# L0+L1 混合检索功能规格 ## 功能概述 实现基于 frontmatter 的精确关键词匹配(L0)+ 向量语义检索(L1)的混合检索系统,为 Agent 提供快速精确的知识库查询能力。 --- ## Spec 1: Frontmatter 解析 ### Requirement 1.1: 支持标准 YAML Frontmatter 格式 **Given** 一个 Markdown 文件包含 frontmatter: ```markdown --- title: 支付网关错误码定义 keywords: [ERR_TIMEOUT, 超时, 支付网关] summary: 记录了支付网关所有核心错误码的含义及排查方向 --- # 正文内容 ``` **When** 调用 FrontmatterParser.parse(content) **Then** 应返回 Frontmatter 对象: - title = "支付网关错误码定义" - keywords = ["ERR_TIMEOUT", "超时", "支付网关"] - summary = "记录了支付网关所有核心错误码的含义及排查方向" **验收标准**: - ✅ 正确解析 title、keywords、summary - ✅ keywords 支持数组格式 - ✅ 忽略预留字段(sections、category 等) --- ### Requirement 1.2: 处理无 Frontmatter 的文件 **Given** 一个 Markdown 文件不包含 frontmatter: ```markdown # 普通文档 这是正文内容。 ``` **When** 调用 FrontmatterParser.parse(content) **Then** 应返回 null **验收标准**: - ✅ hasFrontmatter() 返回 false - ✅ parse() 返回 null - ✅ 不抛出异常 --- ### Requirement 1.3: 处理格式错误的 Frontmatter **Given** 一个 Markdown 文件包含格式错误的 frontmatter: ```markdown --- title: 缺少结束标记 keywords: [ERR_TIMEOUT # 正文 ``` **When** 调用 FrontmatterParser.parse(content) **Then** 应记录警告日志并返回 null **验收标准**: - ✅ 不抛出异常(优雅降级) - ✅ 记录 WARN 级别日志 - ✅ 文档仍可上传(只走 L1) --- ## Spec 2: 文档上传增强 ### Requirement 2.1: 保存原始文件到本地 **Given** 用户上传文件: - file: test-doc.md - category: api **When** 调用 DocumentManagementService.uploadDocument(request) **Then** 应执行以下步骤: 1. ✅ 创建目录:knowledge_base/api/ 2. ✅ 保存文件:knowledge_base/api/test-doc.md 3. ✅ ApiDocument.filePath = "knowledge_base/api/test-doc.md" **验收标准**: - ✅ 文件内容与上传文件一致 - ✅ 目录不存在时自动创建 - ✅ 文件保存失败时抛出异常并回滚事务 --- ### Requirement 2.2: 解析并存储 Frontmatter **Given** 上传的文件包含 frontmatter **When** 调用 DocumentManagementService.uploadDocument(request) **Then** 应执行以下步骤: 1. ✅ 调用 FrontmatterParser.parse() 2. ✅ 将 Frontmatter 对象转为 JSON 字符串 3. ✅ 存入 ApiDocument.metadata 字段 **验收标准**: - ✅ metadata 字段包含完整 frontmatter JSON - ✅ 无 frontmatter 时 metadata = null - ✅ 解析失败时 metadata = null,记录警告 --- ### Requirement 2.3: 更新 L0 索引 **Given** 上传的文件包含有效 frontmatter **When** 文档索引成功(status = INDEXED) **Then** 应调用 KnowledgeIndexService.addToIndex(entry) **验收标准**: - ✅ KnowledgeEntry 包含正确的 filePath、title、keywords、summary - ✅ L0 索引立即可用(启动扫描 + 动态添加) - ✅ 无 frontmatter 的文档不加入 L0 索引 --- ## Spec 3: L0 精确匹配 ### Requirement 3.1: 关键词匹配逻辑 **Given** L0 索引包含文档: - keywords: ["ERR_TIMEOUT", "超时", "支付网关"] **Scenario 3.1.1: 完全匹配** - **When** query = "ERR_TIMEOUT" - **Then** 应命中该文档 **Scenario 3.1.2: 包含匹配** - **When** query = "支付网关超时问题" - **Then** 应命中该文档(query 包含 "支付网关" 和 "超时") **Scenario 3.1.3: 不区分大小写** - **When** query = "err_timeout" - **Then** 应命中该文档 **Scenario 3.1.4: 未匹配** - **When** query = "限流" - **Then** 不应命中该文档 **验收标准**: - ✅ 关键词匹配不区分大小写 - ✅ query 包含任一 keyword 即为匹配 - ✅ 支持部分匹配("支付" 匹配 "支付网关") --- ### Requirement 3.2: 返回匹配结果 **Given** L0 索引包含 3 个文档,query 匹配其中 2 个 **When** 调用 KnowledgeIndexService.exactMatch(query) **Then** 应返回 2 个 KnowledgeEntry **验收标准**: - ✅ 返回所有匹配的文档 - ✅ 按索引顺序返回(启动扫描顺序) - ✅ 空匹配时返回空列表(不返回 null) --- ### Requirement 3.3: 读取文档内容 **Given** 文档路径:knowledge_base/api/test-doc.md **When** 调用 KnowledgeIndexService.readDocument(filePath, 2000) **Then** 应返回文档前 2000 字符 **验收标准**: - ✅ 内容 ≤ 2000 字符时返回完整内容 - ✅ 内容 > 2000 字符时返回前 2000 字符 + "..." - ✅ 文件不存在时记录错误并返回 null --- ## Spec 4: L1 条件调用 ### Requirement 4.1: 高置信度判断 **Scenario 4.1.1: 唯一匹配 = 高置信度** - **Given** L0 匹配结果: 1 个文档 - **When** 调用 LookupKnowledgeTool.lookup(query) - **Then** highConfidence = true,不调用 L1 **Scenario 4.1.2: 多个匹配 = 低置信度** - **Given** L0 匹配结果: 3 个文档 - **When** 调用 LookupKnowledgeTool.lookup(query) - **Then** highConfidence = false,调用 L1 **Scenario 4.1.3: 未匹配 = 低置信度** - **Given** L0 匹配结果: 0 个文档 - **When** 调用 LookupKnowledgeTool.lookup(query) - **Then** highConfidence = false,调用 L1 **验收标准**: - ✅ 唯一匹配时不调用 VectorSearchService - ✅ 多个匹配或未匹配时调用 VectorSearchService - ✅ L1 调用参数:topK=3, category=null --- ## Spec 5: 混合检索结果组装 ### Requirement 5.1: 唯一匹配场景(只返回 L0) **Given** L0 唯一匹配 **When** 调用 LookupKnowledgeTool.lookup("ERR_TIMEOUT") **Then** 应返回: ```json { "found": true, "primary": { "content": "文档前2000字符...", "source": "knowledge_base/api/payment-errors.md", "matchType": "exact_L0", "confidence": "high", "availableSections": null }, "supplement": null } ``` **验收标准**: - ✅ primary 包含 L0 匹配结果 - ✅ supplement = null(未调用 L1) - ✅ confidence = "high" --- ### Requirement 5.2: 多个匹配场景(L0 + L1) **Given** L0 匹配 3 个文档 **When** 调用 LookupKnowledgeTool.lookup("超时") **Then** 应返回: ```json { "found": true, "primary": { "content": "第一个L0匹配文档...", "source": "knowledge_base/api/payment-errors.md", "matchType": "exact_L0", "confidence": "low", "availableSections": null }, "supplement": { "content": "Milvus语义检索片段...", "source": "其他文档路径", "matchType": "semantic_L1" } } ``` **验收标准**: - ✅ primary 包含第一个 L0 匹配结果 - ✅ supplement 包含 L1 Top-1 结果 - ✅ confidence = "low" --- ### Requirement 5.3: 未匹配场景(只返回 L1) **Given** L0 未匹配(0 个结果) **When** 调用 LookupKnowledgeTool.lookup("如何优化性能") **Then** 应返回: ```json { "found": true, "primary": null, "supplement": { "content": "Milvus语义检索片段...", "source": "文档路径", "matchType": "semantic_L1" } } ``` **验收标准**: - ✅ primary = null(L0 未命中) - ✅ supplement 包含 L1 结果 - ✅ found = true(L1 有结果) --- ### Requirement 5.4: 完全未匹配场景 **Given** L0 和 L1 都未匹配 **When** 调用 LookupKnowledgeTool.lookup("完全不存在的内容XYZ") **Then** 应返回: ```json { "found": false, "primary": null, "supplement": null } ``` **验收标准**: - ✅ found = false - ✅ primary 和 supplement 都为 null --- ## Spec 6: 启动扫描 ### Requirement 6.1: 递归扫描 knowledge_base/ **Given** knowledge_base/ 目录结构: ``` knowledge_base/ ├── api/ │ ├── payment.md (有 frontmatter) │ └── order.md (无 frontmatter) ├── domain/ │ └── cache.md (有 frontmatter) └── troubleshooting/ └── timeout.md (有 frontmatter) ``` **When** 应用启动,执行 KnowledgeIndexService.loadIndex() **Then** 应扫描到 4 个 .md 文件,其中 3 个加入 L0 索引 **验收标准**: - ✅ 递归扫描所有子目录 - ✅ 只处理 .md 文件 - ✅ 有 frontmatter 的文档加入索引 - ✅ 无 frontmatter 的文档跳过 - ✅ 启动日志显示索引文档数量 --- ### Requirement 6.2: 目录不存在时自动创建 **Given** knowledge_base/ 目录不存在 **When** 应用启动 **Then** 应自动创建 knowledge_base/ 目录 **验收标准**: - ✅ 目录创建成功 - ✅ 应用正常启动 - ✅ 记录 INFO 日志 --- ### Requirement 6.3: 启动扫描性能 **Given** knowledge_base/ 包含 500 个文档 **When** 应用启动 **Then** 启动扫描应在 5 秒内完成 **验收标准**: - ✅ 启动扫描时间 < 5s - ✅ 不阻塞应用启动 - ✅ 使用 @PostConstruct 异步加载 --- ## Spec 7: Agent 工具集成 ### Requirement 7.1: 工具注册 **Given** LookupKnowledgeTool 使用 @Tool 注解 **When** Agent Framework 初始化 **Then** lookup_knowledge 应自动注册为可用工具 **验收标准**: - ✅ 工具名称:lookup_knowledge - ✅ 工具描述清晰(优先精确匹配,自动补充语义) - ✅ 参数定义:query (必填), section_title (可选) --- ### Requirement 7.2: Agent 调用场景 **Scenario 7.2.1: Agent 查询错误码** - **Given** Agent 诊断时发现错误码 "ERR_TIMEOUT" - **When** Agent 调用 lookup_knowledge("ERR_TIMEOUT") - **Then** 返回错误码定义文档(L0 精确匹配) **Scenario 7.2.2: Agent 查询开放问题** - **Given** Agent 需要了解"缓存优化" - **When** Agent 调用 lookup_knowledge("如何优化缓存") - **Then** 返回语义相关文档(L1 检索) **验收标准**: - ✅ Agent 可以成功调用工具 - ✅ 返回结果符合 Agent 预期格式 - ✅ 工具调用记录到 ToolCall --- ## Spec 8: 文档删除 ### Requirement 8.1: 同步删除 L0 索引 **Given** 文档已加入 L0 索引 **When** 调用 DocumentManagementService.deleteDocument(docId) **Then** 应同步删除: 1. ✅ 本地文件(knowledge_base/{category}/{fileName}) 2. ✅ L0 索引条目 3. ✅ MySQL 元数据(ApiDocument) 4. ✅ Milvus 向量索引 **验收标准**: - ✅ 删除后 L0 查询不再返回该文档 - ✅ 删除后 L1 查询不再返回该文档 - ✅ 本地文件被删除 --- ## Spec 9: 配置管理 ### Requirement 9.1: knowledge.base-path 配置 **Given** application.yml 配置: ```yaml knowledge: base-path: /data/knowledge_base/ ``` **When** KnowledgeIndexService 初始化 **Then** 应使用配置的路径 **验收标准**: - ✅ 支持绝对路径 - ✅ 支持相对路径(相对于应用根目录) - ✅ 未配置时使用默认值:knowledge_base/ --- ## 非功能性规格 ### 性能要求 - L0 查询响应时间:< 10ms(99th percentile) - L0 + L1 组合查询:< 500ms(99th percentile) - 启动扫描时间:< 5s(1000 个文档) - 内存占用:< 10MB(1000 个文档) ### 可用性要求 - L0 索引加载失败不影响应用启动(降级到 L1) - frontmatter 解析失败不影响文档上传 - L1 调用失败时返回 L0 结果 ### 可观测性要求 - 启动扫描:INFO 日志记录文档数量 - L0 匹配:DEBUG 日志记录匹配结果 - L1 条件调用:DEBUG 日志记录调用决策 - 错误场景:ERROR/WARN 日志记录详细信息 --- ## 边界与限制 ### MVP 不支持 - ❌ sections 分段加载(availableSections 返回 null) - ❌ watchdog 热更新(重启生效) - ❌ L0 索引持久化(内存索引) - ❌ 模糊匹配 / 同义词扩展 ### 文件格式限制 - ✅ 仅支持 .md 文件 - ❌ 不支持 .txt、.docx、.pdf ### 索引规模限制 - ⚠️ MVP 推荐 < 1000 个文档 - ⚠️ 超过限制可能导致启动慢或内存占用高