12 KiB
L0+L1 混合检索功能规格
功能概述
实现基于 frontmatter 的精确关键词匹配(L0)+ 向量语义检索(L1)的混合检索系统,为 Agent 提供快速精确的知识库查询能力。
Spec 1: Frontmatter 解析
Requirement 1.1: 支持标准 YAML Frontmatter 格式
Given 一个 Markdown 文件包含 frontmatter:
---
title: 支付网关错误码定义
keywords: [ERR_TIMEOUT, 超时, 支付网关]
summary: 记录了支付网关所有核心错误码的含义及排查方向
---
# 正文内容
When 调用 FrontmatterParser.parse(content)
Then 应返回 Frontmatter 对象:
- title = "支付网关错误码定义"
- keywords = ["ERR_TIMEOUT", "超时", "支付网关"]
- summary = "记录了支付网关所有核心错误码的含义及排查方向"
验收标准:
- ✅ 正确解析 title、keywords、summary
- ✅ keywords 支持数组格式
- ✅ 忽略预留字段(sections、category 等)
Requirement 1.2: 处理无 Frontmatter 的文件
Given 一个 Markdown 文件不包含 frontmatter:
# 普通文档
这是正文内容。
When 调用 FrontmatterParser.parse(content)
Then 应返回 null
验收标准:
- ✅ hasFrontmatter() 返回 false
- ✅ parse() 返回 null
- ✅ 不抛出异常
Requirement 1.3: 处理格式错误的 Frontmatter
Given 一个 Markdown 文件包含格式错误的 frontmatter:
---
title: 缺少结束标记
keywords: [ERR_TIMEOUT
# 正文
When 调用 FrontmatterParser.parse(content)
Then 应记录警告日志并返回 null
验收标准:
- ✅ 不抛出异常(优雅降级)
- ✅ 记录 WARN 级别日志
- ✅ 文档仍可上传(只走 L1)
Spec 2: 文档上传增强
Requirement 2.1: 保存原始文件到本地
Given 用户上传文件:
- file: test-doc.md
- category: api
When 调用 DocumentManagementService.uploadDocument(request)
Then 应执行以下步骤:
- ✅ 创建目录:knowledge_base/api/
- ✅ 保存文件:knowledge_base/api/test-doc.md
- ✅ ApiDocument.filePath = "knowledge_base/api/test-doc.md"
验收标准:
- ✅ 文件内容与上传文件一致
- ✅ 目录不存在时自动创建
- ✅ 文件保存失败时抛出异常并回滚事务
Requirement 2.2: 解析并存储 Frontmatter
Given 上传的文件包含 frontmatter
When 调用 DocumentManagementService.uploadDocument(request)
Then 应执行以下步骤:
- ✅ 调用 FrontmatterParser.parse()
- ✅ 将 Frontmatter 对象转为 JSON 字符串
- ✅ 存入 ApiDocument.metadata 字段
验收标准:
- ✅ metadata 字段包含完整 frontmatter JSON
- ✅ 无 frontmatter 时 metadata = null
- ✅ 解析失败时 metadata = null,记录警告
Requirement 2.3: 更新 L0 索引
Given 上传的文件包含有效 frontmatter
When 文档索引成功(status = INDEXED)
Then 应调用 KnowledgeIndexService.addToIndex(entry)
验收标准:
- ✅ KnowledgeEntry 包含正确的 filePath、title、keywords、summary
- ✅ L0 索引立即可用(启动扫描 + 动态添加)
- ✅ 无 frontmatter 的文档不加入 L0 索引
Spec 3: L0 精确匹配
Requirement 3.1: 关键词匹配逻辑
Given L0 索引包含文档:
- keywords: ["ERR_TIMEOUT", "超时", "支付网关"]
Scenario 3.1.1: 完全匹配
- When query = "ERR_TIMEOUT"
- Then 应命中该文档
Scenario 3.1.2: 包含匹配
- When query = "支付网关超时问题"
- Then 应命中该文档(query 包含 "支付网关" 和 "超时")
Scenario 3.1.3: 不区分大小写
- When query = "err_timeout"
- Then 应命中该文档
Scenario 3.1.4: 未匹配
- When query = "限流"
- Then 不应命中该文档
验收标准:
- ✅ 关键词匹配不区分大小写
- ✅ query 包含任一 keyword 即为匹配
- ✅ 支持部分匹配("支付" 匹配 "支付网关")
Requirement 3.2: 返回匹配结果
Given L0 索引包含 3 个文档,query 匹配其中 2 个
When 调用 KnowledgeIndexService.exactMatch(query)
Then 应返回 2 个 KnowledgeEntry
验收标准:
- ✅ 返回所有匹配的文档
- ✅ 按索引顺序返回(启动扫描顺序)
- ✅ 空匹配时返回空列表(不返回 null)
Requirement 3.3: 读取文档内容
Given 文档路径:knowledge_base/api/test-doc.md
When 调用 KnowledgeIndexService.readDocument(filePath, 2000)
Then 应返回文档前 2000 字符
验收标准:
- ✅ 内容 ≤ 2000 字符时返回完整内容
- ✅ 内容 > 2000 字符时返回前 2000 字符 + "..."
- ✅ 文件不存在时记录错误并返回 null
Spec 4: L1 条件调用
Requirement 4.1: 高置信度判断
Scenario 4.1.1: 唯一匹配 = 高置信度
- Given L0 匹配结果: 1 个文档
- When 调用 LookupKnowledgeTool.lookup(query)
- Then highConfidence = true,不调用 L1
Scenario 4.1.2: 多个匹配 = 低置信度
- Given L0 匹配结果: 3 个文档
- When 调用 LookupKnowledgeTool.lookup(query)
- Then highConfidence = false,调用 L1
Scenario 4.1.3: 未匹配 = 低置信度
- Given L0 匹配结果: 0 个文档
- When 调用 LookupKnowledgeTool.lookup(query)
- Then highConfidence = false,调用 L1
验收标准:
- ✅ 唯一匹配时不调用 VectorSearchService
- ✅ 多个匹配或未匹配时调用 VectorSearchService
- ✅ L1 调用参数:topK=3, category=null
Spec 5: 混合检索结果组装
Requirement 5.1: 唯一匹配场景(只返回 L0)
Given L0 唯一匹配
When 调用 LookupKnowledgeTool.lookup("ERR_TIMEOUT")
Then 应返回:
{
"found": true,
"primary": {
"content": "文档前2000字符...",
"source": "knowledge_base/api/payment-errors.md",
"matchType": "exact_L0",
"confidence": "high",
"availableSections": null
},
"supplement": null
}
验收标准:
- ✅ primary 包含 L0 匹配结果
- ✅ supplement = null(未调用 L1)
- ✅ confidence = "high"
Requirement 5.2: 多个匹配场景(L0 + L1)
Given L0 匹配 3 个文档
When 调用 LookupKnowledgeTool.lookup("超时")
Then 应返回:
{
"found": true,
"primary": {
"content": "第一个L0匹配文档...",
"source": "knowledge_base/api/payment-errors.md",
"matchType": "exact_L0",
"confidence": "low",
"availableSections": null
},
"supplement": {
"content": "Milvus语义检索片段...",
"source": "其他文档路径",
"matchType": "semantic_L1"
}
}
验收标准:
- ✅ primary 包含第一个 L0 匹配结果
- ✅ supplement 包含 L1 Top-1 结果
- ✅ confidence = "low"
Requirement 5.3: 未匹配场景(只返回 L1)
Given L0 未匹配(0 个结果)
When 调用 LookupKnowledgeTool.lookup("如何优化性能")
Then 应返回:
{
"found": true,
"primary": null,
"supplement": {
"content": "Milvus语义检索片段...",
"source": "文档路径",
"matchType": "semantic_L1"
}
}
验收标准:
- ✅ primary = null(L0 未命中)
- ✅ supplement 包含 L1 结果
- ✅ found = true(L1 有结果)
Requirement 5.4: 完全未匹配场景
Given L0 和 L1 都未匹配
When 调用 LookupKnowledgeTool.lookup("完全不存在的内容XYZ")
Then 应返回:
{
"found": false,
"primary": null,
"supplement": null
}
验收标准:
- ✅ found = false
- ✅ primary 和 supplement 都为 null
Spec 6: 启动扫描
Requirement 6.1: 递归扫描 knowledge_base/
Given knowledge_base/ 目录结构:
knowledge_base/
├── api/
│ ├── payment.md (有 frontmatter)
│ └── order.md (无 frontmatter)
├── domain/
│ └── cache.md (有 frontmatter)
└── troubleshooting/
└── timeout.md (有 frontmatter)
When 应用启动,执行 KnowledgeIndexService.loadIndex()
Then 应扫描到 4 个 .md 文件,其中 3 个加入 L0 索引
验收标准:
- ✅ 递归扫描所有子目录
- ✅ 只处理 .md 文件
- ✅ 有 frontmatter 的文档加入索引
- ✅ 无 frontmatter 的文档跳过
- ✅ 启动日志显示索引文档数量
Requirement 6.2: 目录不存在时自动创建
Given knowledge_base/ 目录不存在
When 应用启动
Then 应自动创建 knowledge_base/ 目录
验收标准:
- ✅ 目录创建成功
- ✅ 应用正常启动
- ✅ 记录 INFO 日志
Requirement 6.3: 启动扫描性能
Given knowledge_base/ 包含 500 个文档
When 应用启动
Then 启动扫描应在 5 秒内完成
验收标准:
- ✅ 启动扫描时间 < 5s
- ✅ 不阻塞应用启动
- ✅ 使用 @PostConstruct 异步加载
Spec 7: Agent 工具集成
Requirement 7.1: 工具注册
Given LookupKnowledgeTool 使用 @Tool 注解
When Agent Framework 初始化
Then lookup_knowledge 应自动注册为可用工具
验收标准:
- ✅ 工具名称:lookup_knowledge
- ✅ 工具描述清晰(优先精确匹配,自动补充语义)
- ✅ 参数定义:query (必填), section_title (可选)
Requirement 7.2: Agent 调用场景
Scenario 7.2.1: Agent 查询错误码
- Given Agent 诊断时发现错误码 "ERR_TIMEOUT"
- When Agent 调用 lookup_knowledge("ERR_TIMEOUT")
- Then 返回错误码定义文档(L0 精确匹配)
Scenario 7.2.2: Agent 查询开放问题
- Given Agent 需要了解"缓存优化"
- When Agent 调用 lookup_knowledge("如何优化缓存")
- Then 返回语义相关文档(L1 检索)
验收标准:
- ✅ Agent 可以成功调用工具
- ✅ 返回结果符合 Agent 预期格式
- ✅ 工具调用记录到 ToolCall
Spec 8: 文档删除
Requirement 8.1: 同步删除 L0 索引
Given 文档已加入 L0 索引
When 调用 DocumentManagementService.deleteDocument(docId)
Then 应同步删除:
- ✅ 本地文件(knowledge_base/{category}/{fileName})
- ✅ L0 索引条目
- ✅ MySQL 元数据(ApiDocument)
- ✅ Milvus 向量索引
验收标准:
- ✅ 删除后 L0 查询不再返回该文档
- ✅ 删除后 L1 查询不再返回该文档
- ✅ 本地文件被删除
Spec 9: 配置管理
Requirement 9.1: knowledge.base-path 配置
Given application.yml 配置:
knowledge:
base-path: /data/knowledge_base/
When KnowledgeIndexService 初始化
Then 应使用配置的路径
验收标准:
- ✅ 支持绝对路径
- ✅ 支持相对路径(相对于应用根目录)
- ✅ 未配置时使用默认值:knowledge_base/
非功能性规格
性能要求
- L0 查询响应时间:< 10ms(99th percentile)
- L0 + L1 组合查询:< 500ms(99th percentile)
- 启动扫描时间:< 5s(1000 个文档)
- 内存占用:< 10MB(1000 个文档)
可用性要求
- L0 索引加载失败不影响应用启动(降级到 L1)
- frontmatter 解析失败不影响文档上传
- L1 调用失败时返回 L0 结果
可观测性要求
- 启动扫描:INFO 日志记录文档数量
- L0 匹配:DEBUG 日志记录匹配结果
- L1 条件调用:DEBUG 日志记录调用决策
- 错误场景:ERROR/WARN 日志记录详细信息
边界与限制
MVP 不支持
- ❌ sections 分段加载(availableSections 返回 null)
- ❌ watchdog 热更新(重启生效)
- ❌ L0 索引持久化(内存索引)
- ❌ 模糊匹配 / 同义词扩展
文件格式限制
- ✅ 仅支持 .md 文件
- ❌ 不支持 .txt、.docx、.pdf
索引规模限制
- ⚠️ MVP 推荐 < 1000 个文档
- ⚠️ 超过限制可能导致启动慢或内存占用高