Files
SuperBizAgent-java/openspec/changes/lookup-knowledge-integration/specs/functional-specs.md
T
zhuyongxin d6229f3385 feat(knowledge): 完成 L0+L1 混合检索集成
核心功能:
- 新增 FrontmatterParser 解析 YAML frontmatter
- 新增 KnowledgeIndexService L0 内存索引
- 新增 LookupKnowledgeTool 混合检索工具
- 增强 DocumentManagementService 文件保存和索引同步

技术实现:
- 数据库迁移 V004: api_document.metadata (TEXT)
- 依赖新增: snakeyaml 2.0
- 配置新增: knowledge.base-path
- 可观测性: requestId 追踪 + 性能日志

质量保证:
- 单元测试: 31/31 通过
- 测试覆盖: FrontmatterParser(11), KnowledgeIndexService(13), LookupKnowledgeTool(7)
- 启动验证: L0 索引正常加载

归档文档:
- OpenSpec: openspec/changes/lookup-knowledge-integration/
- devflow 档案: devflow/projects/2026-06-24-lookup-knowledge-integration/
- handoff: handoff/2026-06-24-lookup-knowledge-integration.md
2026-06-24 16:07:10 +08:00

502 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# L0+L1 混合检索功能规格
## 功能概述
实现基于 frontmatter 的精确关键词匹配(L0)+ 向量语义检索(L1)的混合检索系统,为 Agent 提供快速精确的知识库查询能力。
---
## Spec 1: Frontmatter 解析
### Requirement 1.1: 支持标准 YAML Frontmatter 格式
**Given** 一个 Markdown 文件包含 frontmatter:
```markdown
---
title: 支付网关错误码定义
keywords: [ERR_TIMEOUT, 超时, 支付网关]
summary: 记录了支付网关所有核心错误码的含义及排查方向
---
# 正文内容
```
**When** 调用 FrontmatterParser.parse(content)
**Then** 应返回 Frontmatter 对象:
- title = "支付网关错误码定义"
- keywords = ["ERR_TIMEOUT", "超时", "支付网关"]
- summary = "记录了支付网关所有核心错误码的含义及排查方向"
**验收标准**:
- ✅ 正确解析 title、keywords、summary
- ✅ keywords 支持数组格式
- ✅ 忽略预留字段(sections、category 等)
---
### Requirement 1.2: 处理无 Frontmatter 的文件
**Given** 一个 Markdown 文件不包含 frontmatter:
```markdown
# 普通文档
这是正文内容。
```
**When** 调用 FrontmatterParser.parse(content)
**Then** 应返回 null
**验收标准**:
- ✅ hasFrontmatter() 返回 false
- ✅ parse() 返回 null
- ✅ 不抛出异常
---
### Requirement 1.3: 处理格式错误的 Frontmatter
**Given** 一个 Markdown 文件包含格式错误的 frontmatter:
```markdown
---
title: 缺少结束标记
keywords: [ERR_TIMEOUT
# 正文
```
**When** 调用 FrontmatterParser.parse(content)
**Then** 应记录警告日志并返回 null
**验收标准**:
- ✅ 不抛出异常(优雅降级)
- ✅ 记录 WARN 级别日志
- ✅ 文档仍可上传(只走 L1)
---
## Spec 2: 文档上传增强
### Requirement 2.1: 保存原始文件到本地
**Given** 用户上传文件:
- file: test-doc.md
- category: api
**When** 调用 DocumentManagementService.uploadDocument(request)
**Then** 应执行以下步骤:
1. ✅ 创建目录:knowledge_base/api/
2. ✅ 保存文件:knowledge_base/api/test-doc.md
3. ✅ ApiDocument.filePath = "knowledge_base/api/test-doc.md"
**验收标准**:
- ✅ 文件内容与上传文件一致
- ✅ 目录不存在时自动创建
- ✅ 文件保存失败时抛出异常并回滚事务
---
### Requirement 2.2: 解析并存储 Frontmatter
**Given** 上传的文件包含 frontmatter
**When** 调用 DocumentManagementService.uploadDocument(request)
**Then** 应执行以下步骤:
1. ✅ 调用 FrontmatterParser.parse()
2. ✅ 将 Frontmatter 对象转为 JSON 字符串
3. ✅ 存入 ApiDocument.metadata 字段
**验收标准**:
- ✅ metadata 字段包含完整 frontmatter JSON
- ✅ 无 frontmatter 时 metadata = null
- ✅ 解析失败时 metadata = null,记录警告
---
### Requirement 2.3: 更新 L0 索引
**Given** 上传的文件包含有效 frontmatter
**When** 文档索引成功(status = INDEXED)
**Then** 应调用 KnowledgeIndexService.addToIndex(entry)
**验收标准**:
- ✅ KnowledgeEntry 包含正确的 filePath、title、keywords、summary
- ✅ L0 索引立即可用(启动扫描 + 动态添加)
- ✅ 无 frontmatter 的文档不加入 L0 索引
---
## Spec 3: L0 精确匹配
### Requirement 3.1: 关键词匹配逻辑
**Given** L0 索引包含文档:
- keywords: ["ERR_TIMEOUT", "超时", "支付网关"]
**Scenario 3.1.1: 完全匹配**
- **When** query = "ERR_TIMEOUT"
- **Then** 应命中该文档
**Scenario 3.1.2: 包含匹配**
- **When** query = "支付网关超时问题"
- **Then** 应命中该文档(query 包含 "支付网关" 和 "超时")
**Scenario 3.1.3: 不区分大小写**
- **When** query = "err_timeout"
- **Then** 应命中该文档
**Scenario 3.1.4: 未匹配**
- **When** query = "限流"
- **Then** 不应命中该文档
**验收标准**:
- ✅ 关键词匹配不区分大小写
- ✅ query 包含任一 keyword 即为匹配
- ✅ 支持部分匹配("支付" 匹配 "支付网关")
---
### Requirement 3.2: 返回匹配结果
**Given** L0 索引包含 3 个文档,query 匹配其中 2 个
**When** 调用 KnowledgeIndexService.exactMatch(query)
**Then** 应返回 2 个 KnowledgeEntry
**验收标准**:
- ✅ 返回所有匹配的文档
- ✅ 按索引顺序返回(启动扫描顺序)
- ✅ 空匹配时返回空列表(不返回 null)
---
### Requirement 3.3: 读取文档内容
**Given** 文档路径:knowledge_base/api/test-doc.md
**When** 调用 KnowledgeIndexService.readDocument(filePath, 2000)
**Then** 应返回文档前 2000 字符
**验收标准**:
- ✅ 内容 ≤ 2000 字符时返回完整内容
- ✅ 内容 > 2000 字符时返回前 2000 字符 + "..."
- ✅ 文件不存在时记录错误并返回 null
---
## Spec 4: L1 条件调用
### Requirement 4.1: 高置信度判断
**Scenario 4.1.1: 唯一匹配 = 高置信度**
- **Given** L0 匹配结果: 1 个文档
- **When** 调用 LookupKnowledgeTool.lookup(query)
- **Then** highConfidence = true,不调用 L1
**Scenario 4.1.2: 多个匹配 = 低置信度**
- **Given** L0 匹配结果: 3 个文档
- **When** 调用 LookupKnowledgeTool.lookup(query)
- **Then** highConfidence = false,调用 L1
**Scenario 4.1.3: 未匹配 = 低置信度**
- **Given** L0 匹配结果: 0 个文档
- **When** 调用 LookupKnowledgeTool.lookup(query)
- **Then** highConfidence = false,调用 L1
**验收标准**:
- ✅ 唯一匹配时不调用 VectorSearchService
- ✅ 多个匹配或未匹配时调用 VectorSearchService
- ✅ L1 调用参数:topK=3, category=null
---
## Spec 5: 混合检索结果组装
### Requirement 5.1: 唯一匹配场景(只返回 L0)
**Given** L0 唯一匹配
**When** 调用 LookupKnowledgeTool.lookup("ERR_TIMEOUT")
**Then** 应返回:
```json
{
"found": true,
"primary": {
"content": "文档前2000字符...",
"source": "knowledge_base/api/payment-errors.md",
"matchType": "exact_L0",
"confidence": "high",
"availableSections": null
},
"supplement": null
}
```
**验收标准**:
- ✅ primary 包含 L0 匹配结果
- ✅ supplement = null(未调用 L1)
- ✅ confidence = "high"
---
### Requirement 5.2: 多个匹配场景(L0 + L1)
**Given** L0 匹配 3 个文档
**When** 调用 LookupKnowledgeTool.lookup("超时")
**Then** 应返回:
```json
{
"found": true,
"primary": {
"content": "第一个L0匹配文档...",
"source": "knowledge_base/api/payment-errors.md",
"matchType": "exact_L0",
"confidence": "low",
"availableSections": null
},
"supplement": {
"content": "Milvus语义检索片段...",
"source": "其他文档路径",
"matchType": "semantic_L1"
}
}
```
**验收标准**:
- ✅ primary 包含第一个 L0 匹配结果
- ✅ supplement 包含 L1 Top-1 结果
- ✅ confidence = "low"
---
### Requirement 5.3: 未匹配场景(只返回 L1)
**Given** L0 未匹配(0 个结果)
**When** 调用 LookupKnowledgeTool.lookup("如何优化性能")
**Then** 应返回:
```json
{
"found": true,
"primary": null,
"supplement": {
"content": "Milvus语义检索片段...",
"source": "文档路径",
"matchType": "semantic_L1"
}
}
```
**验收标准**:
- ✅ primary = null(L0 未命中)
- ✅ supplement 包含 L1 结果
- ✅ found = true(L1 有结果)
---
### Requirement 5.4: 完全未匹配场景
**Given** L0 和 L1 都未匹配
**When** 调用 LookupKnowledgeTool.lookup("完全不存在的内容XYZ")
**Then** 应返回:
```json
{
"found": false,
"primary": null,
"supplement": null
}
```
**验收标准**:
- ✅ found = false
- ✅ primary 和 supplement 都为 null
---
## Spec 6: 启动扫描
### Requirement 6.1: 递归扫描 knowledge_base/
**Given** knowledge_base/ 目录结构:
```
knowledge_base/
├── api/
│ ├── payment.md (有 frontmatter)
│ └── order.md (无 frontmatter)
├── domain/
│ └── cache.md (有 frontmatter)
└── troubleshooting/
└── timeout.md (有 frontmatter)
```
**When** 应用启动,执行 KnowledgeIndexService.loadIndex()
**Then** 应扫描到 4 个 .md 文件,其中 3 个加入 L0 索引
**验收标准**:
- ✅ 递归扫描所有子目录
- ✅ 只处理 .md 文件
- ✅ 有 frontmatter 的文档加入索引
- ✅ 无 frontmatter 的文档跳过
- ✅ 启动日志显示索引文档数量
---
### Requirement 6.2: 目录不存在时自动创建
**Given** knowledge_base/ 目录不存在
**When** 应用启动
**Then** 应自动创建 knowledge_base/ 目录
**验收标准**:
- ✅ 目录创建成功
- ✅ 应用正常启动
- ✅ 记录 INFO 日志
---
### Requirement 6.3: 启动扫描性能
**Given** knowledge_base/ 包含 500 个文档
**When** 应用启动
**Then** 启动扫描应在 5 秒内完成
**验收标准**:
- ✅ 启动扫描时间 < 5s
- ✅ 不阻塞应用启动
- ✅ 使用 @PostConstruct 异步加载
---
## Spec 7: Agent 工具集成
### Requirement 7.1: 工具注册
**Given** LookupKnowledgeTool 使用 @Tool 注解
**When** Agent Framework 初始化
**Then** lookup_knowledge 应自动注册为可用工具
**验收标准**:
- ✅ 工具名称:lookup_knowledge
- ✅ 工具描述清晰(优先精确匹配,自动补充语义)
- ✅ 参数定义:query (必填), section_title (可选)
---
### Requirement 7.2: Agent 调用场景
**Scenario 7.2.1: Agent 查询错误码**
- **Given** Agent 诊断时发现错误码 "ERR_TIMEOUT"
- **When** Agent 调用 lookup_knowledge("ERR_TIMEOUT")
- **Then** 返回错误码定义文档(L0 精确匹配)
**Scenario 7.2.2: Agent 查询开放问题**
- **Given** Agent 需要了解"缓存优化"
- **When** Agent 调用 lookup_knowledge("如何优化缓存")
- **Then** 返回语义相关文档(L1 检索)
**验收标准**:
- ✅ Agent 可以成功调用工具
- ✅ 返回结果符合 Agent 预期格式
- ✅ 工具调用记录到 ToolCall
---
## Spec 8: 文档删除
### Requirement 8.1: 同步删除 L0 索引
**Given** 文档已加入 L0 索引
**When** 调用 DocumentManagementService.deleteDocument(docId)
**Then** 应同步删除:
1. ✅ 本地文件(knowledge_base/{category}/{fileName})
2. ✅ L0 索引条目
3. ✅ MySQL 元数据(ApiDocument)
4. ✅ Milvus 向量索引
**验收标准**:
- ✅ 删除后 L0 查询不再返回该文档
- ✅ 删除后 L1 查询不再返回该文档
- ✅ 本地文件被删除
---
## Spec 9: 配置管理
### Requirement 9.1: knowledge.base-path 配置
**Given** application.yml 配置:
```yaml
knowledge:
base-path: /data/knowledge_base/
```
**When** KnowledgeIndexService 初始化
**Then** 应使用配置的路径
**验收标准**:
- ✅ 支持绝对路径
- ✅ 支持相对路径(相对于应用根目录)
- ✅ 未配置时使用默认值:knowledge_base/
---
## 非功能性规格
### 性能要求
- L0 查询响应时间:< 10ms(99th percentile)
- L0 + L1 组合查询:< 500ms(99th percentile)
- 启动扫描时间:< 5s(1000 个文档)
- 内存占用:< 10MB(1000 个文档)
### 可用性要求
- L0 索引加载失败不影响应用启动(降级到 L1)
- frontmatter 解析失败不影响文档上传
- L1 调用失败时返回 L0 结果
### 可观测性要求
- 启动扫描:INFO 日志记录文档数量
- L0 匹配:DEBUG 日志记录匹配结果
- L1 条件调用:DEBUG 日志记录调用决策
- 错误场景:ERROR/WARN 日志记录详细信息
---
## 边界与限制
### MVP 不支持
- ❌ sections 分段加载(availableSections 返回 null)
- ❌ watchdog 热更新(重启生效)
- ❌ L0 索引持久化(内存索引)
- ❌ 模糊匹配 / 同义词扩展
### 文件格式限制
- ✅ 仅支持 .md 文件
- ❌ 不支持 .txt、.docx、.pdf
### 索引规模限制
- ⚠️ MVP 推荐 < 1000 个文档
- ⚠️ 超过限制可能导致启动慢或内存占用高