Files
SuperBizAgent-java/openspec/changes/archive/2026-07-05-lookup-knowledge-integration/proposal.md
T

9.0 KiB
Raw Blame History

Proposal: L0+L1 混合检索集成

问题

当前只有 L1 向量语义检索(Milvus + BGE-M3),在遇到精确关键词查询时(如错误码 "ERR_TIMEOUT"、接口名 "PaymentGateway")效率不够高:

  • 需要调用 embedding API 生成向量(约 100-300ms)
  • 语义检索返回相似但可能不精确的结果
  • 无法快速定位已知关键词对应的完整文档

Agent 需要一个"先精确、后语义"的混合检索工具。

建议方案

架构设计:双层检索

lookup_knowledge(query)
    ↓
L0: 精确关键词匹配(内存索引,< 10ms)
    ├─ 匹配成功 + 唯一结果 → 返回完整文档(高置信度)
    └─ 未匹配 或 多个匹配 ↓
L1: 向量语义检索(Milvus,补充上下文)
    └─ 返回 Top-K 相似片段

核心机制:

  1. L0 索引:启动时扫描 knowledge_base/ 目录,解析 Markdown frontmatter,构建内存索引
  2. L1 复用:调用现有 VectorSearchService.searchSimilarDocuments()
  3. 条件调用:L0 唯一匹配时不调用 L1(减少延迟)

1. Frontmatter 规范

所有知识库文档(knowledge_base/ 目录)需在文件头添加 YAML frontmatter:

---
title: 支付网关错误码定义              # 必填
keywords: [ERR_TIMEOUT, 超时, 支付网关]  # 必填,用于精确匹配
summary: 记录了支付网关所有核心错误码的含义及排查方向  # 必填
category: api                          # 可选,与现有 category 对齐
sections:                              # 预留字段(MVP 不实现)
  超时排查: "## 1. 超时类错误"
---

# 文档正文
...

约束:

  • frontmatter 必须在文件最顶部(前面不能有空行)
  • title, keywords, summary 为必填字段
  • 缺少 frontmatter 的文档允许上传,但不参与 L0 索引(只走 L1)

2. 上传流程增强

现有流程:

POST /api/documents/upload
  ↓
DocumentManagementService.uploadDocument()
  ↓
文本提取 → 分块 → 向量化 → Milvus 索引
  ↓
元数据存 MySQL (ApiDocument)

增强后流程:

POST /api/documents/upload
  ↓
1. 文本提取(内存)
2. 保存原始文件到:knowledge_base/{category}/{fileName}
3. 解析 frontmatter(FrontmatterParser)
4. 分块 → 向量化 → Milvus 索引
5. 元数据存 MySQL(ApiDocument.metadata 存储 frontmatter JSON)
6. 更新 L0 内存索引(KnowledgeIndexService)

关键决策(grill 阶段确认):

  • ✅ 保存原始文件到本地(支持 L0 完整读取 + 未来扩展)
  • ✅ metadata 字段:TEXT 类型存储 JSON 字符串
  • ✅ ApiDocument.filePath 存储本地文件路径
  • ✅ L0 高置信度 = 唯一匹配(不调用 L1)
  • 按 category 分类存储:knowledge_base/api/, knowledge_base/domain/, knowledge_base/troubleshooting/

3. L0 索引服务

KnowledgeIndexService:

@Service
public class KnowledgeIndexService {
    // 内存索引结构
    private List<KnowledgeEntry> knowledgeIndex = new ArrayList<>();
    
    // 启动时扫描
    @PostConstruct
    public void loadIndex() {
        // 递归扫描 knowledge_base/
        // 解析 frontmatter
        // 构建内存索引
    }
    
    // L0 精确匹配
    public List<KnowledgeEntry> exactMatch(String query) {
        // 关键词匹配(不区分大小写)
        // 匹配规则:query 包含 keywords 中的任一词
    }
    
    // 读取文档内容
    public String readDocument(String filePath, int maxChars) {
        // 读取文件,返回前 maxChars 字符
    }
}

数据结构:

@Data
public class KnowledgeEntry {
    private String filePath;          // knowledge_base/api/payment-errors.md
    private String title;             // 支付网关错误码定义
    private List<String> keywords;    // [ERR_TIMEOUT, 超时, 支付网关]
    private String summary;           // 一句话摘要
    private String category;          // api
    private Map<String, String> sections;  // 预留字段
}

4. L1 复用

直接调用现有服务:

@Autowired
private VectorSearchService vectorSearchService;

List<VectorSearchService.SearchResult> l1Results = 
    vectorSearchService.searchSimilarDocuments(query, 3, category);

5. 混合检索工具

LookupKnowledgeTool(供 Agent 调用):

@Tool(name = "lookup_knowledge", 
      description = "查询知识库文档。优先精确匹配,自动补充语义相关片段。")
public LookupResult lookup(
    @P("query") String query,
    @P("section_title") String sectionTitle  // 预留参数,MVP 不实现
) {
    // Step 1: L0 精确匹配
    List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
    
    // Step 2: 判断是否高置信度(唯一匹配)
    boolean highConfidence = (l0Matches.size() == 1);
    
    // Step 3: L1 条件调用
    List<SearchResult> l1Results = null;
    if (!highConfidence) {
        l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
    }
    
    // Step 4: 组装结果
    return buildResult(l0Matches, l1Results, highConfidence);
}

返回格式:

{
    "found": true,
    "primary": {
        "content": "文档前2000字符...",
        "source": "knowledge_base/api/payment-errors.md",
        "matchType": "exact_L0",
        "confidence": "high",
        "availableSections": null
    },
    "supplement": {
        "content": "Milvus检索到的相关片段...",
        "source": "其他文档路径",
        "matchType": "semantic_L1"
    }
}

范围

核心功能(MVP)

  1. ✅ FrontmatterParser:解析 YAML frontmatter(使用 snakeyaml)
  2. ✅ KnowledgeIndexService:启动扫描 + 内存索引 + L0 精确匹配
  3. ✅ 上传流程增强:保存本地 + 解析 frontmatter + 更新 L0 索引
  4. ✅ LookupKnowledgeTool:L0 + L1 混合检索 + 条件调用
  5. ✅ ApiDocument.metadata 字段扩展(存储 frontmatter JSON)

预留但不实现

  • ⏸️ sections 分段加载(availableSections 返回 null)
  • ⏸️ watchdog 热更新(重启生效)
  • ⏸️ L0 索引持久化(内存索引,启动扫描)

非目标

  • 不修改现有 VectorSearchService 逻辑
  • 不修改 Milvus 索引结构
  • 不实现文档版本管理
  • 不支持其他文件格式(仅 .md)

技术选型

组件 技术选型 说明
YAML 解析 snakeyaml 2.0 解析 frontmatter
L0 索引 内存 List<KnowledgeEntry> 启动扫描,快速查询
L1 检索 复用 VectorSearchService Milvus + BGE-M3
文件存储 本地文件系统 knowledge_base/{category}/

devflow 上下文约束

必须遵守(来自 phase1-infrastructure):

  • 枚举存储为 VARCHAR,JPA 使用 @Enumerated(EnumType.STRING)
  • Milvus collection 需 loadCollection()
  • 复用现有 VectorSearchService 接口
  • 文档元数据存入 ApiDocument 实体

术语对齐:

  • ApiDocument:文档元数据实体,扩展 metadata 字段存储 frontmatter
  • category:文档分类(api/domain/troubleshooting),与 Phase 1 对齐

关键假设

  1. L0 高置信度定义:唯一匹配

    • 假设:1 个匹配结果即为高置信度,不调用 L1
    • 验证方式:✅ grill 阶段已确认
    • 状态:已验证
  2. knowledge_base/ 目录权限

    • 假设:应用有读写权限
    • 验证方式:启动时创建目录
    • 风险:Docker 部署时路径映射
  3. TEXT 字段存储 JSON

    • 假设:TEXT 类型可存储 JSON 字符串(< 64KB)
    • 验证方式:✅ grill 阶段已确认
    • 状态:已验证

主要风险

风险 1:知识库目录权限问题

  • 影响:无法创建 knowledge_base/ 或保存文件
  • 概率:中(Docker 环境常见)
  • 缓解:启动时检查并创建目录,Docker 部署时正确挂载卷
  • 检测:apply 阶段测试文件保存功能

风险 2:L0 关键词匹配不准确

  • 影响:误匹配或漏匹配
  • 概率:中(依赖 frontmatter 质量)
  • 缓解:frontmatter keywords 需要精心维护,L1 作为兜底
  • 后续:引入模糊匹配或同义词扩展

风险 3:事务一致性(孤儿文件)

  • 影响:文件保存成功但事务回滚,产生孤儿文件
  • 概率:低
  • 缓解:异常时调用 cleanupLocalFile() 清理
  • 检测:集成测试验证

验收标准

功能验收

  1. ✅ 上传带 frontmatter 的 .md 文档成功
  2. ✅ L0 精确匹配:"ERR_TIMEOUT" → 返回完整文档(matchType=exact_L0)
  3. ✅ L0 未匹配:"如何优化性能" → 降级到 L1(matchType=semantic_L1)
  4. ✅ L0 多个匹配:"超时" → 返回 L0 列表 + L1 补充
  5. ✅ 缺少 frontmatter 的文档只走 L1

性能验收

  • L0 查询响应时间 < 10ms
  • L0 + L1 组合查询 < 500ms
  • 启动扫描时间 < 5s(假设 < 1000 个文档)

集成验收

  • Agent 调用 lookup_knowledge("ERR_TIMEOUT") 返回正确文档
  • Agent 调用 lookup_knowledge("支付失败") 返回语义相关文档