# Design: L0+L1 混合检索集成 ## 架构概览 ### 双层检索架构 ``` ┌─────────────────────────────────────────────────────────────┐ │ Agent (ReactAgent) │ └─────────────────────┬───────────────────────────────────────┘ │ 调用 ▼ ┌─────────────────────────────────────────────────────────────┐ │ LookupKnowledgeTool (新增) │ │ - lookup(query, sectionTitle) │ │ - 编排 L0 + L1 检索流程 │ └──────┬──────────────────────────┬───────────────────────────┘ │ │ │ L0 精确匹配 │ L1 语义检索(条件调用) ▼ ▼ ┌──────────────────────┐ ┌──────────────────────────────┐ │ KnowledgeIndexService│ │ VectorSearchService (复用) │ │ (新增) │ │ - searchSimilarDocuments() │ │ - loadIndex() │ │ - Milvus + BGE-M3 │ │ - exactMatch() │ └──────────────────────────────┘ │ - readDocument() │ └──────┬───────────────┘ │ 读取 ▼ ┌──────────────────────────────────────────────────────────────┐ │ knowledge_base/ (本地文件系统) │ │ ├── api/ │ │ ├── domain/ │ │ └── troubleshooting/ │ └──────────────────────────────────────────────────────────────┘ ``` ### 上传流程增强 ``` POST /api/documents/upload │ ▼ DocumentManagementService.uploadDocument() │ ├─ 1. 文件格式验证 ├─ 2. 计算 hash(去重) ├─ 3. 提取文本 (TextExtractorService) │ ├─ 4. 【新增】保存原始文件到本地 │ └─ knowledge_base/{category}/{fileName} │ ├─ 5. 【新增】解析 frontmatter (FrontmatterParser) │ └─ 提取 title, keywords, summary │ ├─ 6. 分块 (DocumentChunkService) ├─ 7. 向量化 + Milvus 索引 (VectorIndexService) │ ├─ 8. 保存元数据到 MySQL (ApiDocument) │ └─ metadata 字段存储 frontmatter JSON │ └─ 9. 【新增】更新 L0 内存索引 └─ KnowledgeIndexService.addToIndex() ``` --- ## 核心组件设计 ### 1. FrontmatterParser(新增) **职责**:解析 Markdown 文件头的 YAML frontmatter **依赖**:snakeyaml 2.0 **接口设计**: ```java package com.superbiz.agent.service; public class FrontmatterParser { /** * 解析 Markdown frontmatter * @param content 完整文件内容 * @return Frontmatter 对象,如果不存在返回 null */ public Frontmatter parse(String content) { // 1. 检查是否以 --- 开头 // 2. 提取 frontmatter 部分(两个 --- 之间) // 3. 使用 Yaml.load() 解析 // 4. 映射到 Frontmatter 对象 } /** * 检查文件是否包含 frontmatter */ public boolean hasFrontmatter(String content) { return content != null && content.trim().startsWith("---"); } } ``` **数据模型**: ```java package com.superbiz.agent.dto; @Data @Builder @NoArgsConstructor @AllArgsConstructor public class Frontmatter { private String title; // 必填 private List keywords; // 必填 private String summary; // 必填 // 预留字段(MVP 不使用) private String category; // 可选 private Map sections; // 可选 private String version; // 可选 private String author; // 可选 private LocalDate lastUpdated; // 可选 } ``` --- ### 2. KnowledgeIndexService(新增) **职责**:L0 精确匹配索引管理 **启动扫描**: ```java @Service public class KnowledgeIndexService { @Value("${knowledge.base-path}") private String knowledgeBasePath; // 从配置文件读取 @Autowired private FrontmatterParser frontmatterParser; // 内存索引 private final List knowledgeIndex = new CopyOnWriteArrayList<>(); @PostConstruct public void loadIndex() { log.info("开始扫描知识库目录: {}", knowledgeBasePath); // 1. 递归扫描 knowledge_base/ // 2. 过滤 .md 文件 // 3. 读取文件内容 // 4. 解析 frontmatter // 5. 构建 KnowledgeEntry // 6. 添加到 knowledgeIndex log.info("知识库索引加载完成,共 {} 个文档", knowledgeIndex.size()); } /** * L0 精确匹配 * @param query 查询关键词 * @return 匹配的文档列表 */ public List exactMatch(String query) { String queryLower = query.toLowerCase(); return knowledgeIndex.stream() .filter(entry -> matchesKeywords(entry, queryLower)) .collect(Collectors.toList()); } private boolean matchesKeywords(KnowledgeEntry entry, String query) { // 关键词匹配(不区分大小写) for (String keyword : entry.getKeywords()) { if (query.contains(keyword.toLowerCase()) || keyword.toLowerCase().contains(query)) { return true; } } return false; } /** * 读取文档内容 * @param filePath 文件路径 * @param maxChars 最大字符数 * @return 文档内容(前 maxChars 字符) */ public String readDocument(String filePath, int maxChars) { try { String content = Files.readString(Paths.get(filePath)); return content.length() > maxChars ? content.substring(0, maxChars) + "..." : content; } catch (IOException e) { log.error("读取文档失败: {}", filePath, e); return null; } } /** * 添加文档到索引(上传时调用) */ public void addToIndex(KnowledgeEntry entry) { knowledgeIndex.add(entry); log.debug("文档已添加到 L0 索引: {}", entry.getTitle()); } /** * 从索引中移除文档(删除时调用) */ public void removeFromIndex(String filePath) { knowledgeIndex.removeIf(e -> e.getFilePath().equals(filePath)); log.debug("文档已从 L0 索引移除: {}", filePath); } } ``` **数据模型**: ```java package com.superbiz.agent.dto; @Data @Builder public class KnowledgeEntry { private String filePath; // knowledge_base/api/payment-errors.md private String title; // 支付网关错误码定义 private List keywords; // [ERR_TIMEOUT, 超时, 支付网关] private String summary; // 一句话摘要 private String category; // api/domain/troubleshooting // 预留字段 private Map sections; } ``` --- ### 3. LookupKnowledgeTool(新增) **职责**:提供给 Agent 的混合检索工具 **实现**: ```java package com.superbiz.agent.tool; @Component public class LookupKnowledgeTool { @Autowired private KnowledgeIndexService knowledgeIndexService; @Autowired private VectorSearchService vectorSearchService; @Tool( name = "lookup_knowledge", description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。" ) public LookupResult lookup( @P("query") String query, @P("section_title") String sectionTitle // 预留参数,MVP 返回 null ) { log.info("收到知识库查询请求: query={}", query); // Step 1: L0 精确匹配 List l0Matches = knowledgeIndexService.exactMatch(query); log.debug("L0 匹配结果: {} 个文档", l0Matches.size()); // Step 2: 判断是否高置信度 boolean highConfidence = (l0Matches.size() == 1); // Step 3: L1 条件调用 List l1Results = null; if (!highConfidence) { log.debug("L0 非唯一匹配,调用 L1 语义检索"); l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null); } // Step 4: 组装结果 return buildResult(l0Matches, l1Results, highConfidence); } private LookupResult buildResult( List l0Matches, List l1Results, boolean highConfidence ) { LookupResult result = new LookupResult(); result.setFound(!l0Matches.isEmpty() || (l1Results != null && !l1Results.isEmpty())); // Primary: L0 结果 if (!l0Matches.isEmpty()) { KnowledgeEntry first = l0Matches.get(0); String content = knowledgeIndexService.readDocument(first.getFilePath(), 2000); result.setPrimary(PrimaryResult.builder() .content(content) .source(first.getFilePath()) .matchType("exact_L0") .confidence(highConfidence ? "high" : "low") .availableSections(null) // MVP 返回 null .build()); } // Supplement: L1 结果 if (l1Results != null && !l1Results.isEmpty()) { VectorSearchService.SearchResult firstL1 = l1Results.get(0); result.setSupplement(SupplementResult.builder() .content(firstL1.getContent()) .source(firstL1.getMetadata()) .matchType("semantic_L1") .build()); } return result; } } ``` **返回模型**: ```java @Data @Builder public class LookupResult { private boolean found; private PrimaryResult primary; private SupplementResult supplement; } @Data @Builder public class PrimaryResult { private String content; private String source; private String matchType; // exact_L0 private String confidence; // high / low private List availableSections; // 预留字段 } @Data @Builder public class SupplementResult { private String content; private String source; private String matchType; // semantic_L1 } ``` --- ### 4. DocumentManagementService(增强) **变更点**: **增加文件保存逻辑**: ```java // 在 uploadDocument() 方法中,提取文本后增加 // 3. 提取文本 String text = textExtractorService.extractText(file, fileName); // 【新增】4. 保存原始文件到本地 String category = request.getCategory() != null ? request.getCategory() : "default"; String localPath = saveToLocal(file, fileName, category); // 【新增】5. 解析 frontmatter Frontmatter frontmatter = null; if (frontmatterParser.hasFrontmatter(text)) { frontmatter = frontmatterParser.parse(text); log.info("解析到 frontmatter: title={}, keywords={}", frontmatter.getTitle(), frontmatter.getKeywords()); } // 6. 分块(继续现有逻辑) List chunks = documentChunkService.chunkDocument(text, fileName); ``` **新增方法**: ```java /** * 保存文件到本地 */ private String saveToLocal(MultipartFile file, String fileName, String category) { try { // 1. 构建目标路径 Path categoryDir = Paths.get(knowledgeBasePath, category); Files.createDirectories(categoryDir); Path targetPath = categoryDir.resolve(fileName); // 2. 保存文件 file.transferTo(targetPath.toFile()); log.info("文件已保存到本地: {}", targetPath); return targetPath.toString(); } catch (IOException e) { throw new DocumentProcessException( fileName, "save-local", "保存文件到本地失败: " + e.getMessage(), e ); } } /** * 清理本地文件(事务回滚时调用) */ private void cleanupLocalFile(String localPath) { if (localPath != null) { try { Files.deleteIfExists(Paths.get(localPath)); log.info("已清理本地文件: {}", localPath); } catch (IOException e) { log.warn("清理本地文件失败: {}", localPath, e); } } } ``` **事务一致性处理**: ```java @Transactional public String uploadDocument(DocumentUploadRequest request) { String localPath = null; try { // ... 提取文本 localPath = saveToLocal(file, fileName, category); // ... frontmatter 解析 // ... 分块、向量化、保存到 MySQL // ... 更新 L0 索引 } catch (Exception e) { // 失败时清理本地文件 cleanupLocalFile(localPath); throw e; } } ``` **更新 ApiDocument 保存**: ```java // 创建文档元数据时增加字段 ApiDocument document = ApiDocument.builder() .docId(docId) .fileName(fileName) .filePath(localPath) // 保存本地路径 .metadata(frontmatter != null ? objectMapper.writeValueAsString(frontmatter) : null) // 存储 frontmatter JSON // ... 其他字段 .build(); ``` **更新 L0 索引**: ```java // 索引成功后,如果有 frontmatter,更新 L0 索引 if (frontmatter != null) { KnowledgeEntry entry = KnowledgeEntry.builder() .filePath(localPath) .title(frontmatter.getTitle()) .keywords(frontmatter.getKeywords()) .summary(frontmatter.getSummary()) .category(category) .build(); knowledgeIndexService.addToIndex(entry); } ``` --- ### 5. ApiDocument 实体扩展 **新增字段**: ```java @Entity @Table(name = "api_document") public class ApiDocument { // ... 现有字段 // 【新增】frontmatter 元数据 @Column(name = "metadata", columnDefinition = "TEXT") private String metadata; // JSON 格式存储 // 【新增】本地文件路径(现有 filePath 字段复用) // 已有:@Column(name = "file_path", length = 512) // private String filePath; } ``` **Flyway 迁移脚本**: ```sql -- V004__add_metadata_to_api_document.sql ALTER TABLE api_document ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)'; ``` --- ## 配置管理 **application.yml 新增配置**: ```yaml # 知识库配置 knowledge: base-path: knowledge_base/ # 知识库根目录 ``` **pom.xml 新增依赖**: ```xml org.yaml snakeyaml 2.0 ``` --- ## 数据流时序图 ### 上传流程时序图 ``` User -> Controller: POST /api/documents/upload Controller -> DocumentManagementService: uploadDocument(request) DocumentManagementService -> TextExtractorService: extractText(file) TextExtractorService --> DocumentManagementService: text DocumentManagementService -> FileSystem: saveToLocal(file, category) FileSystem --> DocumentManagementService: localPath DocumentManagementService -> FrontmatterParser: parse(text) FrontmatterParser --> DocumentManagementService: frontmatter DocumentManagementService -> DocumentChunkService: chunkDocument(text) DocumentChunkService --> DocumentManagementService: chunks DocumentManagementService -> VectorIndexService: indexDocumentChunks(chunks) VectorIndexService -> Milvus: insert vectors Milvus --> VectorIndexService: success DocumentManagementService -> ApiDocumentRepository: save(document) ApiDocumentRepository --> DocumentManagementService: saved DocumentManagementService -> KnowledgeIndexService: addToIndex(entry) KnowledgeIndexService --> DocumentManagementService: indexed DocumentManagementService --> Controller: docId Controller --> User: {"code":200, "data":"doc-id"} ``` ### 查询流程时序图 ``` Agent -> LookupKnowledgeTool: lookup(query) LookupKnowledgeTool -> KnowledgeIndexService: exactMatch(query) KnowledgeIndexService --> LookupKnowledgeTool: l0Matches alt 唯一匹配(高置信度) LookupKnowledgeTool -> KnowledgeIndexService: readDocument(filePath) KnowledgeIndexService -> FileSystem: read file FileSystem --> KnowledgeIndexService: content KnowledgeIndexService --> LookupKnowledgeTool: content else 未匹配或多个匹配(低置信度) LookupKnowledgeTool -> VectorSearchService: searchSimilarDocuments(query) VectorSearchService -> Milvus: search vectors Milvus --> VectorSearchService: l1Results VectorSearchService --> LookupKnowledgeTool: l1Results end LookupKnowledgeTool --> Agent: LookupResult{primary, supplement} ``` --- ## 关键决策记录 ### 决策 1:文件保存策略 - **决策**:保存原始文件到本地文件系统 - **理由**:支持 L0 完整读取 + 未来扩展(版本管理、导出) - **来源**:grill 阶段用户确认 ### 决策 2:metadata 存储方式 - **决策**:TEXT 类型存储 JSON 字符串 - **理由**:简单直接,灵活扩展,无需自定义 JPA Converter - **来源**:grill 阶段用户确认 ### 决策 3:L0 高置信度标准 - **决策**:唯一匹配 = 高置信度,不调用 L1 - **理由**:唯一匹配通常就是用户想要的,调用 L1 只会增加延迟 - **来源**:grill 阶段用户确认 ### 决策 4:knowledge_base/ 路径配置 - **决策**:通过 application.yml 配置,支持环境差异 - **理由**:开发环境和 Docker 环境路径可能不同 - **来源**:grill 阶段用户确认 --- ## 非功能性设计 ### 性能指标 - L0 查询响应时间:< 10ms - L0 + L1 组合查询:< 500ms - 启动扫描时间:< 5s(< 1000 个文档) ### 内存占用 - 单个 KnowledgeEntry:约 1KB(只存储 frontmatter 元数据) - 1000 个文档:约 1MB(启动扫描只读取文件头) - 10000 个文档:约 10MB - **说明**:启动扫描只解析 frontmatter(< 1KB/文档),不读取全文;全文只在查询命中时按需读取 ### 并发安全 - 使用 `CopyOnWriteArrayList` 存储索引(读多写少) - 上传时更新索引(写操作)加锁或使用原子操作 ### 错误处理 - frontmatter 解析失败:记录警告,文档仍可上传(只走 L1) - 文件保存失败:抛出异常,回滚事务 - L0 索引加载失败:记录错误,应用仍可启动(只走 L1) --- ## 测试策略 ### 单元测试 - FrontmatterParser 解析测试(有/无 frontmatter、格式错误) - KnowledgeIndexService 匹配逻辑测试 - LookupKnowledgeTool 条件调用测试 ### 集成测试 - 上传带 frontmatter 的文档 → 验证 L0 索引 - L0 精确匹配 → 验证返回正确文档 - L0 未命中 → 验证降级到 L1 ### 性能测试 - L0 查询响应时间 - 大量文档启动扫描时间 --- ## 实现优先级 ### P0(MVP 必须) 1. FrontmatterParser 2. KnowledgeIndexService(启动扫描 + 精确匹配) 3. DocumentManagementService 增强 4. LookupKnowledgeTool 5. Flyway 迁移脚本 6. 配置管理 ### P1(后续扩展) - sections 分段加载 - watchdog 热更新 - L0 索引持久化 - 模糊匹配 / 同义词扩展