Files
zhuyongxin d6229f3385 feat(knowledge): 完成 L0+L1 混合检索集成
核心功能:
- 新增 FrontmatterParser 解析 YAML frontmatter
- 新增 KnowledgeIndexService L0 内存索引
- 新增 LookupKnowledgeTool 混合检索工具
- 增强 DocumentManagementService 文件保存和索引同步

技术实现:
- 数据库迁移 V004: api_document.metadata (TEXT)
- 依赖新增: snakeyaml 2.0
- 配置新增: knowledge.base-path
- 可观测性: requestId 追踪 + 性能日志

质量保证:
- 单元测试: 31/31 通过
- 测试覆盖: FrontmatterParser(11), KnowledgeIndexService(13), LookupKnowledgeTool(7)
- 启动验证: L0 索引正常加载

归档文档:
- OpenSpec: openspec/changes/lookup-knowledge-integration/
- devflow 档案: devflow/projects/2026-06-24-lookup-knowledge-integration/
- handoff: handoff/2026-06-24-lookup-knowledge-integration.md
2026-06-24 16:07:10 +08:00

20 KiB
Raw Permalink Blame History

Design: L0+L1 混合检索集成

架构概览

双层检索架构

┌─────────────────────────────────────────────────────────────┐
│                      Agent (ReactAgent)                      │
└─────────────────────┬───────────────────────────────────────┘
                      │ 调用
                      ▼
┌─────────────────────────────────────────────────────────────┐
│              LookupKnowledgeTool (新增)                      │
│  - lookup(query, sectionTitle)                              │
│  - 编排 L0 + L1 检索流程                                     │
└──────┬──────────────────────────┬───────────────────────────┘
       │                          │
       │ L0 精确匹配              │ L1 语义检索(条件调用)
       ▼                          ▼
┌──────────────────────┐    ┌──────────────────────────────┐
│ KnowledgeIndexService│    │  VectorSearchService (复用)  │
│  (新增)              │    │  - searchSimilarDocuments()  │
│ - loadIndex()        │    │  - Milvus + BGE-M3          │
│ - exactMatch()       │    └──────────────────────────────┘
│ - readDocument()     │
└──────┬───────────────┘
       │ 读取
       ▼
┌──────────────────────────────────────────────────────────────┐
│              knowledge_base/ (本地文件系统)                   │
│  ├── api/                                                     │
│  ├── domain/                                                  │
│  └── troubleshooting/                                         │
└──────────────────────────────────────────────────────────────┘

上传流程增强

POST /api/documents/upload
    │
    ▼
DocumentManagementService.uploadDocument()
    │
    ├─ 1. 文件格式验证
    ├─ 2. 计算 hash(去重)
    ├─ 3. 提取文本 (TextExtractorService)
    │
    ├─ 4. 【新增】保存原始文件到本地
    │     └─ knowledge_base/{category}/{fileName}
    │
    ├─ 5. 【新增】解析 frontmatter (FrontmatterParser)
    │     └─ 提取 title, keywords, summary
    │
    ├─ 6. 分块 (DocumentChunkService)
    ├─ 7. 向量化 + Milvus 索引 (VectorIndexService)
    │
    ├─ 8. 保存元数据到 MySQL (ApiDocument)
    │     └─ metadata 字段存储 frontmatter JSON
    │
    └─ 9. 【新增】更新 L0 内存索引
          └─ KnowledgeIndexService.addToIndex()

核心组件设计

1. FrontmatterParser(新增)

职责:解析 Markdown 文件头的 YAML frontmatter

依赖:snakeyaml 2.0

接口设计:

package com.superbiz.agent.service;

public class FrontmatterParser {
    
    /**
     * 解析 Markdown frontmatter
     * @param content 完整文件内容
     * @return Frontmatter 对象,如果不存在返回 null
     */
    public Frontmatter parse(String content) {
        // 1. 检查是否以 --- 开头
        // 2. 提取 frontmatter 部分(两个 --- 之间)
        // 3. 使用 Yaml.load() 解析
        // 4. 映射到 Frontmatter 对象
    }
    
    /**
     * 检查文件是否包含 frontmatter
     */
    public boolean hasFrontmatter(String content) {
        return content != null && content.trim().startsWith("---");
    }
}

数据模型:

package com.superbiz.agent.dto;

@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class Frontmatter {
    private String title;              // 必填
    private List<String> keywords;     // 必填
    private String summary;            // 必填
    
    // 预留字段(MVP 不使用)
    private String category;           // 可选
    private Map<String, String> sections;  // 可选
    private String version;            // 可选
    private String author;             // 可选
    private LocalDate lastUpdated;     // 可选
}

2. KnowledgeIndexService(新增)

职责:L0 精确匹配索引管理

启动扫描:

@Service
public class KnowledgeIndexService {
    
    @Value("${knowledge.base-path}")
    private String knowledgeBasePath;  // 从配置文件读取
    
    @Autowired
    private FrontmatterParser frontmatterParser;
    
    // 内存索引
    private final List<KnowledgeEntry> knowledgeIndex = 
        new CopyOnWriteArrayList<>();
    
    @PostConstruct
    public void loadIndex() {
        log.info("开始扫描知识库目录: {}", knowledgeBasePath);
        
        // 1. 递归扫描 knowledge_base/
        // 2. 过滤 .md 文件
        // 3. 读取文件内容
        // 4. 解析 frontmatter
        // 5. 构建 KnowledgeEntry
        // 6. 添加到 knowledgeIndex
        
        log.info("知识库索引加载完成,共 {} 个文档", knowledgeIndex.size());
    }
    
    /**
     * L0 精确匹配
     * @param query 查询关键词
     * @return 匹配的文档列表
     */
    public List<KnowledgeEntry> exactMatch(String query) {
        String queryLower = query.toLowerCase();
        
        return knowledgeIndex.stream()
            .filter(entry -> matchesKeywords(entry, queryLower))
            .collect(Collectors.toList());
    }
    
    private boolean matchesKeywords(KnowledgeEntry entry, String query) {
        // 关键词匹配(不区分大小写)
        for (String keyword : entry.getKeywords()) {
            if (query.contains(keyword.toLowerCase()) || 
                keyword.toLowerCase().contains(query)) {
                return true;
            }
        }
        return false;
    }
    
    /**
     * 读取文档内容
     * @param filePath 文件路径
     * @param maxChars 最大字符数
     * @return 文档内容(前 maxChars 字符)
     */
    public String readDocument(String filePath, int maxChars) {
        try {
            String content = Files.readString(Paths.get(filePath));
            return content.length() > maxChars ? 
                content.substring(0, maxChars) + "..." : content;
        } catch (IOException e) {
            log.error("读取文档失败: {}", filePath, e);
            return null;
        }
    }
    
    /**
     * 添加文档到索引(上传时调用)
     */
    public void addToIndex(KnowledgeEntry entry) {
        knowledgeIndex.add(entry);
        log.debug("文档已添加到 L0 索引: {}", entry.getTitle());
    }
    
    /**
     * 从索引中移除文档(删除时调用)
     */
    public void removeFromIndex(String filePath) {
        knowledgeIndex.removeIf(e -> e.getFilePath().equals(filePath));
        log.debug("文档已从 L0 索引移除: {}", filePath);
    }
}

数据模型:

package com.superbiz.agent.dto;

@Data
@Builder
public class KnowledgeEntry {
    private String filePath;          // knowledge_base/api/payment-errors.md
    private String title;             // 支付网关错误码定义
    private List<String> keywords;    // [ERR_TIMEOUT, 超时, 支付网关]
    private String summary;           // 一句话摘要
    private String category;          // api/domain/troubleshooting
    
    // 预留字段
    private Map<String, String> sections;
}

3. LookupKnowledgeTool(新增)

职责:提供给 Agent 的混合检索工具

实现:

package com.superbiz.agent.tool;

@Component
public class LookupKnowledgeTool {
    
    @Autowired
    private KnowledgeIndexService knowledgeIndexService;
    
    @Autowired
    private VectorSearchService vectorSearchService;
    
    @Tool(
        name = "lookup_knowledge",
        description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。"
    )
    public LookupResult lookup(
        @P("query") String query,
        @P("section_title") String sectionTitle  // 预留参数,MVP 返回 null
    ) {
        log.info("收到知识库查询请求: query={}", query);
        
        // Step 1: L0 精确匹配
        List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
        log.debug("L0 匹配结果: {} 个文档", l0Matches.size());
        
        // Step 2: 判断是否高置信度
        boolean highConfidence = (l0Matches.size() == 1);
        
        // Step 3: L1 条件调用
        List<VectorSearchService.SearchResult> l1Results = null;
        if (!highConfidence) {
            log.debug("L0 非唯一匹配,调用 L1 语义检索");
            l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
        }
        
        // Step 4: 组装结果
        return buildResult(l0Matches, l1Results, highConfidence);
    }
    
    private LookupResult buildResult(
        List<KnowledgeEntry> l0Matches,
        List<VectorSearchService.SearchResult> l1Results,
        boolean highConfidence
    ) {
        LookupResult result = new LookupResult();
        result.setFound(!l0Matches.isEmpty() || (l1Results != null && !l1Results.isEmpty()));
        
        // Primary: L0 结果
        if (!l0Matches.isEmpty()) {
            KnowledgeEntry first = l0Matches.get(0);
            String content = knowledgeIndexService.readDocument(first.getFilePath(), 2000);
            
            result.setPrimary(PrimaryResult.builder()
                .content(content)
                .source(first.getFilePath())
                .matchType("exact_L0")
                .confidence(highConfidence ? "high" : "low")
                .availableSections(null)  // MVP 返回 null
                .build());
        }
        
        // Supplement: L1 结果
        if (l1Results != null && !l1Results.isEmpty()) {
            VectorSearchService.SearchResult firstL1 = l1Results.get(0);
            result.setSupplement(SupplementResult.builder()
                .content(firstL1.getContent())
                .source(firstL1.getMetadata())
                .matchType("semantic_L1")
                .build());
        }
        
        return result;
    }
}

返回模型:

@Data
@Builder
public class LookupResult {
    private boolean found;
    private PrimaryResult primary;
    private SupplementResult supplement;
}

@Data
@Builder
public class PrimaryResult {
    private String content;
    private String source;
    private String matchType;  // exact_L0
    private String confidence; // high / low
    private List<String> availableSections;  // 预留字段
}

@Data
@Builder
public class SupplementResult {
    private String content;
    private String source;
    private String matchType;  // semantic_L1
}

4. DocumentManagementService(增强)

变更点:

增加文件保存逻辑:

// 在 uploadDocument() 方法中,提取文本后增加
// 3. 提取文本
String text = textExtractorService.extractText(file, fileName);

// 【新增】4. 保存原始文件到本地
String category = request.getCategory() != null ? request.getCategory() : "default";
String localPath = saveToLocal(file, fileName, category);

// 【新增】5. 解析 frontmatter
Frontmatter frontmatter = null;
if (frontmatterParser.hasFrontmatter(text)) {
    frontmatter = frontmatterParser.parse(text);
    log.info("解析到 frontmatter: title={}, keywords={}", 
        frontmatter.getTitle(), frontmatter.getKeywords());
}

// 6. 分块(继续现有逻辑)
List<DocumentChunk> chunks = documentChunkService.chunkDocument(text, fileName);

新增方法:

/**
 * 保存文件到本地
 */
private String saveToLocal(MultipartFile file, String fileName, String category) {
    try {
        // 1. 构建目标路径
        Path categoryDir = Paths.get(knowledgeBasePath, category);
        Files.createDirectories(categoryDir);
        
        Path targetPath = categoryDir.resolve(fileName);
        
        // 2. 保存文件
        file.transferTo(targetPath.toFile());
        
        log.info("文件已保存到本地: {}", targetPath);
        return targetPath.toString();
        
    } catch (IOException e) {
        throw new DocumentProcessException(
            fileName, "save-local",
            "保存文件到本地失败: " + e.getMessage(), e
        );
    }
}

/**
 * 清理本地文件(事务回滚时调用)
 */
private void cleanupLocalFile(String localPath) {
    if (localPath != null) {
        try {
            Files.deleteIfExists(Paths.get(localPath));
            log.info("已清理本地文件: {}", localPath);
        } catch (IOException e) {
            log.warn("清理本地文件失败: {}", localPath, e);
        }
    }
}

事务一致性处理:

@Transactional
public String uploadDocument(DocumentUploadRequest request) {
    String localPath = null;
    try {
        // ... 提取文本
        localPath = saveToLocal(file, fileName, category);
        // ... frontmatter 解析
        // ... 分块、向量化、保存到 MySQL
        // ... 更新 L0 索引
        
    } catch (Exception e) {
        // 失败时清理本地文件
        cleanupLocalFile(localPath);
        throw e;
    }
}

更新 ApiDocument 保存:

// 创建文档元数据时增加字段
ApiDocument document = ApiDocument.builder()
    .docId(docId)
    .fileName(fileName)
    .filePath(localPath)  // 保存本地路径
    .metadata(frontmatter != null ? 
        objectMapper.writeValueAsString(frontmatter) : null)  // 存储 frontmatter JSON
    // ... 其他字段
    .build();

更新 L0 索引:

// 索引成功后,如果有 frontmatter,更新 L0 索引
if (frontmatter != null) {
    KnowledgeEntry entry = KnowledgeEntry.builder()
        .filePath(localPath)
        .title(frontmatter.getTitle())
        .keywords(frontmatter.getKeywords())
        .summary(frontmatter.getSummary())
        .category(category)
        .build();
    
    knowledgeIndexService.addToIndex(entry);
}

5. ApiDocument 实体扩展

新增字段:

@Entity
@Table(name = "api_document")
public class ApiDocument {
    // ... 现有字段
    
    // 【新增】frontmatter 元数据
    @Column(name = "metadata", columnDefinition = "TEXT")
    private String metadata;  // JSON 格式存储
    
    // 【新增】本地文件路径(现有 filePath 字段复用)
    // 已有:@Column(name = "file_path", length = 512)
    //      private String filePath;
}

Flyway 迁移脚本:

-- V004__add_metadata_to_api_document.sql
ALTER TABLE api_document 
ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)';

配置管理

application.yml 新增配置:

# 知识库配置
knowledge:
  base-path: knowledge_base/  # 知识库根目录

pom.xml 新增依赖:

<!-- YAML 解析 -->
<dependency>
    <groupId>org.yaml</groupId>
    <artifactId>snakeyaml</artifactId>
    <version>2.0</version>
</dependency>

数据流时序图

上传流程时序图

User -> Controller: POST /api/documents/upload
Controller -> DocumentManagementService: uploadDocument(request)
DocumentManagementService -> TextExtractorService: extractText(file)
TextExtractorService --> DocumentManagementService: text

DocumentManagementService -> FileSystem: saveToLocal(file, category)
FileSystem --> DocumentManagementService: localPath

DocumentManagementService -> FrontmatterParser: parse(text)
FrontmatterParser --> DocumentManagementService: frontmatter

DocumentManagementService -> DocumentChunkService: chunkDocument(text)
DocumentChunkService --> DocumentManagementService: chunks

DocumentManagementService -> VectorIndexService: indexDocumentChunks(chunks)
VectorIndexService -> Milvus: insert vectors
Milvus --> VectorIndexService: success

DocumentManagementService -> ApiDocumentRepository: save(document)
ApiDocumentRepository --> DocumentManagementService: saved

DocumentManagementService -> KnowledgeIndexService: addToIndex(entry)
KnowledgeIndexService --> DocumentManagementService: indexed

DocumentManagementService --> Controller: docId
Controller --> User: {"code":200, "data":"doc-id"}

查询流程时序图

Agent -> LookupKnowledgeTool: lookup(query)
LookupKnowledgeTool -> KnowledgeIndexService: exactMatch(query)
KnowledgeIndexService --> LookupKnowledgeTool: l0Matches

alt 唯一匹配(高置信度)
    LookupKnowledgeTool -> KnowledgeIndexService: readDocument(filePath)
    KnowledgeIndexService -> FileSystem: read file
    FileSystem --> KnowledgeIndexService: content
    KnowledgeIndexService --> LookupKnowledgeTool: content
else 未匹配或多个匹配(低置信度)
    LookupKnowledgeTool -> VectorSearchService: searchSimilarDocuments(query)
    VectorSearchService -> Milvus: search vectors
    Milvus --> VectorSearchService: l1Results
    VectorSearchService --> LookupKnowledgeTool: l1Results
end

LookupKnowledgeTool --> Agent: LookupResult{primary, supplement}

关键决策记录

决策 1:文件保存策略

  • 决策:保存原始文件到本地文件系统
  • 理由:支持 L0 完整读取 + 未来扩展(版本管理、导出)
  • 来源:grill 阶段用户确认

决策 2:metadata 存储方式

  • 决策:TEXT 类型存储 JSON 字符串
  • 理由:简单直接,灵活扩展,无需自定义 JPA Converter
  • 来源:grill 阶段用户确认

决策 3:L0 高置信度标准

  • 决策:唯一匹配 = 高置信度,不调用 L1
  • 理由:唯一匹配通常就是用户想要的,调用 L1 只会增加延迟
  • 来源:grill 阶段用户确认

决策 4:knowledge_base/ 路径配置

  • 决策:通过 application.yml 配置,支持环境差异
  • 理由:开发环境和 Docker 环境路径可能不同
  • 来源:grill 阶段用户确认

非功能性设计

性能指标

  • L0 查询响应时间:< 10ms
  • L0 + L1 组合查询:< 500ms
  • 启动扫描时间:< 5s(< 1000 个文档)

内存占用

  • 单个 KnowledgeEntry:约 1KB(只存储 frontmatter 元数据)
  • 1000 个文档:约 1MB(启动扫描只读取文件头)
  • 10000 个文档:约 10MB
  • 说明:启动扫描只解析 frontmatter(< 1KB/文档),不读取全文;全文只在查询命中时按需读取

并发安全

  • 使用 CopyOnWriteArrayList 存储索引(读多写少)
  • 上传时更新索引(写操作)加锁或使用原子操作

错误处理

  • frontmatter 解析失败:记录警告,文档仍可上传(只走 L1)
  • 文件保存失败:抛出异常,回滚事务
  • L0 索引加载失败:记录错误,应用仍可启动(只走 L1)

测试策略

单元测试

  • FrontmatterParser 解析测试(有/无 frontmatter、格式错误)
  • KnowledgeIndexService 匹配逻辑测试
  • LookupKnowledgeTool 条件调用测试

集成测试

  • 上传带 frontmatter 的文档 → 验证 L0 索引
  • L0 精确匹配 → 验证返回正确文档
  • L0 未命中 → 验证降级到 L1

性能测试

  • L0 查询响应时间
  • 大量文档启动扫描时间

实现优先级

P0(MVP 必须)

  1. FrontmatterParser
  2. KnowledgeIndexService(启动扫描 + 精确匹配)
  3. DocumentManagementService 增强
  4. LookupKnowledgeTool
  5. Flyway 迁移脚本
  6. 配置管理

P1(后续扩展)

  • sections 分段加载
  • watchdog 热更新
  • L0 索引持久化
  • 模糊匹配 / 同义词扩展