feat(knowledge): 完成 L0+L1 混合检索集成

核心功能:
- 新增 FrontmatterParser 解析 YAML frontmatter
- 新增 KnowledgeIndexService L0 内存索引
- 新增 LookupKnowledgeTool 混合检索工具
- 增强 DocumentManagementService 文件保存和索引同步

技术实现:
- 数据库迁移 V004: api_document.metadata (TEXT)
- 依赖新增: snakeyaml 2.0
- 配置新增: knowledge.base-path
- 可观测性: requestId 追踪 + 性能日志

质量保证:
- 单元测试: 31/31 通过
- 测试覆盖: FrontmatterParser(11), KnowledgeIndexService(13), LookupKnowledgeTool(7)
- 启动验证: L0 索引正常加载

归档文档:
- OpenSpec: openspec/changes/lookup-knowledge-integration/
- devflow 档案: devflow/projects/2026-06-24-lookup-knowledge-integration/
- handoff: handoff/2026-06-24-lookup-knowledge-integration.md
This commit is contained in:
zhuyongxin
2026-06-24 16:07:10 +08:00
parent c86045b33f
commit d6229f3385
32 changed files with 5396 additions and 67 deletions
@@ -0,0 +1,657 @@
# Design: L0+L1 混合检索集成
## 架构概览
### 双层检索架构
```
┌─────────────────────────────────────────────────────────────┐
│ Agent (ReactAgent) │
└─────────────────────┬───────────────────────────────────────┘
│ 调用
▼
┌─────────────────────────────────────────────────────────────┐
│ LookupKnowledgeTool (新增) │
│ - lookup(query, sectionTitle) │
│ - 编排 L0 + L1 检索流程 │
└──────┬──────────────────────────┬───────────────────────────┘
│ │
│ L0 精确匹配 │ L1 语义检索(条件调用)
▼ ▼
┌──────────────────────┐ ┌──────────────────────────────┐
│ KnowledgeIndexService│ │ VectorSearchService (复用) │
│ (新增) │ │ - searchSimilarDocuments() │
│ - loadIndex() │ │ - Milvus + BGE-M3 │
│ - exactMatch() │ └──────────────────────────────┘
│ - readDocument() │
└──────┬───────────────┘
│ 读取
▼
┌──────────────────────────────────────────────────────────────┐
│ knowledge_base/ (本地文件系统) │
│ ├── api/ │
│ ├── domain/ │
│ └── troubleshooting/ │
└──────────────────────────────────────────────────────────────┘
```
### 上传流程增强
```
POST /api/documents/upload
│
▼
DocumentManagementService.uploadDocument()
│
├─ 1. 文件格式验证
├─ 2. 计算 hash(去重)
├─ 3. 提取文本 (TextExtractorService)
│
├─ 4. 【新增】保存原始文件到本地
│ └─ knowledge_base/{category}/{fileName}
│
├─ 5. 【新增】解析 frontmatter (FrontmatterParser)
│ └─ 提取 title, keywords, summary
│
├─ 6. 分块 (DocumentChunkService)
├─ 7. 向量化 + Milvus 索引 (VectorIndexService)
│
├─ 8. 保存元数据到 MySQL (ApiDocument)
│ └─ metadata 字段存储 frontmatter JSON
│
└─ 9. 【新增】更新 L0 内存索引
└─ KnowledgeIndexService.addToIndex()
```
---
## 核心组件设计
### 1. FrontmatterParser(新增)
**职责**:解析 Markdown 文件头的 YAML frontmatter
**依赖**:snakeyaml 2.0
**接口设计**:
```java
package com.superbiz.agent.service;
public class FrontmatterParser {
/**
* 解析 Markdown frontmatter
* @param content 完整文件内容
* @return Frontmatter 对象,如果不存在返回 null
*/
public Frontmatter parse(String content) {
// 1. 检查是否以 --- 开头
// 2. 提取 frontmatter 部分(两个 --- 之间)
// 3. 使用 Yaml.load() 解析
// 4. 映射到 Frontmatter 对象
}
/**
* 检查文件是否包含 frontmatter
*/
public boolean hasFrontmatter(String content) {
return content != null && content.trim().startsWith("---");
}
}
```
**数据模型**:
```java
package com.superbiz.agent.dto;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class Frontmatter {
private String title; // 必填
private List<String> keywords; // 必填
private String summary; // 必填
// 预留字段(MVP 不使用)
private String category; // 可选
private Map<String, String> sections; // 可选
private String version; // 可选
private String author; // 可选
private LocalDate lastUpdated; // 可选
}
```
---
### 2. KnowledgeIndexService(新增)
**职责**:L0 精确匹配索引管理
**启动扫描**:
```java
@Service
public class KnowledgeIndexService {
@Value("${knowledge.base-path}")
private String knowledgeBasePath; // 从配置文件读取
@Autowired
private FrontmatterParser frontmatterParser;
// 内存索引
private final List<KnowledgeEntry> knowledgeIndex =
new CopyOnWriteArrayList<>();
@PostConstruct
public void loadIndex() {
log.info("开始扫描知识库目录: {}", knowledgeBasePath);
// 1. 递归扫描 knowledge_base/
// 2. 过滤 .md 文件
// 3. 读取文件内容
// 4. 解析 frontmatter
// 5. 构建 KnowledgeEntry
// 6. 添加到 knowledgeIndex
log.info("知识库索引加载完成,共 {} 个文档", knowledgeIndex.size());
}
/**
* L0 精确匹配
* @param query 查询关键词
* @return 匹配的文档列表
*/
public List<KnowledgeEntry> exactMatch(String query) {
String queryLower = query.toLowerCase();
return knowledgeIndex.stream()
.filter(entry -> matchesKeywords(entry, queryLower))
.collect(Collectors.toList());
}
private boolean matchesKeywords(KnowledgeEntry entry, String query) {
// 关键词匹配(不区分大小写)
for (String keyword : entry.getKeywords()) {
if (query.contains(keyword.toLowerCase()) ||
keyword.toLowerCase().contains(query)) {
return true;
}
}
return false;
}
/**
* 读取文档内容
* @param filePath 文件路径
* @param maxChars 最大字符数
* @return 文档内容(前 maxChars 字符)
*/
public String readDocument(String filePath, int maxChars) {
try {
String content = Files.readString(Paths.get(filePath));
return content.length() > maxChars ?
content.substring(0, maxChars) + "..." : content;
} catch (IOException e) {
log.error("读取文档失败: {}", filePath, e);
return null;
}
}
/**
* 添加文档到索引(上传时调用)
*/
public void addToIndex(KnowledgeEntry entry) {
knowledgeIndex.add(entry);
log.debug("文档已添加到 L0 索引: {}", entry.getTitle());
}
/**
* 从索引中移除文档(删除时调用)
*/
public void removeFromIndex(String filePath) {
knowledgeIndex.removeIf(e -> e.getFilePath().equals(filePath));
log.debug("文档已从 L0 索引移除: {}", filePath);
}
}
```
**数据模型**:
```java
package com.superbiz.agent.dto;
@Data
@Builder
public class KnowledgeEntry {
private String filePath; // knowledge_base/api/payment-errors.md
private String title; // 支付网关错误码定义
private List<String> keywords; // [ERR_TIMEOUT, 超时, 支付网关]
private String summary; // 一句话摘要
private String category; // api/domain/troubleshooting
// 预留字段
private Map<String, String> sections;
}
```
---
### 3. LookupKnowledgeTool(新增)
**职责**:提供给 Agent 的混合检索工具
**实现**:
```java
package com.superbiz.agent.tool;
@Component
public class LookupKnowledgeTool {
@Autowired
private KnowledgeIndexService knowledgeIndexService;
@Autowired
private VectorSearchService vectorSearchService;
@Tool(
name = "lookup_knowledge",
description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。"
)
public LookupResult lookup(
@P("query") String query,
@P("section_title") String sectionTitle // 预留参数,MVP 返回 null
) {
log.info("收到知识库查询请求: query={}", query);
// Step 1: L0 精确匹配
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
log.debug("L0 匹配结果: {} 个文档", l0Matches.size());
// Step 2: 判断是否高置信度
boolean highConfidence = (l0Matches.size() == 1);
// Step 3: L1 条件调用
List<VectorSearchService.SearchResult> l1Results = null;
if (!highConfidence) {
log.debug("L0 非唯一匹配,调用 L1 语义检索");
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
}
// Step 4: 组装结果
return buildResult(l0Matches, l1Results, highConfidence);
}
private LookupResult buildResult(
List<KnowledgeEntry> l0Matches,
List<VectorSearchService.SearchResult> l1Results,
boolean highConfidence
) {
LookupResult result = new LookupResult();
result.setFound(!l0Matches.isEmpty() || (l1Results != null && !l1Results.isEmpty()));
// Primary: L0 结果
if (!l0Matches.isEmpty()) {
KnowledgeEntry first = l0Matches.get(0);
String content = knowledgeIndexService.readDocument(first.getFilePath(), 2000);
result.setPrimary(PrimaryResult.builder()
.content(content)
.source(first.getFilePath())
.matchType("exact_L0")
.confidence(highConfidence ? "high" : "low")
.availableSections(null) // MVP 返回 null
.build());
}
// Supplement: L1 结果
if (l1Results != null && !l1Results.isEmpty()) {
VectorSearchService.SearchResult firstL1 = l1Results.get(0);
result.setSupplement(SupplementResult.builder()
.content(firstL1.getContent())
.source(firstL1.getMetadata())
.matchType("semantic_L1")
.build());
}
return result;
}
}
```
**返回模型**:
```java
@Data
@Builder
public class LookupResult {
private boolean found;
private PrimaryResult primary;
private SupplementResult supplement;
}
@Data
@Builder
public class PrimaryResult {
private String content;
private String source;
private String matchType; // exact_L0
private String confidence; // high / low
private List<String> availableSections; // 预留字段
}
@Data
@Builder
public class SupplementResult {
private String content;
private String source;
private String matchType; // semantic_L1
}
```
---
### 4. DocumentManagementService(增强)
**变更点**:
**增加文件保存逻辑**:
```java
// 在 uploadDocument() 方法中,提取文本后增加
// 3. 提取文本
String text = textExtractorService.extractText(file, fileName);
// 【新增】4. 保存原始文件到本地
String category = request.getCategory() != null ? request.getCategory() : "default";
String localPath = saveToLocal(file, fileName, category);
// 【新增】5. 解析 frontmatter
Frontmatter frontmatter = null;
if (frontmatterParser.hasFrontmatter(text)) {
frontmatter = frontmatterParser.parse(text);
log.info("解析到 frontmatter: title={}, keywords={}",
frontmatter.getTitle(), frontmatter.getKeywords());
}
// 6. 分块(继续现有逻辑)
List<DocumentChunk> chunks = documentChunkService.chunkDocument(text, fileName);
```
**新增方法**:
```java
/**
* 保存文件到本地
*/
private String saveToLocal(MultipartFile file, String fileName, String category) {
try {
// 1. 构建目标路径
Path categoryDir = Paths.get(knowledgeBasePath, category);
Files.createDirectories(categoryDir);
Path targetPath = categoryDir.resolve(fileName);
// 2. 保存文件
file.transferTo(targetPath.toFile());
log.info("文件已保存到本地: {}", targetPath);
return targetPath.toString();
} catch (IOException e) {
throw new DocumentProcessException(
fileName, "save-local",
"保存文件到本地失败: " + e.getMessage(), e
);
}
}
/**
* 清理本地文件(事务回滚时调用)
*/
private void cleanupLocalFile(String localPath) {
if (localPath != null) {
try {
Files.deleteIfExists(Paths.get(localPath));
log.info("已清理本地文件: {}", localPath);
} catch (IOException e) {
log.warn("清理本地文件失败: {}", localPath, e);
}
}
}
```
**事务一致性处理**:
```java
@Transactional
public String uploadDocument(DocumentUploadRequest request) {
String localPath = null;
try {
// ... 提取文本
localPath = saveToLocal(file, fileName, category);
// ... frontmatter 解析
// ... 分块、向量化、保存到 MySQL
// ... 更新 L0 索引
} catch (Exception e) {
// 失败时清理本地文件
cleanupLocalFile(localPath);
throw e;
}
}
```
**更新 ApiDocument 保存**:
```java
// 创建文档元数据时增加字段
ApiDocument document = ApiDocument.builder()
.docId(docId)
.fileName(fileName)
.filePath(localPath) // 保存本地路径
.metadata(frontmatter != null ?
objectMapper.writeValueAsString(frontmatter) : null) // 存储 frontmatter JSON
// ... 其他字段
.build();
```
**更新 L0 索引**:
```java
// 索引成功后,如果有 frontmatter,更新 L0 索引
if (frontmatter != null) {
KnowledgeEntry entry = KnowledgeEntry.builder()
.filePath(localPath)
.title(frontmatter.getTitle())
.keywords(frontmatter.getKeywords())
.summary(frontmatter.getSummary())
.category(category)
.build();
knowledgeIndexService.addToIndex(entry);
}
```
---
### 5. ApiDocument 实体扩展
**新增字段**:
```java
@Entity
@Table(name = "api_document")
public class ApiDocument {
// ... 现有字段
// 【新增】frontmatter 元数据
@Column(name = "metadata", columnDefinition = "TEXT")
private String metadata; // JSON 格式存储
// 【新增】本地文件路径(现有 filePath 字段复用)
// 已有:@Column(name = "file_path", length = 512)
// private String filePath;
}
```
**Flyway 迁移脚本**:
```sql
-- V004__add_metadata_to_api_document.sql
ALTER TABLE api_document
ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)';
```
---
## 配置管理
**application.yml 新增配置**:
```yaml
# 知识库配置
knowledge:
base-path: knowledge_base/ # 知识库根目录
```
**pom.xml 新增依赖**:
```xml
<!-- YAML 解析 -->
<dependency>
<groupId>org.yaml</groupId>
<artifactId>snakeyaml</artifactId>
<version>2.0</version>
</dependency>
```
---
## 数据流时序图
### 上传流程时序图
```
User -> Controller: POST /api/documents/upload
Controller -> DocumentManagementService: uploadDocument(request)
DocumentManagementService -> TextExtractorService: extractText(file)
TextExtractorService --> DocumentManagementService: text
DocumentManagementService -> FileSystem: saveToLocal(file, category)
FileSystem --> DocumentManagementService: localPath
DocumentManagementService -> FrontmatterParser: parse(text)
FrontmatterParser --> DocumentManagementService: frontmatter
DocumentManagementService -> DocumentChunkService: chunkDocument(text)
DocumentChunkService --> DocumentManagementService: chunks
DocumentManagementService -> VectorIndexService: indexDocumentChunks(chunks)
VectorIndexService -> Milvus: insert vectors
Milvus --> VectorIndexService: success
DocumentManagementService -> ApiDocumentRepository: save(document)
ApiDocumentRepository --> DocumentManagementService: saved
DocumentManagementService -> KnowledgeIndexService: addToIndex(entry)
KnowledgeIndexService --> DocumentManagementService: indexed
DocumentManagementService --> Controller: docId
Controller --> User: {"code":200, "data":"doc-id"}
```
### 查询流程时序图
```
Agent -> LookupKnowledgeTool: lookup(query)
LookupKnowledgeTool -> KnowledgeIndexService: exactMatch(query)
KnowledgeIndexService --> LookupKnowledgeTool: l0Matches
alt 唯一匹配(高置信度)
LookupKnowledgeTool -> KnowledgeIndexService: readDocument(filePath)
KnowledgeIndexService -> FileSystem: read file
FileSystem --> KnowledgeIndexService: content
KnowledgeIndexService --> LookupKnowledgeTool: content
else 未匹配或多个匹配(低置信度)
LookupKnowledgeTool -> VectorSearchService: searchSimilarDocuments(query)
VectorSearchService -> Milvus: search vectors
Milvus --> VectorSearchService: l1Results
VectorSearchService --> LookupKnowledgeTool: l1Results
end
LookupKnowledgeTool --> Agent: LookupResult{primary, supplement}
```
---
## 关键决策记录
### 决策 1:文件保存策略
- **决策**:保存原始文件到本地文件系统
- **理由**:支持 L0 完整读取 + 未来扩展(版本管理、导出)
- **来源**:grill 阶段用户确认
### 决策 2:metadata 存储方式
- **决策**:TEXT 类型存储 JSON 字符串
- **理由**:简单直接,灵活扩展,无需自定义 JPA Converter
- **来源**:grill 阶段用户确认
### 决策 3:L0 高置信度标准
- **决策**:唯一匹配 = 高置信度,不调用 L1
- **理由**:唯一匹配通常就是用户想要的,调用 L1 只会增加延迟
- **来源**:grill 阶段用户确认
### 决策 4:knowledge_base/ 路径配置
- **决策**:通过 application.yml 配置,支持环境差异
- **理由**:开发环境和 Docker 环境路径可能不同
- **来源**:grill 阶段用户确认
---
## 非功能性设计
### 性能指标
- L0 查询响应时间:< 10ms
- L0 + L1 组合查询:< 500ms
- 启动扫描时间:< 5s(< 1000 个文档)
### 内存占用
- 单个 KnowledgeEntry:约 1KB(只存储 frontmatter 元数据)
- 1000 个文档:约 1MB(启动扫描只读取文件头)
- 10000 个文档:约 10MB
- **说明**:启动扫描只解析 frontmatter(< 1KB/文档),不读取全文;全文只在查询命中时按需读取
### 并发安全
- 使用 `CopyOnWriteArrayList` 存储索引(读多写少)
- 上传时更新索引(写操作)加锁或使用原子操作
### 错误处理
- frontmatter 解析失败:记录警告,文档仍可上传(只走 L1)
- 文件保存失败:抛出异常,回滚事务
- L0 索引加载失败:记录错误,应用仍可启动(只走 L1)
---
## 测试策略
### 单元测试
- FrontmatterParser 解析测试(有/无 frontmatter、格式错误)
- KnowledgeIndexService 匹配逻辑测试
- LookupKnowledgeTool 条件调用测试
### 集成测试
- 上传带 frontmatter 的文档 → 验证 L0 索引
- L0 精确匹配 → 验证返回正确文档
- L0 未命中 → 验证降级到 L1
### 性能测试
- L0 查询响应时间
- 大量文档启动扫描时间
---
## 实现优先级
### P0(MVP 必须)
1. FrontmatterParser
2. KnowledgeIndexService(启动扫描 + 精确匹配)
3. DocumentManagementService 增强
4. LookupKnowledgeTool
5. Flyway 迁移脚本
6. 配置管理
### P1(后续扩展)
- sections 分段加载
- watchdog 热更新
- L0 索引持久化
- 模糊匹配 / 同义词扩展