282 lines
9.0 KiB
Markdown
282 lines
9.0 KiB
Markdown
# Proposal: L0+L1 混合检索集成
|
||
|
||
## 问题
|
||
|
||
当前只有 L1 向量语义检索(Milvus + BGE-M3),在遇到精确关键词查询时(如错误码 "ERR_TIMEOUT"、接口名 "PaymentGateway")效率不够高:
|
||
- 需要调用 embedding API 生成向量(约 100-300ms)
|
||
- 语义检索返回相似但可能不精确的结果
|
||
- 无法快速定位已知关键词对应的完整文档
|
||
|
||
Agent 需要一个"先精确、后语义"的混合检索工具。
|
||
|
||
## 建议方案
|
||
|
||
### 架构设计:双层检索
|
||
|
||
```
|
||
lookup_knowledge(query)
|
||
↓
|
||
L0: 精确关键词匹配(内存索引,< 10ms)
|
||
├─ 匹配成功 + 唯一结果 → 返回完整文档(高置信度)
|
||
└─ 未匹配 或 多个匹配 ↓
|
||
L1: 向量语义检索(Milvus,补充上下文)
|
||
└─ 返回 Top-K 相似片段
|
||
```
|
||
|
||
**核心机制**:
|
||
1. **L0 索引**:启动时扫描 `knowledge_base/` 目录,解析 Markdown frontmatter,构建内存索引
|
||
2. **L1 复用**:调用现有 `VectorSearchService.searchSimilarDocuments()`
|
||
3. **条件调用**:L0 唯一匹配时不调用 L1(减少延迟)
|
||
|
||
### 1. Frontmatter 规范
|
||
|
||
所有知识库文档(`knowledge_base/` 目录)需在文件头添加 YAML frontmatter:
|
||
|
||
```yaml
|
||
---
|
||
title: 支付网关错误码定义 # 必填
|
||
keywords: [ERR_TIMEOUT, 超时, 支付网关] # 必填,用于精确匹配
|
||
summary: 记录了支付网关所有核心错误码的含义及排查方向 # 必填
|
||
category: api # 可选,与现有 category 对齐
|
||
sections: # 预留字段(MVP 不实现)
|
||
超时排查: "## 1. 超时类错误"
|
||
---
|
||
|
||
# 文档正文
|
||
...
|
||
```
|
||
|
||
**约束**:
|
||
- frontmatter 必须在文件最顶部(前面不能有空行)
|
||
- `title`, `keywords`, `summary` 为必填字段
|
||
- 缺少 frontmatter 的文档允许上传,但不参与 L0 索引(只走 L1)
|
||
|
||
### 2. 上传流程增强
|
||
|
||
**现有流程**:
|
||
```
|
||
POST /api/documents/upload
|
||
↓
|
||
DocumentManagementService.uploadDocument()
|
||
↓
|
||
文本提取 → 分块 → 向量化 → Milvus 索引
|
||
↓
|
||
元数据存 MySQL (ApiDocument)
|
||
```
|
||
|
||
**增强后流程**:
|
||
```
|
||
POST /api/documents/upload
|
||
↓
|
||
1. 文本提取(内存)
|
||
2. 保存原始文件到:knowledge_base/{category}/{fileName}
|
||
3. 解析 frontmatter(FrontmatterParser)
|
||
4. 分块 → 向量化 → Milvus 索引
|
||
5. 元数据存 MySQL(ApiDocument.metadata 存储 frontmatter JSON)
|
||
6. 更新 L0 内存索引(KnowledgeIndexService)
|
||
```
|
||
|
||
**关键决策**(grill 阶段确认):
|
||
- ✅ 保存原始文件到本地(支持 L0 完整读取 + 未来扩展)
|
||
- ✅ metadata 字段:TEXT 类型存储 JSON 字符串
|
||
- ✅ ApiDocument.filePath 存储本地文件路径
|
||
- ✅ L0 高置信度 = 唯一匹配(不调用 L1)
|
||
- 按 category 分类存储:`knowledge_base/api/`, `knowledge_base/domain/`, `knowledge_base/troubleshooting/`
|
||
|
||
### 3. L0 索引服务
|
||
|
||
**KnowledgeIndexService**:
|
||
```java
|
||
@Service
|
||
public class KnowledgeIndexService {
|
||
// 内存索引结构
|
||
private List<KnowledgeEntry> knowledgeIndex = new ArrayList<>();
|
||
|
||
// 启动时扫描
|
||
@PostConstruct
|
||
public void loadIndex() {
|
||
// 递归扫描 knowledge_base/
|
||
// 解析 frontmatter
|
||
// 构建内存索引
|
||
}
|
||
|
||
// L0 精确匹配
|
||
public List<KnowledgeEntry> exactMatch(String query) {
|
||
// 关键词匹配(不区分大小写)
|
||
// 匹配规则:query 包含 keywords 中的任一词
|
||
}
|
||
|
||
// 读取文档内容
|
||
public String readDocument(String filePath, int maxChars) {
|
||
// 读取文件,返回前 maxChars 字符
|
||
}
|
||
}
|
||
```
|
||
|
||
**数据结构**:
|
||
```java
|
||
@Data
|
||
public class KnowledgeEntry {
|
||
private String filePath; // knowledge_base/api/payment-errors.md
|
||
private String title; // 支付网关错误码定义
|
||
private List<String> keywords; // [ERR_TIMEOUT, 超时, 支付网关]
|
||
private String summary; // 一句话摘要
|
||
private String category; // api
|
||
private Map<String, String> sections; // 预留字段
|
||
}
|
||
```
|
||
|
||
### 4. L1 复用
|
||
|
||
直接调用现有服务:
|
||
```java
|
||
@Autowired
|
||
private VectorSearchService vectorSearchService;
|
||
|
||
List<VectorSearchService.SearchResult> l1Results =
|
||
vectorSearchService.searchSimilarDocuments(query, 3, category);
|
||
```
|
||
|
||
### 5. 混合检索工具
|
||
|
||
**LookupKnowledgeTool**(供 Agent 调用):
|
||
```java
|
||
@Tool(name = "lookup_knowledge",
|
||
description = "查询知识库文档。优先精确匹配,自动补充语义相关片段。")
|
||
public LookupResult lookup(
|
||
@P("query") String query,
|
||
@P("section_title") String sectionTitle // 预留参数,MVP 不实现
|
||
) {
|
||
// Step 1: L0 精确匹配
|
||
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
|
||
|
||
// Step 2: 判断是否高置信度(唯一匹配)
|
||
boolean highConfidence = (l0Matches.size() == 1);
|
||
|
||
// Step 3: L1 条件调用
|
||
List<SearchResult> l1Results = null;
|
||
if (!highConfidence) {
|
||
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
|
||
}
|
||
|
||
// Step 4: 组装结果
|
||
return buildResult(l0Matches, l1Results, highConfidence);
|
||
}
|
||
```
|
||
|
||
**返回格式**:
|
||
```json
|
||
{
|
||
"found": true,
|
||
"primary": {
|
||
"content": "文档前2000字符...",
|
||
"source": "knowledge_base/api/payment-errors.md",
|
||
"matchType": "exact_L0",
|
||
"confidence": "high",
|
||
"availableSections": null
|
||
},
|
||
"supplement": {
|
||
"content": "Milvus检索到的相关片段...",
|
||
"source": "其他文档路径",
|
||
"matchType": "semantic_L1"
|
||
}
|
||
}
|
||
```
|
||
|
||
## 范围
|
||
|
||
### 核心功能(MVP)
|
||
1. ✅ FrontmatterParser:解析 YAML frontmatter(使用 snakeyaml)
|
||
2. ✅ KnowledgeIndexService:启动扫描 + 内存索引 + L0 精确匹配
|
||
3. ✅ 上传流程增强:保存本地 + 解析 frontmatter + 更新 L0 索引
|
||
4. ✅ LookupKnowledgeTool:L0 + L1 混合检索 + 条件调用
|
||
5. ✅ ApiDocument.metadata 字段扩展(存储 frontmatter JSON)
|
||
|
||
### 预留但不实现
|
||
- ⏸️ sections 分段加载(`availableSections` 返回 null)
|
||
- ⏸️ watchdog 热更新(重启生效)
|
||
- ⏸️ L0 索引持久化(内存索引,启动扫描)
|
||
|
||
## 非目标
|
||
|
||
- 不修改现有 VectorSearchService 逻辑
|
||
- 不修改 Milvus 索引结构
|
||
- 不实现文档版本管理
|
||
- 不支持其他文件格式(仅 .md)
|
||
|
||
## 技术选型
|
||
|
||
| 组件 | 技术选型 | 说明 |
|
||
|------|---------|------|
|
||
| YAML 解析 | snakeyaml 2.0 | 解析 frontmatter |
|
||
| L0 索引 | 内存 `List<KnowledgeEntry>` | 启动扫描,快速查询 |
|
||
| L1 检索 | 复用 VectorSearchService | Milvus + BGE-M3 |
|
||
| 文件存储 | 本地文件系统 | `knowledge_base/{category}/` |
|
||
|
||
## devflow 上下文约束
|
||
|
||
**必须遵守**(来自 phase1-infrastructure):
|
||
- 枚举存储为 VARCHAR,JPA 使用 `@Enumerated(EnumType.STRING)`
|
||
- Milvus collection 需 `loadCollection()`
|
||
- 复用现有 `VectorSearchService` 接口
|
||
- 文档元数据存入 `ApiDocument` 实体
|
||
|
||
**术语对齐**:
|
||
- `ApiDocument`:文档元数据实体,扩展 `metadata` 字段存储 frontmatter
|
||
- `category`:文档分类(api/domain/troubleshooting),与 Phase 1 对齐
|
||
|
||
### 关键假设
|
||
|
||
1. **L0 高置信度定义:唯一匹配**
|
||
- 假设:1 个匹配结果即为高置信度,不调用 L1
|
||
- 验证方式:✅ grill 阶段已确认
|
||
- 状态:已验证
|
||
|
||
2. **knowledge_base/ 目录权限**
|
||
- 假设:应用有读写权限
|
||
- 验证方式:启动时创建目录
|
||
- 风险:Docker 部署时路径映射
|
||
|
||
3. **TEXT 字段存储 JSON**
|
||
- 假设:TEXT 类型可存储 JSON 字符串(< 64KB)
|
||
- 验证方式:✅ grill 阶段已确认
|
||
- 状态:已验证
|
||
|
||
## 主要风险
|
||
|
||
### 风险 1:知识库目录权限问题
|
||
- **影响**:无法创建 knowledge_base/ 或保存文件
|
||
- **概率**:中(Docker 环境常见)
|
||
- **缓解**:启动时检查并创建目录,Docker 部署时正确挂载卷
|
||
- **检测**:apply 阶段测试文件保存功能
|
||
|
||
### 风险 2:L0 关键词匹配不准确
|
||
- **影响**:误匹配或漏匹配
|
||
- **概率**:中(依赖 frontmatter 质量)
|
||
- **缓解**:frontmatter keywords 需要精心维护,L1 作为兜底
|
||
- **后续**:引入模糊匹配或同义词扩展
|
||
|
||
### 风险 3:事务一致性(孤儿文件)
|
||
- **影响**:文件保存成功但事务回滚,产生孤儿文件
|
||
- **概率**:低
|
||
- **缓解**:异常时调用 cleanupLocalFile() 清理
|
||
- **检测**:集成测试验证
|
||
|
||
## 验收标准
|
||
|
||
### 功能验收
|
||
1. ✅ 上传带 frontmatter 的 .md 文档成功
|
||
2. ✅ L0 精确匹配:"ERR_TIMEOUT" → 返回完整文档(matchType=exact_L0)
|
||
3. ✅ L0 未匹配:"如何优化性能" → 降级到 L1(matchType=semantic_L1)
|
||
4. ✅ L0 多个匹配:"超时" → 返回 L0 列表 + L1 补充
|
||
5. ✅ 缺少 frontmatter 的文档只走 L1
|
||
|
||
### 性能验收
|
||
- L0 查询响应时间 < 10ms
|
||
- L0 + L1 组合查询 < 500ms
|
||
- 启动扫描时间 < 5s(假设 < 1000 个文档)
|
||
|
||
### 集成验收
|
||
- Agent 调用 `lookup_knowledge("ERR_TIMEOUT")` 返回正确文档
|
||
- Agent 调用 `lookup_knowledge("支付失败")` 返回语义相关文档
|