docs: archive historical openspec changes
This commit is contained in:
@@ -0,0 +1,281 @@
|
||||
# Proposal: L0+L1 混合检索集成
|
||||
|
||||
## 问题
|
||||
|
||||
当前只有 L1 向量语义检索(Milvus + BGE-M3),在遇到精确关键词查询时(如错误码 "ERR_TIMEOUT"、接口名 "PaymentGateway")效率不够高:
|
||||
- 需要调用 embedding API 生成向量(约 100-300ms)
|
||||
- 语义检索返回相似但可能不精确的结果
|
||||
- 无法快速定位已知关键词对应的完整文档
|
||||
|
||||
Agent 需要一个"先精确、后语义"的混合检索工具。
|
||||
|
||||
## 建议方案
|
||||
|
||||
### 架构设计:双层检索
|
||||
|
||||
```
|
||||
lookup_knowledge(query)
|
||||
↓
|
||||
L0: 精确关键词匹配(内存索引,< 10ms)
|
||||
├─ 匹配成功 + 唯一结果 → 返回完整文档(高置信度)
|
||||
└─ 未匹配 或 多个匹配 ↓
|
||||
L1: 向量语义检索(Milvus,补充上下文)
|
||||
└─ 返回 Top-K 相似片段
|
||||
```
|
||||
|
||||
**核心机制**:
|
||||
1. **L0 索引**:启动时扫描 `knowledge_base/` 目录,解析 Markdown frontmatter,构建内存索引
|
||||
2. **L1 复用**:调用现有 `VectorSearchService.searchSimilarDocuments()`
|
||||
3. **条件调用**:L0 唯一匹配时不调用 L1(减少延迟)
|
||||
|
||||
### 1. Frontmatter 规范
|
||||
|
||||
所有知识库文档(`knowledge_base/` 目录)需在文件头添加 YAML frontmatter:
|
||||
|
||||
```yaml
|
||||
---
|
||||
title: 支付网关错误码定义 # 必填
|
||||
keywords: [ERR_TIMEOUT, 超时, 支付网关] # 必填,用于精确匹配
|
||||
summary: 记录了支付网关所有核心错误码的含义及排查方向 # 必填
|
||||
category: api # 可选,与现有 category 对齐
|
||||
sections: # 预留字段(MVP 不实现)
|
||||
超时排查: "## 1. 超时类错误"
|
||||
---
|
||||
|
||||
# 文档正文
|
||||
...
|
||||
```
|
||||
|
||||
**约束**:
|
||||
- frontmatter 必须在文件最顶部(前面不能有空行)
|
||||
- `title`, `keywords`, `summary` 为必填字段
|
||||
- 缺少 frontmatter 的文档允许上传,但不参与 L0 索引(只走 L1)
|
||||
|
||||
### 2. 上传流程增强
|
||||
|
||||
**现有流程**:
|
||||
```
|
||||
POST /api/documents/upload
|
||||
↓
|
||||
DocumentManagementService.uploadDocument()
|
||||
↓
|
||||
文本提取 → 分块 → 向量化 → Milvus 索引
|
||||
↓
|
||||
元数据存 MySQL (ApiDocument)
|
||||
```
|
||||
|
||||
**增强后流程**:
|
||||
```
|
||||
POST /api/documents/upload
|
||||
↓
|
||||
1. 文本提取(内存)
|
||||
2. 保存原始文件到:knowledge_base/{category}/{fileName}
|
||||
3. 解析 frontmatter(FrontmatterParser)
|
||||
4. 分块 → 向量化 → Milvus 索引
|
||||
5. 元数据存 MySQL(ApiDocument.metadata 存储 frontmatter JSON)
|
||||
6. 更新 L0 内存索引(KnowledgeIndexService)
|
||||
```
|
||||
|
||||
**关键决策**(grill 阶段确认):
|
||||
- ✅ 保存原始文件到本地(支持 L0 完整读取 + 未来扩展)
|
||||
- ✅ metadata 字段:TEXT 类型存储 JSON 字符串
|
||||
- ✅ ApiDocument.filePath 存储本地文件路径
|
||||
- ✅ L0 高置信度 = 唯一匹配(不调用 L1)
|
||||
- 按 category 分类存储:`knowledge_base/api/`, `knowledge_base/domain/`, `knowledge_base/troubleshooting/`
|
||||
|
||||
### 3. L0 索引服务
|
||||
|
||||
**KnowledgeIndexService**:
|
||||
```java
|
||||
@Service
|
||||
public class KnowledgeIndexService {
|
||||
// 内存索引结构
|
||||
private List<KnowledgeEntry> knowledgeIndex = new ArrayList<>();
|
||||
|
||||
// 启动时扫描
|
||||
@PostConstruct
|
||||
public void loadIndex() {
|
||||
// 递归扫描 knowledge_base/
|
||||
// 解析 frontmatter
|
||||
// 构建内存索引
|
||||
}
|
||||
|
||||
// L0 精确匹配
|
||||
public List<KnowledgeEntry> exactMatch(String query) {
|
||||
// 关键词匹配(不区分大小写)
|
||||
// 匹配规则:query 包含 keywords 中的任一词
|
||||
}
|
||||
|
||||
// 读取文档内容
|
||||
public String readDocument(String filePath, int maxChars) {
|
||||
// 读取文件,返回前 maxChars 字符
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**数据结构**:
|
||||
```java
|
||||
@Data
|
||||
public class KnowledgeEntry {
|
||||
private String filePath; // knowledge_base/api/payment-errors.md
|
||||
private String title; // 支付网关错误码定义
|
||||
private List<String> keywords; // [ERR_TIMEOUT, 超时, 支付网关]
|
||||
private String summary; // 一句话摘要
|
||||
private String category; // api
|
||||
private Map<String, String> sections; // 预留字段
|
||||
}
|
||||
```
|
||||
|
||||
### 4. L1 复用
|
||||
|
||||
直接调用现有服务:
|
||||
```java
|
||||
@Autowired
|
||||
private VectorSearchService vectorSearchService;
|
||||
|
||||
List<VectorSearchService.SearchResult> l1Results =
|
||||
vectorSearchService.searchSimilarDocuments(query, 3, category);
|
||||
```
|
||||
|
||||
### 5. 混合检索工具
|
||||
|
||||
**LookupKnowledgeTool**(供 Agent 调用):
|
||||
```java
|
||||
@Tool(name = "lookup_knowledge",
|
||||
description = "查询知识库文档。优先精确匹配,自动补充语义相关片段。")
|
||||
public LookupResult lookup(
|
||||
@P("query") String query,
|
||||
@P("section_title") String sectionTitle // 预留参数,MVP 不实现
|
||||
) {
|
||||
// Step 1: L0 精确匹配
|
||||
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
|
||||
|
||||
// Step 2: 判断是否高置信度(唯一匹配)
|
||||
boolean highConfidence = (l0Matches.size() == 1);
|
||||
|
||||
// Step 3: L1 条件调用
|
||||
List<SearchResult> l1Results = null;
|
||||
if (!highConfidence) {
|
||||
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
|
||||
}
|
||||
|
||||
// Step 4: 组装结果
|
||||
return buildResult(l0Matches, l1Results, highConfidence);
|
||||
}
|
||||
```
|
||||
|
||||
**返回格式**:
|
||||
```json
|
||||
{
|
||||
"found": true,
|
||||
"primary": {
|
||||
"content": "文档前2000字符...",
|
||||
"source": "knowledge_base/api/payment-errors.md",
|
||||
"matchType": "exact_L0",
|
||||
"confidence": "high",
|
||||
"availableSections": null
|
||||
},
|
||||
"supplement": {
|
||||
"content": "Milvus检索到的相关片段...",
|
||||
"source": "其他文档路径",
|
||||
"matchType": "semantic_L1"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 范围
|
||||
|
||||
### 核心功能(MVP)
|
||||
1. ✅ FrontmatterParser:解析 YAML frontmatter(使用 snakeyaml)
|
||||
2. ✅ KnowledgeIndexService:启动扫描 + 内存索引 + L0 精确匹配
|
||||
3. ✅ 上传流程增强:保存本地 + 解析 frontmatter + 更新 L0 索引
|
||||
4. ✅ LookupKnowledgeTool:L0 + L1 混合检索 + 条件调用
|
||||
5. ✅ ApiDocument.metadata 字段扩展(存储 frontmatter JSON)
|
||||
|
||||
### 预留但不实现
|
||||
- ⏸️ sections 分段加载(`availableSections` 返回 null)
|
||||
- ⏸️ watchdog 热更新(重启生效)
|
||||
- ⏸️ L0 索引持久化(内存索引,启动扫描)
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不修改现有 VectorSearchService 逻辑
|
||||
- 不修改 Milvus 索引结构
|
||||
- 不实现文档版本管理
|
||||
- 不支持其他文件格式(仅 .md)
|
||||
|
||||
## 技术选型
|
||||
|
||||
| 组件 | 技术选型 | 说明 |
|
||||
|------|---------|------|
|
||||
| YAML 解析 | snakeyaml 2.0 | 解析 frontmatter |
|
||||
| L0 索引 | 内存 `List<KnowledgeEntry>` | 启动扫描,快速查询 |
|
||||
| L1 检索 | 复用 VectorSearchService | Milvus + BGE-M3 |
|
||||
| 文件存储 | 本地文件系统 | `knowledge_base/{category}/` |
|
||||
|
||||
## devflow 上下文约束
|
||||
|
||||
**必须遵守**(来自 phase1-infrastructure):
|
||||
- 枚举存储为 VARCHAR,JPA 使用 `@Enumerated(EnumType.STRING)`
|
||||
- Milvus collection 需 `loadCollection()`
|
||||
- 复用现有 `VectorSearchService` 接口
|
||||
- 文档元数据存入 `ApiDocument` 实体
|
||||
|
||||
**术语对齐**:
|
||||
- `ApiDocument`:文档元数据实体,扩展 `metadata` 字段存储 frontmatter
|
||||
- `category`:文档分类(api/domain/troubleshooting),与 Phase 1 对齐
|
||||
|
||||
### 关键假设
|
||||
|
||||
1. **L0 高置信度定义:唯一匹配**
|
||||
- 假设:1 个匹配结果即为高置信度,不调用 L1
|
||||
- 验证方式:✅ grill 阶段已确认
|
||||
- 状态:已验证
|
||||
|
||||
2. **knowledge_base/ 目录权限**
|
||||
- 假设:应用有读写权限
|
||||
- 验证方式:启动时创建目录
|
||||
- 风险:Docker 部署时路径映射
|
||||
|
||||
3. **TEXT 字段存储 JSON**
|
||||
- 假设:TEXT 类型可存储 JSON 字符串(< 64KB)
|
||||
- 验证方式:✅ grill 阶段已确认
|
||||
- 状态:已验证
|
||||
|
||||
## 主要风险
|
||||
|
||||
### 风险 1:知识库目录权限问题
|
||||
- **影响**:无法创建 knowledge_base/ 或保存文件
|
||||
- **概率**:中(Docker 环境常见)
|
||||
- **缓解**:启动时检查并创建目录,Docker 部署时正确挂载卷
|
||||
- **检测**:apply 阶段测试文件保存功能
|
||||
|
||||
### 风险 2:L0 关键词匹配不准确
|
||||
- **影响**:误匹配或漏匹配
|
||||
- **概率**:中(依赖 frontmatter 质量)
|
||||
- **缓解**:frontmatter keywords 需要精心维护,L1 作为兜底
|
||||
- **后续**:引入模糊匹配或同义词扩展
|
||||
|
||||
### 风险 3:事务一致性(孤儿文件)
|
||||
- **影响**:文件保存成功但事务回滚,产生孤儿文件
|
||||
- **概率**:低
|
||||
- **缓解**:异常时调用 cleanupLocalFile() 清理
|
||||
- **检测**:集成测试验证
|
||||
|
||||
## 验收标准
|
||||
|
||||
### 功能验收
|
||||
1. ✅ 上传带 frontmatter 的 .md 文档成功
|
||||
2. ✅ L0 精确匹配:"ERR_TIMEOUT" → 返回完整文档(matchType=exact_L0)
|
||||
3. ✅ L0 未匹配:"如何优化性能" → 降级到 L1(matchType=semantic_L1)
|
||||
4. ✅ L0 多个匹配:"超时" → 返回 L0 列表 + L1 补充
|
||||
5. ✅ 缺少 frontmatter 的文档只走 L1
|
||||
|
||||
### 性能验收
|
||||
- L0 查询响应时间 < 10ms
|
||||
- L0 + L1 组合查询 < 500ms
|
||||
- 启动扫描时间 < 5s(假设 < 1000 个文档)
|
||||
|
||||
### 集成验收
|
||||
- Agent 调用 `lookup_knowledge("ERR_TIMEOUT")` 返回正确文档
|
||||
- Agent 调用 `lookup_knowledge("支付失败")` 返回语义相关文档
|
||||
Reference in New Issue
Block a user