docs: archive historical openspec changes

This commit is contained in:
aruo
2026-07-05 14:04:32 +08:00
parent 63b62b28a2
commit b22f2d22c8
11 changed files with 0 additions and 0 deletions
@@ -0,0 +1,281 @@
# Proposal: L0+L1 混合检索集成
## 问题
当前只有 L1 向量语义检索(Milvus + BGE-M3),在遇到精确关键词查询时(如错误码 "ERR_TIMEOUT"、接口名 "PaymentGateway")效率不够高:
- 需要调用 embedding API 生成向量(约 100-300ms)
- 语义检索返回相似但可能不精确的结果
- 无法快速定位已知关键词对应的完整文档
Agent 需要一个"先精确、后语义"的混合检索工具。
## 建议方案
### 架构设计:双层检索
```
lookup_knowledge(query)
↓
L0: 精确关键词匹配(内存索引,< 10ms)
├─ 匹配成功 + 唯一结果 → 返回完整文档(高置信度)
└─ 未匹配 或 多个匹配 ↓
L1: 向量语义检索(Milvus,补充上下文)
└─ 返回 Top-K 相似片段
```
**核心机制**:
1. **L0 索引**:启动时扫描 `knowledge_base/` 目录,解析 Markdown frontmatter,构建内存索引
2. **L1 复用**:调用现有 `VectorSearchService.searchSimilarDocuments()`
3. **条件调用**:L0 唯一匹配时不调用 L1(减少延迟)
### 1. Frontmatter 规范
所有知识库文档(`knowledge_base/` 目录)需在文件头添加 YAML frontmatter:
```yaml
---
title: 支付网关错误码定义 # 必填
keywords: [ERR_TIMEOUT, 超时, 支付网关] # 必填,用于精确匹配
summary: 记录了支付网关所有核心错误码的含义及排查方向 # 必填
category: api # 可选,与现有 category 对齐
sections: # 预留字段(MVP 不实现)
超时排查: "## 1. 超时类错误"
---
# 文档正文
...
```
**约束**:
- frontmatter 必须在文件最顶部(前面不能有空行)
- `title`, `keywords`, `summary` 为必填字段
- 缺少 frontmatter 的文档允许上传,但不参与 L0 索引(只走 L1)
### 2. 上传流程增强
**现有流程**:
```
POST /api/documents/upload
↓
DocumentManagementService.uploadDocument()
↓
文本提取 → 分块 → 向量化 → Milvus 索引
↓
元数据存 MySQL (ApiDocument)
```
**增强后流程**:
```
POST /api/documents/upload
↓
1. 文本提取(内存)
2. 保存原始文件到:knowledge_base/{category}/{fileName}
3. 解析 frontmatter(FrontmatterParser)
4. 分块 → 向量化 → Milvus 索引
5. 元数据存 MySQL(ApiDocument.metadata 存储 frontmatter JSON)
6. 更新 L0 内存索引(KnowledgeIndexService)
```
**关键决策**(grill 阶段确认):
- ✅ 保存原始文件到本地(支持 L0 完整读取 + 未来扩展)
- ✅ metadata 字段:TEXT 类型存储 JSON 字符串
- ✅ ApiDocument.filePath 存储本地文件路径
- ✅ L0 高置信度 = 唯一匹配(不调用 L1)
- 按 category 分类存储:`knowledge_base/api/`, `knowledge_base/domain/`, `knowledge_base/troubleshooting/`
### 3. L0 索引服务
**KnowledgeIndexService**:
```java
@Service
public class KnowledgeIndexService {
// 内存索引结构
private List<KnowledgeEntry> knowledgeIndex = new ArrayList<>();
// 启动时扫描
@PostConstruct
public void loadIndex() {
// 递归扫描 knowledge_base/
// 解析 frontmatter
// 构建内存索引
}
// L0 精确匹配
public List<KnowledgeEntry> exactMatch(String query) {
// 关键词匹配(不区分大小写)
// 匹配规则:query 包含 keywords 中的任一词
}
// 读取文档内容
public String readDocument(String filePath, int maxChars) {
// 读取文件,返回前 maxChars 字符
}
}
```
**数据结构**:
```java
@Data
public class KnowledgeEntry {
private String filePath; // knowledge_base/api/payment-errors.md
private String title; // 支付网关错误码定义
private List<String> keywords; // [ERR_TIMEOUT, 超时, 支付网关]
private String summary; // 一句话摘要
private String category; // api
private Map<String, String> sections; // 预留字段
}
```
### 4. L1 复用
直接调用现有服务:
```java
@Autowired
private VectorSearchService vectorSearchService;
List<VectorSearchService.SearchResult> l1Results =
vectorSearchService.searchSimilarDocuments(query, 3, category);
```
### 5. 混合检索工具
**LookupKnowledgeTool**(供 Agent 调用):
```java
@Tool(name = "lookup_knowledge",
description = "查询知识库文档。优先精确匹配,自动补充语义相关片段。")
public LookupResult lookup(
@P("query") String query,
@P("section_title") String sectionTitle // 预留参数,MVP 不实现
) {
// Step 1: L0 精确匹配
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
// Step 2: 判断是否高置信度(唯一匹配)
boolean highConfidence = (l0Matches.size() == 1);
// Step 3: L1 条件调用
List<SearchResult> l1Results = null;
if (!highConfidence) {
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
}
// Step 4: 组装结果
return buildResult(l0Matches, l1Results, highConfidence);
}
```
**返回格式**:
```json
{
"found": true,
"primary": {
"content": "文档前2000字符...",
"source": "knowledge_base/api/payment-errors.md",
"matchType": "exact_L0",
"confidence": "high",
"availableSections": null
},
"supplement": {
"content": "Milvus检索到的相关片段...",
"source": "其他文档路径",
"matchType": "semantic_L1"
}
}
```
## 范围
### 核心功能(MVP)
1. ✅ FrontmatterParser:解析 YAML frontmatter(使用 snakeyaml)
2. ✅ KnowledgeIndexService:启动扫描 + 内存索引 + L0 精确匹配
3. ✅ 上传流程增强:保存本地 + 解析 frontmatter + 更新 L0 索引
4. ✅ LookupKnowledgeTool:L0 + L1 混合检索 + 条件调用
5. ✅ ApiDocument.metadata 字段扩展(存储 frontmatter JSON)
### 预留但不实现
- ⏸️ sections 分段加载(`availableSections` 返回 null)
- ⏸️ watchdog 热更新(重启生效)
- ⏸️ L0 索引持久化(内存索引,启动扫描)
## 非目标
- 不修改现有 VectorSearchService 逻辑
- 不修改 Milvus 索引结构
- 不实现文档版本管理
- 不支持其他文件格式(仅 .md)
## 技术选型
| 组件 | 技术选型 | 说明 |
|------|---------|------|
| YAML 解析 | snakeyaml 2.0 | 解析 frontmatter |
| L0 索引 | 内存 `List<KnowledgeEntry>` | 启动扫描,快速查询 |
| L1 检索 | 复用 VectorSearchService | Milvus + BGE-M3 |
| 文件存储 | 本地文件系统 | `knowledge_base/{category}/` |
## devflow 上下文约束
**必须遵守**(来自 phase1-infrastructure):
- 枚举存储为 VARCHAR,JPA 使用 `@Enumerated(EnumType.STRING)`
- Milvus collection 需 `loadCollection()`
- 复用现有 `VectorSearchService` 接口
- 文档元数据存入 `ApiDocument` 实体
**术语对齐**:
- `ApiDocument`:文档元数据实体,扩展 `metadata` 字段存储 frontmatter
- `category`:文档分类(api/domain/troubleshooting),与 Phase 1 对齐
### 关键假设
1. **L0 高置信度定义:唯一匹配**
- 假设:1 个匹配结果即为高置信度,不调用 L1
- 验证方式:✅ grill 阶段已确认
- 状态:已验证
2. **knowledge_base/ 目录权限**
- 假设:应用有读写权限
- 验证方式:启动时创建目录
- 风险:Docker 部署时路径映射
3. **TEXT 字段存储 JSON**
- 假设:TEXT 类型可存储 JSON 字符串(< 64KB)
- 验证方式:✅ grill 阶段已确认
- 状态:已验证
## 主要风险
### 风险 1:知识库目录权限问题
- **影响**:无法创建 knowledge_base/ 或保存文件
- **概率**:中(Docker 环境常见)
- **缓解**:启动时检查并创建目录,Docker 部署时正确挂载卷
- **检测**:apply 阶段测试文件保存功能
### 风险 2:L0 关键词匹配不准确
- **影响**:误匹配或漏匹配
- **概率**:中(依赖 frontmatter 质量)
- **缓解**:frontmatter keywords 需要精心维护,L1 作为兜底
- **后续**:引入模糊匹配或同义词扩展
### 风险 3:事务一致性(孤儿文件)
- **影响**:文件保存成功但事务回滚,产生孤儿文件
- **概率**:低
- **缓解**:异常时调用 cleanupLocalFile() 清理
- **检测**:集成测试验证
## 验收标准
### 功能验收
1. ✅ 上传带 frontmatter 的 .md 文档成功
2. ✅ L0 精确匹配:"ERR_TIMEOUT" → 返回完整文档(matchType=exact_L0)
3. ✅ L0 未匹配:"如何优化性能" → 降级到 L1(matchType=semantic_L1)
4. ✅ L0 多个匹配:"超时" → 返回 L0 列表 + L1 补充
5. ✅ 缺少 frontmatter 的文档只走 L1
### 性能验收
- L0 查询响应时间 < 10ms
- L0 + L1 组合查询 < 500ms
- 启动扫描时间 < 5s(假设 < 1000 个文档)
### 集成验收
- Agent 调用 `lookup_knowledge("ERR_TIMEOUT")` 返回正确文档
- Agent 调用 `lookup_knowledge("支付失败")` 返回语义相关文档