Files

658 lines
20 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Design: L0+L1 混合检索集成
## 架构概览
### 双层检索架构
```
┌─────────────────────────────────────────────────────────────┐
│ Agent (ReactAgent) │
└─────────────────────┬───────────────────────────────────────┘
│ 调用
▼
┌─────────────────────────────────────────────────────────────┐
│ LookupKnowledgeTool (新增) │
│ - lookup(query, sectionTitle) │
│ - 编排 L0 + L1 检索流程 │
└──────┬──────────────────────────┬───────────────────────────┘
│ │
│ L0 精确匹配 │ L1 语义检索(条件调用)
▼ ▼
┌──────────────────────┐ ┌──────────────────────────────┐
│ KnowledgeIndexService│ │ VectorSearchService (复用) │
│ (新增) │ │ - searchSimilarDocuments() │
│ - loadIndex() │ │ - Milvus + BGE-M3 │
│ - exactMatch() │ └──────────────────────────────┘
│ - readDocument() │
└──────┬───────────────┘
│ 读取
▼
┌──────────────────────────────────────────────────────────────┐
│ knowledge_base/ (本地文件系统) │
│ ├── api/ │
│ ├── domain/ │
│ └── troubleshooting/ │
└──────────────────────────────────────────────────────────────┘
```
### 上传流程增强
```
POST /api/documents/upload
│
▼
DocumentManagementService.uploadDocument()
│
├─ 1. 文件格式验证
├─ 2. 计算 hash(去重)
├─ 3. 提取文本 (TextExtractorService)
│
├─ 4. 【新增】保存原始文件到本地
│ └─ knowledge_base/{category}/{fileName}
│
├─ 5. 【新增】解析 frontmatter (FrontmatterParser)
│ └─ 提取 title, keywords, summary
│
├─ 6. 分块 (DocumentChunkService)
├─ 7. 向量化 + Milvus 索引 (VectorIndexService)
│
├─ 8. 保存元数据到 MySQL (ApiDocument)
│ └─ metadata 字段存储 frontmatter JSON
│
└─ 9. 【新增】更新 L0 内存索引
└─ KnowledgeIndexService.addToIndex()
```
---
## 核心组件设计
### 1. FrontmatterParser(新增)
**职责**:解析 Markdown 文件头的 YAML frontmatter
**依赖**:snakeyaml 2.0
**接口设计**:
```java
package com.superbiz.agent.service;
public class FrontmatterParser {
/**
* 解析 Markdown frontmatter
* @param content 完整文件内容
* @return Frontmatter 对象,如果不存在返回 null
*/
public Frontmatter parse(String content) {
// 1. 检查是否以 --- 开头
// 2. 提取 frontmatter 部分(两个 --- 之间)
// 3. 使用 Yaml.load() 解析
// 4. 映射到 Frontmatter 对象
}
/**
* 检查文件是否包含 frontmatter
*/
public boolean hasFrontmatter(String content) {
return content != null && content.trim().startsWith("---");
}
}
```
**数据模型**:
```java
package com.superbiz.agent.dto;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class Frontmatter {
private String title; // 必填
private List<String> keywords; // 必填
private String summary; // 必填
// 预留字段(MVP 不使用)
private String category; // 可选
private Map<String, String> sections; // 可选
private String version; // 可选
private String author; // 可选
private LocalDate lastUpdated; // 可选
}
```
---
### 2. KnowledgeIndexService(新增)
**职责**:L0 精确匹配索引管理
**启动扫描**:
```java
@Service
public class KnowledgeIndexService {
@Value("${knowledge.base-path}")
private String knowledgeBasePath; // 从配置文件读取
@Autowired
private FrontmatterParser frontmatterParser;
// 内存索引
private final List<KnowledgeEntry> knowledgeIndex =
new CopyOnWriteArrayList<>();
@PostConstruct
public void loadIndex() {
log.info("开始扫描知识库目录: {}", knowledgeBasePath);
// 1. 递归扫描 knowledge_base/
// 2. 过滤 .md 文件
// 3. 读取文件内容
// 4. 解析 frontmatter
// 5. 构建 KnowledgeEntry
// 6. 添加到 knowledgeIndex
log.info("知识库索引加载完成,共 {} 个文档", knowledgeIndex.size());
}
/**
* L0 精确匹配
* @param query 查询关键词
* @return 匹配的文档列表
*/
public List<KnowledgeEntry> exactMatch(String query) {
String queryLower = query.toLowerCase();
return knowledgeIndex.stream()
.filter(entry -> matchesKeywords(entry, queryLower))
.collect(Collectors.toList());
}
private boolean matchesKeywords(KnowledgeEntry entry, String query) {
// 关键词匹配(不区分大小写)
for (String keyword : entry.getKeywords()) {
if (query.contains(keyword.toLowerCase()) ||
keyword.toLowerCase().contains(query)) {
return true;
}
}
return false;
}
/**
* 读取文档内容
* @param filePath 文件路径
* @param maxChars 最大字符数
* @return 文档内容(前 maxChars 字符)
*/
public String readDocument(String filePath, int maxChars) {
try {
String content = Files.readString(Paths.get(filePath));
return content.length() > maxChars ?
content.substring(0, maxChars) + "..." : content;
} catch (IOException e) {
log.error("读取文档失败: {}", filePath, e);
return null;
}
}
/**
* 添加文档到索引(上传时调用)
*/
public void addToIndex(KnowledgeEntry entry) {
knowledgeIndex.add(entry);
log.debug("文档已添加到 L0 索引: {}", entry.getTitle());
}
/**
* 从索引中移除文档(删除时调用)
*/
public void removeFromIndex(String filePath) {
knowledgeIndex.removeIf(e -> e.getFilePath().equals(filePath));
log.debug("文档已从 L0 索引移除: {}", filePath);
}
}
```
**数据模型**:
```java
package com.superbiz.agent.dto;
@Data
@Builder
public class KnowledgeEntry {
private String filePath; // knowledge_base/api/payment-errors.md
private String title; // 支付网关错误码定义
private List<String> keywords; // [ERR_TIMEOUT, 超时, 支付网关]
private String summary; // 一句话摘要
private String category; // api/domain/troubleshooting
// 预留字段
private Map<String, String> sections;
}
```
---
### 3. LookupKnowledgeTool(新增)
**职责**:提供给 Agent 的混合检索工具
**实现**:
```java
package com.superbiz.agent.tool;
@Component
public class LookupKnowledgeTool {
@Autowired
private KnowledgeIndexService knowledgeIndexService;
@Autowired
private VectorSearchService vectorSearchService;
@Tool(
name = "lookup_knowledge",
description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。"
)
public LookupResult lookup(
@P("query") String query,
@P("section_title") String sectionTitle // 预留参数,MVP 返回 null
) {
log.info("收到知识库查询请求: query={}", query);
// Step 1: L0 精确匹配
List<KnowledgeEntry> l0Matches = knowledgeIndexService.exactMatch(query);
log.debug("L0 匹配结果: {} 个文档", l0Matches.size());
// Step 2: 判断是否高置信度
boolean highConfidence = (l0Matches.size() == 1);
// Step 3: L1 条件调用
List<VectorSearchService.SearchResult> l1Results = null;
if (!highConfidence) {
log.debug("L0 非唯一匹配,调用 L1 语义检索");
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, null);
}
// Step 4: 组装结果
return buildResult(l0Matches, l1Results, highConfidence);
}
private LookupResult buildResult(
List<KnowledgeEntry> l0Matches,
List<VectorSearchService.SearchResult> l1Results,
boolean highConfidence
) {
LookupResult result = new LookupResult();
result.setFound(!l0Matches.isEmpty() || (l1Results != null && !l1Results.isEmpty()));
// Primary: L0 结果
if (!l0Matches.isEmpty()) {
KnowledgeEntry first = l0Matches.get(0);
String content = knowledgeIndexService.readDocument(first.getFilePath(), 2000);
result.setPrimary(PrimaryResult.builder()
.content(content)
.source(first.getFilePath())
.matchType("exact_L0")
.confidence(highConfidence ? "high" : "low")
.availableSections(null) // MVP 返回 null
.build());
}
// Supplement: L1 结果
if (l1Results != null && !l1Results.isEmpty()) {
VectorSearchService.SearchResult firstL1 = l1Results.get(0);
result.setSupplement(SupplementResult.builder()
.content(firstL1.getContent())
.source(firstL1.getMetadata())
.matchType("semantic_L1")
.build());
}
return result;
}
}
```
**返回模型**:
```java
@Data
@Builder
public class LookupResult {
private boolean found;
private PrimaryResult primary;
private SupplementResult supplement;
}
@Data
@Builder
public class PrimaryResult {
private String content;
private String source;
private String matchType; // exact_L0
private String confidence; // high / low
private List<String> availableSections; // 预留字段
}
@Data
@Builder
public class SupplementResult {
private String content;
private String source;
private String matchType; // semantic_L1
}
```
---
### 4. DocumentManagementService(增强)
**变更点**:
**增加文件保存逻辑**:
```java
// 在 uploadDocument() 方法中,提取文本后增加
// 3. 提取文本
String text = textExtractorService.extractText(file, fileName);
// 【新增】4. 保存原始文件到本地
String category = request.getCategory() != null ? request.getCategory() : "default";
String localPath = saveToLocal(file, fileName, category);
// 【新增】5. 解析 frontmatter
Frontmatter frontmatter = null;
if (frontmatterParser.hasFrontmatter(text)) {
frontmatter = frontmatterParser.parse(text);
log.info("解析到 frontmatter: title={}, keywords={}",
frontmatter.getTitle(), frontmatter.getKeywords());
}
// 6. 分块(继续现有逻辑)
List<DocumentChunk> chunks = documentChunkService.chunkDocument(text, fileName);
```
**新增方法**:
```java
/**
* 保存文件到本地
*/
private String saveToLocal(MultipartFile file, String fileName, String category) {
try {
// 1. 构建目标路径
Path categoryDir = Paths.get(knowledgeBasePath, category);
Files.createDirectories(categoryDir);
Path targetPath = categoryDir.resolve(fileName);
// 2. 保存文件
file.transferTo(targetPath.toFile());
log.info("文件已保存到本地: {}", targetPath);
return targetPath.toString();
} catch (IOException e) {
throw new DocumentProcessException(
fileName, "save-local",
"保存文件到本地失败: " + e.getMessage(), e
);
}
}
/**
* 清理本地文件(事务回滚时调用)
*/
private void cleanupLocalFile(String localPath) {
if (localPath != null) {
try {
Files.deleteIfExists(Paths.get(localPath));
log.info("已清理本地文件: {}", localPath);
} catch (IOException e) {
log.warn("清理本地文件失败: {}", localPath, e);
}
}
}
```
**事务一致性处理**:
```java
@Transactional
public String uploadDocument(DocumentUploadRequest request) {
String localPath = null;
try {
// ... 提取文本
localPath = saveToLocal(file, fileName, category);
// ... frontmatter 解析
// ... 分块、向量化、保存到 MySQL
// ... 更新 L0 索引
} catch (Exception e) {
// 失败时清理本地文件
cleanupLocalFile(localPath);
throw e;
}
}
```
**更新 ApiDocument 保存**:
```java
// 创建文档元数据时增加字段
ApiDocument document = ApiDocument.builder()
.docId(docId)
.fileName(fileName)
.filePath(localPath) // 保存本地路径
.metadata(frontmatter != null ?
objectMapper.writeValueAsString(frontmatter) : null) // 存储 frontmatter JSON
// ... 其他字段
.build();
```
**更新 L0 索引**:
```java
// 索引成功后,如果有 frontmatter,更新 L0 索引
if (frontmatter != null) {
KnowledgeEntry entry = KnowledgeEntry.builder()
.filePath(localPath)
.title(frontmatter.getTitle())
.keywords(frontmatter.getKeywords())
.summary(frontmatter.getSummary())
.category(category)
.build();
knowledgeIndexService.addToIndex(entry);
}
```
---
### 5. ApiDocument 实体扩展
**新增字段**:
```java
@Entity
@Table(name = "api_document")
public class ApiDocument {
// ... 现有字段
// 【新增】frontmatter 元数据
@Column(name = "metadata", columnDefinition = "TEXT")
private String metadata; // JSON 格式存储
// 【新增】本地文件路径(现有 filePath 字段复用)
// 已有:@Column(name = "file_path", length = 512)
// private String filePath;
}
```
**Flyway 迁移脚本**:
```sql
-- V004__add_metadata_to_api_document.sql
ALTER TABLE api_document
ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)';
```
---
## 配置管理
**application.yml 新增配置**:
```yaml
# 知识库配置
knowledge:
base-path: knowledge_base/ # 知识库根目录
```
**pom.xml 新增依赖**:
```xml
<!-- YAML 解析 -->
<dependency>
<groupId>org.yaml</groupId>
<artifactId>snakeyaml</artifactId>
<version>2.0</version>
</dependency>
```
---
## 数据流时序图
### 上传流程时序图
```
User -> Controller: POST /api/documents/upload
Controller -> DocumentManagementService: uploadDocument(request)
DocumentManagementService -> TextExtractorService: extractText(file)
TextExtractorService --> DocumentManagementService: text
DocumentManagementService -> FileSystem: saveToLocal(file, category)
FileSystem --> DocumentManagementService: localPath
DocumentManagementService -> FrontmatterParser: parse(text)
FrontmatterParser --> DocumentManagementService: frontmatter
DocumentManagementService -> DocumentChunkService: chunkDocument(text)
DocumentChunkService --> DocumentManagementService: chunks
DocumentManagementService -> VectorIndexService: indexDocumentChunks(chunks)
VectorIndexService -> Milvus: insert vectors
Milvus --> VectorIndexService: success
DocumentManagementService -> ApiDocumentRepository: save(document)
ApiDocumentRepository --> DocumentManagementService: saved
DocumentManagementService -> KnowledgeIndexService: addToIndex(entry)
KnowledgeIndexService --> DocumentManagementService: indexed
DocumentManagementService --> Controller: docId
Controller --> User: {"code":200, "data":"doc-id"}
```
### 查询流程时序图
```
Agent -> LookupKnowledgeTool: lookup(query)
LookupKnowledgeTool -> KnowledgeIndexService: exactMatch(query)
KnowledgeIndexService --> LookupKnowledgeTool: l0Matches
alt 唯一匹配(高置信度)
LookupKnowledgeTool -> KnowledgeIndexService: readDocument(filePath)
KnowledgeIndexService -> FileSystem: read file
FileSystem --> KnowledgeIndexService: content
KnowledgeIndexService --> LookupKnowledgeTool: content
else 未匹配或多个匹配(低置信度)
LookupKnowledgeTool -> VectorSearchService: searchSimilarDocuments(query)
VectorSearchService -> Milvus: search vectors
Milvus --> VectorSearchService: l1Results
VectorSearchService --> LookupKnowledgeTool: l1Results
end
LookupKnowledgeTool --> Agent: LookupResult{primary, supplement}
```
---
## 关键决策记录
### 决策 1:文件保存策略
- **决策**:保存原始文件到本地文件系统
- **理由**:支持 L0 完整读取 + 未来扩展(版本管理、导出)
- **来源**:grill 阶段用户确认
### 决策 2:metadata 存储方式
- **决策**:TEXT 类型存储 JSON 字符串
- **理由**:简单直接,灵活扩展,无需自定义 JPA Converter
- **来源**:grill 阶段用户确认
### 决策 3:L0 高置信度标准
- **决策**:唯一匹配 = 高置信度,不调用 L1
- **理由**:唯一匹配通常就是用户想要的,调用 L1 只会增加延迟
- **来源**:grill 阶段用户确认
### 决策 4:knowledge_base/ 路径配置
- **决策**:通过 application.yml 配置,支持环境差异
- **理由**:开发环境和 Docker 环境路径可能不同
- **来源**:grill 阶段用户确认
---
## 非功能性设计
### 性能指标
- L0 查询响应时间:< 10ms
- L0 + L1 组合查询:< 500ms
- 启动扫描时间:< 5s(< 1000 个文档)
### 内存占用
- 单个 KnowledgeEntry:约 1KB(只存储 frontmatter 元数据)
- 1000 个文档:约 1MB(启动扫描只读取文件头)
- 10000 个文档:约 10MB
- **说明**:启动扫描只解析 frontmatter(< 1KB/文档),不读取全文;全文只在查询命中时按需读取
### 并发安全
- 使用 `CopyOnWriteArrayList` 存储索引(读多写少)
- 上传时更新索引(写操作)加锁或使用原子操作
### 错误处理
- frontmatter 解析失败:记录警告,文档仍可上传(只走 L1)
- 文件保存失败:抛出异常,回滚事务
- L0 索引加载失败:记录错误,应用仍可启动(只走 L1)
---
## 测试策略
### 单元测试
- FrontmatterParser 解析测试(有/无 frontmatter、格式错误)
- KnowledgeIndexService 匹配逻辑测试
- LookupKnowledgeTool 条件调用测试
### 集成测试
- 上传带 frontmatter 的文档 → 验证 L0 索引
- L0 精确匹配 → 验证返回正确文档
- L0 未命中 → 验证降级到 L1
### 性能测试
- L0 查询响应时间
- 大量文档启动扫描时间
---
## 实现优先级
### P0(MVP 必须)
1. FrontmatterParser
2. KnowledgeIndexService(启动扫描 + 精确匹配)
3. DocumentManagementService 增强
4. LookupKnowledgeTool
5. Flyway 迁移脚本
6. 配置管理
### P1(后续扩展)
- sections 分段加载
- watchdog 热更新
- L0 索引持久化
- 模糊匹配 / 同义词扩展