feat(knowledge): 完整实现知识库初始化 - 包含 Milvus 向量索引
## 核心改动
在上一版本基础上,补充完整的 Milvus (L1) 向量索引功能。
### 新增依赖注入
```java
@Autowired
private DocumentChunkService documentChunkService;
@Autowired
private VectorIndexService vectorIndexService;
@Autowired
private VectorEmbeddingService vectorEmbeddingService;
```
### 完整的数据流
```
knowledge_base/*.md
↓ 1. 扫描 & 解析 frontmatter
↓ 2. 保存到 MySQL (api_document)
↓ 3. 提取正文 & 文档分块
↓ 4. 生成向量并索引到 Milvus
↓ 5. 加入 L0 内存索引
完成 (L0 + L1 双层索引)
```
### 关键代码
```java
// 1. 提取正文(去除 frontmatter)
String body = extractBody(content);
// 2. 文档分块
List<DocumentChunk> chunks = documentChunkService.chunkDocument(body, relativePath);
// 3. 上传到 Milvus
vectorIndexService.indexDocumentChunks(document.getDocId(), chunks, category);
// 4. 更新状态
document.setStatus("INDEXED");
document.setChunkCount(chunks.size());
```
### 错误处理
- Milvus 索引失败时:
- 更新文档状态为 FAILED
- 记录错误信息到 error_message 字段
- 继续处理下一个文档(不中断整个流程)
### 响应示例
```json
{
"success": true,
"scanned": 6,
"inserted": 6,
"failed": 0,
"details": {
"api/payment-errors.md": "导入成功(L0+L1)"
}
}
```
### 数据库字段
新增:
- `chunk_count`:分块数量
- `error_message`:错误信息(失败时)
## 验证步骤
```bash
# 1. 启动应用(确保 Milvus 已运行)
mvn spring-boot:run
# 2. 初始化知识库
curl -X POST http://localhost:9900/api/knowledge/init
# 3. 验证结果
# - MySQL: 检查 api_document 表
# - Milvus: 检查 knowledge_base_collection
# - L0: 日志显示"知识库索引加载完成,共 6 个文档"
# 4. 测试 L1 语义检索
# lookup_knowledge("支付为什么会失败")
# 应该返回 semantic_L1 结果
```
## 文档更新
- 更新使用文档,删除"暂未实现 L1"的说明
- 添加 Milvus 数据结构说明
- 添加 Milvus 相关错误处理
This commit is contained in:
@@ -2,14 +2,14 @@
|
|||||||
|
|
||||||
## 概述
|
## 概述
|
||||||
|
|
||||||
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0)。
|
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0 + L1)。
|
||||||
|
|
||||||
**功能特点**:
|
**功能特点**:
|
||||||
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
|
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
|
||||||
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
|
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
|
||||||
3. ✅ **数据入库**:保存文档元数据到 MySQL
|
3. ✅ **数据入库**:保存文档元数据到 MySQL
|
||||||
4. ✅ **L0 索引**:自动加入内存精确匹配索引
|
4. ✅ **L0 索引**:自动加入内存精确匹配索引
|
||||||
5. ⏳ **L1 索引**:暂未实现,需要后续通过独立索引任务完成
|
5. ✅ **L1 索引**:文档分块并上传到 Milvus 向量数据库
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -46,12 +46,12 @@ curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
|||||||
"inserted": 6,
|
"inserted": 6,
|
||||||
"failed": 0,
|
"failed": 0,
|
||||||
"details": {
|
"details": {
|
||||||
"api/payment-errors.md": "导入成功(L0)",
|
"api/payment-errors.md": "导入成功(L0+L1)",
|
||||||
"domain/spring-ai-tool-best-practices.md": "导入成功(L0)",
|
"domain/spring-ai-tool-best-practices.md": "导入成功(L0+L1)",
|
||||||
"infrastructure/flyway-best-practices.md": "导入成功(L0)",
|
"infrastructure/flyway-best-practices.md": "导入成功(L0+L1)",
|
||||||
"infrastructure/mysql-connection-pool.md": "导入成功(L0)",
|
"infrastructure/mysql-connection-pool.md": "导入成功(L0+L1)",
|
||||||
"infrastructure/redis-config.md": "导入成功(L0)",
|
"infrastructure/redis-config.md": "导入成功(L0+L1)",
|
||||||
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0)"
|
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0+L1)"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -82,7 +82,7 @@ curl http://localhost:9900/api/knowledge/stats
|
|||||||
{
|
{
|
||||||
"success": true,
|
"success": true,
|
||||||
"totalDocuments": 6,
|
"totalDocuments": 6,
|
||||||
"totalVectors": 0,
|
"totalVectors": 48,
|
||||||
"categories": {
|
"categories": {
|
||||||
"api": 1,
|
"api": 1,
|
||||||
"domain": 1,
|
"domain": 1,
|
||||||
@@ -94,7 +94,7 @@ curl http://localhost:9900/api/knowledge/stats
|
|||||||
|
|
||||||
**字段说明**:
|
**字段说明**:
|
||||||
- `totalDocuments`:数据库中的文档总数
|
- `totalDocuments`:数据库中的文档总数
|
||||||
- `totalVectors`:Milvus 中的向量总数(当前为 0,未实现)
|
- `totalVectors`:Milvus 中的向量总数(chunk 数量)
|
||||||
- `categories`:按分类统计的文档数量
|
- `categories`:按分类统计的文档数量
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -198,6 +198,28 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
|||||||
|
|
||||||
## 数据存储
|
## 数据存储
|
||||||
|
|
||||||
|
### 完整的数据流
|
||||||
|
|
||||||
|
```
|
||||||
|
knowledge_base/*.md
|
||||||
|
↓ 1. 扫描
|
||||||
|
KnowledgeBaseInitService
|
||||||
|
↓ 2. 解析 frontmatter
|
||||||
|
Frontmatter (title, keywords, summary)
|
||||||
|
↓ 3. 保存到数据库
|
||||||
|
MySQL (api_document)
|
||||||
|
↓ 4. 提取正文 & 分块
|
||||||
|
DocumentChunkService
|
||||||
|
↓ 5. 生成向量
|
||||||
|
VectorEmbeddingService
|
||||||
|
↓ 6. 索引到 Milvus
|
||||||
|
Milvus (L1 向量索引)
|
||||||
|
↓ 7. 加入内存索引
|
||||||
|
KnowledgeIndexService (L0)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
### 数据库表结构(api_document)
|
### 数据库表结构(api_document)
|
||||||
|
|
||||||
| 字段 | 类型 | 说明 | 示例 |
|
| 字段 | 类型 | 说明 | 示例 |
|
||||||
@@ -207,7 +229,9 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
|||||||
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
|
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
|
||||||
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
|
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
|
||||||
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
|
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
|
||||||
| `status` | VARCHAR(16) | 状态 | INDEXED |
|
| `status` | VARCHAR(16) | 状态 | INDEXED / FAILED |
|
||||||
|
| `chunk_count` | INT | 分块数量 | 8 |
|
||||||
|
| `error_message` | TEXT | 错误信息 | null |
|
||||||
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
|
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
|
||||||
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
|
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
|
||||||
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
|
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
|
||||||
@@ -225,6 +249,26 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
### Milvus 向量索引
|
||||||
|
|
||||||
|
每个文档会被分块(chunk)并生成向量,存储到 Milvus 集合中:
|
||||||
|
|
||||||
|
**Collection**: `knowledge_base_collection`
|
||||||
|
|
||||||
|
**字段**:
|
||||||
|
- `doc_id`:文档 ID
|
||||||
|
- `chunk_id`:分块 ID
|
||||||
|
- `chunk_text`:分块文本内容
|
||||||
|
- `embedding`:768 维向量
|
||||||
|
- `category`:文档分类
|
||||||
|
- `file_path`:文件路径
|
||||||
|
|
||||||
|
**分块策略**:
|
||||||
|
- Chunk Size:根据 `DocumentChunkConfig` 配置(默认 500 token)
|
||||||
|
- Overlap:重叠区域(默认 50 token)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## L0 内存索引
|
## L0 内存索引
|
||||||
|
|
||||||
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
|
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
|
||||||
@@ -305,6 +349,61 @@ category: api
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
### 问题 4: Milvus 连接失败
|
||||||
|
|
||||||
|
**症状**:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": true,
|
||||||
|
"scanned": 6,
|
||||||
|
"inserted": 0,
|
||||||
|
"failed": 6,
|
||||||
|
"details": {
|
||||||
|
"api/payment-errors.md": "Milvus 索引失败: Connection refused"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**原因**:
|
||||||
|
- Milvus 服务未启动
|
||||||
|
- 网络连接问题
|
||||||
|
- 配置错误
|
||||||
|
|
||||||
|
**解决**:
|
||||||
|
```bash
|
||||||
|
# 检查 Milvus 是否运行
|
||||||
|
docker ps | grep milvus
|
||||||
|
|
||||||
|
# 检查配置
|
||||||
|
grep milvus application.yml
|
||||||
|
|
||||||
|
# 启动 Milvus
|
||||||
|
docker-compose up -d milvus-standalone
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 问题 5: 文档分块失败
|
||||||
|
|
||||||
|
**症状**:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"details": {
|
||||||
|
"test/large-doc.md": "Milvus 索引失败: Document too large"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**原因**:
|
||||||
|
- 文档内容过大
|
||||||
|
- 分块配置不当
|
||||||
|
|
||||||
|
**解决**:
|
||||||
|
- 检查 `DocumentChunkConfig` 配置
|
||||||
|
- 调整 chunk size 和 overlap
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
#### 3. 文档缺少标题
|
#### 3. 文档缺少标题
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -318,38 +417,6 @@ category: api
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 后续扩展(L1 向量索引)
|
|
||||||
|
|
||||||
当前版本暂未实现 L1 向量索引(Milvus),计划后续扩展:
|
|
||||||
|
|
||||||
### 扩展方案
|
|
||||||
|
|
||||||
1. **独立索引任务**:
|
|
||||||
```bash
|
|
||||||
POST /api/knowledge/build-vectors
|
|
||||||
```
|
|
||||||
- 读取数据库中所有文档
|
|
||||||
- 调用 `DocumentManagementService` 处理分块
|
|
||||||
- 上传到 Milvus
|
|
||||||
|
|
||||||
2. **或者修改当前接口**:
|
|
||||||
- 在 `initializeKnowledgeBase` 中调用文档分块和向量索引
|
|
||||||
- 需要处理大文件的分块逻辑
|
|
||||||
|
|
||||||
### 验证 L1 的方法(未来)
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# 1. 调用 L1 索引构建
|
|
||||||
curl -X POST http://localhost:9900/api/knowledge/build-vectors
|
|
||||||
|
|
||||||
# 2. 查询统计信息
|
|
||||||
curl http://localhost:9900/api/knowledge/stats
|
|
||||||
|
|
||||||
# 3. 确认 totalVectors > 0
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 最佳实践
|
## 最佳实践
|
||||||
|
|
||||||
### ✅ 推荐做法
|
### ✅ 推荐做法
|
||||||
|
|||||||
@@ -4,6 +4,7 @@ import com.superbiz.agent.domain.entity.ApiDocument;
|
|||||||
import com.superbiz.agent.repository.ApiDocumentRepository;
|
import com.superbiz.agent.repository.ApiDocumentRepository;
|
||||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||||
import com.superbiz.agent.dto.Frontmatter;
|
import com.superbiz.agent.dto.Frontmatter;
|
||||||
|
import com.superbiz.agent.dto.DocumentChunk;
|
||||||
import lombok.Data;
|
import lombok.Data;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
@@ -37,6 +38,15 @@ public class KnowledgeBaseInitService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private FrontmatterParser frontmatterParser;
|
private FrontmatterParser frontmatterParser;
|
||||||
|
|
||||||
|
@Autowired
|
||||||
|
private DocumentChunkService documentChunkService;
|
||||||
|
|
||||||
|
@Autowired
|
||||||
|
private VectorIndexService vectorIndexService;
|
||||||
|
|
||||||
|
@Autowired
|
||||||
|
private VectorEmbeddingService vectorEmbeddingService;
|
||||||
|
|
||||||
@Autowired
|
@Autowired
|
||||||
private KnowledgeIndexService knowledgeIndexService;
|
private KnowledgeIndexService knowledgeIndexService;
|
||||||
|
|
||||||
@@ -112,6 +122,35 @@ public class KnowledgeBaseInitService {
|
|||||||
// 保存到数据库
|
// 保存到数据库
|
||||||
ApiDocument document = saveToDatabase(relativePath, title, summary, category, content, keywords);
|
ApiDocument document = saveToDatabase(relativePath, title, summary, category, content, keywords);
|
||||||
|
|
||||||
|
// 提取文档正文(去除 frontmatter)
|
||||||
|
String body = extractBody(content);
|
||||||
|
|
||||||
|
// 文档分块
|
||||||
|
List<DocumentChunk> chunks = documentChunkService.chunkDocument(body, relativePath);
|
||||||
|
logger.debug("文档分块完成: {} -> {} 个 chunk", relativePath, chunks.size());
|
||||||
|
|
||||||
|
// 上传到 Milvus
|
||||||
|
try {
|
||||||
|
vectorIndexService.indexDocumentChunks(document.getDocId(), chunks, category);
|
||||||
|
|
||||||
|
document.setStatus("INDEXED");
|
||||||
|
document.setChunkCount(chunks.size());
|
||||||
|
apiDocumentRepository.save(document);
|
||||||
|
|
||||||
|
logger.info("文档已索引到 Milvus: {} (docId={}, chunks={})",
|
||||||
|
title, document.getDocId(), chunks.size());
|
||||||
|
} catch (Exception e) {
|
||||||
|
logger.error("上传到 Milvus 失败: {}", relativePath, e);
|
||||||
|
|
||||||
|
document.setStatus("FAILED");
|
||||||
|
document.setErrorMessage(e.getMessage());
|
||||||
|
apiDocumentRepository.save(document);
|
||||||
|
|
||||||
|
result.incrementFailed();
|
||||||
|
result.addDetail(relativePath, "Milvus 索引失败: " + e.getMessage());
|
||||||
|
continue; // 跳过该文档,继续处理下一个
|
||||||
|
}
|
||||||
|
|
||||||
// 添加到 L0 内存索引
|
// 添加到 L0 内存索引
|
||||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||||
.filePath(relativePath)
|
.filePath(relativePath)
|
||||||
@@ -122,12 +161,9 @@ public class KnowledgeBaseInitService {
|
|||||||
.build();
|
.build();
|
||||||
knowledgeIndexService.addToIndex(entry);
|
knowledgeIndexService.addToIndex(entry);
|
||||||
|
|
||||||
// TODO: 上传到 Milvus (L1) - 需要通过独立的索引任务完成
|
|
||||||
// 当前版本只处理数据库入库和 L0 索引
|
|
||||||
|
|
||||||
result.incrementInserted();
|
result.incrementInserted();
|
||||||
result.addDetail(relativePath, "导入成功(L0)");
|
result.addDetail(relativePath, "导入成功(L0+L1)");
|
||||||
logger.info("文档导入成功: {} -> {} (L0 索引已更新)", relativePath, title);
|
logger.info("文档导入成功: {} -> {} (L0+L1 索引已更新)", relativePath, title);
|
||||||
|
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
logger.error("处理文档失败: {}", relativePath, e);
|
logger.error("处理文档失败: {}", relativePath, e);
|
||||||
|
|||||||
Reference in New Issue
Block a user