feat(knowledge): 完整实现知识库初始化 - 包含 Milvus 向量索引
## 核心改动
在上一版本基础上,补充完整的 Milvus (L1) 向量索引功能。
### 新增依赖注入
```java
@Autowired
private DocumentChunkService documentChunkService;
@Autowired
private VectorIndexService vectorIndexService;
@Autowired
private VectorEmbeddingService vectorEmbeddingService;
```
### 完整的数据流
```
knowledge_base/*.md
↓ 1. 扫描 & 解析 frontmatter
↓ 2. 保存到 MySQL (api_document)
↓ 3. 提取正文 & 文档分块
↓ 4. 生成向量并索引到 Milvus
↓ 5. 加入 L0 内存索引
完成 (L0 + L1 双层索引)
```
### 关键代码
```java
// 1. 提取正文(去除 frontmatter)
String body = extractBody(content);
// 2. 文档分块
List<DocumentChunk> chunks = documentChunkService.chunkDocument(body, relativePath);
// 3. 上传到 Milvus
vectorIndexService.indexDocumentChunks(document.getDocId(), chunks, category);
// 4. 更新状态
document.setStatus("INDEXED");
document.setChunkCount(chunks.size());
```
### 错误处理
- Milvus 索引失败时:
- 更新文档状态为 FAILED
- 记录错误信息到 error_message 字段
- 继续处理下一个文档(不中断整个流程)
### 响应示例
```json
{
"success": true,
"scanned": 6,
"inserted": 6,
"failed": 0,
"details": {
"api/payment-errors.md": "导入成功(L0+L1)"
}
}
```
### 数据库字段
新增:
- `chunk_count`:分块数量
- `error_message`:错误信息(失败时)
## 验证步骤
```bash
# 1. 启动应用(确保 Milvus 已运行)
mvn spring-boot:run
# 2. 初始化知识库
curl -X POST http://localhost:9900/api/knowledge/init
# 3. 验证结果
# - MySQL: 检查 api_document 表
# - Milvus: 检查 knowledge_base_collection
# - L0: 日志显示"知识库索引加载完成,共 6 个文档"
# 4. 测试 L1 语义检索
# lookup_knowledge("支付为什么会失败")
# 应该返回 semantic_L1 结果
```
## 文档更新
- 更新使用文档,删除"暂未实现 L1"的说明
- 添加 Milvus 数据结构说明
- 添加 Milvus 相关错误处理
This commit is contained in:
@@ -2,14 +2,14 @@
|
||||
|
||||
## 概述
|
||||
|
||||
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0)。
|
||||
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0 + L1)。
|
||||
|
||||
**功能特点**:
|
||||
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
|
||||
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
|
||||
3. ✅ **数据入库**:保存文档元数据到 MySQL
|
||||
4. ✅ **L0 索引**:自动加入内存精确匹配索引
|
||||
5. ⏳ **L1 索引**:暂未实现,需要后续通过独立索引任务完成
|
||||
5. ✅ **L1 索引**:文档分块并上传到 Milvus 向量数据库
|
||||
|
||||
---
|
||||
|
||||
@@ -46,12 +46,12 @@ curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
"inserted": 6,
|
||||
"failed": 0,
|
||||
"details": {
|
||||
"api/payment-errors.md": "导入成功(L0)",
|
||||
"domain/spring-ai-tool-best-practices.md": "导入成功(L0)",
|
||||
"infrastructure/flyway-best-practices.md": "导入成功(L0)",
|
||||
"infrastructure/mysql-connection-pool.md": "导入成功(L0)",
|
||||
"infrastructure/redis-config.md": "导入成功(L0)",
|
||||
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0)"
|
||||
"api/payment-errors.md": "导入成功(L0+L1)",
|
||||
"domain/spring-ai-tool-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/flyway-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/mysql-connection-pool.md": "导入成功(L0+L1)",
|
||||
"infrastructure/redis-config.md": "导入成功(L0+L1)",
|
||||
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0+L1)"
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -82,7 +82,7 @@ curl http://localhost:9900/api/knowledge/stats
|
||||
{
|
||||
"success": true,
|
||||
"totalDocuments": 6,
|
||||
"totalVectors": 0,
|
||||
"totalVectors": 48,
|
||||
"categories": {
|
||||
"api": 1,
|
||||
"domain": 1,
|
||||
@@ -94,7 +94,7 @@ curl http://localhost:9900/api/knowledge/stats
|
||||
|
||||
**字段说明**:
|
||||
- `totalDocuments`:数据库中的文档总数
|
||||
- `totalVectors`:Milvus 中的向量总数(当前为 0,未实现)
|
||||
- `totalVectors`:Milvus 中的向量总数(chunk 数量)
|
||||
- `categories`:按分类统计的文档数量
|
||||
|
||||
---
|
||||
@@ -198,6 +198,28 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
||||
|
||||
## 数据存储
|
||||
|
||||
### 完整的数据流
|
||||
|
||||
```
|
||||
knowledge_base/*.md
|
||||
↓ 1. 扫描
|
||||
KnowledgeBaseInitService
|
||||
↓ 2. 解析 frontmatter
|
||||
Frontmatter (title, keywords, summary)
|
||||
↓ 3. 保存到数据库
|
||||
MySQL (api_document)
|
||||
↓ 4. 提取正文 & 分块
|
||||
DocumentChunkService
|
||||
↓ 5. 生成向量
|
||||
VectorEmbeddingService
|
||||
↓ 6. 索引到 Milvus
|
||||
Milvus (L1 向量索引)
|
||||
↓ 7. 加入内存索引
|
||||
KnowledgeIndexService (L0)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 数据库表结构(api_document)
|
||||
|
||||
| 字段 | 类型 | 说明 | 示例 |
|
||||
@@ -207,7 +229,9 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
||||
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
|
||||
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
|
||||
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
|
||||
| `status` | VARCHAR(16) | 状态 | INDEXED |
|
||||
| `status` | VARCHAR(16) | 状态 | INDEXED / FAILED |
|
||||
| `chunk_count` | INT | 分块数量 | 8 |
|
||||
| `error_message` | TEXT | 错误信息 | null |
|
||||
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
|
||||
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
|
||||
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
|
||||
@@ -225,6 +249,26 @@ if (!force && existingFilePaths.contains(relativePath)) {
|
||||
|
||||
---
|
||||
|
||||
### Milvus 向量索引
|
||||
|
||||
每个文档会被分块(chunk)并生成向量,存储到 Milvus 集合中:
|
||||
|
||||
**Collection**: `knowledge_base_collection`
|
||||
|
||||
**字段**:
|
||||
- `doc_id`:文档 ID
|
||||
- `chunk_id`:分块 ID
|
||||
- `chunk_text`:分块文本内容
|
||||
- `embedding`:768 维向量
|
||||
- `category`:文档分类
|
||||
- `file_path`:文件路径
|
||||
|
||||
**分块策略**:
|
||||
- Chunk Size:根据 `DocumentChunkConfig` 配置(默认 500 token)
|
||||
- Overlap:重叠区域(默认 50 token)
|
||||
|
||||
---
|
||||
|
||||
## L0 内存索引
|
||||
|
||||
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
|
||||
@@ -305,6 +349,61 @@ category: api
|
||||
|
||||
---
|
||||
|
||||
### 问题 4: Milvus 连接失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"scanned": 6,
|
||||
"inserted": 0,
|
||||
"failed": 6,
|
||||
"details": {
|
||||
"api/payment-errors.md": "Milvus 索引失败: Connection refused"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- Milvus 服务未启动
|
||||
- 网络连接问题
|
||||
- 配置错误
|
||||
|
||||
**解决**:
|
||||
```bash
|
||||
# 检查 Milvus 是否运行
|
||||
docker ps | grep milvus
|
||||
|
||||
# 检查配置
|
||||
grep milvus application.yml
|
||||
|
||||
# 启动 Milvus
|
||||
docker-compose up -d milvus-standalone
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 问题 5: 文档分块失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"details": {
|
||||
"test/large-doc.md": "Milvus 索引失败: Document too large"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- 文档内容过大
|
||||
- 分块配置不当
|
||||
|
||||
**解决**:
|
||||
- 检查 `DocumentChunkConfig` 配置
|
||||
- 调整 chunk size 和 overlap
|
||||
|
||||
---
|
||||
|
||||
#### 3. 文档缺少标题
|
||||
```json
|
||||
{
|
||||
@@ -318,38 +417,6 @@ category: api
|
||||
|
||||
---
|
||||
|
||||
## 后续扩展(L1 向量索引)
|
||||
|
||||
当前版本暂未实现 L1 向量索引(Milvus),计划后续扩展:
|
||||
|
||||
### 扩展方案
|
||||
|
||||
1. **独立索引任务**:
|
||||
```bash
|
||||
POST /api/knowledge/build-vectors
|
||||
```
|
||||
- 读取数据库中所有文档
|
||||
- 调用 `DocumentManagementService` 处理分块
|
||||
- 上传到 Milvus
|
||||
|
||||
2. **或者修改当前接口**:
|
||||
- 在 `initializeKnowledgeBase` 中调用文档分块和向量索引
|
||||
- 需要处理大文件的分块逻辑
|
||||
|
||||
### 验证 L1 的方法(未来)
|
||||
|
||||
```bash
|
||||
# 1. 调用 L1 索引构建
|
||||
curl -X POST http://localhost:9900/api/knowledge/build-vectors
|
||||
|
||||
# 2. 查询统计信息
|
||||
curl http://localhost:9900/api/knowledge/stats
|
||||
|
||||
# 3. 确认 totalVectors > 0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 最佳实践
|
||||
|
||||
### ✅ 推荐做法
|
||||
|
||||
Reference in New Issue
Block a user