zhuyongxin
|
ed7efc58b7
|
feat(rag): close eval pipeline with live snapshots
|
2026-07-06 21:39:27 +08:00 |
|
zhuyongxin
|
1ff7f09d25
|
fix chat session traces and document paths
|
2026-07-03 13:53:16 +08:00 |
|
zhuyongxin
|
bb44140901
|
feat(knowledge): 会话级去重 + 知识域地图注入 Planner 解决 ISS-001 重复检索
- RetrievedDocTracker: sessionId → Set<filePath> 会话级去重,LookupKnowledgeTool Step 5 过滤已检索文档
- KnowledgeDomainService: 域级聚合,LLM 生成 when_to_retrieve,构建 knowledge map YAML
- DocumentFieldEnricher: 上传时 LLM 补全 covers + whenToRetrieve(含同域文档排除上下文)
- KnowledgeDomain entity + V009 迁移: 域级元数据持久化,避免重启重复 LLM 调用
- ChatService: 注入 knowledge map 到 Planner prompt,会话结束时清理去重状态
- KnowledgeIndexService: 手写 JSON 解析替换为 Jackson ObjectMapper,启动时补建缺失域记录
- chat-planner-prompt: 新增知识库检索规则(按域 when_to_retrieve 判断,每域最多一次检索)
- doc-field-enricher-prompt / domain-summary-prompt: 外部化 LLM 提示词
|
2026-07-01 10:47:46 +08:00 |
|
zhuyongxin
|
91931363d4
|
refactor(knowledge): 重构 FaultCategory 枚举为文档分类
## 改动内容
### 1. 重构 FaultCategory 枚举
**修改前**:故障类别枚举
```java
EXTERNAL_API("外部接口调用失败"),
INTERNAL_ERROR("系统内部错误"),
DATABASE("数据库问题"),
...
```
**修改后**:文档分类枚举
```java
API("API 接口文档"),
INFRASTRUCTURE("基础设施文档"),
DOMAIN("领域业务文档"),
TROUBLESHOOTING("故障排查文档"),
GENERAL("通用文档");
```
### 2. 新增 fromString 映射方法
```java
public static FaultCategory fromString(String category) {
switch (category.toLowerCase()) {
case "api": return API;
case "infrastructure": return INFRASTRUCTURE;
case "domain": return DOMAIN;
case "troubleshooting": return TROUBLESHOOTING;
default: return GENERAL;
}
}
```
### 3. 更新所有引用
- `ApiDocument`: 默认值 EXTERNAL_API → GENERAL
- `DocumentManagementService`: 默认值 EXTERNAL_API → GENERAL
- `KnowledgeBaseInitService`: 使用 FaultCategory.fromString() 映射
### 4. 字段映射关系
| Frontmatter | 数据库字段 | 枚举值 | 说明 |
|-------------|-----------|--------|------|
| `category: "api"` | `fault_category` | API | API 接口文档 |
| `category: "infrastructure"` | `fault_category` | INFRASTRUCTURE | 基础设施文档 |
| `category: "domain"` | `fault_category` | DOMAIN | 领域业务文档 |
| `category: "troubleshooting"` | `fault_category` | TROUBLESHOOTING | 故障排查文档 |
| `category: "xxx"` | `fault_category` | GENERAL | 默认/其他 |
## 数据库影响
**不需要修改数据库结构**:
- `fault_category` 字段仍然是 VARCHAR(32)
- 只是存储的值从 `EXTERNAL_API` 变为 `API`, `INFRASTRUCTURE` 等
**已存在的数据**:
- 旧数据中的 `EXTERNAL_API` 仍可以正常读取(枚举向后兼容)
- 新导入的文档会使用新的枚举值
## 验证
```bash
# 1. 重新初始化
curl -X POST http://localhost:9900/api/knowledge/init?force=true
# 2. 查询统计
curl http://localhost:9900/api/knowledge/stats
# 3. 响应
{
"categories": {
"API": 1,
"INFRASTRUCTURE": 3,
"DOMAIN": 1,
"TROUBLESHOOTING": 1
}
}
```
## 数据库查询
```sql
SELECT fault_category, COUNT(*)
FROM api_document
GROUP BY fault_category;
-- 结果
API | 1
INFRASTRUCTURE | 3
DOMAIN | 1
TROUBLESHOOTING | 1
```
|
2026-06-25 14:23:57 +08:00 |
|
zhuyongxin
|
d6229f3385
|
feat(knowledge): 完成 L0+L1 混合检索集成
核心功能:
- 新增 FrontmatterParser 解析 YAML frontmatter
- 新增 KnowledgeIndexService L0 内存索引
- 新增 LookupKnowledgeTool 混合检索工具
- 增强 DocumentManagementService 文件保存和索引同步
技术实现:
- 数据库迁移 V004: api_document.metadata (TEXT)
- 依赖新增: snakeyaml 2.0
- 配置新增: knowledge.base-path
- 可观测性: requestId 追踪 + 性能日志
质量保证:
- 单元测试: 31/31 通过
- 测试覆盖: FrontmatterParser(11), KnowledgeIndexService(13), LookupKnowledgeTool(7)
- 启动验证: L0 索引正常加载
归档文档:
- OpenSpec: openspec/changes/lookup-knowledge-integration/
- devflow 档案: devflow/projects/2026-06-24-lookup-knowledge-integration/
- handoff: handoff/2026-06-24-lookup-knowledge-integration.md
|
2026-06-24 16:07:10 +08:00 |
|
zhuyongxin
|
24101a8d66
|
feat(phase1): 支持上传时指定文档类别
功能增强:
- DocumentController 新增 category 参数
POST /api/documents/upload?category=api
- DocumentUploadRequest 新增 category 字段
- 支持用户指定:api、domain、troubleshoot 等
- 默认值:upload(未指定时)
- VectorIndexService.indexDocumentChunks 接收 category
- 将用户指定的类别存入 Milvus metadata
- metadata.category = 用户指定值 或 "upload"
使用示例:
```bash
# 上传 API 文档
curl -X POST /api/documents/upload \
-F "file=@redis-api.md" \
-F "category=api"
# 上传领域知识文档
curl -X POST /api/documents/upload \
-F "file=@cache-theory.md" \
-F "category=domain"
# 检索时按类别过滤
searchSimilarDocuments("Redis接口", 5, "api")
```
完整流程:
1. 文件索引:自动从路径提取(aiops-docs/api/ → "api")
2. 用户上传:从接口参数获取(category=api)
3. 检索时:可按类别过滤(category 参数)
编译验证:BUILD SUCCESS
|
2026-06-23 16:43:47 +08:00 |
|
zhuyongxin
|
4ef8d87961
|
feat(phase1): 实现文档分块向量化索引
Task 5.6: 向量化索引实现
- VectorIndexService 新增方法:
- indexDocumentChunks(docId, chunks): 索引文档分块到 Milvus
- deleteDocumentChunks(docId): 删除文档的所有向量
- buildDocumentMetadata(): 构建文档元数据(区分文件索引)
核心流程:
1. 上传时:文本提取 → 分块 → 向量化 → 存入 Milvus + MySQL
2. 检索时:问题向量化 → Milvus 语义检索 → 返回相似文档
3. 删除时:删除元数据 + 删除向量索引
实现细节:
- 复用 indexSingleFile 的向量化逻辑
- metadata.docId 标识文档来源(区分 upload: 和 file:)
- 删除表达式:metadata["docId"] == "xxx"
- 自动去重:上传前删除旧向量数据
DocumentManagementService 完整实现:
- uploadDocument: 完整向量化流程(移除 TODO)
- deleteDocument: 同步删除向量索引(移除 TODO)
编译验证:BUILD SUCCESS
Progress: 32/34 tasks completed (94%)
|
2026-06-23 16:08:50 +08:00 |
|
zhuyongxin
|
e76d4ce48f
|
feat(phase1): 完成文档查询和删除接口
Task 5.4: 文档查询接口
- DocumentManagementService 新增查询方法:
- queryDocumentById: 根据 docId 查询单个文档
- queryDocumentsByStatus: 根据状态查询(分页)
- queryDocumentsByFaultSource: 根据故障源查询
- convertToResponse: 实体转 DTO 工具方法
- DocumentController 新增 RESTful 接口:
- GET /api/documents/{docId}
- GET /api/documents/status/{status}?page=0&size=20
- GET /api/documents/faultSource/{faultSource}
Task 5.5: 文档删除接口
- DocumentManagementService 新增删除方法:
- deleteDocument: 删除文档元数据
- TODO: 向量索引删除待实现
- DocumentController 新增删除接口:
- DELETE /api/documents/{docId}
功能特性:
- 统一异常处理:DocumentProcessException
- 统一响应格式:Result<T>
- 分页支持:Page/PageRequest
- 事务支持:@Transactional
编译验证:BUILD SUCCESS
Progress: 28/33 tasks completed (85%)
|
2026-06-23 15:48:04 +08:00 |
|
zhuyongxin
|
f446290d0f
|
feat(phase1): 完成文档上传接口
Task 5.3: 文档上传接口
- 创建 DocumentManagementService:文档上传核心逻辑
- 文件格式验证(仅 .md/.txt)
- 文件 hash 计算与去重检查
- 文本提取与分块处理
- 文档元数据持久化(ApiDocument)
- 向量化索引标记为 TODO(待补充)
- 创建 DocumentController:RESTful 上传接口
- POST /api/documents/upload
- 支持参数:file, faultCategory, faultSource, apiName, version
- 返回:文档 docId
功能特性:
- MD5 hash 去重:防止重复上传
- 事务支持:元数据与索引状态一致性
- 异常处理:DocumentProcessException 统一封装
- 分块配置:使用 DocumentChunkConfig 默认配置
待补充:
- TODO: VectorIndexService.indexDocumentChunks() 实现
- 当前文档状态直接标记为 INDEXED
编译验证:BUILD SUCCESS
Progress: 26/33 tasks completed (79%)
|
2026-06-23 15:28:02 +08:00 |
|