Files
SuperBizAgent-java/.docs/knowledge-observability.md
T
zhuyongxin d6229f3385 feat(knowledge): 完成 L0+L1 混合检索集成
核心功能:
- 新增 FrontmatterParser 解析 YAML frontmatter
- 新增 KnowledgeIndexService L0 内存索引
- 新增 LookupKnowledgeTool 混合检索工具
- 增强 DocumentManagementService 文件保存和索引同步

技术实现:
- 数据库迁移 V004: api_document.metadata (TEXT)
- 依赖新增: snakeyaml 2.0
- 配置新增: knowledge.base-path
- 可观测性: requestId 追踪 + 性能日志

质量保证:
- 单元测试: 31/31 通过
- 测试覆盖: FrontmatterParser(11), KnowledgeIndexService(13), LookupKnowledgeTool(7)
- 启动验证: L0 索引正常加载

归档文档:
- OpenSpec: openspec/changes/lookup-knowledge-integration/
- devflow 档案: devflow/projects/2026-06-24-lookup-knowledge-integration/
- handoff: handoff/2026-06-24-lookup-knowledge-integration.md
2026-06-24 16:07:10 +08:00

215 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 知识库检索可观测性指南
## 日志层次
### INFO 级别 - 关键业务流程
适用于生产环境监控,记录关键决策点和业务指标。
#### LookupKnowledgeTool(知识库查询)
```
[requestId] 收到知识库查询请求: query=ERR_TIMEOUT
[requestId] L0精确匹配完成: matches=1, time=2ms
[requestId] L0非唯一匹配,触发L1语义检索
[requestId] L1语义检索完成: matches=3, time=450ms
[requestId] 查询完成: found=true, hasL0=true, hasL1=false, confidence=high, totalTime=455ms
```
**关键指标**:
- `requestId`: 追踪单次查询的完整流程
- `matches`: L0/L1 匹配数量
- `time`: 各阶段耗时(ms)
- `confidence`: 置信度(high/low)
- `totalTime`: 端到端总耗时
#### DocumentManagementService(文档上传)
```
开始上传文档: fileName=payment-errors.md, size=1024 bytes
解析到frontmatter: title=支付网关错误码, keywords=[ERR_TIMEOUT, 超时], time=5ms
文档分块完成: fileName=payment-errors.md, chunks=3, time=12ms
文档向量索引完成: docId=abc123, category=api, time=850ms
文档已加入L0索引: docId=abc123, title=支付网关错误码
文档上传完成: docId=abc123, fileName=payment-errors.md, hasFrontmatter=true, totalTime=920ms
```
**关键指标**:
- `docId`: 文档唯一标识
- `hasFrontmatter`: 是否包含元数据
- `chunks`: 分块数量
- `totalTime`: 上传总耗时
#### KnowledgeIndexService(启动扫描)
```
开始扫描知识库目录: knowledge_base/
知识库索引加载完成,共 5 个文档
```
### DEBUG 级别 - 详细诊断信息
适用于开发和调试,记录详细的执行细节。
```
[requestId] 置信度判断: highConfidence=true, reason=唯一匹配
[requestId] L0唯一匹配,跳过L1检索
L0结果已构建: source=knowledge_base/api/payment-errors.md, contentLength=1024
L0精确匹配: query=ERR_TIMEOUT, matches=1, indexSize=5, time=1ms
文档已加入索引: title=支付网关错误码, filePath=knowledge_base\api\payment-errors.md
```
### WARN 级别 - 异常但可恢复
```
文档已存在: hash=abc123def, docId=xyz789
Frontmatter序列化失败
L0匹配但文件读取失败: knowledge_base/api/missing.md
```
### ERROR 级别 - 严重错误
```
文档上传失败: fileName=test.md
知识库索引加载失败
文档索引失败: docId=abc123
```
---
## 可观测性场景
### 场景 1: 追踪单次查询
**目标**:查看某次查询的完整流程
**步骤**:
1. 从日志中提取 `requestId`(8位UUID)
2. 使用 requestId 过滤所有相关日志
**示例**:
```bash
grep "[a1b2c3d4]" logs/application.log
```
**输出**:
```
[a1b2c3d4] 收到知识库查询请求: query=超时
[a1b2c3d4] L0精确匹配完成: matches=2, time=3ms
[a1b2c3d4] 置信度判断: highConfidence=false, reason=多个或零个匹配
[a1b2c3d4] L0非唯一匹配,触发L1语义检索
[a1b2c3d4] L1语义检索完成: matches=3, time=420ms
[a1b2c3d4] 查询完成: found=true, hasL0=true, hasL1=true, confidence=low, totalTime=425ms
```
---
### 场景 2: 性能监控
**目标**:监控 L0/L1 检索性能
**关键指标**:
- L0 耗时:通常 < 10ms
- L1 耗时:通常 200-500ms
- 总耗时:通常 < 1s
**异常识别**:
```bash
# 查找慢查询(总耗时 > 1000ms)
grep "totalTime=" logs/application.log | awk -F'totalTime=' '{print $2}' | awk -F'ms' '{if ($1 > 1000) print}'
```
---
### 场景 3: L0 命中率分析
**目标**:统计 L0 精确匹配效果
**指标**:
- 唯一匹配率(高置信度)
- 多个匹配率(低置信度)
- 未命中率(需要 L1)
**统计脚本**:
```bash
# 统计 L0 匹配情况
grep "L0精确匹配完成" logs/application.log | \
awk -F'matches=' '{print $2}' | \
awk -F',' '{print $1}' | \
sort | uniq -c
```
---
### 场景 4: 文档上传监控
**目标**:监控文档上传流程
**关键检查点**:
1. Frontmatter 解析成功率
2. 向量索引耗时
3. L0 索引更新
**查询**:
```bash
# 查找上传失败的文档
grep "文档上传失败" logs/application-error.log
# 统计 frontmatter 解析率
grep "hasFrontmatter=" logs/application.log | \
awk -F'hasFrontmatter=' '{print $2}' | \
awk -F',' '{print $1}' | \
sort | uniq -c
```
---
### 场景 5: Agent 工具调用链
**目标**:观测 Agent 如何使用 lookup_knowledge 工具
**配置**(application.yml):
```yaml
logging:
level:
org.springframework.ai: DEBUG
com.superbiz.agent.tool: INFO
```
**日志示例**:
```
[Agent] Calling tool: lookup_knowledge with query=ERR_TIMEOUT
[a1b2c3d4] 收到知识库查询请求: query=ERR_TIMEOUT
[a1b2c3d4] L0精确匹配完成: matches=1, time=2ms
[a1b2c3d4] 查询完成: found=true, confidence=high, totalTime=5ms
[Agent] Tool returned: {"found":true,"primary":{"content":"...","confidence":"high"}}
```
---
## 日志分析最佳实践
### 1. 使用结构化查询
```bash
# 按 requestId 分组统计耗时
grep "查询完成" logs/application.log | \
awk -F'totalTime=' '{print $2}' | \
awk -F'ms' '{sum+=$1; count++} END {print "平均耗时:", sum/count, "ms"}'
```
### 2. 监控关键指标
- L0 索引大小(启动时)
- L0 平均耗时
- L1 调用频率
- 高置信度比例
### 3. 告警规则
- 总耗时 > 2s
- L0 索引加载失败
- 文档上传失败率 > 10%
---
## MVP 阶段限制
当前日志为轻量级实现,**不包含**:
- ❌ 结构化日志(JSON格式)
- ❌ 指标收集(Micrometer/Prometheus)
- ❌ 分布式追踪(Zipkin/Skywalking)
- ❌ 独立日志文件
- ❌ 实时监控面板
**后续增强方向**:
1. 引入 Micrometer 指标
2. 配置独立的 knowledge-lookup.log
3. 集成 APM 工具
4. 添加 Grafana 监控面板