# 知识库检索可观测性指南 ## 日志层次 ### INFO 级别 - 关键业务流程 适用于生产环境监控,记录关键决策点和业务指标。 #### LookupKnowledgeTool(知识库查询) ``` [requestId] 收到知识库查询请求: query=ERR_TIMEOUT [requestId] L0精确匹配完成: matches=1, time=2ms [requestId] L0非唯一匹配,触发L1语义检索 [requestId] L1语义检索完成: matches=3, time=450ms [requestId] 查询完成: found=true, hasL0=true, hasL1=false, confidence=high, totalTime=455ms ``` **关键指标**: - `requestId`: 追踪单次查询的完整流程 - `matches`: L0/L1 匹配数量 - `time`: 各阶段耗时(ms) - `confidence`: 置信度(high/low) - `totalTime`: 端到端总耗时 #### DocumentManagementService(文档上传) ``` 开始上传文档: fileName=payment-errors.md, size=1024 bytes 解析到frontmatter: title=支付网关错误码, keywords=[ERR_TIMEOUT, 超时], time=5ms 文档分块完成: fileName=payment-errors.md, chunks=3, time=12ms 文档向量索引完成: docId=abc123, category=api, time=850ms 文档已加入L0索引: docId=abc123, title=支付网关错误码 文档上传完成: docId=abc123, fileName=payment-errors.md, hasFrontmatter=true, totalTime=920ms ``` **关键指标**: - `docId`: 文档唯一标识 - `hasFrontmatter`: 是否包含元数据 - `chunks`: 分块数量 - `totalTime`: 上传总耗时 #### KnowledgeIndexService(启动扫描) ``` 开始扫描知识库目录: knowledge_base/ 知识库索引加载完成,共 5 个文档 ``` ### DEBUG 级别 - 详细诊断信息 适用于开发和调试,记录详细的执行细节。 ``` [requestId] 置信度判断: highConfidence=true, reason=唯一匹配 [requestId] L0唯一匹配,跳过L1检索 L0结果已构建: source=knowledge_base/api/payment-errors.md, contentLength=1024 L0精确匹配: query=ERR_TIMEOUT, matches=1, indexSize=5, time=1ms 文档已加入索引: title=支付网关错误码, filePath=knowledge_base\api\payment-errors.md ``` ### WARN 级别 - 异常但可恢复 ``` 文档已存在: hash=abc123def, docId=xyz789 Frontmatter序列化失败 L0匹配但文件读取失败: knowledge_base/api/missing.md ``` ### ERROR 级别 - 严重错误 ``` 文档上传失败: fileName=test.md 知识库索引加载失败 文档索引失败: docId=abc123 ``` --- ## 可观测性场景 ### 场景 1: 追踪单次查询 **目标**:查看某次查询的完整流程 **步骤**: 1. 从日志中提取 `requestId`(8位UUID) 2. 使用 requestId 过滤所有相关日志 **示例**: ```bash grep "[a1b2c3d4]" logs/application.log ``` **输出**: ``` [a1b2c3d4] 收到知识库查询请求: query=超时 [a1b2c3d4] L0精确匹配完成: matches=2, time=3ms [a1b2c3d4] 置信度判断: highConfidence=false, reason=多个或零个匹配 [a1b2c3d4] L0非唯一匹配,触发L1语义检索 [a1b2c3d4] L1语义检索完成: matches=3, time=420ms [a1b2c3d4] 查询完成: found=true, hasL0=true, hasL1=true, confidence=low, totalTime=425ms ``` --- ### 场景 2: 性能监控 **目标**:监控 L0/L1 检索性能 **关键指标**: - L0 耗时:通常 < 10ms - L1 耗时:通常 200-500ms - 总耗时:通常 < 1s **异常识别**: ```bash # 查找慢查询(总耗时 > 1000ms) grep "totalTime=" logs/application.log | awk -F'totalTime=' '{print $2}' | awk -F'ms' '{if ($1 > 1000) print}' ``` --- ### 场景 3: L0 命中率分析 **目标**:统计 L0 精确匹配效果 **指标**: - 唯一匹配率(高置信度) - 多个匹配率(低置信度) - 未命中率(需要 L1) **统计脚本**: ```bash # 统计 L0 匹配情况 grep "L0精确匹配完成" logs/application.log | \ awk -F'matches=' '{print $2}' | \ awk -F',' '{print $1}' | \ sort | uniq -c ``` --- ### 场景 4: 文档上传监控 **目标**:监控文档上传流程 **关键检查点**: 1. Frontmatter 解析成功率 2. 向量索引耗时 3. L0 索引更新 **查询**: ```bash # 查找上传失败的文档 grep "文档上传失败" logs/application-error.log # 统计 frontmatter 解析率 grep "hasFrontmatter=" logs/application.log | \ awk -F'hasFrontmatter=' '{print $2}' | \ awk -F',' '{print $1}' | \ sort | uniq -c ``` --- ### 场景 5: Agent 工具调用链 **目标**:观测 Agent 如何使用 lookup_knowledge 工具 **配置**(application.yml): ```yaml logging: level: org.springframework.ai: DEBUG com.superbiz.agent.tool: INFO ``` **日志示例**: ``` [Agent] Calling tool: lookup_knowledge with query=ERR_TIMEOUT [a1b2c3d4] 收到知识库查询请求: query=ERR_TIMEOUT [a1b2c3d4] L0精确匹配完成: matches=1, time=2ms [a1b2c3d4] 查询完成: found=true, confidence=high, totalTime=5ms [Agent] Tool returned: {"found":true,"primary":{"content":"...","confidence":"high"}} ``` --- ## 日志分析最佳实践 ### 1. 使用结构化查询 ```bash # 按 requestId 分组统计耗时 grep "查询完成" logs/application.log | \ awk -F'totalTime=' '{print $2}' | \ awk -F'ms' '{sum+=$1; count++} END {print "平均耗时:", sum/count, "ms"}' ``` ### 2. 监控关键指标 - L0 索引大小(启动时) - L0 平均耗时 - L1 调用频率 - 高置信度比例 ### 3. 告警规则 - 总耗时 > 2s - L0 索引加载失败 - 文档上传失败率 > 10% --- ## MVP 阶段限制 当前日志为轻量级实现,**不包含**: - ❌ 结构化日志(JSON格式) - ❌ 指标收集(Micrometer/Prometheus) - ❌ 分布式追踪(Zipkin/Skywalking) - ❌ 独立日志文件 - ❌ 实时监控面板 **后续增强方向**: 1. 引入 Micrometer 指标 2. 配置独立的 knowledge-lookup.log 3. 集成 APM 工具 4. 添加 Grafana 监控面板