Files
SuperBizAgent-java/mvp/issues/rag-l1-score-calibration.md
T

50 lines
1.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RAG L1 分数阈值未校准
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:向量搜索、相关性判断、工具调用记录
---
## 现象
L1 语义检索使用 Milvus 向量距离后,会做相关性归一和阈值判断。但当前阈值更偏经验值,没有基于真实查询集和真实分数分布做校准。
由于当前使用 L2 距离,不同 embedding 模型、不同语料密度、不同 query 长度都会影响分数分布。
---
## 当前实现
- `VectorSearchService` 使用 query embedding 搜索 Milvus。
- Milvus metric type 为 `L2`。
- `LookupKnowledgeTool` 会把 L2 score 转成 normalized relevance。
- 阈值没有配套评测集或分布统计。
---
## 影响
- 阈值过松时,低相关片段会进入 Agent 上下文。
- 阈值过紧时,正确片段可能被过滤掉。
- 面试中如果被追问“为什么这个阈值合理”,当前只能回答是 MVP 经验值。
---
## 建议修复
1. 固化一组 RAG 回归查询集,覆盖告警、数据库、流程规范、AIOps 诊断等场景。
2. 记录每次 topK 的原始 L2 score、归一化分数、最终是否采纳。
3. 统计正例和负例分布,确定阈值区间。
4. 将阈值配置化,并在 README 或 issue 中记录选择依据。
5. 后续引入 reranker 后,L1 阈值可以从“最终判断”退化为“粗召回过滤”。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/resources/application.yml`