feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,25 @@
|
||||
# Evidence: rag-quality-score-unify
|
||||
|
||||
## Code (pre-change)
|
||||
|
||||
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后并行 dense 回填 L2;BM25-only → `bm25_only_no_dense` + maxL2。
|
||||
- `KnowledgeEvidencePostProcessor`:一律 `normalizeL2(score)` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序;PRECISE 需 `hasHintSupport`。
|
||||
|
||||
## User decisions (grill)
|
||||
|
||||
| ID | 结论 |
|
||||
|---|---|
|
||||
| Q1 | 一级 label 仅 dense/hybrid;bm25_only 不作正式 label |
|
||||
| Q2 | 后处理去掉 contains 加分改序,保 originalRank |
|
||||
| Q3 | 唯一 toQualityScore;后处理统一 |
|
||||
| Q4 | dense mode 保留作对照 |
|
||||
| Q7 | 接受 relevance_level / retry 分布变化 |
|
||||
| Q8 | hybrid quality = **纯 rank 映射** |
|
||||
|
||||
## Post-change anchors
|
||||
|
||||
- `RetrievalScoreLabels` / `RetrievalScoreNormalizer`
|
||||
- `MilvusHybridKnowledgeStore`(无 L2 overwrite / 无 bm25_only 发射)
|
||||
- `KnowledgeEvidencePostProcessor`(rank sort + explain-only L0 overlap)
|
||||
- OpenSpec delta:`rag-retrieval-quality-score`
|
||||
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
|
||||
Reference in New Issue
Block a user