feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,21 @@
|
||||
# Brief: rag-quality-score-unify
|
||||
|
||||
## Background
|
||||
|
||||
真 BM25 hybrid(dense + BM25 + RRF)已上线,但后处理仍把 hybrid 结果伪装成 L2 做 `normalizeL2`,并用 L0 domain/entity/keyword contains 加分改序。排序权威与质量闸门分裂,词面信号被 BM25 与后处理双重计分。
|
||||
|
||||
## Goals
|
||||
|
||||
- 一级 `scoreLabel` 仅 `dense` | `hybrid`
|
||||
- 唯一 `toQualityScore`;后处理 label-agnostic
|
||||
- 排序主序 = 检索 `originalRank`;去掉关键词 boost 改序
|
||||
- hybrid quality = 本轮 rank 纯映射(不做 max(rank, denseSim)、不为闸门回填 L2)
|
||||
- 保留 `mode=dense` 作同库召回对照;线上默认 hybrid
|
||||
|
||||
## Scope
|
||||
|
||||
内部 RAG:store 发射、normalizer、evidence post-process、单测、架构文档 §6。
|
||||
|
||||
## Non-goals
|
||||
|
||||
精排 / query rewrite / 邻块、schema rebuild、改 Agent ACI 字段名、删除 dense 对照 mode。
|
||||
Reference in New Issue
Block a user