feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,20 @@
|
||||
# Brief: rag-quality-score-unify
|
||||
|
||||
## Background
|
||||
|
||||
True BM25 hybrid retrieval is live, but post-processing still normalizes as if every score were dense L2 and re-ranks with L0 keyword contains boosts. That splits ranking authority from quality gates and double-counts lexical signal.
|
||||
|
||||
## Goals
|
||||
|
||||
- Unify score labels to `dense` | `hybrid`.
|
||||
- Single `toQualityScore`; label-agnostic post-process.
|
||||
- Preserve retrieval rank; remove boost re-rank.
|
||||
- Hybrid quality = pure rank mapping (slice 1).
|
||||
|
||||
## Scope
|
||||
|
||||
Internal RAG pipeline: store emission, normalizer, evidence post-process, tests, architecture docs.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Fine re-rankers, query rewrite, neighbor chunks, schema rebuild, ACI field renames, removing dense comparison mode.
|
||||
Reference in New Issue
Block a user