feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -0,0 +1,20 @@
# Brief: rag-quality-score-unify
## Background
True BM25 hybrid retrieval is live, but post-processing still normalizes as if every score were dense L2 and re-ranks with L0 keyword contains boosts. That splits ranking authority from quality gates and double-counts lexical signal.
## Goals
- Unify score labels to `dense` | `hybrid`.
- Single `toQualityScore`; label-agnostic post-process.
- Preserve retrieval rank; remove boost re-rank.
- Hybrid quality = pure rank mapping (slice 1).
## Scope
Internal RAG pipeline: store emission, normalizer, evidence post-process, tests, architecture docs.
## Non-goals
Fine re-rankers, query rewrite, neighbor chunks, schema rebuild, ACI field renames, removing dense comparison mode.