feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -0,0 +1,16 @@
# Brief: rag-eval-hybrid-baseline
## Background
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
## Goals (knife-1 only)
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
- Fixtures record `searchMode` / `kbScope`.
- README documents hybrid-era offline vs live loop.
- Best-effort live seed + regenerate fixtures + update baseline.
## Non-goals
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.