feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,16 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
|
||||
|
||||
## Goals (knife-1 only)
|
||||
|
||||
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
|
||||
- Fixtures record `searchMode` / `kbScope`.
|
||||
- README documents hybrid-era offline vs live loop.
|
||||
- Best-effort live seed + regenerate fixtures + update baseline.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
|
||||
Reference in New Issue
Block a user