feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,13 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval structure is correct but generator/docs/fixtures predate hybrid search mode.
|
||||
|
||||
## Goals
|
||||
|
||||
Knife 1 only: `search.mode` wiring, fixture meta, README, best-effort fixture refresh.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dual dense/hybrid fixture trees; golden mustNot/chunk/level hard gates; new frameworks.
|
||||
Reference in New Issue
Block a user