feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,18 @@
|
||||
# Decisions: rag-eval-hybrid-baseline(最终版)
|
||||
|
||||
## Process
|
||||
|
||||
sm-flow standard-lean: Discover → Commit → Apply → Archive.
|
||||
|
||||
## Key decisions
|
||||
|
||||
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
|
||||
2. Fixture meta: `searchMode`, `kbScope` when set.
|
||||
3. Knife-2 (dual fixtures / mustNot golden) deferred.
|
||||
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
|
||||
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
|
||||
|
||||
## Trade-offs
|
||||
|
||||
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
|
||||
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
|
||||
Reference in New Issue
Block a user