Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
18 lines
705 B
Markdown
18 lines
705 B
Markdown
# Evidence: rag-eval-hybrid-baseline
|
|
|
|
## Pre-change
|
|
|
|
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
|
|
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
|
|
- Offline eval already supported Hit levels, recall@K, baseline diff.
|
|
|
|
## User decisions
|
|
|
|
- Scope: knife-1 only (no dual fixture dirs).
|
|
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
|
|
|
|
## Apply-discovered
|
|
|
|
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
|
|
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
|