Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
3.4 KiB
3.4 KiB
Change: Align RAG offline eval with hybrid + qualityScore era
Why
eval/rag-retrieval already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
- Snapshot generator still passes dead
retrieval.vector-store.mode=spring|sdk. - Fixtures lack
searchMode/ scope meta; content still shows boost-stylehitReasonsand old score story. - No first-class dense vs hybrid fixture split for recall comparison.
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
Without this, offline eval cannot gate the current main path (retrieval.search.mode=hybrid, V2 store, qualityScore post-process).
What Changes
Knife 1 (must) — make offline eval reflect current main path
-
Generator wiring
- Replace
retrieval.vector-store.modewithretrieval.search.mode(hybriddefault;denseallowed). - Keep
-Dretrieval.kb-scope=rag-eval(or configurable). - Update
scripts/generate_rag_lookup_snapshots.ps1and any Java system-property docs/comments.
- Replace
-
Fixture meta
- Each fixture SHALL record at least:
caseId,query,retrievedAt,searchMode,kbScope(when set), pluslookupResultpayload. - Snapshot writer emits current
LookupResultshape (evidence identity fields if already present on blocks).
- Each fixture SHALL record at least:
-
Refresh path
- Document and support: prepare seed → generate fixtures (hybrid) →
eval_rag_retrieval.py→ update baseline. - Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
- Document and support: prepare seed → generate fixtures (hybrid) →
-
Docs
- Update
eval/rag-retrieval/README.mdto hybrid/quality narrative; remove spring vector-store as default.
- Update
Out of scope this change (confirmed grill)
- Knife 2:
fixtures/hybridvsfixtures/densedual layout and comparison report. - Golden extensions:
tags/mustNot*/ chunk keys /expectedRelevanceLevelhard gates.
Non-goals
- Rewriting eval into a new framework or LLM-as-judge.
- Full threshold calibration productization.
- Neighbor chunks / query rewrite / cross-encoder.
- Changing production retrieval code paths (except eval generator test harness props).
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (wiring+docs still complete; refresh recorded as unverified).
- Dense/hybrid dual fixture directories (later change).
Context constraints
- Continues
rag-quality-score-unify,rag-bm25-hybrid-drop-sdk,rag-chunk-evidence-identity-dedup. - Offline checker must remain dependency-free (no Milvus/LLM in
eval_rag_retrieval.py). - Seed isolation via
kb_scope=rag-evalstays.
Impact
- Interface: L1/L2 docs + eval artifacts only; no Agent ACI change.
- Risk: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
- Scale: micro→standard lean — multi-file scripts/docs/fixtures; no production architecture change. Use standard artifacts for clarity (
design+specs+tasks).
Success
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
- Fixtures carry searchMode meta.
- README describes correct offline/live loop.
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
- (If knife 2) dual fixture roots documented and runnable.