# Change: Align RAG offline eval with hybrid + qualityScore era ## Why `eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative: - Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`. - Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story. - No first-class dense vs hybrid fixture split for recall comparison. - Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out. Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process). ## What Changes ### Knife 1 (must) — make offline eval reflect current main path 1. **Generator wiring** - Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed). - Keep `-Dretrieval.kb-scope=rag-eval` (or configurable). - Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments. 2. **Fixture meta** - Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload. - Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks). 3. **Refresh path** - Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline. - Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance. 4. **Docs** - Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default. ### Out of scope this change (confirmed grill) - Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report. - Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates. ## Non-goals - Rewriting eval into a new framework or LLM-as-judge. - Full threshold calibration productization. - Neighbor chunks / query rewrite / cross-encoder. - Changing production retrieval code paths (except eval generator test harness props). - Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified). - Dense/hybrid dual fixture directories (later change). ## Context constraints - Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`. - Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`). - Seed isolation via `kb_scope=rag-eval` stays. ## Impact - **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change. - **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review. - **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`). ## Success - Generator defaults to hybrid search mode; dead vector-store mode flag gone. - Fixtures carry searchMode meta. - README describes correct offline/live loop. - Offline eval runs on refreshed or existing fixtures without requiring removed config keys. - (If knife 2) dual fixture roots documented and runnable.