Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
67 lines
3.4 KiB
Markdown
67 lines
3.4 KiB
Markdown
# Change: Align RAG offline eval with hybrid + qualityScore era
|
||
|
||
## Why
|
||
|
||
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
|
||
|
||
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
|
||
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
|
||
- No first-class dense vs hybrid fixture split for recall comparison.
|
||
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
|
||
|
||
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
|
||
|
||
## What Changes
|
||
|
||
### Knife 1 (must) — make offline eval reflect current main path
|
||
|
||
1. **Generator wiring**
|
||
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
|
||
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
|
||
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
|
||
|
||
2. **Fixture meta**
|
||
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
|
||
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
|
||
|
||
3. **Refresh path**
|
||
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
|
||
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
|
||
|
||
4. **Docs**
|
||
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
|
||
|
||
### Out of scope this change (confirmed grill)
|
||
|
||
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
|
||
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
|
||
|
||
## Non-goals
|
||
|
||
- Rewriting eval into a new framework or LLM-as-judge.
|
||
- Full threshold calibration productization.
|
||
- Neighbor chunks / query rewrite / cross-encoder.
|
||
- Changing production retrieval code paths (except eval generator test harness props).
|
||
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
|
||
- Dense/hybrid dual fixture directories (later change).
|
||
|
||
## Context constraints
|
||
|
||
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
|
||
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
|
||
- Seed isolation via `kb_scope=rag-eval` stays.
|
||
|
||
## Impact
|
||
|
||
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
|
||
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
|
||
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
|
||
|
||
## Success
|
||
|
||
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
|
||
- Fixtures carry searchMode meta.
|
||
- README describes correct offline/live loop.
|
||
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
|
||
- (If knife 2) dual fixture roots documented and runnable.
|