feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -0,0 +1,66 @@
# Change: Align RAG offline eval with hybrid + qualityScore era
## Why
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
- No first-class dense vs hybrid fixture split for recall comparison.
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
## What Changes
### Knife 1 (must) — make offline eval reflect current main path
1. **Generator wiring**
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
2. **Fixture meta**
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
3. **Refresh path**
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
4. **Docs**
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
### Out of scope this change (confirmed grill)
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
## Non-goals
- Rewriting eval into a new framework or LLM-as-judge.
- Full threshold calibration productization.
- Neighbor chunks / query rewrite / cross-encoder.
- Changing production retrieval code paths (except eval generator test harness props).
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
- Dense/hybrid dual fixture directories (later change).
## Context constraints
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
- Seed isolation via `kb_scope=rag-eval` stays.
## Impact
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
## Success
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
- Fixtures carry searchMode meta.
- README describes correct offline/live loop.
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
- (If knife 2) dual fixture roots documented and runnable.