Files
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

3.4 KiB
Raw Permalink Blame History

Change: Align RAG offline eval with hybrid + qualityScore era

Why

eval/rag-retrieval already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:

  • Snapshot generator still passes dead retrieval.vector-store.mode=spring|sdk.
  • Fixtures lack searchMode / scope meta; content still shows boost-style hitReasons and old score story.
  • No first-class dense vs hybrid fixture split for recall comparison.
  • Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.

Without this, offline eval cannot gate the current main path (retrieval.search.mode=hybrid, V2 store, qualityScore post-process).

What Changes

Knife 1 (must) — make offline eval reflect current main path

  1. Generator wiring

    • Replace retrieval.vector-store.mode with retrieval.search.mode (hybrid default; dense allowed).
    • Keep -Dretrieval.kb-scope=rag-eval (or configurable).
    • Update scripts/generate_rag_lookup_snapshots.ps1 and any Java system-property docs/comments.
  2. Fixture meta

    • Each fixture SHALL record at least: caseId, query, retrievedAt, searchMode, kbScope (when set), plus lookupResult payload.
    • Snapshot writer emits current LookupResult shape (evidence identity fields if already present on blocks).
  3. Refresh path

    • Document and support: prepare seed → generate fixtures (hybrid) → eval_rag_retrieval.py → update baseline.
    • Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
  4. Docs

    • Update eval/rag-retrieval/README.md to hybrid/quality narrative; remove spring vector-store as default.

Out of scope this change (confirmed grill)

  • Knife 2: fixtures/hybrid vs fixtures/dense dual layout and comparison report.
  • Golden extensions: tags / mustNot* / chunk keys / expectedRelevanceLevel hard gates.

Non-goals

  • Rewriting eval into a new framework or LLM-as-judge.
  • Full threshold calibration productization.
  • Neighbor chunks / query rewrite / cross-encoder.
  • Changing production retrieval code paths (except eval generator test harness props).
  • Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (wiring+docs still complete; refresh recorded as unverified).
  • Dense/hybrid dual fixture directories (later change).

Context constraints

  • Continues rag-quality-score-unify, rag-bm25-hybrid-drop-sdk, rag-chunk-evidence-identity-dedup.
  • Offline checker must remain dependency-free (no Milvus/LLM in eval_rag_retrieval.py).
  • Seed isolation via kb_scope=rag-eval stays.

Impact

  • Interface: L1/L2 docs + eval artifacts only; no Agent ACI change.
  • Risk: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
  • Scale: micro→standard lean — multi-file scripts/docs/fixtures; no production architecture change. Use standard artifacts for clarity (design + specs + tasks).

Success

  • Generator defaults to hybrid search mode; dead vector-store mode flag gone.
  • Fixtures carry searchMode meta.
  • README describes correct offline/live loop.
  • Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
  • (If knife 2) dual fixture roots documented and runnable.