Files
SuperBizAgent-java/openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline/design.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

2.7 KiB
Raw Blame History

Design: rag-eval-hybrid-baseline

Context

Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (retrieval.search.mode) via MilvusHybridKnowledgeStore. Snapshot generation still injects removed retrieval.vector-store.mode.

Goals / Non-Goals

Goals: Wire snapshot generation to search.mode; emit fixture meta (searchMode, kbScope); document hybrid-era loop; refresh fixtures/baseline when env allows.

Non-Goals: Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.

Decisions

D1 — Replace vector-store mode with search mode

Before After
-Dretrieval.vector-store.mode=spring|sdk -Dretrieval.search.mode=hybrid|dense
PS1 param VectorStoreMode SearchMode default hybrid

Java snapshot test does not need a Spring bean switch: LookupKnowledgeTool already honors global retrieval.search.mode via VectorSearchService. Only system property / process config must set the property before context loads (Maven -D + optional properties on @SpringBootTest if required).

D2 — Fixture meta minimum

caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
  • searchMode: actual mode used for generation.
  • kbScope: from -Dretrieval.kb-scope when non-empty.
  • Offline evaluator MAY ignore unknown meta fields (backward compatible).

D3 — LookupResult payload

Continue serializing full LookupResult from tool. Prefer preserving any new block fields (docId, evidenceKey, scoreLabel) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).

D4 — Acceptance if live refresh fails

Must deliver: ps1, test meta emission, README.
Should attempt: seed + generate + eval.
If blocked: do not fail the change; record commands and gap in acceptance/devflow.

D5 — Baseline update policy

When fixtures refresh successfully: run offline eval; if intentional behavior change, update reports/baseline.* with diff review. Do not force green by weakening golden without note.

Risks

Risk Mitigation
Env cannot refresh fixtures Q2: wiring-first acceptance
Old fixtures fail offline after code drift Document; refresh when possible; optional temporary note in README
@SpringBootTest ignores late -D for some props Set search.mode via test properties default hybrid + override from system property if needed

Interface impact

L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.

Audit

Eval-only pipeline; no new runtime module. Couples to existing LookupKnowledgeTool and config keys only.