Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
2.7 KiB
Design: rag-eval-hybrid-baseline
Context
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (retrieval.search.mode) via MilvusHybridKnowledgeStore. Snapshot generation still injects removed retrieval.vector-store.mode.
Goals / Non-Goals
Goals: Wire snapshot generation to search.mode; emit fixture meta (searchMode, kbScope); document hybrid-era loop; refresh fixtures/baseline when env allows.
Non-Goals: Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
Decisions
D1 — Replace vector-store mode with search mode
| Before | After |
|---|---|
-Dretrieval.vector-store.mode=spring|sdk |
-Dretrieval.search.mode=hybrid|dense |
PS1 param VectorStoreMode |
SearchMode default hybrid |
Java snapshot test does not need a Spring bean switch: LookupKnowledgeTool already honors global retrieval.search.mode via VectorSearchService. Only system property / process config must set the property before context loads (Maven -D + optional properties on @SpringBootTest if required).
D2 — Fixture meta minimum
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
searchMode: actual mode used for generation.kbScope: from-Dretrieval.kb-scopewhen non-empty.- Offline evaluator MAY ignore unknown meta fields (backward compatible).
D3 — LookupResult payload
Continue serializing full LookupResult from tool. Prefer preserving any new block fields (docId, evidenceKey, scoreLabel) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
D4 — Acceptance if live refresh fails
Must deliver: ps1, test meta emission, README.
Should attempt: seed + generate + eval.
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
D5 — Baseline update policy
When fixtures refresh successfully: run offline eval; if intentional behavior change, update reports/baseline.* with diff review. Do not force green by weakening golden without note.
Risks
| Risk | Mitigation |
|---|---|
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
@SpringBootTest ignores late -D for some props |
Set search.mode via test properties default hybrid + override from system property if needed |
Interface impact
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
Audit
Eval-only pipeline; no new runtime module. Couples to existing LookupKnowledgeTool and config keys only.