Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
63 lines
2.7 KiB
Markdown
63 lines
2.7 KiB
Markdown
# Design: rag-eval-hybrid-baseline
|
||
|
||
## Context
|
||
|
||
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
|
||
|
||
## Goals / Non-Goals
|
||
|
||
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
|
||
|
||
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
|
||
|
||
## Decisions
|
||
|
||
### D1 — Replace vector-store mode with search mode
|
||
|
||
| Before | After |
|
||
|--------|--------|
|
||
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
|
||
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
|
||
|
||
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
|
||
|
||
### D2 — Fixture meta minimum
|
||
|
||
```text
|
||
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
|
||
```
|
||
|
||
- `searchMode`: actual mode used for generation.
|
||
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
|
||
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
|
||
|
||
### D3 — LookupResult payload
|
||
|
||
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
|
||
|
||
### D4 — Acceptance if live refresh fails
|
||
|
||
Must deliver: ps1, test meta emission, README.
|
||
Should attempt: seed + generate + eval.
|
||
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
|
||
|
||
### D5 — Baseline update policy
|
||
|
||
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
|
||
|
||
## Risks
|
||
|
||
| Risk | Mitigation |
|
||
|------|------------|
|
||
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
|
||
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
|
||
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
|
||
|
||
## Interface impact
|
||
|
||
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
|
||
|
||
## Audit
|
||
|
||
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
|