feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# Design: rag-eval-hybrid-baseline
|
||||
|
||||
## Context
|
||||
|
||||
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
|
||||
|
||||
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1 — Replace vector-store mode with search mode
|
||||
|
||||
| Before | After |
|
||||
|--------|--------|
|
||||
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
|
||||
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
|
||||
|
||||
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
|
||||
|
||||
### D2 — Fixture meta minimum
|
||||
|
||||
```text
|
||||
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
|
||||
```
|
||||
|
||||
- `searchMode`: actual mode used for generation.
|
||||
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
|
||||
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
|
||||
|
||||
### D3 — LookupResult payload
|
||||
|
||||
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
|
||||
|
||||
### D4 — Acceptance if live refresh fails
|
||||
|
||||
Must deliver: ps1, test meta emission, README.
|
||||
Should attempt: seed + generate + eval.
|
||||
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
|
||||
|
||||
### D5 — Baseline update policy
|
||||
|
||||
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Mitigation |
|
||||
|------|------------|
|
||||
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
|
||||
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
|
||||
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
|
||||
|
||||
## Interface impact
|
||||
|
||||
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
|
||||
|
||||
## Audit
|
||||
|
||||
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
|
||||
Reference in New Issue
Block a user