feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -0,0 +1,62 @@
# Design: rag-eval-hybrid-baseline
## Context
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
## Goals / Non-Goals
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
## Decisions
### D1 — Replace vector-store mode with search mode
| Before | After |
|--------|--------|
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
### D2 — Fixture meta minimum
```text
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
```
- `searchMode`: actual mode used for generation.
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
### D3 — LookupResult payload
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
### D4 — Acceptance if live refresh fails
Must deliver: ps1, test meta emission, README.
Should attempt: seed + generate + eval.
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
### D5 — Baseline update policy
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
## Risks
| Risk | Mitigation |
|------|------------|
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
## Interface impact
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
## Audit
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.