feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
+89 -143
View File
@@ -1,38 +1,65 @@
# RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor.
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
The baseline is intentionally narrower than full diagnosis evaluation. It checks
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
evidence keywords before changing L0 behavior, query augmentation, evidence
post-processing, or Spring AI VectorStore integration.
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
Related design notes:
- `docs/RAG-Hybrid质量分与后处理.md`
- `docs/RAG-Agent如何读relevance_level.md`
- `mvp/architecture/RAG知识检索架构.md` §6
## Offline vs live
| Layer | What | Needs live stack? |
|-------|------|-------------------|
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
Daily CI / local quick check: **offline only**.
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
## Layout
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
fixtures/*.json Saved retrieval fixtures for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/baseline-diff.* Optional diff reports
reports/live-post-reindex.* Optional live acceptance reports
cases/golden-cases.json Fixed queries + expectations
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
reports/baseline.json|md Last accepted offline report
reports/baseline-diff.* Optional diff vs previous report
```
## Seed Docs + Import/Reindex
## Fixture shape (minimum)
The live-tool eval uses canonical seed documents so the real
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
whatever ad hoc documents happen to exist in the local knowledge base.
```text
caseId
query
retrievedAt
searchMode # hybrid | dense (required on newly generated fixtures)
kbScope # e.g. rag-eval when generation used a scope
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
```
Seed documents live in:
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
Older fixtures may omit `searchMode`; regenerate to attach meta.
## Seed docs + import
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
```text
eval/rag-retrieval/seed-docs/*.md
```
Each seed doc uses frontmatter fields that are propagated into vector metadata:
Frontmatter example:
```yaml
source: mysql-connection-pool
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
Import or reindex the seed docs through the real upload pipeline:
Import via real upload pipeline:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
deletes the existing document with the same `source`/`docId`, uploads the seed
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
Milvus chunks.
Isolation:
`kb_scope` isolates eval data:
- App default may leave `retrieval.kb-scope` empty (all docs).
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
without `kb_scope` remain searchable;
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
vector retrieval both use only the canonical eval seed docs;
- the fallback retry skips only the L0 category filter, not the `kb_scope`
boundary.
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
L0, and document enrichment; only the Markdown body is chunked and embedded.
This keeps controlled L0 decoys from becoming semantically relevant just because
their frontmatter keywords matched the query.
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
## Run
From the repository root:
## Offline run (no live stack)
```bash
python scripts/eval_rag_retrieval.py
```
Custom paths are also supported:
Custom paths:
```bash
python scripts/eval_rag_retrieval.py \
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Generate Fixtures From LookupKnowledgeTool
Use the snapshot generator when fixtures should reflect the real
`LookupKnowledgeTool` pipeline:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
For the intended live loop, run seed import first:
## Generate fixtures (live stack)
```powershell
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
python scripts\eval_rag_retrieval.py
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
```
The script runs a Spring test harness:
```text
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
```
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
bean, calls `lookupKnowledge(query)` for each case, writes
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
Custom paths are supported:
Dense baseline snapshot (same seed, comparison only):
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 `
-Cases eval\rag-retrieval\cases\golden-cases.json `
-Fixtures eval\rag-retrieval\fixtures `
-RetrievedAt 2026-07-06T00:00:00Z
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
```
The generator is disabled in normal test runs. It only executes when
`rag.snapshot.enabled=true` is provided because it writes repository files and
depends on the configured runtime retrieval stack.
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
If generated fixtures fail the offline baseline, treat that as a real alignment
signal: either the golden expectations need to be adjusted to the current
knowledge base, or the knowledge base/indexing path needs to be fixed.
## Modular RAG Contract
Fixtures must use the current `lookupResult` shape, which mirrors the
`lookup_knowledge` output:
Maven equivalent:
```text
lookupResult.evidenceBlocks
lookupResult.contextPack
lookupResult.retrievalTrace
lookupResult.rerankTrace
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
-Drag.snapshot.enabled=true \
-Dretrieval.kb-scope=rag-eval \
-Dretrieval.search.mode=hybrid \
test
```
Golden cases can assert both retrieval quality and pipeline behavior:
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
## Golden assertions
Supported expectation fields include:
- `expectedSources` / `expectedDocIds`
- `expectedBreadcrumbs`
- `expectedKeywords`
- `expectedBreadcrumbs` / `expectedKeywords`
- `expectedSelectedAttempt`
- `expectedFallbackReason`
- `expectedFallbackReasons`
- `expectedFallbackReason` / `expectedFallbackReasons`
- `expectedEvidenceStatus`
- `expectedContextSources`
- `expectedRerankTopSource`
This lets the baseline catch regressions such as losing the expected evidence
source, skipping context packing, changing the selected retrieval attempt, or
breaking the filtered-vector to unfiltered-retry fallback.
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
## Baseline Diff
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
To compare a freshly generated report against an existing baseline:
## Baseline diff
```bash
python scripts/eval_rag_retrieval.py \
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
```
The diff reports aggregate regressions and case-level changes for:
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
Non-zero exit on case failure or regression in diff mode.
- pass rate, recall@K, strong hit rate, miss count
- pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
## Hit levels
The command exits non-zero when a case fails or the diff contains a regression.
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
- `medium`: expected document found, coverage incomplete
- `weak`: keyword hit without expected document
- `miss`: neither
## Hit Levels
`Recall@K` counts `strong` + `medium`.
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
- `weak`: expected evidence keyword is found, but expected document is missing.
- `miss`: expected document and expected evidence are not found.
## Optional live smoke (post-reindex)
`Recall@K` counts `strong` and `medium` as retrieved.
## Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
or the Spring Boot application. It is a regression harness for retrieval behavior,
not a claim that live production retrieval accuracy is complete.
## Live Post-Reindex Acceptance
When embedding input changes, existing vectors do not update by themselves. For
example, after adding `title` and `breadcrumb` to the embedding text, the live
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
semantic signal.
Use this optional live acceptance flow after the application is running and the
knowledge base has been reindexed:
After reindex, with app up:
```bash
python scripts/eval_rag_live_acceptance.py
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
```
Custom service URL and output paths are supported:
```bash
python scripts/eval_rag_live_acceptance.py \
--base-url http://127.0.0.1:9900 \
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
```
The script calls:
```text
GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
results, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.