feat(rag): close eval pipeline with live snapshots

This commit is contained in:
zhuyongxin
2026-07-06 21:39:27 +08:00
parent cf3333d607
commit ed7efc58b7
47 changed files with 2613 additions and 177 deletions
+149 -2
View File
@@ -12,12 +12,59 @@ post-processing, or Spring AI VectorStore integration.
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
fixtures/*.json Saved retrieval fixtures for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/baseline-diff.* Optional diff reports
reports/live-post-reindex.* Optional live acceptance reports
```
## Seed Docs + Import/Reindex
The live-tool eval uses canonical seed documents so the real
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
whatever ad hoc documents happen to exist in the local knowledge base.
Seed documents live in:
```text
eval/rag-retrieval/seed-docs/*.md
```
Each seed doc uses frontmatter fields that are propagated into vector metadata:
```yaml
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
Import or reindex the seed docs through the real upload pipeline:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
deletes the existing document with the same `source`/`docId`, uploads the seed
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
Milvus chunks.
`kb_scope` isolates eval data:
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
without `kb_scope` remain searchable;
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
vector retrieval both use only the canonical eval seed docs;
- the fallback retry skips only the L0 category filter, not the `kb_scope`
boundary.
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
L0, and document enrichment; only the Markdown body is chunked and embedded.
This keeps controlled L0 decoys from becoming semantically relevant just because
their frontmatter keywords matched the query.
## Run
From the repository root:
@@ -36,6 +83,106 @@ python scripts/eval_rag_retrieval.py \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Generate Fixtures From LookupKnowledgeTool
Use the snapshot generator when fixtures should reflect the real
`LookupKnowledgeTool` pipeline:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
For the intended live loop, run seed import first:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
python scripts\eval_rag_retrieval.py
```
The script runs a Spring test harness:
```text
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
```
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
bean, calls `lookupKnowledge(query)` for each case, writes
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
Custom paths are supported:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 `
-Cases eval\rag-retrieval\cases\golden-cases.json `
-Fixtures eval\rag-retrieval\fixtures `
-RetrievedAt 2026-07-06T00:00:00Z
```
The generator is disabled in normal test runs. It only executes when
`rag.snapshot.enabled=true` is provided because it writes repository files and
depends on the configured runtime retrieval stack.
If generated fixtures fail the offline baseline, treat that as a real alignment
signal: either the golden expectations need to be adjusted to the current
knowledge base, or the knowledge base/indexing path needs to be fixed.
## Modular RAG Contract
Fixtures must use the current `lookupResult` shape, which mirrors the
`lookup_knowledge` output:
```text
lookupResult.evidenceBlocks
lookupResult.contextPack
lookupResult.retrievalTrace
lookupResult.rerankTrace
```
Golden cases can assert both retrieval quality and pipeline behavior:
- `expectedSources` / `expectedDocIds`
- `expectedBreadcrumbs`
- `expectedKeywords`
- `expectedSelectedAttempt`
- `expectedFallbackReason`
- `expectedFallbackReasons`
- `expectedEvidenceStatus`
- `expectedContextSources`
- `expectedRerankTopSource`
This lets the baseline catch regressions such as losing the expected evidence
source, skipping context packing, changing the selected retrieval attempt, or
breaking the filtered-vector to unfiltered-retry fallback.
## Baseline Diff
To compare a freshly generated report against an existing baseline:
```bash
python scripts/eval_rag_retrieval.py \
--json-report eval/rag-retrieval/reports/current.json \
--markdown-report eval/rag-retrieval/reports/current.md \
--compare-to eval/rag-retrieval/reports/baseline.json \
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
```
The diff reports aggregate regressions and case-level changes for:
- pass rate, recall@K, strong hit rate, miss count
- pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
The command exits non-zero when a case fails or the diff contains a regression.
## Hit Levels
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
@@ -81,6 +228,6 @@ GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
candidates, breadcrumb, score labels, and raw response fields. This is a live
results, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.