feat(rag): close eval pipeline with live snapshots

This commit is contained in:
zhuyongxin
2026-07-06 21:39:27 +08:00
parent cf3333d607
commit ed7efc58b7
47 changed files with 2613 additions and 177 deletions
@@ -0,0 +1,78 @@
# Acceptance: rag-eval-pipeline-closure
## Status
Archived.
## Acceptance Criteria
| Item | Status | Notes |
|---|---|---|
| Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. |
| LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. |
| Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. |
| Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. |
| Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. |
| Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. |
| Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. |
| Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. |
| Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. |
| Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. |
## Verification
```powershell
python scripts\eval_rag_retrieval.py
```
Result: passed. 7 cases, passRate=1.0, recall@5=1.0.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test
```
Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=<temp-fixtures>" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test
python scripts\eval_rag_retrieval.py --fixtures <temp-fixtures> --json-report <temp-current.json> --markdown-report <temp-current.md>
```
Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1.
```powershell
python scripts\eval_rag_retrieval.py --json-report <temp-current.json> --markdown-report <temp-current.md> --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report <temp-diff.json> --diff-markdown-report <temp-diff.md>
```
Result: passed. regressions=0.
```powershell
mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test
```
Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled.
```powershell
$null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1))
$null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1))
```
Result: PowerShell syntax OK.
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack.
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z -SkipEval
```
Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0.
```powershell
git diff --check
```
Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.
@@ -0,0 +1,21 @@
# Brief: rag-eval-pipeline-closure
## Background
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
## Goals
1. Reuse the existing offline RAG retrieval baseline.
2. Extend it to support modular `LookupResult` fixtures.
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
4. Add a RAG baseline diff path for regression detection.
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
6. Document how RAG baseline and diagnosis baseline form a quality loop.
## Non-Goals
- No new production API.
- No LLM-as-judge scoring.
- No production API behavior changes.
- No replacement for diagnosis eval.
@@ -0,0 +1,57 @@
# Decisions: rag-eval-pipeline-closure
## D1. Reuse the existing evaluator
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
## D2. Use LookupResult as the only fixture contract
Decision: support `lookupResult` only.
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
## D3. Make modular assertions opt-in per case
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
## D4. Diff remains deterministic
Decision: RAG diff compares report fields only and does not call live services or models.
Reason: this keeps it suitable for local regression checks and CI-style gates.
## D5. Isolate live eval docs with kb_scope
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
## D6. Import seed docs through the real upload pipeline
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
## D7. Strip frontmatter before chunk embedding
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
## D8. Treat retry behavior as the stable fallback contract
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
## D9. Default live snapshots to Spring AI VectorStore
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.
@@ -0,0 +1,87 @@
# Evidence: rag-eval-pipeline-closure
## Changed Assets
- `scripts/eval_rag_retrieval.py`
- `eval/rag-retrieval/cases/golden-cases.json`
- `eval/rag-retrieval/fixtures/*.json`
- `eval/rag-retrieval/reports/baseline.json`
- `eval/rag-retrieval/reports/baseline.md`
- `eval/rag-retrieval/README.md`
- `mvp/architecture/rag-eval-closure.md`
- `src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java`
- `src/test/java/com/superbiz/agent/eval/RagEvalSeedImporterTest.java`
- `scripts/generate_rag_lookup_snapshots.ps1`
- `scripts/prepare_rag_eval_seed.ps1`
- `eval/rag-retrieval/seed-docs/*.md`
- `src/main/java/com/superbiz/agent/dto/Frontmatter.java`
- `src/main/java/com/superbiz/agent/dto/KnowledgeEntry.java`
- `src/main/java/com/superbiz/agent/service/FrontmatterParser.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/resources/application.yml`
## Baseline Result
```text
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
## Regression Signals
The evaluator now fails on:
- missing expected source
- missing breadcrumb or evidence keyword
- non-`lookupResult` fixture
- selected attempt mismatch
- fallback reason mismatch
- fallback reason outside accepted values
- evidence status mismatch
- missing context source
- rerank top source mismatch
The snapshot generator now provides:
- real `LookupKnowledgeTool` invocation
- one fixture per golden case
- explicit opt-in through `rag.snapshot.enabled=true`
- optional post-generation baseline evaluation
- scoped retrieval through `retrieval.kb-scope=rag-eval`
- Spring AI VectorStore by default through `retrieval.vector-store.mode=spring`
- scoped L0 hints through the same `retrieval.kb-scope`
The seed importer now provides:
- canonical eval docs under `eval/rag-retrieval/seed-docs`
- real `DocumentManagementService` import/reindex
- stable `source`/`docId` metadata
- `kb_scope=rag-eval` isolation from local non-eval documents
- frontmatter stripping before chunk embedding
- an over-filter decoy seed doc for fallback-path evaluation
The diff now detects:
- aggregate pass/recall regression
- case pass regression
- hit-level regression
- first-rank regression
- selected attempt/fallback/evidence/rerank changes
## Final Spring Live Snapshot
```text
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
Key fallback trace:
```text
selectedAttempt=UNFILTERED_VECTOR_RETRY
fallbackReason=filtered_vector_no_evidence
rerankTopSource=rag-l0-filter-fallback
```