# Acceptance: rag-eval-pipeline-closure ## Status Archived. ## Acceptance Criteria | Item | Status | Notes | |---|---|---| | Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. | | LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. | | Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. | | Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. | | Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. | | Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. | | Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. | | Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. | | Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. | | Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. | ## Verification ```powershell python scripts\eval_rag_retrieval.py ``` Result: passed. 7 cases, passRate=1.0, recall@5=1.0. ```powershell mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test ``` Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided. ```powershell mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test python scripts\eval_rag_retrieval.py --fixtures --json-report --markdown-report ``` Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1. ```powershell python scripts\eval_rag_retrieval.py --json-report --markdown-report --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report --diff-markdown-report ``` Result: passed. regressions=0. ```powershell mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test ``` Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled. ```powershell $null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1)) $null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1)) ``` Result: PowerShell syntax OK. ```powershell .\scripts\prepare_rag_eval_seed.ps1 ``` Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack. ```powershell .\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures -RetrievedAt 2026-07-06T00:00:00Z -SkipEval ``` Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0. ```powershell git diff --check ``` Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.