# Decisions: rag-eval-pipeline-closure ## D1. Reuse the existing evaluator Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator. Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path. ## D2. Use LookupResult as the only fixture contract Decision: support `lookupResult` only. Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks. ## D3. Make modular assertions opt-in per case Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`. Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field. ## D4. Diff remains deterministic Decision: RAG diff compares report fields only and does not call live services or models. Reason: this keeps it suitable for local regression checks and CI-style gates. ## D5. Isolate live eval docs with kb_scope Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents. Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data. Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval. ## D6. Import seed docs through the real upload pipeline Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`. Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative. ## D7. Strip frontmatter before chunk embedding Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding. Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases. ## D8. Treat retry behavior as the stable fallback contract Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source. Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence. ## D9. Default live snapshots to Spring AI VectorStore Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`. Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.