Files
SuperBizAgent-java/devflow/projects/2026-07-06-rag-eval-pipeline-closure/decisions.md
T

58 lines
3.4 KiB
Markdown

# Decisions: rag-eval-pipeline-closure
## D1. Reuse the existing evaluator
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
## D2. Use LookupResult as the only fixture contract
Decision: support `lookupResult` only.
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
## D3. Make modular assertions opt-in per case
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
## D4. Diff remains deterministic
Decision: RAG diff compares report fields only and does not call live services or models.
Reason: this keeps it suitable for local regression checks and CI-style gates.
## D5. Isolate live eval docs with kb_scope
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
## D6. Import seed docs through the real upload pipeline
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
## D7. Strip frontmatter before chunk embedding
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
## D8. Treat retry behavior as the stable fallback contract
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
## D9. Default live snapshots to Spring AI VectorStore
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.