Files

3.4 KiB

Decisions: rag-eval-pipeline-closure

D1. Reuse the existing evaluator

Decision: extend scripts/eval_rag_retrieval.py instead of creating a second evaluator.

Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.

D2. Use LookupResult as the only fixture contract

Decision: support lookupResult only.

Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.

D3. Make modular assertions opt-in per case

Decision: use fields such as expectedSelectedAttempt, expectedFallbackReason/expectedFallbackReasons, expectedEvidenceStatus, expectedContextSources, and expectedRerankTopSource.

Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.

D4. Diff remains deterministic

Decision: RAG diff compares report fields only and does not call live services or models.

Reason: this keeps it suitable for local regression checks and CI-style gates.

D5. Isolate live eval docs with kb_scope

Decision: add kb_scope metadata and use rag-eval for canonical eval seed documents.

Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real LookupKnowledgeTool snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.

Default runtime keeps retrieval.kb-scope empty so legacy documents without kb_scope remain searchable. Eval scripts pass -Dretrieval.kb-scope=rag-eval. The same scope applies to L0 query hints and L1 vector retrieval.

D6. Import seed docs through the real upload pipeline

Decision: seed docs are imported by RagEvalSeedImporterTest through DocumentManagementService.uploadDocument.

Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.

D7. Strip frontmatter before chunk embedding

Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.

Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.

D8. Treat retry behavior as the stable fallback contract

Decision: the fallback golden case accepts both filtered_vector_low_quality and filtered_vector_no_evidence, while still requiring selectedAttempt=UNFILTERED_VECTOR_RETRY, expected evidence source, context packing, and rerank top source.

Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps kb_scope, retries the original query, and returns the correct evidence.

D9. Default live snapshots to Spring AI VectorStore

Decision: generate_rag_lookup_snapshots.ps1 defaults to retrieval.vector-store.mode=spring.

Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through -VectorStoreMode sdk for comparison.