feat(rag): close eval pipeline with live snapshots
This commit is contained in:
@@ -0,0 +1,21 @@
|
||||
# Brief: rag-eval-pipeline-closure
|
||||
|
||||
## Background
|
||||
|
||||
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Reuse the existing offline RAG retrieval baseline.
|
||||
2. Extend it to support modular `LookupResult` fixtures.
|
||||
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
|
||||
4. Add a RAG baseline diff path for regression detection.
|
||||
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
|
||||
6. Document how RAG baseline and diagnosis baseline form a quality loop.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new production API.
|
||||
- No LLM-as-judge scoring.
|
||||
- No production API behavior changes.
|
||||
- No replacement for diagnosis eval.
|
||||
Reference in New Issue
Block a user