1.7 KiB
1.7 KiB
RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor.
The baseline is intentionally narrower than full diagnosis evaluation. It checks whether fixed retrieval queries can recover expected documents, breadcrumbs, and evidence keywords before changing L0 behavior, query augmentation, evidence post-processing, or Spring AI VectorStore integration.
Layout
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
Run
From the repository root:
python scripts/eval_rag_retrieval.py
Custom paths are also supported:
python scripts/eval_rag_retrieval.py \
--cases eval/rag-retrieval/cases/golden-cases.json \
--fixtures eval/rag-retrieval/fixtures \
--json-report eval/rag-retrieval/reports/baseline.json \
--markdown-report eval/rag-retrieval/reports/baseline.md
Hit Levels
strong: expected document is found and breadcrumb or evidence keyword coverage is satisfied.medium: expected document is found, but breadcrumb or keyword coverage is incomplete.weak: expected evidence keyword is found, but expected document is missing.miss: expected document and expected evidence are not found.
Recall@K counts strong and medium as retrieved.
Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM, or the Spring Boot application. It is a regression harness for retrieval behavior, not a claim that live production retrieval accuracy is complete.