52 lines
1.7 KiB
Markdown
52 lines
1.7 KiB
Markdown
# RAG Retrieval Baseline
|
|
|
|
This directory contains the offline retrieval baseline for the RAG refactor.
|
|
|
|
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
|
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
|
evidence keywords before changing L0 behavior, query augmentation, evidence
|
|
post-processing, or Spring AI VectorStore integration.
|
|
|
|
## Layout
|
|
|
|
```text
|
|
eval/rag-retrieval/
|
|
cases/golden-cases.json Fixed retrieval golden cases
|
|
fixtures/*.json Saved retrieval candidates for each case
|
|
reports/baseline.json Machine-readable baseline report
|
|
reports/baseline.md Human-readable baseline report
|
|
```
|
|
|
|
## Run
|
|
|
|
From the repository root:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py
|
|
```
|
|
|
|
Custom paths are also supported:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py \
|
|
--cases eval/rag-retrieval/cases/golden-cases.json \
|
|
--fixtures eval/rag-retrieval/fixtures \
|
|
--json-report eval/rag-retrieval/reports/baseline.json \
|
|
--markdown-report eval/rag-retrieval/reports/baseline.md
|
|
```
|
|
|
|
## Hit Levels
|
|
|
|
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
|
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
|
- `weak`: expected evidence keyword is found, but expected document is missing.
|
|
- `miss`: expected document and expected evidence are not found.
|
|
|
|
`Recall@K` counts `strong` and `medium` as retrieved.
|
|
|
|
## Scope
|
|
|
|
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
|
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
|
not a claim that live production retrieval accuracy is complete.
|