87 lines
2.8 KiB
Markdown
87 lines
2.8 KiB
Markdown
# RAG Retrieval Baseline
|
|
|
|
This directory contains the offline retrieval baseline for the RAG refactor.
|
|
|
|
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
|
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
|
evidence keywords before changing L0 behavior, query augmentation, evidence
|
|
post-processing, or Spring AI VectorStore integration.
|
|
|
|
## Layout
|
|
|
|
```text
|
|
eval/rag-retrieval/
|
|
cases/golden-cases.json Fixed retrieval golden cases
|
|
fixtures/*.json Saved retrieval candidates for each case
|
|
reports/baseline.json Machine-readable baseline report
|
|
reports/baseline.md Human-readable baseline report
|
|
reports/live-post-reindex.* Optional live acceptance reports
|
|
```
|
|
|
|
## Run
|
|
|
|
From the repository root:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py
|
|
```
|
|
|
|
Custom paths are also supported:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py \
|
|
--cases eval/rag-retrieval/cases/golden-cases.json \
|
|
--fixtures eval/rag-retrieval/fixtures \
|
|
--json-report eval/rag-retrieval/reports/baseline.json \
|
|
--markdown-report eval/rag-retrieval/reports/baseline.md
|
|
```
|
|
|
|
## Hit Levels
|
|
|
|
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
|
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
|
- `weak`: expected evidence keyword is found, but expected document is missing.
|
|
- `miss`: expected document and expected evidence are not found.
|
|
|
|
`Recall@K` counts `strong` and `medium` as retrieved.
|
|
|
|
## Scope
|
|
|
|
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
|
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
|
not a claim that live production retrieval accuracy is complete.
|
|
|
|
## Live Post-Reindex Acceptance
|
|
|
|
When embedding input changes, existing vectors do not update by themselves. For
|
|
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
|
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
|
semantic signal.
|
|
|
|
Use this optional live acceptance flow after the application is running and the
|
|
knowledge base has been reindexed:
|
|
|
|
```bash
|
|
python scripts/eval_rag_live_acceptance.py
|
|
```
|
|
|
|
Custom service URL and output paths are supported:
|
|
|
|
```bash
|
|
python scripts/eval_rag_live_acceptance.py \
|
|
--base-url http://127.0.0.1:9900 \
|
|
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
|
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
|
```
|
|
|
|
The script calls:
|
|
|
|
```text
|
|
GET /api/search/similar
|
|
```
|
|
|
|
It writes JSON and Markdown reports with query, topK, result count, top
|
|
candidates, breadcrumb, score labels, and raw response fields. This is a live
|
|
smoke check for environment readiness and post-reindex behavior; it does not
|
|
replace the deterministic offline baseline above.
|