Files
SuperBizAgent-java/eval/rag-retrieval/README.md
T

2.8 KiB

RAG Retrieval Baseline

This directory contains the offline retrieval baseline for the RAG refactor.

The baseline is intentionally narrower than full diagnosis evaluation. It checks whether fixed retrieval queries can recover expected documents, breadcrumbs, and evidence keywords before changing L0 behavior, query augmentation, evidence post-processing, or Spring AI VectorStore integration.

Layout

eval/rag-retrieval/
  cases/golden-cases.json      Fixed retrieval golden cases
  fixtures/*.json              Saved retrieval candidates for each case
  reports/baseline.json        Machine-readable baseline report
  reports/baseline.md          Human-readable baseline report
  reports/live-post-reindex.*  Optional live acceptance reports

Run

From the repository root:

python scripts/eval_rag_retrieval.py

Custom paths are also supported:

python scripts/eval_rag_retrieval.py \
  --cases eval/rag-retrieval/cases/golden-cases.json \
  --fixtures eval/rag-retrieval/fixtures \
  --json-report eval/rag-retrieval/reports/baseline.json \
  --markdown-report eval/rag-retrieval/reports/baseline.md

Hit Levels

  • strong: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
  • medium: expected document is found, but breadcrumb or keyword coverage is incomplete.
  • weak: expected evidence keyword is found, but expected document is missing.
  • miss: expected document and expected evidence are not found.

Recall@K counts strong and medium as retrieved.

Scope

This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM, or the Spring Boot application. It is a regression harness for retrieval behavior, not a claim that live production retrieval accuracy is complete.

Live Post-Reindex Acceptance

When embedding input changes, existing vectors do not update by themselves. For example, after adding title and breadcrumb to the embedding text, the live Milvus/Zilliz collection must be reindexed before retrieval can reflect that new semantic signal.

Use this optional live acceptance flow after the application is running and the knowledge base has been reindexed:

python scripts/eval_rag_live_acceptance.py

Custom service URL and output paths are supported:

python scripts/eval_rag_live_acceptance.py \
  --base-url http://127.0.0.1:9900 \
  --json-report eval/rag-retrieval/reports/live-post-reindex.json \
  --markdown-report eval/rag-retrieval/reports/live-post-reindex.md

The script calls:

GET /api/search/similar

It writes JSON and Markdown reports with query, topK, result count, top candidates, breadcrumb, score labels, and raw response fields. This is a live smoke check for environment readiness and post-reindex behavior; it does not replace the deterministic offline baseline above.