Files

7.5 KiB

RAG Retrieval Baseline

This directory contains the offline retrieval baseline for the RAG refactor.

The baseline is intentionally narrower than full diagnosis evaluation. It checks whether fixed retrieval queries can recover expected documents, breadcrumbs, and evidence keywords before changing L0 behavior, query augmentation, evidence post-processing, or Spring AI VectorStore integration.

Layout

eval/rag-retrieval/
  cases/golden-cases.json      Fixed retrieval golden cases
  seed-docs/*.md               Canonical docs imported into the live KB for real-tool eval
  fixtures/*.json              Saved retrieval fixtures for each case
  reports/baseline.json        Machine-readable baseline report
  reports/baseline.md          Human-readable baseline report
  reports/baseline-diff.*      Optional diff reports
  reports/live-post-reindex.*  Optional live acceptance reports

Seed Docs + Import/Reindex

The live-tool eval uses canonical seed documents so the real LookupKnowledgeTool can retrieve stable evidence from MySQL/Milvus instead of whatever ad hoc documents happen to exist in the local knowledge base.

Seed documents live in:

eval/rag-retrieval/seed-docs/*.md

Each seed doc uses frontmatter fields that are propagated into vector metadata:

source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval

Import or reindex the seed docs through the real upload pipeline:

.\scripts\prepare_rag_eval_seed.ps1

The script runs RagEvalSeedImporterTest with rag.seed.enabled=true. It deletes the existing document with the same source/docId, uploads the seed doc through DocumentManagementService, updates DB metadata and L0, and rebuilds Milvus chunks.

kb_scope isolates eval data:

  • default application config leaves retrieval.kb-scope empty, so legacy docs without kb_scope remain searchable;
  • eval scripts pass -Dretrieval.kb-scope=rag-eval, so L0 query hints and L1 vector retrieval both use only the canonical eval seed docs;
  • the fallback retry skips only the L0 category filter, not the kb_scope boundary.

Frontmatter is not embedded as chunk content during upload. It feeds metadata, L0, and document enrichment; only the Markdown body is chunked and embedded. This keeps controlled L0 decoys from becoming semantically relevant just because their frontmatter keywords matched the query.

Run

From the repository root:

python scripts/eval_rag_retrieval.py

Custom paths are also supported:

python scripts/eval_rag_retrieval.py \
  --cases eval/rag-retrieval/cases/golden-cases.json \
  --fixtures eval/rag-retrieval/fixtures \
  --json-report eval/rag-retrieval/reports/baseline.json \
  --markdown-report eval/rag-retrieval/reports/baseline.md

Generate Fixtures From LookupKnowledgeTool

Use the snapshot generator when fixtures should reflect the real LookupKnowledgeTool pipeline:

.\scripts\generate_rag_lookup_snapshots.ps1

For the intended live loop, run seed import first:

.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
python scripts\eval_rag_retrieval.py

The script runs a Spring test harness:

mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test

The generator reads golden-cases.json, injects the real LookupKnowledgeTool bean, calls lookupKnowledge(query) for each case, writes fixtures/{caseId}.json, and then runs eval_rag_retrieval.py unless -SkipEval is provided. It defaults to Spring AI VectorStore mode; pass -VectorStoreMode sdk only when intentionally comparing the legacy SDK path.

Custom paths are supported:

.\scripts\generate_rag_lookup_snapshots.ps1 `
  -Cases eval\rag-retrieval\cases\golden-cases.json `
  -Fixtures eval\rag-retrieval\fixtures `
  -RetrievedAt 2026-07-06T00:00:00Z

The generator is disabled in normal test runs. It only executes when rag.snapshot.enabled=true is provided because it writes repository files and depends on the configured runtime retrieval stack.

If generated fixtures fail the offline baseline, treat that as a real alignment signal: either the golden expectations need to be adjusted to the current knowledge base, or the knowledge base/indexing path needs to be fixed.

Modular RAG Contract

Fixtures must use the current lookupResult shape, which mirrors the lookup_knowledge output:

lookupResult.evidenceBlocks
lookupResult.contextPack
lookupResult.retrievalTrace
lookupResult.rerankTrace

Golden cases can assert both retrieval quality and pipeline behavior:

  • expectedSources / expectedDocIds
  • expectedBreadcrumbs
  • expectedKeywords
  • expectedSelectedAttempt
  • expectedFallbackReason
  • expectedFallbackReasons
  • expectedEvidenceStatus
  • expectedContextSources
  • expectedRerankTopSource

This lets the baseline catch regressions such as losing the expected evidence source, skipping context packing, changing the selected retrieval attempt, or breaking the filtered-vector to unfiltered-retry fallback.

Baseline Diff

To compare a freshly generated report against an existing baseline:

python scripts/eval_rag_retrieval.py \
  --json-report eval/rag-retrieval/reports/current.json \
  --markdown-report eval/rag-retrieval/reports/current.md \
  --compare-to eval/rag-retrieval/reports/baseline.json \
  --diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
  --diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md

The diff reports aggregate regressions and case-level changes for:

  • pass rate, recall@K, strong hit rate, miss count
  • pass state
  • hit level
  • first expected rank
  • selected attempt
  • fallback reason
  • evidence status
  • rerank top source

The command exits non-zero when a case fails or the diff contains a regression.

Hit Levels

  • strong: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
  • medium: expected document is found, but breadcrumb or keyword coverage is incomplete.
  • weak: expected evidence keyword is found, but expected document is missing.
  • miss: expected document and expected evidence are not found.

Recall@K counts strong and medium as retrieved.

Scope

This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM, or the Spring Boot application. It is a regression harness for retrieval behavior, not a claim that live production retrieval accuracy is complete.

Live Post-Reindex Acceptance

When embedding input changes, existing vectors do not update by themselves. For example, after adding title and breadcrumb to the embedding text, the live Milvus/Zilliz collection must be reindexed before retrieval can reflect that new semantic signal.

Use this optional live acceptance flow after the application is running and the knowledge base has been reindexed:

python scripts/eval_rag_live_acceptance.py

Custom service URL and output paths are supported:

python scripts/eval_rag_live_acceptance.py \
  --base-url http://127.0.0.1:9900 \
  --json-report eval/rag-retrieval/reports/live-post-reindex.json \
  --markdown-report eval/rag-retrieval/reports/live-post-reindex.md

The script calls:

GET /api/search/similar

It writes JSON and Markdown reports with query, topK, result count, top results, breadcrumb, score labels, and raw response fields. This is a live smoke check for environment readiness and post-reindex behavior; it does not replace the deterministic offline baseline above.