234 lines
7.5 KiB
Markdown
234 lines
7.5 KiB
Markdown
# RAG Retrieval Baseline
|
|
|
|
This directory contains the offline retrieval baseline for the RAG refactor.
|
|
|
|
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
|
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
|
evidence keywords before changing L0 behavior, query augmentation, evidence
|
|
post-processing, or Spring AI VectorStore integration.
|
|
|
|
## Layout
|
|
|
|
```text
|
|
eval/rag-retrieval/
|
|
cases/golden-cases.json Fixed retrieval golden cases
|
|
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
|
fixtures/*.json Saved retrieval fixtures for each case
|
|
reports/baseline.json Machine-readable baseline report
|
|
reports/baseline.md Human-readable baseline report
|
|
reports/baseline-diff.* Optional diff reports
|
|
reports/live-post-reindex.* Optional live acceptance reports
|
|
```
|
|
|
|
## Seed Docs + Import/Reindex
|
|
|
|
The live-tool eval uses canonical seed documents so the real
|
|
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
|
whatever ad hoc documents happen to exist in the local knowledge base.
|
|
|
|
Seed documents live in:
|
|
|
|
```text
|
|
eval/rag-retrieval/seed-docs/*.md
|
|
```
|
|
|
|
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
|
|
|
```yaml
|
|
source: mysql-connection-pool
|
|
breadcrumb: Database > MySQL > Connection Pool
|
|
kb_scope: rag-eval
|
|
```
|
|
|
|
Import or reindex the seed docs through the real upload pipeline:
|
|
|
|
```powershell
|
|
.\scripts\prepare_rag_eval_seed.ps1
|
|
```
|
|
|
|
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
|
deletes the existing document with the same `source`/`docId`, uploads the seed
|
|
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
|
Milvus chunks.
|
|
|
|
`kb_scope` isolates eval data:
|
|
|
|
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
|
without `kb_scope` remain searchable;
|
|
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
|
vector retrieval both use only the canonical eval seed docs;
|
|
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
|
boundary.
|
|
|
|
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
|
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
|
This keeps controlled L0 decoys from becoming semantically relevant just because
|
|
their frontmatter keywords matched the query.
|
|
|
|
## Run
|
|
|
|
From the repository root:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py
|
|
```
|
|
|
|
Custom paths are also supported:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py \
|
|
--cases eval/rag-retrieval/cases/golden-cases.json \
|
|
--fixtures eval/rag-retrieval/fixtures \
|
|
--json-report eval/rag-retrieval/reports/baseline.json \
|
|
--markdown-report eval/rag-retrieval/reports/baseline.md
|
|
```
|
|
|
|
## Generate Fixtures From LookupKnowledgeTool
|
|
|
|
Use the snapshot generator when fixtures should reflect the real
|
|
`LookupKnowledgeTool` pipeline:
|
|
|
|
```powershell
|
|
.\scripts\generate_rag_lookup_snapshots.ps1
|
|
```
|
|
|
|
For the intended live loop, run seed import first:
|
|
|
|
```powershell
|
|
.\scripts\prepare_rag_eval_seed.ps1
|
|
.\scripts\generate_rag_lookup_snapshots.ps1
|
|
python scripts\eval_rag_retrieval.py
|
|
```
|
|
|
|
The script runs a Spring test harness:
|
|
|
|
```text
|
|
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
|
```
|
|
|
|
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
|
bean, calls `lookupKnowledge(query)` for each case, writes
|
|
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
|
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
|
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
|
|
|
Custom paths are supported:
|
|
|
|
```powershell
|
|
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
|
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
|
-Fixtures eval\rag-retrieval\fixtures `
|
|
-RetrievedAt 2026-07-06T00:00:00Z
|
|
```
|
|
|
|
The generator is disabled in normal test runs. It only executes when
|
|
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
|
depends on the configured runtime retrieval stack.
|
|
|
|
If generated fixtures fail the offline baseline, treat that as a real alignment
|
|
signal: either the golden expectations need to be adjusted to the current
|
|
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
|
|
|
## Modular RAG Contract
|
|
|
|
Fixtures must use the current `lookupResult` shape, which mirrors the
|
|
`lookup_knowledge` output:
|
|
|
|
```text
|
|
lookupResult.evidenceBlocks
|
|
lookupResult.contextPack
|
|
lookupResult.retrievalTrace
|
|
lookupResult.rerankTrace
|
|
```
|
|
|
|
Golden cases can assert both retrieval quality and pipeline behavior:
|
|
|
|
- `expectedSources` / `expectedDocIds`
|
|
- `expectedBreadcrumbs`
|
|
- `expectedKeywords`
|
|
- `expectedSelectedAttempt`
|
|
- `expectedFallbackReason`
|
|
- `expectedFallbackReasons`
|
|
- `expectedEvidenceStatus`
|
|
- `expectedContextSources`
|
|
- `expectedRerankTopSource`
|
|
|
|
This lets the baseline catch regressions such as losing the expected evidence
|
|
source, skipping context packing, changing the selected retrieval attempt, or
|
|
breaking the filtered-vector to unfiltered-retry fallback.
|
|
|
|
## Baseline Diff
|
|
|
|
To compare a freshly generated report against an existing baseline:
|
|
|
|
```bash
|
|
python scripts/eval_rag_retrieval.py \
|
|
--json-report eval/rag-retrieval/reports/current.json \
|
|
--markdown-report eval/rag-retrieval/reports/current.md \
|
|
--compare-to eval/rag-retrieval/reports/baseline.json \
|
|
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
|
|
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
|
```
|
|
|
|
The diff reports aggregate regressions and case-level changes for:
|
|
|
|
- pass rate, recall@K, strong hit rate, miss count
|
|
- pass state
|
|
- hit level
|
|
- first expected rank
|
|
- selected attempt
|
|
- fallback reason
|
|
- evidence status
|
|
- rerank top source
|
|
|
|
The command exits non-zero when a case fails or the diff contains a regression.
|
|
|
|
## Hit Levels
|
|
|
|
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
|
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
|
- `weak`: expected evidence keyword is found, but expected document is missing.
|
|
- `miss`: expected document and expected evidence are not found.
|
|
|
|
`Recall@K` counts `strong` and `medium` as retrieved.
|
|
|
|
## Scope
|
|
|
|
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
|
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
|
not a claim that live production retrieval accuracy is complete.
|
|
|
|
## Live Post-Reindex Acceptance
|
|
|
|
When embedding input changes, existing vectors do not update by themselves. For
|
|
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
|
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
|
semantic signal.
|
|
|
|
Use this optional live acceptance flow after the application is running and the
|
|
knowledge base has been reindexed:
|
|
|
|
```bash
|
|
python scripts/eval_rag_live_acceptance.py
|
|
```
|
|
|
|
Custom service URL and output paths are supported:
|
|
|
|
```bash
|
|
python scripts/eval_rag_live_acceptance.py \
|
|
--base-url http://127.0.0.1:9900 \
|
|
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
|
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
|
```
|
|
|
|
The script calls:
|
|
|
|
```text
|
|
GET /api/search/similar
|
|
```
|
|
|
|
It writes JSON and Markdown reports with query, topK, result count, top
|
|
results, breadcrumb, score labels, and raw response fields. This is a live
|
|
smoke check for environment readiness and post-reindex behavior; it does not
|
|
replace the deterministic offline baseline above.
|