Files
SuperBizAgent-java/eval/rag-retrieval/README.md
zhuyongxin 584639fa2a docs(mvp): move engineering notes under mvp/engineering
Relocate RAG and diagnosis decision/E2E writeups from docs/ root into
mvp/engineering so architecture, issues, and engineering narrative stay
together. Update indexes and cross-links; leave docs/learning as legacy.
2026-07-29 10:49:45 +08:00

6.0 KiB
Raw Permalink Blame History

RAG Retrieval Baseline

Offline regression harness for lookup_knowledge after hybrid retrieval + qualityScore post-process.

It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is not a full diagnosis-agent E2E.

Production knowledge path: MilvusHybridKnowledgeStore with retrieval.search.mode=hybrid (dense+BM25+RRF).
mode=dense remains a same-collection baseline for recall comparison (not a second index).

Related design notes:

  • mvp/engineering/rag/RAG-Hybrid质量分与后处理.md
  • mvp/engineering/rag/RAG-Agent如何读relevance_level.md
  • mvp/architecture/RAG知识检索架构.md §6

Offline vs live

Layer What Needs live stack?
Offline fixtures/*.json × golden-cases.json → pass/fail + baseline diff No (no Milvus/LLM/Boot)
Snapshot generate Real LookupKnowledgeTool writes fixtures Yes (embedding + Milvus + DB/L0 as configured)
Live smoke optional eval_rag_live_acceptance.py Yes (running app)

Daily CI / local quick check: offline only.
After changing retrieval, indexing, or search mode: regenerate fixtures, then offline eval, then update baseline if the diff is intentional.

Layout

eval/rag-retrieval/
  cases/golden-cases.json      Fixed queries + expectations
  seed-docs/*.md               Canonical docs for live snapshot (kb_scope: rag-eval)
  fixtures/*.json              Frozen lookupResult snapshots (+ searchMode meta)
  reports/baseline.json|md     Last accepted offline report
  reports/baseline-diff.*      Optional diff vs previous report

Fixture shape (minimum)

caseId
query
retrievedAt
searchMode          # hybrid | dense  (required on newly generated fixtures)
kbScope             # e.g. rag-eval when generation used a scope
lookupResult        # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …

Offline eval ignores unknown top-level meta and does not full-JSON-compare.
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.

Older fixtures may omit searchMode; regenerate to attach meta.

Seed docs + import

Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:

eval/rag-retrieval/seed-docs/*.md

Frontmatter example:

source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval

Import via real upload pipeline:

.\scripts\prepare_rag_eval_seed.ps1

Isolation:

  • App default may leave retrieval.kb-scope empty (all docs).
  • Eval generation passes -Dretrieval.kb-scope=rag-eval.
  • Category-filter fallback retries without L0 category filter only; kb_scope still applies.

Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).

Seeds must live in the current hybrid collection schema (milvus.collection, default biz). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.

Offline run (no live stack)

python scripts/eval_rag_retrieval.py

Custom paths:

python scripts/eval_rag_retrieval.py \
  --cases eval/rag-retrieval/cases/golden-cases.json \
  --fixtures eval/rag-retrieval/fixtures \
  --json-report eval/rag-retrieval/reports/baseline.json \
  --markdown-report eval/rag-retrieval/reports/baseline.md

Generate fixtures (live stack)

.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval

Dense baseline snapshot (same seed, comparison only):

.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval

(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)

Maven equivalent:

mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
  -Drag.snapshot.enabled=true \
  -Dretrieval.kb-scope=rag-eval \
  -Dretrieval.search.mode=hybrid \
  test

Generator is off in normal tests; only runs when rag.snapshot.enabled=true (writes files).

If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline with an explicit reason — do not silently overwrite.

Golden assertions

Supported expectation fields include:

  • expectedSources / expectedDocIds
  • expectedBreadcrumbs / expectedKeywords
  • expectedSelectedAttempt
  • expectedFallbackReason / expectedFallbackReasons
  • expectedEvidenceStatus
  • expectedContextSources
  • expectedRerankTopSource

Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.

Note: relevance_level is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).

Baseline diff

python scripts/eval_rag_retrieval.py \
  --json-report eval/rag-retrieval/reports/current.json \
  --markdown-report eval/rag-retrieval/reports/current.md \
  --compare-to eval/rag-retrieval/reports/baseline.json \
  --diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
  --diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md

Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
Non-zero exit on case failure or regression in diff mode.

Hit levels

  • strong: expected document found and breadcrumb or keyword coverage OK
  • medium: expected document found, coverage incomplete
  • weak: keyword hit without expected document
  • miss: neither

Recall@K counts strong + medium.

Optional live smoke (post-reindex)

After reindex, with app up:

python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900

Calls GET /api/search/similar. Environment smoke only — does not replace offline baseline.