Files
SuperBizAgent-java/eval/rag-retrieval/README.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

6.0 KiB
Raw Blame History

RAG Retrieval Baseline

Offline regression harness for lookup_knowledge after hybrid retrieval + qualityScore post-process.

It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is not a full diagnosis-agent E2E.

Production knowledge path: MilvusHybridKnowledgeStore with retrieval.search.mode=hybrid (dense+BM25+RRF).
mode=dense remains a same-collection baseline for recall comparison (not a second index).

Related design notes:

  • docs/RAG-Hybrid质量分与后处理.md
  • docs/RAG-Agent如何读relevance_level.md
  • mvp/architecture/RAG知识检索架构.md §6

Offline vs live

Layer What Needs live stack?
Offline fixtures/*.json × golden-cases.json → pass/fail + baseline diff No (no Milvus/LLM/Boot)
Snapshot generate Real LookupKnowledgeTool writes fixtures Yes (embedding + Milvus + DB/L0 as configured)
Live smoke optional eval_rag_live_acceptance.py Yes (running app)

Daily CI / local quick check: offline only.
After changing retrieval, indexing, or search mode: regenerate fixtures, then offline eval, then update baseline if the diff is intentional.

Layout

eval/rag-retrieval/
  cases/golden-cases.json      Fixed queries + expectations
  seed-docs/*.md               Canonical docs for live snapshot (kb_scope: rag-eval)
  fixtures/*.json              Frozen lookupResult snapshots (+ searchMode meta)
  reports/baseline.json|md     Last accepted offline report
  reports/baseline-diff.*      Optional diff vs previous report

Fixture shape (minimum)

caseId
query
retrievedAt
searchMode          # hybrid | dense  (required on newly generated fixtures)
kbScope             # e.g. rag-eval when generation used a scope
lookupResult        # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …

Offline eval ignores unknown top-level meta and does not full-JSON-compare.
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.

Older fixtures may omit searchMode; regenerate to attach meta.

Seed docs + import

Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:

eval/rag-retrieval/seed-docs/*.md

Frontmatter example:

source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval

Import via real upload pipeline:

.\scripts\prepare_rag_eval_seed.ps1

Isolation:

  • App default may leave retrieval.kb-scope empty (all docs).
  • Eval generation passes -Dretrieval.kb-scope=rag-eval.
  • Category-filter fallback retries without L0 category filter only; kb_scope still applies.

Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).

Seeds must live in the current hybrid collection schema (milvus.collection, default biz). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.

Offline run (no live stack)

python scripts/eval_rag_retrieval.py

Custom paths:

python scripts/eval_rag_retrieval.py \
  --cases eval/rag-retrieval/cases/golden-cases.json \
  --fixtures eval/rag-retrieval/fixtures \
  --json-report eval/rag-retrieval/reports/baseline.json \
  --markdown-report eval/rag-retrieval/reports/baseline.md

Generate fixtures (live stack)

.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval

Dense baseline snapshot (same seed, comparison only):

.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval

(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)

Maven equivalent:

mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
  -Drag.snapshot.enabled=true \
  -Dretrieval.kb-scope=rag-eval \
  -Dretrieval.search.mode=hybrid \
  test

Generator is off in normal tests; only runs when rag.snapshot.enabled=true (writes files).

If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline with an explicit reason — do not silently overwrite.

Golden assertions

Supported expectation fields include:

  • expectedSources / expectedDocIds
  • expectedBreadcrumbs / expectedKeywords
  • expectedSelectedAttempt
  • expectedFallbackReason / expectedFallbackReasons
  • expectedEvidenceStatus
  • expectedContextSources
  • expectedRerankTopSource

Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.

Note: relevance_level is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).

Baseline diff

python scripts/eval_rag_retrieval.py \
  --json-report eval/rag-retrieval/reports/current.json \
  --markdown-report eval/rag-retrieval/reports/current.md \
  --compare-to eval/rag-retrieval/reports/baseline.json \
  --diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
  --diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md

Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
Non-zero exit on case failure or regression in diff mode.

Hit levels

  • strong: expected document found and breadcrumb or keyword coverage OK
  • medium: expected document found, coverage incomplete
  • weak: keyword hit without expected document
  • miss: neither

Recall@K counts strong + medium.

Optional live smoke (post-reindex)

After reindex, with app up:

python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900

Calls GET /api/search/similar. Environment smoke only — does not replace offline baseline.