Files
SuperBizAgent-java/eval/rag-retrieval/README.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

180 lines
6.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RAG Retrieval Baseline
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
Related design notes:
- `docs/RAG-Hybrid质量分与后处理.md`
- `docs/RAG-Agent如何读relevance_level.md`
- `mvp/architecture/RAG知识检索架构.md` §6
## Offline vs live
| Layer | What | Needs live stack? |
|-------|------|-------------------|
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
Daily CI / local quick check: **offline only**.
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
## Layout
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed queries + expectations
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
reports/baseline.json|md Last accepted offline report
reports/baseline-diff.* Optional diff vs previous report
```
## Fixture shape (minimum)
```text
caseId
query
retrievedAt
searchMode # hybrid | dense (required on newly generated fixtures)
kbScope # e.g. rag-eval when generation used a scope
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
```
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
Older fixtures may omit `searchMode`; regenerate to attach meta.
## Seed docs + import
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
```text
eval/rag-retrieval/seed-docs/*.md
```
Frontmatter example:
```yaml
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
Import via real upload pipeline:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
Isolation:
- App default may leave `retrieval.kb-scope` empty (all docs).
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
## Offline run (no live stack)
```bash
python scripts/eval_rag_retrieval.py
```
Custom paths:
```bash
python scripts/eval_rag_retrieval.py \
--cases eval/rag-retrieval/cases/golden-cases.json \
--fixtures eval/rag-retrieval/fixtures \
--json-report eval/rag-retrieval/reports/baseline.json \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Generate fixtures (live stack)
```powershell
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
```
Dense baseline snapshot (same seed, comparison only):
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
```
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
Maven equivalent:
```text
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
-Drag.snapshot.enabled=true \
-Dretrieval.kb-scope=rag-eval \
-Dretrieval.search.mode=hybrid \
test
```
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
## Golden assertions
Supported expectation fields include:
- `expectedSources` / `expectedDocIds`
- `expectedBreadcrumbs` / `expectedKeywords`
- `expectedSelectedAttempt`
- `expectedFallbackReason` / `expectedFallbackReasons`
- `expectedEvidenceStatus`
- `expectedContextSources`
- `expectedRerankTopSource`
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
## Baseline diff
```bash
python scripts/eval_rag_retrieval.py \
--json-report eval/rag-retrieval/reports/current.json \
--markdown-report eval/rag-retrieval/reports/current.md \
--compare-to eval/rag-retrieval/reports/baseline.json \
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
```
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
Non-zero exit on case failure or regression in diff mode.
## Hit levels
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
- `medium`: expected document found, coverage incomplete
- `weak`: keyword hit without expected document
- `miss`: neither
`Recall@K` counts `strong` + `medium`.
## Optional live smoke (post-reindex)
After reindex, with app up:
```bash
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
```
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.