Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
180 lines
6.0 KiB
Markdown
180 lines
6.0 KiB
Markdown
# RAG Retrieval Baseline
|
||
|
||
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
|
||
|
||
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
|
||
|
||
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
|
||
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
|
||
|
||
Related design notes:
|
||
|
||
- `docs/RAG-Hybrid质量分与后处理.md`
|
||
- `docs/RAG-Agent如何读relevance_level.md`
|
||
- `mvp/architecture/RAG知识检索架构.md` §6
|
||
|
||
## Offline vs live
|
||
|
||
| Layer | What | Needs live stack? |
|
||
|-------|------|-------------------|
|
||
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
|
||
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
|
||
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
|
||
|
||
Daily CI / local quick check: **offline only**.
|
||
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
|
||
|
||
## Layout
|
||
|
||
```text
|
||
eval/rag-retrieval/
|
||
cases/golden-cases.json Fixed queries + expectations
|
||
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
|
||
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
|
||
reports/baseline.json|md Last accepted offline report
|
||
reports/baseline-diff.* Optional diff vs previous report
|
||
```
|
||
|
||
## Fixture shape (minimum)
|
||
|
||
```text
|
||
caseId
|
||
query
|
||
retrievedAt
|
||
searchMode # hybrid | dense (required on newly generated fixtures)
|
||
kbScope # e.g. rag-eval when generation used a scope
|
||
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
|
||
```
|
||
|
||
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
|
||
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
|
||
|
||
Older fixtures may omit `searchMode`; regenerate to attach meta.
|
||
|
||
## Seed docs + import
|
||
|
||
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
|
||
|
||
```text
|
||
eval/rag-retrieval/seed-docs/*.md
|
||
```
|
||
|
||
Frontmatter example:
|
||
|
||
```yaml
|
||
source: mysql-connection-pool
|
||
breadcrumb: Database > MySQL > Connection Pool
|
||
kb_scope: rag-eval
|
||
```
|
||
|
||
Import via real upload pipeline:
|
||
|
||
```powershell
|
||
.\scripts\prepare_rag_eval_seed.ps1
|
||
```
|
||
|
||
Isolation:
|
||
|
||
- App default may leave `retrieval.kb-scope` empty (all docs).
|
||
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
|
||
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
|
||
|
||
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
|
||
|
||
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
|
||
|
||
## Offline run (no live stack)
|
||
|
||
```bash
|
||
python scripts/eval_rag_retrieval.py
|
||
```
|
||
|
||
Custom paths:
|
||
|
||
```bash
|
||
python scripts/eval_rag_retrieval.py \
|
||
--cases eval/rag-retrieval/cases/golden-cases.json \
|
||
--fixtures eval/rag-retrieval/fixtures \
|
||
--json-report eval/rag-retrieval/reports/baseline.json \
|
||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||
```
|
||
|
||
## Generate fixtures (live stack)
|
||
|
||
```powershell
|
||
.\scripts\prepare_rag_eval_seed.ps1
|
||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
|
||
```
|
||
|
||
Dense baseline snapshot (same seed, comparison only):
|
||
|
||
```powershell
|
||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
|
||
```
|
||
|
||
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
|
||
|
||
Maven equivalent:
|
||
|
||
```text
|
||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
|
||
-Drag.snapshot.enabled=true \
|
||
-Dretrieval.kb-scope=rag-eval \
|
||
-Dretrieval.search.mode=hybrid \
|
||
test
|
||
```
|
||
|
||
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
|
||
|
||
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
|
||
|
||
## Golden assertions
|
||
|
||
Supported expectation fields include:
|
||
|
||
- `expectedSources` / `expectedDocIds`
|
||
- `expectedBreadcrumbs` / `expectedKeywords`
|
||
- `expectedSelectedAttempt`
|
||
- `expectedFallbackReason` / `expectedFallbackReasons`
|
||
- `expectedEvidenceStatus`
|
||
- `expectedContextSources`
|
||
- `expectedRerankTopSource`
|
||
|
||
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
|
||
|
||
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
|
||
|
||
## Baseline diff
|
||
|
||
```bash
|
||
python scripts/eval_rag_retrieval.py \
|
||
--json-report eval/rag-retrieval/reports/current.json \
|
||
--markdown-report eval/rag-retrieval/reports/current.md \
|
||
--compare-to eval/rag-retrieval/reports/baseline.json \
|
||
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
|
||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||
```
|
||
|
||
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
|
||
Non-zero exit on case failure or regression in diff mode.
|
||
|
||
## Hit levels
|
||
|
||
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
|
||
- `medium`: expected document found, coverage incomplete
|
||
- `weak`: keyword hit without expected document
|
||
- `miss`: neither
|
||
|
||
`Recall@K` counts `strong` + `medium`.
|
||
|
||
## Optional live smoke (post-reindex)
|
||
|
||
After reindex, with app up:
|
||
|
||
```bash
|
||
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
|
||
```
|
||
|
||
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
|