feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
+89
-143
@@ -1,38 +1,65 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
This directory contains the offline retrieval baseline for the RAG refactor.
|
||||
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
|
||||
|
||||
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
||||
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
||||
evidence keywords before changing L0 behavior, query augmentation, evidence
|
||||
post-processing, or Spring AI VectorStore integration.
|
||||
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
|
||||
|
||||
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
|
||||
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
|
||||
|
||||
Related design notes:
|
||||
|
||||
- `docs/RAG-Hybrid质量分与后处理.md`
|
||||
- `docs/RAG-Agent如何读relevance_level.md`
|
||||
- `mvp/architecture/RAG知识检索架构.md` §6
|
||||
|
||||
## Offline vs live
|
||||
|
||||
| Layer | What | Needs live stack? |
|
||||
|-------|------|-------------------|
|
||||
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
|
||||
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
|
||||
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
|
||||
|
||||
Daily CI / local quick check: **offline only**.
|
||||
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
|
||||
|
||||
## Layout
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
||||
fixtures/*.json Saved retrieval fixtures for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/baseline-diff.* Optional diff reports
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
cases/golden-cases.json Fixed queries + expectations
|
||||
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
|
||||
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
|
||||
reports/baseline.json|md Last accepted offline report
|
||||
reports/baseline-diff.* Optional diff vs previous report
|
||||
```
|
||||
|
||||
## Seed Docs + Import/Reindex
|
||||
## Fixture shape (minimum)
|
||||
|
||||
The live-tool eval uses canonical seed documents so the real
|
||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
||||
```text
|
||||
caseId
|
||||
query
|
||||
retrievedAt
|
||||
searchMode # hybrid | dense (required on newly generated fixtures)
|
||||
kbScope # e.g. rag-eval when generation used a scope
|
||||
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
|
||||
```
|
||||
|
||||
Seed documents live in:
|
||||
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
|
||||
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
|
||||
|
||||
Older fixtures may omit `searchMode`; regenerate to attach meta.
|
||||
|
||||
## Seed docs + import
|
||||
|
||||
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
```
|
||||
|
||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
||||
Frontmatter example:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
Import or reindex the seed docs through the real upload pipeline:
|
||||
Import via real upload pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
||||
Milvus chunks.
|
||||
Isolation:
|
||||
|
||||
`kb_scope` isolates eval data:
|
||||
- App default may leave `retrieval.kb-scope` empty (all docs).
|
||||
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
|
||||
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
|
||||
|
||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
||||
without `kb_scope` remain searchable;
|
||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
||||
vector retrieval both use only the canonical eval seed docs;
|
||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
||||
boundary.
|
||||
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
|
||||
|
||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
||||
their frontmatter keywords matched the query.
|
||||
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
## Offline run (no live stack)
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Custom paths are also supported:
|
||||
Custom paths:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Generate Fixtures From LookupKnowledgeTool
|
||||
|
||||
Use the snapshot generator when fixtures should reflect the real
|
||||
`LookupKnowledgeTool` pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
For the intended live loop, run seed import first:
|
||||
## Generate fixtures (live stack)
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
python scripts\eval_rag_retrieval.py
|
||||
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
|
||||
```
|
||||
|
||||
The script runs a Spring test harness:
|
||||
|
||||
```text
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
||||
```
|
||||
|
||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
||||
|
||||
Custom paths are supported:
|
||||
Dense baseline snapshot (same seed, comparison only):
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
||||
-Fixtures eval\rag-retrieval\fixtures `
|
||||
-RetrievedAt 2026-07-06T00:00:00Z
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
|
||||
```
|
||||
|
||||
The generator is disabled in normal test runs. It only executes when
|
||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
||||
depends on the configured runtime retrieval stack.
|
||||
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
|
||||
|
||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
||||
signal: either the golden expectations need to be adjusted to the current
|
||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
||||
|
||||
## Modular RAG Contract
|
||||
|
||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
||||
`lookup_knowledge` output:
|
||||
Maven equivalent:
|
||||
|
||||
```text
|
||||
lookupResult.evidenceBlocks
|
||||
lookupResult.contextPack
|
||||
lookupResult.retrievalTrace
|
||||
lookupResult.rerankTrace
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
|
||||
-Drag.snapshot.enabled=true \
|
||||
-Dretrieval.kb-scope=rag-eval \
|
||||
-Dretrieval.search.mode=hybrid \
|
||||
test
|
||||
```
|
||||
|
||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
||||
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
|
||||
|
||||
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
|
||||
|
||||
## Golden assertions
|
||||
|
||||
Supported expectation fields include:
|
||||
|
||||
- `expectedSources` / `expectedDocIds`
|
||||
- `expectedBreadcrumbs`
|
||||
- `expectedKeywords`
|
||||
- `expectedBreadcrumbs` / `expectedKeywords`
|
||||
- `expectedSelectedAttempt`
|
||||
- `expectedFallbackReason`
|
||||
- `expectedFallbackReasons`
|
||||
- `expectedFallbackReason` / `expectedFallbackReasons`
|
||||
- `expectedEvidenceStatus`
|
||||
- `expectedContextSources`
|
||||
- `expectedRerankTopSource`
|
||||
|
||||
This lets the baseline catch regressions such as losing the expected evidence
|
||||
source, skipping context packing, changing the selected retrieval attempt, or
|
||||
breaking the filtered-vector to unfiltered-retry fallback.
|
||||
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
|
||||
|
||||
## Baseline Diff
|
||||
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
|
||||
|
||||
To compare a freshly generated report against an existing baseline:
|
||||
## Baseline diff
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
|
||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||
```
|
||||
|
||||
The diff reports aggregate regressions and case-level changes for:
|
||||
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
|
||||
Non-zero exit on case failure or regression in diff mode.
|
||||
|
||||
- pass rate, recall@K, strong hit rate, miss count
|
||||
- pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
## Hit levels
|
||||
|
||||
The command exits non-zero when a case fails or the diff contains a regression.
|
||||
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
|
||||
- `medium`: expected document found, coverage incomplete
|
||||
- `weak`: keyword hit without expected document
|
||||
- `miss`: neither
|
||||
|
||||
## Hit Levels
|
||||
`Recall@K` counts `strong` + `medium`.
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
||||
- `weak`: expected evidence keyword is found, but expected document is missing.
|
||||
- `miss`: expected document and expected evidence are not found.
|
||||
## Optional live smoke (post-reindex)
|
||||
|
||||
`Recall@K` counts `strong` and `medium` as retrieved.
|
||||
|
||||
## Scope
|
||||
|
||||
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
||||
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
||||
not a claim that live production retrieval accuracy is complete.
|
||||
|
||||
## Live Post-Reindex Acceptance
|
||||
|
||||
When embedding input changes, existing vectors do not update by themselves. For
|
||||
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
||||
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
||||
semantic signal.
|
||||
|
||||
Use this optional live acceptance flow after the application is running and the
|
||||
knowledge base has been reindexed:
|
||||
After reindex, with app up:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
|
||||
```
|
||||
|
||||
Custom service URL and output paths are supported:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py \
|
||||
--base-url http://127.0.0.1:9900 \
|
||||
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
||||
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The script calls:
|
||||
|
||||
```text
|
||||
GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
|
||||
|
||||
Reference in New Issue
Block a user