Files
SuperBizAgent-java/mvp/architecture/archive/2026-07-22-legacy/rag-eval-closure.md
T

184 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RAG 评测闭环架构
**更新日期**:2026-07-06
本文记录当前 RAG 质量闭环。它的目标不是证明检索“永远正确”,而是让每次改 `lookup_knowledge`、L0 hint、向量召回、post-retrieval、rerank 或 context packing 时,都能得到可重复的回归信号。
## 1. 闭环分层
```text
RAG pipeline change
-> LookupKnowledgeTool snapshot generation
-> offline RAG retrieval baseline
-> RAG baseline diff
-> diagnosis eval baseline
-> diagnosis baseline diff
-> accept / fix / archive
```
| 层级 | 位置 | 作用 |
|---|---|---|
| RAG retrieval baseline | `eval/rag-retrieval/` | 检查固定 query 是否命中期望证据、路径和 fallback |
| RAG baseline diff | `scripts/eval_rag_retrieval.py --compare-to ...` | 对比当前报告和旧基线,输出 regression/change |
| Diagnosis eval baseline | `mvp/eval/` | 检查 Agent 最终诊断 trace、报告和证据行为 |
| Live acceptance | `scripts/eval_rag_live_acceptance.py` | 在应用和向量库运行后做真实环境 smoke check |
## 2. Offline RAG Baseline
核心资产:
```text
eval/rag-retrieval/cases/golden-cases.json
eval/rag-retrieval/fixtures/*.json
eval/rag-retrieval/reports/baseline.json
eval/rag-retrieval/reports/baseline.md
scripts/eval_rag_retrieval.py
scripts/generate_rag_lookup_snapshots.ps1
src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java
```
运行:
```powershell
python scripts\eval_rag_retrieval.py
```
该 baseline 完全离线,不依赖 MySQL、Redis、Milvus、LLM 或 Spring Boot。它适合在改 RAG 代码后快速判断:
- 期望 source 是否仍在 topK 内。
- breadcrumb 和 evidence keyword 是否仍能覆盖。
- `LookupResult` 是否仍包含 `evidenceBlocks/contextPack/retrievalTrace/rerankTrace`。
- selected attempt 是否符合预期。
- fallback reason 是否符合预期。
- context pack 是否包含期望 source。
- rerank top source 是否稳定。
## 3. 模块化输出契约
fixture 必须使用当前模块化格式:
```json
{
"lookupResult": {
"evidenceBlocks": [],
"contextPack": {},
"retrievalTrace": {},
"rerankTrace": {}
}
}
```
当前 golden cases 直接以模块化格式为唯一契约,因为这个版本的目标是验证完整 RAG pipeline,而不只是验证候选召回。
## 4. Fallback Case
当前 baseline 增加了 `chat-l0-filter-fallback`:
```text
FILTERED_VECTOR low quality or no evidence
-> UNFILTERED_VECTOR_RETRY
-> fallbackReason = filtered_vector_low_quality | filtered_vector_no_evidence
```
这个 case 固化了 MVP 版本的降级策略:如果经过 L0 filter 后 L1 低质量或没有证据,就跳过 L0 filter,用原始 query 再做一次无过滤向量检索。不同向量后端对“低质量候选”和“无候选”的边界可能不同,所以 golden case 允许两个 fallback reason,但强制要求 retry 行为和最终证据正确。
## 5. Diff 闭环
生成当前报告并与旧基线对比:
```powershell
python scripts\eval_rag_retrieval.py `
--json-report eval\rag-retrieval\reports\current.json `
--markdown-report eval\rag-retrieval\reports\current.md `
--compare-to eval\rag-retrieval\reports\baseline.json `
--diff-json-report eval\rag-retrieval\reports\baseline-diff.json `
--diff-markdown-report eval\rag-retrieval\reports\baseline-diff.md
```
diff 会检查:
- pass rate
- recall@K
- strong hit rate
- miss count
- case pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
当 case 失败或 diff 出现 regression 时,脚本会返回非 0 退出码,可作为本地质量门禁或 CI 门禁。
## 6. 与 Diagnosis Eval 的关系
RAG baseline 解决的是“证据有没有被正确检索、处理和打包”。
Diagnosis eval 解决的是“Agent 有没有把证据用于最终诊断,并保持 trace 可解释”。
两者不是替代关系:
- 改 RAG pipeline:先跑 RAG baseline,再跑相关 Agent 测试。
- 改 prompt、Agent 编排、Verifier:重点跑 diagnosis eval。
- 改 embedding 输入、reindex、向量库配置:跑 RAG baseline + live acceptance。
## 7. 面试表达
可以概括为:
> 我没有只做一个 RAG 调用,而是把 RAG 拆成 Query Transform、Retrieval、Post-Retrieval、Rerank、Context Packing,并为它建设了离线 golden cases、baseline report、baseline diff 和上层 diagnosis eval,形成可回放、可对比、可回归的 Agent 质量闭环。
## 8. LookupKnowledgeTool Snapshot
真实工具快照生成命令:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
该命令默认使用 `retrieval.vector-store.mode=spring`,通过 `RagLookupSnapshotGeneratorTest` 启动 Spring test context,注入真实 `LookupKnowledgeTool` bean,对 `golden-cases.json` 中每个 query 调用 `lookupKnowledge(query)`,并把返回的 `LookupResult` 写入 `eval/rag-retrieval/fixtures/{caseId}.json`。
普通测试不会执行快照生成器;只有显式传入 `rag.snapshot.enabled=true` 时才会写 fixture。
## 9. Seed Docs And Scope Isolation
Live `LookupKnowledgeTool` snapshots are only stable if the expected documents
exist in the real knowledge base and vector index. The eval loop therefore adds
a canonical seed layer:
```text
eval/rag-retrieval/seed-docs/*.md
-> scripts/prepare_rag_eval_seed.ps1
-> RagEvalSeedImporterTest
-> DocumentManagementService.uploadDocument
-> api_document metadata + L0 index + Milvus chunks
```
仓库内还保留一份 `knowledge_base/rag-eval/` 镜像,方便直接查看和提交 eval 知识库文档。它们放在单独目录下,避免和 `knowledge_base/api`、`knowledge_base/infrastructure` 等业务知识目录混在一起;检索 category 仍由 frontmatter 中的 `category` 决定。
Seed frontmatter includes:
```yaml
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
`source` becomes the stable `docId` when it fits the DB column, and is also
written to vector metadata as `_source` and `source`. `breadcrumb` is copied into
chunk metadata so evidence blocks can keep a stable path. `kb_scope` isolates
eval documents from local production documents.
Default runtime behavior keeps `retrieval.kb-scope` empty, so existing documents
without `kb_scope` are still searchable. Eval scripts pass
`-Dretrieval.kb-scope=rag-eval`, so L0 query hints, the filtered attempt, and
the unfiltered retry stay inside the eval corpus while the retry still skips the
L0 category filter.
Frontmatter is used for DB metadata, L0 hints, and vector metadata. It is
stripped before document chunking so embedding content represents the Markdown
body, not the YAML control plane. This is important for fallback eval: a decoy
document may intentionally match L0 keywords, but its body should remain low
quality evidence so the retry path can be exercised.