feat(rag): close eval pipeline with live snapshots

This commit is contained in:
zhuyongxin
2026-07-06 21:39:27 +08:00
parent cf3333d607
commit ed7efc58b7
47 changed files with 2613 additions and 177 deletions
+7 -5
View File
@@ -18,6 +18,7 @@
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
@@ -37,8 +38,9 @@ SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
7. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
8. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
9. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
10. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
11. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+4 -4
View File
@@ -37,7 +37,7 @@ flowchart TD
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
SpringTry -->|success| Results["SearchResult list"]
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
Mode -->|spring / spring-ai| SpringOnly["Spring AI VectorStore only"]
Mode -->|sdk| SdkOnly["Milvus SDK only"]
SpringOnly --> Results
@@ -66,7 +66,7 @@ Agent Executor
-> mode=auto
-> Spring AI VectorStore
-> fallback: Milvus SDK
-> mode=spring-ai
-> mode=spring / spring-ai
-> Spring AI VectorStore only
-> mode=sdk
-> Milvus SDK only
@@ -134,7 +134,7 @@ Executor -> LookupKnowledgeTool -> VectorSearchService
`VectorSearchService` 是当前检索门面:
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
- `spring-ai`:只走 Spring AI VectorStore。
- `spring` / `spring-ai`:只走 Spring AI VectorStore。
- `sdk`:只走原 Milvus SDK。
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
@@ -374,7 +374,7 @@ RAG 架构变更必须先过评测,再认为可合入主链路。
- `lookup_knowledge` 保持显式 Agent Tool。
- L0 降级为 domain/entity hint。
- L1 默认执行语义检索。
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
- `VectorSearchService` 支持 `auto`、`spring`/`spring-ai`、`sdk` 三种模式。
- Spring AI VectorStore 成为读取主路径。
- Milvus SDK fallback 保留。
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
+181
View File
@@ -0,0 +1,181 @@
# RAG 评测闭环架构
**更新日期**:2026-07-06
本文记录当前 RAG 质量闭环。它的目标不是证明检索“永远正确”,而是让每次改 `lookup_knowledge`、L0 hint、向量召回、post-retrieval、rerank 或 context packing 时,都能得到可重复的回归信号。
## 1. 闭环分层
```text
RAG pipeline change
-> LookupKnowledgeTool snapshot generation
-> offline RAG retrieval baseline
-> RAG baseline diff
-> diagnosis eval baseline
-> diagnosis baseline diff
-> accept / fix / archive
```
| 层级 | 位置 | 作用 |
|---|---|---|
| RAG retrieval baseline | `eval/rag-retrieval/` | 检查固定 query 是否命中期望证据、路径和 fallback |
| RAG baseline diff | `scripts/eval_rag_retrieval.py --compare-to ...` | 对比当前报告和旧基线,输出 regression/change |
| Diagnosis eval baseline | `mvp/eval/` | 检查 Agent 最终诊断 trace、报告和证据行为 |
| Live acceptance | `scripts/eval_rag_live_acceptance.py` | 在应用和向量库运行后做真实环境 smoke check |
## 2. Offline RAG Baseline
核心资产:
```text
eval/rag-retrieval/cases/golden-cases.json
eval/rag-retrieval/fixtures/*.json
eval/rag-retrieval/reports/baseline.json
eval/rag-retrieval/reports/baseline.md
scripts/eval_rag_retrieval.py
scripts/generate_rag_lookup_snapshots.ps1
src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java
```
运行:
```powershell
python scripts\eval_rag_retrieval.py
```
该 baseline 完全离线,不依赖 MySQL、Redis、Milvus、LLM 或 Spring Boot。它适合在改 RAG 代码后快速判断:
- 期望 source 是否仍在 topK 内。
- breadcrumb 和 evidence keyword 是否仍能覆盖。
- `LookupResult` 是否仍包含 `evidenceBlocks/contextPack/retrievalTrace/rerankTrace`。
- selected attempt 是否符合预期。
- fallback reason 是否符合预期。
- context pack 是否包含期望 source。
- rerank top source 是否稳定。
## 3. 模块化输出契约
fixture 必须使用当前模块化格式:
```json
{
"lookupResult": {
"evidenceBlocks": [],
"contextPack": {},
"retrievalTrace": {},
"rerankTrace": {}
}
}
```
当前 golden cases 直接以模块化格式为唯一契约,因为这个版本的目标是验证完整 RAG pipeline,而不只是验证候选召回。
## 4. Fallback Case
当前 baseline 增加了 `chat-l0-filter-fallback`:
```text
FILTERED_VECTOR low quality or no evidence
-> UNFILTERED_VECTOR_RETRY
-> fallbackReason = filtered_vector_low_quality | filtered_vector_no_evidence
```
这个 case 固化了 MVP 版本的降级策略:如果经过 L0 filter 后 L1 低质量或没有证据,就跳过 L0 filter,用原始 query 再做一次无过滤向量检索。不同向量后端对“低质量候选”和“无候选”的边界可能不同,所以 golden case 允许两个 fallback reason,但强制要求 retry 行为和最终证据正确。
## 5. Diff 闭环
生成当前报告并与旧基线对比:
```powershell
python scripts\eval_rag_retrieval.py `
--json-report eval\rag-retrieval\reports\current.json `
--markdown-report eval\rag-retrieval\reports\current.md `
--compare-to eval\rag-retrieval\reports\baseline.json `
--diff-json-report eval\rag-retrieval\reports\baseline-diff.json `
--diff-markdown-report eval\rag-retrieval\reports\baseline-diff.md
```
diff 会检查:
- pass rate
- recall@K
- strong hit rate
- miss count
- case pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
当 case 失败或 diff 出现 regression 时,脚本会返回非 0 退出码,可作为本地质量门禁或 CI 门禁。
## 6. 与 Diagnosis Eval 的关系
RAG baseline 解决的是“证据有没有被正确检索、处理和打包”。
Diagnosis eval 解决的是“Agent 有没有把证据用于最终诊断,并保持 trace 可解释”。
两者不是替代关系:
- 改 RAG pipeline:先跑 RAG baseline,再跑相关 Agent 测试。
- 改 prompt、Agent 编排、Verifier:重点跑 diagnosis eval。
- 改 embedding 输入、reindex、向量库配置:跑 RAG baseline + live acceptance。
## 7. 面试表达
可以概括为:
> 我没有只做一个 RAG 调用,而是把 RAG 拆成 Query Transform、Retrieval、Post-Retrieval、Rerank、Context Packing,并为它建设了离线 golden cases、baseline report、baseline diff 和上层 diagnosis eval,形成可回放、可对比、可回归的 Agent 质量闭环。
## 8. LookupKnowledgeTool Snapshot
真实工具快照生成命令:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
该命令默认使用 `retrieval.vector-store.mode=spring`,通过 `RagLookupSnapshotGeneratorTest` 启动 Spring test context,注入真实 `LookupKnowledgeTool` bean,对 `golden-cases.json` 中每个 query 调用 `lookupKnowledge(query)`,并把返回的 `LookupResult` 写入 `eval/rag-retrieval/fixtures/{caseId}.json`。
普通测试不会执行快照生成器;只有显式传入 `rag.snapshot.enabled=true` 时才会写 fixture。
## 9. Seed Docs And Scope Isolation
Live `LookupKnowledgeTool` snapshots are only stable if the expected documents
exist in the real knowledge base and vector index. The eval loop therefore adds
a canonical seed layer:
```text
eval/rag-retrieval/seed-docs/*.md
-> scripts/prepare_rag_eval_seed.ps1
-> RagEvalSeedImporterTest
-> DocumentManagementService.uploadDocument
-> api_document metadata + L0 index + Milvus chunks
```
Seed frontmatter includes:
```yaml
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
`source` becomes the stable `docId` when it fits the DB column, and is also
written to vector metadata as `_source` and `source`. `breadcrumb` is copied into
chunk metadata so evidence blocks can keep a stable path. `kb_scope` isolates
eval documents from local production documents.
Default runtime behavior keeps `retrieval.kb-scope` empty, so existing documents
without `kb_scope` are still searchable. Eval scripts pass
`-Dretrieval.kb-scope=rag-eval`, so L0 query hints, the filtered attempt, and
the unfiltered retry stay inside the eval corpus while the retry still skips the
L0 category filter.
Frontmatter is used for DB metadata, L0 hints, and vector metadata. It is
stripped before document chunking so embedding content represents the Markdown
body, not the YAML control plane. This is important for fallback eval: a decoy
document may intentionally match L0 keywords, but its body should remain low
quality evidence so the retry path can be exercised.
+2 -2
View File
@@ -30,7 +30,7 @@ flowchart TD
L1 --> Mode{"retrieval.vector-store.mode"}
Mode -->|auto| Spring["Spring AI VectorStore"]
Spring -->|failure| SDK["Milvus SDK fallback"]
Mode -->|spring-ai| Spring
Mode -->|spring / spring-ai| Spring
Mode -->|sdk| SDK
Spring --> Candidates["L1 candidates"]
@@ -89,7 +89,7 @@ L1 通过 `VectorSearchService` 调度,支持三种模式:
| 模式 | 行为 | 用途 |
|---|---|---|
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
| `spring` / `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
### Spring AI VectorStore 路径