feat(rag): close eval pipeline with live snapshots
This commit is contained in:
@@ -18,6 +18,7 @@
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||
@@ -37,8 +38,9 @@ SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat
|
||||
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
7. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
8. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
9. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
10. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
11. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
|
||||
@@ -37,7 +37,7 @@ flowchart TD
|
||||
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
|
||||
SpringTry -->|success| Results["SearchResult list"]
|
||||
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
|
||||
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||
Mode -->|spring / spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||
Mode -->|sdk| SdkOnly["Milvus SDK only"]
|
||||
|
||||
SpringOnly --> Results
|
||||
@@ -66,7 +66,7 @@ Agent Executor
|
||||
-> mode=auto
|
||||
-> Spring AI VectorStore
|
||||
-> fallback: Milvus SDK
|
||||
-> mode=spring-ai
|
||||
-> mode=spring / spring-ai
|
||||
-> Spring AI VectorStore only
|
||||
-> mode=sdk
|
||||
-> Milvus SDK only
|
||||
@@ -134,7 +134,7 @@ Executor -> LookupKnowledgeTool -> VectorSearchService
|
||||
`VectorSearchService` 是当前检索门面:
|
||||
|
||||
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
|
||||
- `spring-ai`:只走 Spring AI VectorStore。
|
||||
- `spring` / `spring-ai`:只走 Spring AI VectorStore。
|
||||
- `sdk`:只走原 Milvus SDK。
|
||||
|
||||
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
|
||||
@@ -374,7 +374,7 @@ RAG 架构变更必须先过评测,再认为可合入主链路。
|
||||
- `lookup_knowledge` 保持显式 Agent Tool。
|
||||
- L0 降级为 domain/entity hint。
|
||||
- L1 默认执行语义检索。
|
||||
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
|
||||
- `VectorSearchService` 支持 `auto`、`spring`/`spring-ai`、`sdk` 三种模式。
|
||||
- Spring AI VectorStore 成为读取主路径。
|
||||
- Milvus SDK fallback 保留。
|
||||
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
|
||||
|
||||
@@ -0,0 +1,181 @@
|
||||
# RAG 评测闭环架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
|
||||
本文记录当前 RAG 质量闭环。它的目标不是证明检索“永远正确”,而是让每次改 `lookup_knowledge`、L0 hint、向量召回、post-retrieval、rerank 或 context packing 时,都能得到可重复的回归信号。
|
||||
|
||||
## 1. 闭环分层
|
||||
|
||||
```text
|
||||
RAG pipeline change
|
||||
-> LookupKnowledgeTool snapshot generation
|
||||
-> offline RAG retrieval baseline
|
||||
-> RAG baseline diff
|
||||
-> diagnosis eval baseline
|
||||
-> diagnosis baseline diff
|
||||
-> accept / fix / archive
|
||||
```
|
||||
|
||||
| 层级 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| RAG retrieval baseline | `eval/rag-retrieval/` | 检查固定 query 是否命中期望证据、路径和 fallback |
|
||||
| RAG baseline diff | `scripts/eval_rag_retrieval.py --compare-to ...` | 对比当前报告和旧基线,输出 regression/change |
|
||||
| Diagnosis eval baseline | `mvp/eval/` | 检查 Agent 最终诊断 trace、报告和证据行为 |
|
||||
| Live acceptance | `scripts/eval_rag_live_acceptance.py` | 在应用和向量库运行后做真实环境 smoke check |
|
||||
|
||||
## 2. Offline RAG Baseline
|
||||
|
||||
核心资产:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/cases/golden-cases.json
|
||||
eval/rag-retrieval/fixtures/*.json
|
||||
eval/rag-retrieval/reports/baseline.json
|
||||
eval/rag-retrieval/reports/baseline.md
|
||||
scripts/eval_rag_retrieval.py
|
||||
scripts/generate_rag_lookup_snapshots.ps1
|
||||
src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java
|
||||
```
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
该 baseline 完全离线,不依赖 MySQL、Redis、Milvus、LLM 或 Spring Boot。它适合在改 RAG 代码后快速判断:
|
||||
|
||||
- 期望 source 是否仍在 topK 内。
|
||||
- breadcrumb 和 evidence keyword 是否仍能覆盖。
|
||||
- `LookupResult` 是否仍包含 `evidenceBlocks/contextPack/retrievalTrace/rerankTrace`。
|
||||
- selected attempt 是否符合预期。
|
||||
- fallback reason 是否符合预期。
|
||||
- context pack 是否包含期望 source。
|
||||
- rerank top source 是否稳定。
|
||||
|
||||
## 3. 模块化输出契约
|
||||
|
||||
fixture 必须使用当前模块化格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookupResult": {
|
||||
"evidenceBlocks": [],
|
||||
"contextPack": {},
|
||||
"retrievalTrace": {},
|
||||
"rerankTrace": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
当前 golden cases 直接以模块化格式为唯一契约,因为这个版本的目标是验证完整 RAG pipeline,而不只是验证候选召回。
|
||||
|
||||
## 4. Fallback Case
|
||||
|
||||
当前 baseline 增加了 `chat-l0-filter-fallback`:
|
||||
|
||||
```text
|
||||
FILTERED_VECTOR low quality or no evidence
|
||||
-> UNFILTERED_VECTOR_RETRY
|
||||
-> fallbackReason = filtered_vector_low_quality | filtered_vector_no_evidence
|
||||
```
|
||||
|
||||
这个 case 固化了 MVP 版本的降级策略:如果经过 L0 filter 后 L1 低质量或没有证据,就跳过 L0 filter,用原始 query 再做一次无过滤向量检索。不同向量后端对“低质量候选”和“无候选”的边界可能不同,所以 golden case 允许两个 fallback reason,但强制要求 retry 行为和最终证据正确。
|
||||
|
||||
## 5. Diff 闭环
|
||||
|
||||
生成当前报告并与旧基线对比:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py `
|
||||
--json-report eval\rag-retrieval\reports\current.json `
|
||||
--markdown-report eval\rag-retrieval\reports\current.md `
|
||||
--compare-to eval\rag-retrieval\reports\baseline.json `
|
||||
--diff-json-report eval\rag-retrieval\reports\baseline-diff.json `
|
||||
--diff-markdown-report eval\rag-retrieval\reports\baseline-diff.md
|
||||
```
|
||||
|
||||
diff 会检查:
|
||||
|
||||
- pass rate
|
||||
- recall@K
|
||||
- strong hit rate
|
||||
- miss count
|
||||
- case pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
|
||||
当 case 失败或 diff 出现 regression 时,脚本会返回非 0 退出码,可作为本地质量门禁或 CI 门禁。
|
||||
|
||||
## 6. 与 Diagnosis Eval 的关系
|
||||
|
||||
RAG baseline 解决的是“证据有没有被正确检索、处理和打包”。
|
||||
|
||||
Diagnosis eval 解决的是“Agent 有没有把证据用于最终诊断,并保持 trace 可解释”。
|
||||
|
||||
两者不是替代关系:
|
||||
|
||||
- 改 RAG pipeline:先跑 RAG baseline,再跑相关 Agent 测试。
|
||||
- 改 prompt、Agent 编排、Verifier:重点跑 diagnosis eval。
|
||||
- 改 embedding 输入、reindex、向量库配置:跑 RAG baseline + live acceptance。
|
||||
|
||||
## 7. 面试表达
|
||||
|
||||
可以概括为:
|
||||
|
||||
> 我没有只做一个 RAG 调用,而是把 RAG 拆成 Query Transform、Retrieval、Post-Retrieval、Rerank、Context Packing,并为它建设了离线 golden cases、baseline report、baseline diff 和上层 diagnosis eval,形成可回放、可对比、可回归的 Agent 质量闭环。
|
||||
|
||||
## 8. LookupKnowledgeTool Snapshot
|
||||
|
||||
真实工具快照生成命令:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
该命令默认使用 `retrieval.vector-store.mode=spring`,通过 `RagLookupSnapshotGeneratorTest` 启动 Spring test context,注入真实 `LookupKnowledgeTool` bean,对 `golden-cases.json` 中每个 query 调用 `lookupKnowledge(query)`,并把返回的 `LookupResult` 写入 `eval/rag-retrieval/fixtures/{caseId}.json`。
|
||||
|
||||
普通测试不会执行快照生成器;只有显式传入 `rag.snapshot.enabled=true` 时才会写 fixture。
|
||||
|
||||
## 9. Seed Docs And Scope Isolation
|
||||
|
||||
Live `LookupKnowledgeTool` snapshots are only stable if the expected documents
|
||||
exist in the real knowledge base and vector index. The eval loop therefore adds
|
||||
a canonical seed layer:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
-> scripts/prepare_rag_eval_seed.ps1
|
||||
-> RagEvalSeedImporterTest
|
||||
-> DocumentManagementService.uploadDocument
|
||||
-> api_document metadata + L0 index + Milvus chunks
|
||||
```
|
||||
|
||||
Seed frontmatter includes:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
`source` becomes the stable `docId` when it fits the DB column, and is also
|
||||
written to vector metadata as `_source` and `source`. `breadcrumb` is copied into
|
||||
chunk metadata so evidence blocks can keep a stable path. `kb_scope` isolates
|
||||
eval documents from local production documents.
|
||||
|
||||
Default runtime behavior keeps `retrieval.kb-scope` empty, so existing documents
|
||||
without `kb_scope` are still searchable. Eval scripts pass
|
||||
`-Dretrieval.kb-scope=rag-eval`, so L0 query hints, the filtered attempt, and
|
||||
the unfiltered retry stay inside the eval corpus while the retry still skips the
|
||||
L0 category filter.
|
||||
|
||||
Frontmatter is used for DB metadata, L0 hints, and vector metadata. It is
|
||||
stripped before document chunking so embedding content represents the Markdown
|
||||
body, not the YAML control plane. This is important for fallback eval: a decoy
|
||||
document may intentionally match L0 keywords, but its body should remain low
|
||||
quality evidence so the retry path can be exercised.
|
||||
@@ -30,7 +30,7 @@ flowchart TD
|
||||
L1 --> Mode{"retrieval.vector-store.mode"}
|
||||
Mode -->|auto| Spring["Spring AI VectorStore"]
|
||||
Spring -->|failure| SDK["Milvus SDK fallback"]
|
||||
Mode -->|spring-ai| Spring
|
||||
Mode -->|spring / spring-ai| Spring
|
||||
Mode -->|sdk| SDK
|
||||
|
||||
Spring --> Candidates["L1 candidates"]
|
||||
@@ -89,7 +89,7 @@ L1 通过 `VectorSearchService` 调度,支持三种模式:
|
||||
| 模式 | 行为 | 用途 |
|
||||
|---|---|---|
|
||||
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
|
||||
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||
| `spring` / `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
|
||||
|
||||
### Spring AI VectorStore 路径
|
||||
|
||||
Reference in New Issue
Block a user