feat(rag): close eval pipeline with live snapshots
This commit is contained in:
@@ -5,6 +5,7 @@
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
# Acceptance: rag-eval-pipeline-closure
|
||||
|
||||
## Status
|
||||
|
||||
Archived.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
| Item | Status | Notes |
|
||||
|---|---|---|
|
||||
| Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. |
|
||||
| LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. |
|
||||
| Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. |
|
||||
| Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. |
|
||||
| Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. |
|
||||
| Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. |
|
||||
| Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. |
|
||||
| Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. |
|
||||
| Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. |
|
||||
| Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. |
|
||||
|
||||
## Verification
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Result: passed. 7 cases, passRate=1.0, recall@5=1.0.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test
|
||||
```
|
||||
|
||||
Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=<temp-fixtures>" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test
|
||||
python scripts\eval_rag_retrieval.py --fixtures <temp-fixtures> --json-report <temp-current.json> --markdown-report <temp-current.md>
|
||||
```
|
||||
|
||||
Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1.
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py --json-report <temp-current.json> --markdown-report <temp-current.md> --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report <temp-diff.json> --diff-markdown-report <temp-diff.md>
|
||||
```
|
||||
|
||||
Result: passed. regressions=0.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test
|
||||
```
|
||||
|
||||
Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled.
|
||||
|
||||
```powershell
|
||||
$null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1))
|
||||
$null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1))
|
||||
```
|
||||
|
||||
Result: PowerShell syntax OK.
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack.
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z -SkipEval
|
||||
```
|
||||
|
||||
Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0.
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Brief: rag-eval-pipeline-closure
|
||||
|
||||
## Background
|
||||
|
||||
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Reuse the existing offline RAG retrieval baseline.
|
||||
2. Extend it to support modular `LookupResult` fixtures.
|
||||
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
|
||||
4. Add a RAG baseline diff path for regression detection.
|
||||
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
|
||||
6. Document how RAG baseline and diagnosis baseline form a quality loop.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new production API.
|
||||
- No LLM-as-judge scoring.
|
||||
- No production API behavior changes.
|
||||
- No replacement for diagnosis eval.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Decisions: rag-eval-pipeline-closure
|
||||
|
||||
## D1. Reuse the existing evaluator
|
||||
|
||||
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
|
||||
|
||||
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
|
||||
|
||||
## D2. Use LookupResult as the only fixture contract
|
||||
|
||||
Decision: support `lookupResult` only.
|
||||
|
||||
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
|
||||
|
||||
## D3. Make modular assertions opt-in per case
|
||||
|
||||
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
|
||||
|
||||
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
|
||||
|
||||
## D4. Diff remains deterministic
|
||||
|
||||
Decision: RAG diff compares report fields only and does not call live services or models.
|
||||
|
||||
Reason: this keeps it suitable for local regression checks and CI-style gates.
|
||||
|
||||
## D5. Isolate live eval docs with kb_scope
|
||||
|
||||
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
|
||||
|
||||
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
|
||||
|
||||
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
|
||||
|
||||
## D6. Import seed docs through the real upload pipeline
|
||||
|
||||
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
|
||||
|
||||
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
|
||||
|
||||
## D7. Strip frontmatter before chunk embedding
|
||||
|
||||
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
|
||||
|
||||
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
|
||||
|
||||
## D8. Treat retry behavior as the stable fallback contract
|
||||
|
||||
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
|
||||
|
||||
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
|
||||
|
||||
## D9. Default live snapshots to Spring AI VectorStore
|
||||
|
||||
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
|
||||
|
||||
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.
|
||||
@@ -0,0 +1,87 @@
|
||||
# Evidence: rag-eval-pipeline-closure
|
||||
|
||||
## Changed Assets
|
||||
|
||||
- `scripts/eval_rag_retrieval.py`
|
||||
- `eval/rag-retrieval/cases/golden-cases.json`
|
||||
- `eval/rag-retrieval/fixtures/*.json`
|
||||
- `eval/rag-retrieval/reports/baseline.json`
|
||||
- `eval/rag-retrieval/reports/baseline.md`
|
||||
- `eval/rag-retrieval/README.md`
|
||||
- `mvp/architecture/rag-eval-closure.md`
|
||||
- `src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/RagEvalSeedImporterTest.java`
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`
|
||||
- `scripts/prepare_rag_eval_seed.ps1`
|
||||
- `eval/rag-retrieval/seed-docs/*.md`
|
||||
- `src/main/java/com/superbiz/agent/dto/Frontmatter.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/KnowledgeEntry.java`
|
||||
- `src/main/java/com/superbiz/agent/service/FrontmatterParser.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||
- `src/main/resources/application.yml`
|
||||
|
||||
## Baseline Result
|
||||
|
||||
```text
|
||||
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
```
|
||||
|
||||
## Regression Signals
|
||||
|
||||
The evaluator now fails on:
|
||||
|
||||
- missing expected source
|
||||
- missing breadcrumb or evidence keyword
|
||||
- non-`lookupResult` fixture
|
||||
- selected attempt mismatch
|
||||
- fallback reason mismatch
|
||||
- fallback reason outside accepted values
|
||||
- evidence status mismatch
|
||||
- missing context source
|
||||
- rerank top source mismatch
|
||||
|
||||
The snapshot generator now provides:
|
||||
|
||||
- real `LookupKnowledgeTool` invocation
|
||||
- one fixture per golden case
|
||||
- explicit opt-in through `rag.snapshot.enabled=true`
|
||||
- optional post-generation baseline evaluation
|
||||
- scoped retrieval through `retrieval.kb-scope=rag-eval`
|
||||
- Spring AI VectorStore by default through `retrieval.vector-store.mode=spring`
|
||||
- scoped L0 hints through the same `retrieval.kb-scope`
|
||||
|
||||
The seed importer now provides:
|
||||
|
||||
- canonical eval docs under `eval/rag-retrieval/seed-docs`
|
||||
- real `DocumentManagementService` import/reindex
|
||||
- stable `source`/`docId` metadata
|
||||
- `kb_scope=rag-eval` isolation from local non-eval documents
|
||||
- frontmatter stripping before chunk embedding
|
||||
- an over-filter decoy seed doc for fallback-path evaluation
|
||||
|
||||
The diff now detects:
|
||||
|
||||
- aggregate pass/recall regression
|
||||
- case pass regression
|
||||
- hit-level regression
|
||||
- first-rank regression
|
||||
- selected attempt/fallback/evidence/rerank changes
|
||||
|
||||
## Final Spring Live Snapshot
|
||||
|
||||
```text
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z
|
||||
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
```
|
||||
|
||||
Key fallback trace:
|
||||
|
||||
```text
|
||||
selectedAttempt=UNFILTERED_VECTOR_RETRY
|
||||
fallbackReason=filtered_vector_no_evidence
|
||||
rerankTopSource=rag-l0-filter-fallback
|
||||
```
|
||||
Reference in New Issue
Block a user