feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1 @@
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
Committed OpenSpec
|
||||
change: rag-eval-hybrid-baseline
|
||||
committed_at: 2026-07-28
|
||||
scale: standard-lean
|
||||
scope: knife-1-only
|
||||
acceptance: wiring-required; fixture-refresh-best-effort
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-28
|
||||
@@ -0,0 +1,13 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval structure is correct but generator/docs/fixtures predate hybrid search mode.
|
||||
|
||||
## Goals
|
||||
|
||||
Knife 1 only: `search.mode` wiring, fixture meta, README, best-effort fixture refresh.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dual dense/hybrid fixture trees; golden mustNot/chunk/level hard gates; new frameworks.
|
||||
@@ -0,0 +1,111 @@
|
||||
# Decisions — rag-eval-hybrid-baseline
|
||||
|
||||
## sm-flow meta
|
||||
|
||||
- **Checkpoint**: Discover (in progress)
|
||||
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
|
||||
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
|
||||
- **Slug**: `rag-eval-hybrid-baseline`
|
||||
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
|
||||
|
||||
## Clarify summary
|
||||
|
||||
| Item | Content |
|
||||
|------|---------|
|
||||
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
|
||||
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
|
||||
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
|
||||
| Non-goals | New framework, LLM judge, prod retrieval redesign |
|
||||
|
||||
## Context summary
|
||||
|
||||
| Source | Conclusion | Into OpenSpec |
|
||||
|--------|------------|---------------|
|
||||
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
|
||||
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
|
||||
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
|
||||
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
|
||||
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
|
||||
|
||||
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
|
||||
|
||||
## Question pool (grill)
|
||||
|
||||
| ID | Dim | Mode | Question | Status |
|
||||
|----|-----|------|----------|--------|
|
||||
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
|
||||
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
|
||||
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
|
||||
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
|
||||
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
|
||||
|
||||
### Q3–Q5 evidence
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
|
||||
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
|
||||
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
|
||||
|
||||
### Q1 用户确认
|
||||
|
||||
- **选择**: 仅第一刀(推荐)
|
||||
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
|
||||
|
||||
### Q2 用户确认
|
||||
|
||||
- **选择**: 接线优先,刷新可记未验证
|
||||
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
|
||||
|
||||
---
|
||||
|
||||
## Discover status
|
||||
|
||||
- [x] clarify
|
||||
- [x] context
|
||||
- [x] propose (`proposal.md`)
|
||||
- [x] grill (Q1–Q5 closed)
|
||||
|
||||
**Discover checkpoint: 完成。**
|
||||
|
||||
---
|
||||
|
||||
## Commit checkpoint
|
||||
|
||||
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
|
||||
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
|
||||
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
|
||||
- **Gate**: `.committed` written
|
||||
|
||||
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
|
||||
|
||||
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
|
||||
|
||||
---
|
||||
|
||||
## Apply checkpoint
|
||||
|
||||
- **Capability**: openspec-apply-change + Committed tasks
|
||||
- **Authorization**: user「实现」
|
||||
|
||||
### Delivered (knife-1)
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
|
||||
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
|
||||
- `eval/rag-retrieval/README.md` hybrid-era docs
|
||||
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
|
||||
|
||||
### Apply-discovered regression + fix
|
||||
|
||||
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
|
||||
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
|
||||
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
|
||||
|
||||
### Commands run
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
```
|
||||
|
||||
**Apply checkpoint: 完成。** Ready for Archive when user requests.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Design: rag-eval-hybrid-baseline
|
||||
|
||||
## Context
|
||||
|
||||
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
|
||||
|
||||
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1 — Replace vector-store mode with search mode
|
||||
|
||||
| Before | After |
|
||||
|--------|--------|
|
||||
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
|
||||
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
|
||||
|
||||
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
|
||||
|
||||
### D2 — Fixture meta minimum
|
||||
|
||||
```text
|
||||
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
|
||||
```
|
||||
|
||||
- `searchMode`: actual mode used for generation.
|
||||
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
|
||||
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
|
||||
|
||||
### D3 — LookupResult payload
|
||||
|
||||
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
|
||||
|
||||
### D4 — Acceptance if live refresh fails
|
||||
|
||||
Must deliver: ps1, test meta emission, README.
|
||||
Should attempt: seed + generate + eval.
|
||||
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
|
||||
|
||||
### D5 — Baseline update policy
|
||||
|
||||
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Mitigation |
|
||||
|------|------------|
|
||||
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
|
||||
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
|
||||
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
|
||||
|
||||
## Interface impact
|
||||
|
||||
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
|
||||
|
||||
## Audit
|
||||
|
||||
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
|
||||
@@ -0,0 +1,66 @@
|
||||
# Change: Align RAG offline eval with hybrid + qualityScore era
|
||||
|
||||
## Why
|
||||
|
||||
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
|
||||
|
||||
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
|
||||
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
|
||||
- No first-class dense vs hybrid fixture split for recall comparison.
|
||||
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
|
||||
|
||||
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
|
||||
|
||||
## What Changes
|
||||
|
||||
### Knife 1 (must) — make offline eval reflect current main path
|
||||
|
||||
1. **Generator wiring**
|
||||
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
|
||||
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
|
||||
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
|
||||
|
||||
2. **Fixture meta**
|
||||
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
|
||||
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
|
||||
|
||||
3. **Refresh path**
|
||||
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
|
||||
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
|
||||
|
||||
4. **Docs**
|
||||
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
|
||||
|
||||
### Out of scope this change (confirmed grill)
|
||||
|
||||
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
|
||||
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Rewriting eval into a new framework or LLM-as-judge.
|
||||
- Full threshold calibration productization.
|
||||
- Neighbor chunks / query rewrite / cross-encoder.
|
||||
- Changing production retrieval code paths (except eval generator test harness props).
|
||||
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
|
||||
- Dense/hybrid dual fixture directories (later change).
|
||||
|
||||
## Context constraints
|
||||
|
||||
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
|
||||
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
|
||||
- Seed isolation via `kb_scope=rag-eval` stays.
|
||||
|
||||
## Impact
|
||||
|
||||
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
|
||||
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
|
||||
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
|
||||
|
||||
## Success
|
||||
|
||||
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
|
||||
- Fixtures carry searchMode meta.
|
||||
- README describes correct offline/live loop.
|
||||
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
|
||||
- (If knife 2) dual fixture roots documented and runnable.
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
# rag-eval-offline-baseline Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Snapshot generation SHALL use retrieval search mode
|
||||
|
||||
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
|
||||
|
||||
#### Scenario: Default hybrid generation
|
||||
|
||||
- **WHEN** the snapshot generator is invoked with default parameters
|
||||
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
|
||||
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
|
||||
|
||||
#### Scenario: Dense mode override for comparison runs
|
||||
|
||||
- **WHEN** the operator sets search mode to `dense`
|
||||
- **THEN** fixture generation SHALL use dense retrieval for that run
|
||||
|
||||
### Requirement: Generated fixtures SHALL record search meta
|
||||
|
||||
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
|
||||
|
||||
#### Scenario: Meta fields present
|
||||
|
||||
- **WHEN** a fixture is written for a golden case
|
||||
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
|
||||
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
|
||||
|
||||
### Requirement: Offline evaluation SHALL remain dependency-free
|
||||
|
||||
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
|
||||
|
||||
#### Scenario: Offline eval without live stack
|
||||
|
||||
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
|
||||
- **THEN** it SHALL produce pass/fail results using fixture contents only
|
||||
|
||||
### Requirement: Eval documentation SHALL describe the hybrid-era loop
|
||||
|
||||
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
|
||||
|
||||
#### Scenario: README main path
|
||||
|
||||
- **WHEN** an engineer follows the eval README happy path
|
||||
- **THEN** the documented default generation mode SHALL be hybrid search mode
|
||||
@@ -0,0 +1,34 @@
|
||||
# Tasks: rag-eval-hybrid-baseline
|
||||
|
||||
## 1. Generator wiring
|
||||
|
||||
- [x] 1.1 Update `scripts/generate_rag_lookup_snapshots.ps1`: replace `VectorStoreMode` / `vector-store.mode` with `SearchMode` default `hybrid` and `-Dretrieval.search.mode=...`
|
||||
- [x] 1.2 Keep `-Dretrieval.kb-scope` (default `rag-eval`); document `-SearchMode dense` override
|
||||
- [x] 1.3 Ensure snapshot test picks up `retrieval.search.mode` (system property and/or `@SpringBootTest` properties)
|
||||
|
||||
## 2. Fixture meta
|
||||
|
||||
- [x] 2.1 `RagLookupSnapshotGeneratorTest` writes `searchMode` and `kbScope` (when set) on each fixture
|
||||
- [x] 2.2 Confirm offline `eval_rag_retrieval.py` still loads fixtures (ignore extra meta)
|
||||
|
||||
## 3. Docs
|
||||
|
||||
- [x] 3.1 Rewrite `eval/rag-retrieval/README.md` hybrid-era loop; remove spring vector-store as default generation path
|
||||
- [x] 3.2 Note offline vs live responsibilities; point to qualityScore/hybrid main path briefly
|
||||
|
||||
## 4. Refresh attempt (best-effort)
|
||||
|
||||
- [x] 4.1 Attempt `prepare_rag_eval_seed` + hybrid snapshot generate + offline eval when environment allows
|
||||
- [x] 4.2 On success: update fixtures and `reports/baseline.*` if needed after diff review
|
||||
- [x] 4.3 On failure: record exact commands, error summary, and “unverified refresh” in change decisions/acceptance notes — do not block wiring delivery
|
||||
|
||||
## 5. Verify
|
||||
|
||||
- [x] 5.1 Static: grep shows no required `vector-store.mode` in snapshot generator path
|
||||
- [x] 5.2 Offline: `python scripts/eval_rag_retrieval.py` runs on committed fixtures (pass or documented baseline drift)
|
||||
|
||||
## 6. Apply-discovered fix (hybrid quality gate)
|
||||
|
||||
- [x] 6.1 Hybrid attaches optional `denseDistance` without overwriting RRF order/scoreLabel
|
||||
- [x] 6.2 `toQualityScore(hybrid)` prefers dense L2 for absolute gates; rank fallback if no dense
|
||||
- [x] 6.3 Regenerate fixtures; L0 filter fallback case green again; baseline updated
|
||||
Reference in New Issue
Block a user