feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -3,7 +3,7 @@
## Context
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
- Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1.
- Delivery baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1.
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
@@ -9,7 +9,7 @@
As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.
This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checklist.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
This change is **Delivery 1** from `docs/Milvus-Hybrid接入清单.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
## What Changes
@@ -47,7 +47,7 @@ This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checkl
- Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
- Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
- Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids
- Docs baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1
- Docs baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1
## Risks
@@ -0,0 +1,6 @@
Committed OpenSpec
change: rag-eval-hybrid-baseline
committed_at: 2026-07-28
scale: standard-lean
scope: knife-1-only
acceptance: wiring-required; fixture-refresh-best-effort
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-28
@@ -0,0 +1,13 @@
# Brief: rag-eval-hybrid-baseline
## Background
Offline RAG eval structure is correct but generator/docs/fixtures predate hybrid search mode.
## Goals
Knife 1 only: `search.mode` wiring, fixture meta, README, best-effort fixture refresh.
## Non-goals
Dual dense/hybrid fixture trees; golden mustNot/chunk/level hard gates; new frameworks.
@@ -0,0 +1,111 @@
# Decisions — rag-eval-hybrid-baseline
## sm-flow meta
- **Checkpoint**: Discover (in progress)
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
- **Slug**: `rag-eval-hybrid-baseline`
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
## Clarify summary
| Item | Content |
|------|---------|
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
| Non-goals | New framework, LLM judge, prod retrieval redesign |
## Context summary
| Source | Conclusion | Into OpenSpec |
|--------|------------|---------------|
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
## Question pool (grill)
| ID | Dim | Mode | Question | Status |
|----|-----|------|----------|--------|
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
### Q3–Q5 evidence
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
### Q1 用户确认
- **选择**: 仅第一刀(推荐)
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
### Q2 用户确认
- **选择**: 接线优先,刷新可记未验证
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
---
## Discover status
- [x] clarify
- [x] context
- [x] propose (`proposal.md`)
- [x] grill (Q1–Q5 closed)
**Discover checkpoint: 完成。**
---
## Commit checkpoint
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
- **Gate**: `.committed` written
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
---
## Apply checkpoint
- **Capability**: openspec-apply-change + Committed tasks
- **Authorization**: user「实现」
### Delivered (knife-1)
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
- `eval/rag-retrieval/README.md` hybrid-era docs
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
### Apply-discovered regression + fix
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
### Commands run
```text
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
```
**Apply checkpoint: 完成。** Ready for Archive when user requests.
@@ -0,0 +1,62 @@
# Design: rag-eval-hybrid-baseline
## Context
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
## Goals / Non-Goals
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
## Decisions
### D1 — Replace vector-store mode with search mode
| Before | After |
|--------|--------|
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
### D2 — Fixture meta minimum
```text
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
```
- `searchMode`: actual mode used for generation.
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
### D3 — LookupResult payload
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
### D4 — Acceptance if live refresh fails
Must deliver: ps1, test meta emission, README.
Should attempt: seed + generate + eval.
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
### D5 — Baseline update policy
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
## Risks
| Risk | Mitigation |
|------|------------|
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
## Interface impact
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
## Audit
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
@@ -0,0 +1,66 @@
# Change: Align RAG offline eval with hybrid + qualityScore era
## Why
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
- No first-class dense vs hybrid fixture split for recall comparison.
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
## What Changes
### Knife 1 (must) — make offline eval reflect current main path
1. **Generator wiring**
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
2. **Fixture meta**
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
3. **Refresh path**
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
4. **Docs**
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
### Out of scope this change (confirmed grill)
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
## Non-goals
- Rewriting eval into a new framework or LLM-as-judge.
- Full threshold calibration productization.
- Neighbor chunks / query rewrite / cross-encoder.
- Changing production retrieval code paths (except eval generator test harness props).
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
- Dense/hybrid dual fixture directories (later change).
## Context constraints
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
- Seed isolation via `kb_scope=rag-eval` stays.
## Impact
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
## Success
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
- Fixtures carry searchMode meta.
- README describes correct offline/live loop.
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
- (If knife 2) dual fixture roots documented and runnable.
@@ -0,0 +1,50 @@
# rag-eval-offline-baseline Specification
## Purpose
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
## ADDED Requirements
### Requirement: Snapshot generation SHALL use retrieval search mode
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
#### Scenario: Default hybrid generation
- **WHEN** the snapshot generator is invoked with default parameters
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
#### Scenario: Dense mode override for comparison runs
- **WHEN** the operator sets search mode to `dense`
- **THEN** fixture generation SHALL use dense retrieval for that run
### Requirement: Generated fixtures SHALL record search meta
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
#### Scenario: Meta fields present
- **WHEN** a fixture is written for a golden case
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
### Requirement: Offline evaluation SHALL remain dependency-free
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
#### Scenario: Offline eval without live stack
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
- **THEN** it SHALL produce pass/fail results using fixture contents only
### Requirement: Eval documentation SHALL describe the hybrid-era loop
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
#### Scenario: README main path
- **WHEN** an engineer follows the eval README happy path
- **THEN** the documented default generation mode SHALL be hybrid search mode
@@ -0,0 +1,34 @@
# Tasks: rag-eval-hybrid-baseline
## 1. Generator wiring
- [x] 1.1 Update `scripts/generate_rag_lookup_snapshots.ps1`: replace `VectorStoreMode` / `vector-store.mode` with `SearchMode` default `hybrid` and `-Dretrieval.search.mode=...`
- [x] 1.2 Keep `-Dretrieval.kb-scope` (default `rag-eval`); document `-SearchMode dense` override
- [x] 1.3 Ensure snapshot test picks up `retrieval.search.mode` (system property and/or `@SpringBootTest` properties)
## 2. Fixture meta
- [x] 2.1 `RagLookupSnapshotGeneratorTest` writes `searchMode` and `kbScope` (when set) on each fixture
- [x] 2.2 Confirm offline `eval_rag_retrieval.py` still loads fixtures (ignore extra meta)
## 3. Docs
- [x] 3.1 Rewrite `eval/rag-retrieval/README.md` hybrid-era loop; remove spring vector-store as default generation path
- [x] 3.2 Note offline vs live responsibilities; point to qualityScore/hybrid main path briefly
## 4. Refresh attempt (best-effort)
- [x] 4.1 Attempt `prepare_rag_eval_seed` + hybrid snapshot generate + offline eval when environment allows
- [x] 4.2 On success: update fixtures and `reports/baseline.*` if needed after diff review
- [x] 4.3 On failure: record exact commands, error summary, and “unverified refresh” in change decisions/acceptance notes — do not block wiring delivery
## 5. Verify
- [x] 5.1 Static: grep shows no required `vector-store.mode` in snapshot generator path
- [x] 5.2 Offline: `python scripts/eval_rag_retrieval.py` runs on committed fixtures (pass or documented baseline drift)
## 6. Apply-discovered fix (hybrid quality gate)
- [x] 6.1 Hybrid attaches optional `denseDistance` without overwriting RRF order/scoreLabel
- [x] 6.2 `toQualityScore(hybrid)` prefers dense L2 for absolute gates; rank fallback if no dense
- [x] 6.3 Regenerate fixtures; L0 filter fallback case green again; baseline updated
@@ -0,0 +1,6 @@
Committed OpenSpec
change: rag-quality-score-unify
committed_at: 2026-07-28
scale: standard
interface_impact: L2
gate: proposal+design+specs+tasks+cross-artifact+audit
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-28
@@ -0,0 +1,20 @@
# Brief: rag-quality-score-unify
## Background
True BM25 hybrid retrieval is live, but post-processing still normalizes as if every score were dense L2 and re-ranks with L0 keyword contains boosts. That splits ranking authority from quality gates and double-counts lexical signal.
## Goals
- Unify score labels to `dense` | `hybrid`.
- Single `toQualityScore`; label-agnostic post-process.
- Preserve retrieval rank; remove boost re-rank.
- Hybrid quality = pure rank mapping (slice 1).
## Scope
Internal RAG pipeline: store emission, normalizer, evidence post-process, tests, architecture docs.
## Non-goals
Fine re-rankers, query rewrite, neighbor chunks, schema rebuild, ACI field renames, removing dense comparison mode.
@@ -0,0 +1,136 @@
# Decisions — rag-quality-score-unify
## sm-flow meta
- **Checkpoint**: Discover(clarify + context + propose + grill)
- **Scale**: standard
- **Capability**: sm-flow 内置协议;openspec CLI `new change`;grill 使用内置协议(conversation-confirmed + evidence-driven),标注 fallback:未调用外部 `grill-with-docs` skill 文件执行器
- **Slug**: `rag-quality-score-unify`
- **OpenSpec path**: `openspec/changes/rag-quality-score-unify/`
## Clarify summary
| 项 | 内容 |
|---|---|
| 问题 | hybrid 已 RRF 融合,后处理仍 L2 伪装 + 关键词 boost 改序,质量信号不统一 |
| 期望 | label 仅 dense/hybrid;toQualityScore 唯一归一化;后处理保 rank、去 boost 改序 |
| 影响代码 | `MilvusHybridKnowledgeStore`, `VectorSearchService`, `KnowledgeEvidencePostProcessor`, retrieval 包新类, DTO 注释, 测试, 架构文档 |
| 非目标 | 精排/rewrite/邻块、删 dense mode、改 ACI 字段结构、改 schema |
## Context summary (devflow)
| 来源 | 结论 | 需进 OpenSpec |
|---|---|---|
| `devflow/index.md` | rag-chunk-identity / bm25-hybrid / hybrid-rrf 均 archived | 是:承接不回退 |
| `rag-bm25-hybrid-drop-sdk/decisions.md` | 曾要求 dense L2 enrichment 兼容阈值 | **是:本 change 废止该 decision** |
| glossary | lookup_knowledge 为证据工具;不在此改 ACI 主结构 | 是:非目标 |
| 架构文档 §6.0 | mode dense=对照,hybrid=主路径 | 是:保留 |
**index 使用状态**: 已命中相关 RAG 条目。
## Question pool (grill)
| ID | 维度 | 模式 | 问题 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | user-interview(对话已确认) | 一级 scoreLabel 是否只有 dense/hybrid,bm25_only 不作正式 label? | **已确认** |
| Q2 | 边界 | user-interview(对话已确认) | 后处理是否去掉关键词 contains 加分改序,仅保 originalRank? | **已确认** |
| Q3 | 边界 | user-interview(对话已确认) | 归一化是否唯一 toQualityScore;后处理 label-agnostic? | **已确认** |
| Q4 | 验收 | user-interview(对话已确认) | dense mode 保留作召回对照;主路径 hybrid? | **已确认** |
| Q5 | 技术 | evidence-driven | 当前代码是否仍 L2 回填 + boost 重排? | **已查证** |
| Q6 | 技术 | evidence-driven | Agent ACI 是否暴露 scoreLabel? | **已查证** |
| Q7 | 验收 | user-interview(对话已确认) | 接受 relevance_level / retry 分布变化? | **已确认** |
| Q8 | 边界 | user-interview | hybrid quality 切片 1 是否采用**纯 rank 映射**(不做 max(rank,denseSim))? | **已确认** |
### Q1–Q4, Q7 用户确认摘录(本会话)
- Label:「应该只有 hybrid 和 dense」「bm25_only 不是第三种」→ 同意收成两种。
- 归一化:「抽取抽象转换…后处理抽象统一」→ 同意。
- 后处理:「关键词打分不合理」「可以,就按照这个」(去 boost 改序 + 归一化一起做)。
- mode:保留 dense 作对照,写入架构 §6.0。
- 行为变化:讨论中已说明 hybrid 顺序/等级/retry 会变,用户要求按该方案实施(经 sm-flow)。
### Q5 evidence-driven
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后仍 dense 回填 L2 / `bm25_only_no_dense`。
- `KnowledgeEvidencePostProcessor.score`:`normalizeL2` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序。
- `LookupKnowledgeTool`:`isLowQuality` 看 `topSimilarity`(来自 baseScore)。
### Q6 evidence-driven
- Agent 主契约 `RagToolResult` / projector 暴露 evidence 列表与 relevance_level,不依赖 scoreLabel 字符串;改 label 为内部/L2 影响。
### Q8 用户确认(2026-07-28)
- **问题原文**: hybrid 的 qualityScore(切片 1)采用哪种映射?
- **用户选择**: 纯 rank 映射(推荐)
- **确认状态**: 已确认
- **实现约束**: `toQualityScore(hybrid)` = `rankToQuality(originalRank, batchSize)`;不看 RRF 原分量纲;不做 `max(rank, denseSim)`;不在 hybrid 路径为质量闸门再查/回填 dense L2。
---
## Discover status
- [x] clarify
- [x] context
- [x] propose (`proposal.md`)
- [x] grill 完成(Q1–Q8 均已关闭)
**Discover checkpoint: 完成。**
---
## Commit checkpoint
### Capability
- specify: sm-flow 内置 + openspec status/instructions(fallback:按 template 手写 design/specs/tasks)
- audit: sm-flow 内置协议(未调用外部 zoom-out)
- commit gate: 文件完整性 + 一致性检查后写入 `.committed`
### Cross-artifact 对齐
| 链路 | 状态 |
|---|---|
| brief/proposal 目标范围 → design | 已对齐 |
| design 决策(label/normalizer/保序/去 boost/纯 rank)→ specs | 已对齐 |
| specs 可观察行为 → tasks 可执行切片 | 已对齐 |
| decisions Q1–Q8 → proposal/design/specs | 已对齐 |
### Audit(≤5 句)
1. 链路仍是 Tool→Retriever→Store→Post→Pack→Project,无新外部系统。
2. 分数所有权上收 store 发射 + normalizer;后处理只裁剪与质量闸门。
3. 废止 bm25-hybrid 的「dense L2 enrichment」决策,属有意行为变化(L2 接口影响)。
4. 风险主要是 hybrid 序数 quality 与阈值标定,已记入 design Risks。
5. 不触及 Agent ACI 字段名与 Milvus schema。
### Commit gate checklist
- [x] proposal / design / specs / tasks / brief 存在
- [x] 核心概念在 design 有对应
- [x] design 关键决策在 tasks 有任务
- [x] tasks 可验证(checkbox 纵向切片)
- [x] 无未确认 user-interview
- [x] `.committed` 已创建
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
**下一步**: 等待用户明确授权 **Apply**(例如「开始 apply / 实现」)。未授权前不改业务接线代码。
---
## Apply checkpoint
- **Capability**: openspec-apply-change + Committed OpenSpec tasks
- **授权**: 用户「实现」
- **完成**: tasks.md 全部勾选
- **验证**:
- `RetrievalScoreNormalizerTest` 4 passed
- `KnowledgeEvidencePostProcessorTest` 6 passed
- `LookupKnowledgeToolTest` 7 passed
- `VectorSearchServiceTest` 2 passed
- `VectorKnowledgeSearchAdapterHybridTest` 1 passed
- **已知限制**: hybrid quality 为本轮 rank 序数映射,跨 query 绝对值不可比;阈值可能需后续标定
- **行为变化**: 已落地(去 L2 回填、去 boost 改序、label dense/hybrid)
**Apply checkpoint: 完成。** 可进入 Archive(需用户确认是否 archive OpenSpec)。
@@ -0,0 +1,134 @@
# Design: rag-quality-score-unify
## Context
- Knowledge path already uses single `MilvusHybridKnowledgeStore` (MilvusClientV2) with dense + BM25 + RRF.
- Chunk-level `evidenceKey` dedup and `retrieve-k` / `return-n` are landed.
- Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
- Prior design in `rag-bm25-hybrid-drop-sdk` required dense L2 enrichment for threshold compatibility — **this change supersedes that decision**.
Stakeholders: `lookup_knowledge` internal pipeline; Agent ACI field *names* unchanged; operators comparing `retrieval.search.mode=dense|hybrid`.
## Goals / Non-Goals
**Goals:**
1. First-class `scoreLabel` values: only `dense` | `hybrid` (aliases canonicalize).
2. Single `toQualityScore` adapter; post-process is label-agnostic.
3. Preserve retrieval `originalRank` as sort authority; remove keyword/domain boost re-ranking.
4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
5. Keep `mode=dense` for offline recall comparison; production default remains hybrid.
**Non-Goals:**
- Cross-encoder / query rewrite / neighbor chunks.
- Schema rebuild or collection rename.
- Changing Agent-facing ACI JSON field names.
- Configurable `max(rank, denseSim)` quality (future).
## Decisions
### D1 — Two labels only
| label | `score` meaning | quality mapping |
|---|---|---|
| `dense` | L2 distance (smaller better) | `1 - clamp(l2)/maxL2Distance` |
| `hybrid` | engine fused score optional in `rawScore`; **not** used as L2 | `rankToQuality(originalRank, batchSize)` |
Canonicalize legacy strings: `l2_distance`→dense; `rrf_fused` / `bm25_only_*`→hybrid.
**Why not keep `bm25_only`:** it is not a search mode; it was a L2-fake patch. Hybrid path hits are all `hybrid`.
### D2 — Stop dense L2 overwrite on hybrid hits
`searchHybrid` SHALL:
1. Run `hybridSearch` + RRFRanker.
2. Emit hits in RRF order with `scoreLabel=hybrid`, `originalRank=1..n`.
3. Set `rawScore` from engine when present; `score` MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption.
4. MUST NOT set `bm25_only_no_dense` or force `score=maxL2Distance` for threshold faking.
5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).
`searchDense` SHALL emit `scoreLabel=dense` and L2 in `score`.
### D3 — `RetrievalScoreNormalizer` is the only label branch
```text
qualityScore = RetrievalScoreNormalizer.toQualityScore(
scoreLabel, score, originalRank, batchSize, maxL2Distance)
```
- Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
- Dense: existing L2 formula (behavior parity for dense mode).
Post-processor, `isLowQuality`, and `relevance_level` consume **only** `qualityScore` (exposed today as `baseScore` / `topSimilarity` fields for minimal DTO churn).
### D4 — Post-process sort and boosts
```text
order = originalRank ASC, then stable evidenceKey
// NO finalScore = base + 0.15 domain + ...
```
- Remove additive boosts from sort key and from `finalScore` used for ordering.
- Optional: if L0 hint string matches, append explanatory `hitReasons` only (e.g. `l0_keyword_overlap`) — zero score delta.
- Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
- `relevance_level`: compare top `qualityScore` to existing thresholds; **remove** `hasHintSupport` gate for PRECISE.
- `isLowQuality` / category unfiltered retry: unchanged control flow, new score semantics.
### D5 — Trace fields
- `RerankTrace` may keep `baseScore`/`finalScore` names but both equal `qualityScore` when boosts are zero; `boostReasons` empty or explanation-only reasons without `:+0.xx` score deltas.
- Prefer renaming comments to “quality trace”; no Agent contract change required.
### D6 — WIP files
Workspace may contain draft `RetrievalScoreLabels` / `RetrievalScoreNormalizer`. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.
## Interface impact
- **Level: L2** — internal DTO semantics (`score`, `scoreLabel`), post-process ordering and relevance distribution.
- Agent ACI: field names stable; `relevance_level` *values distribution* may change (accepted behavioral change).
## Data flow (target)
```text
VectorSearchService (mode dense|hybrid)
-> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
-> KnowledgeDocumentRetriever candidates
-> KnowledgeEvidencePostProcessor
for each: qualityScore = toQualityScore(...)
sort by originalRank
dedup / caps / return-n
relevance + topSimilarity from qualityScore
-> pack / assemble / project
```
## Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute |
| Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” |
| PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains |
| Tests assert boost re-order | Update `LookupKnowledgeToolTest.rerankUsesHintMatches...` |
| Old label strings in sidecar/eval | canonicalize in normalizer |
## Migration Plan
1. Deploy code; no Milvus schema migration.
2. Default `retrieval.search.mode=hybrid` unchanged.
3. Rollback: revert change; old L2-enrichment behavior returns.
4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).
## Open Questions
- None for slice-1 (Q8 confirmed: pure rank).
- Follow-up: threshold calibration after live traces; optional denseSim blend.
## Audit notes (inline)
Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
No new cross-module lifecycle. Couples only internal RAG pipeline.
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.
@@ -0,0 +1,80 @@
# Change: Unify RAG quality score (dense/hybrid labels) and stop keyword boost re-rank
## Why
BM25 hybrid 已在库内完成 dense + BM25 + RRF 融合,但后处理仍:
1. 把 hybrid 结果**伪装成 L2** 再 `normalizeL2`(含 `bm25_only_no_dense` 弱分占位);
2. 用 L0 domain/entity/keyword **contains 加分改主序**。
这导致:排序信号与质量闸门分裂;词面信号被 BM25 与后处理**双重计分**;「词面热、语义冷」的片段可能被抬到前面;hybrid 的 RRF 序被冲掉。
需要统一:**检索负责序,后处理只做 quality 归一化 + 裁剪装配**。
## What Changes
### 检索层(`MilvusHybridKnowledgeStore` / `VectorSearchService`)
- 一级 `scoreLabel` 仅两种:`dense` | `hybrid`(与 `retrieval.search.mode` 对齐)。
- **废弃**正式一级 label:`l2_distance` / `rrf_fused` / `bm25_only_no_dense`(可读兼容映射到 dense/hybrid)。
- `mode=dense`:`score` = L2 距离,`label=dense`,`originalRank` = ANN 序。
- `mode=hybrid`:`label=hybrid`;**不再**用 dense L2 覆盖主 `score`;**不再**对 BM25-only 伪造 maxL2;`originalRank` = RRF 返回序;`rawScore` 可保留引擎融合分。
- `mode=dense|hybrid` **保留**:hybrid 为线上主路径;dense 为同库对照/评测(已写入架构 §6.0)。
### 归一化(新)
- 新增唯一转换点 `RetrievalScoreNormalizer.toQualityScore(label, score, rank, batchSize, maxL2)` → `qualityScore ∈ [0,1]`(越大越好)。
- `dense`:`1 - clamp(L2)/maxL2Distance`
- `hybrid`:按 **rank** 映射(本轮 batch 线性),不把 RRF 原分当 L2 套公式。
- Label 差异**只**在此消化。
### 后处理(`KnowledgeEvidencePostProcessor`)
- **统一流程**,只消费 `qualityScore` + `originalRank`(label-agnostic)。
- **排序主序 = `originalRank` 升序**(保检索序);去掉 domain/entity/keyword/source_type **加分改序**。
- L0 contains 匹配若保留,仅写入 `hitReasons` / trace 解释,**不参与 sort key、不加 finalScore**。
- `relevance_level` / `isLowQuality` / category unfiltered retry:只看 top `qualityScore` 与既有阈值;**不再**要求 `hasHintSupport` 才能 PRECISE。
- 保留:evidenceKey 去重、`max-chunks-per-document`、`return-n`、excerpt 截断、EvidenceBlock 装配。
### 文档 / 测试
- 更新 `mvp/architecture/RAG知识检索架构.md` §6 分数与后处理约定。
- 单测:dense 路径 quality 与现 L2 归一化一致;hybrid 保序且不被 keyword 打乱;无 `bm25_only` 一级 label;旧 label 别名可 canonicalize。
## Non-goals
- Cross-encoder / listwise 精排、query rewrite、邻块扩展。
- 删除 `mode=dense` 对照开关。
- 改变 Agent 可见 ACI 字段结构(`evidence[]` / `relevance_level` 枚举名可不变;**分布会变**)。
- 修改 Milvus schema / 强制全量 rebuild(本 change 不改 collection 结构)。
- 上线可配 `max(rank, denseSim)` 混合 quality(可后续迭代;本 change 切片 1 用纯 rank 映射 hybrid)。
## Context constraints (from devflow)
- 承接 archived:`rag-chunk-evidence-identity-dedup`、`rag-bm25-hybrid-drop-sdk`、`rag-hybrid-search-rrf`、`modular-rag-pipeline`。
- 单一知识后端仍为 `MilvusHybridKnowledgeStore`(V2);不得恢复 sdk/spring 主路径路由。
- Agent 投影仍不暴露 raw fused score / 完整 contextPack 作为主契约(内部 LookupResult/trace 可保留调试字段)。
- 历史 decision「Dense L2 enrichment for threshold compatibility」**本 change 有意废止**,改为 qualityScore 统一闸门。
## Impact
- **行为变化(对内检索质量语义)**:
- hybrid 下证据顺序更贴近 RRF;
- 词面命中不再被后处理 contains 二次抬序;
- `relevance_level` 与 unfiltered retry 触发分布可能变化;
- BM25-only 命中不再被标成 quality≈0。
- **接口影响**:L2(内部 DTO/注释/scoreLabel 字符串约定);Agent ACI 字段名不变。
- **风险**:hybrid rank→quality 为序数映射,绝对值不跨 query 可比;阈值 0.75/0.5 可能需后续观测再调(本 change 先沿用配置项)。
## Scale
- **standard**(多文件、有意行为变化、需 design + specs + tasks + 测试)。
## Depends on
- 已落地 hybrid schema + chunk evidenceKey(archived changes 如上)。
- 对话已确认的设计口径(见 change `decisions.md`)。
## WIP note
- 工作区可能已有未接线的 `RetrievalScoreLabels` / `RetrievalScoreNormalizer` 草稿文件;apply 阶段以 **Committed OpenSpec** 为准接入或改写,不视为已完成实现。
@@ -0,0 +1,94 @@
# rag-retrieval-quality-score Specification
## Purpose
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
## ADDED Requirements
### Requirement: Primary score labels SHALL be only dense or hybrid
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
#### Scenario: Dense mode labels hits as dense
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
- **AND** `score` SHALL be the dense L2 distance
#### Scenario: Hybrid mode labels hits as hybrid
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
#### Scenario: Legacy aliases canonicalize
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
#### Scenario: No L2 enrichment overwrite
- **WHEN** hybrid search completes
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
### Requirement: Quality score SHALL be produced by a single normalizer
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
#### Scenario: Dense quality from L2
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
#### Scenario: Hybrid quality from rank
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
### Requirement: Post-process SHALL preserve retrieval rank order
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
#### Scenario: Keyword overlap does not promote lower rank
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
- **AND** B matches more L0 keywords via string contains than A
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
#### Scenario: Structural caps still apply after rank order
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
- **AND** `rag.return-n` SHALL still bound total blocks
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
#### Scenario: Low quality uses top qualityScore
- **WHEN** post-process finishes with at least one evidence block
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
#### Scenario: PRECISE does not require hint support
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
### Requirement: Dense mode remains available for recall comparison
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
#### Scenario: Mode dense still callable
- **WHEN** mode is `dense`
- **THEN** search SHALL call dense ANN only and label hits `dense`
@@ -0,0 +1,39 @@
# Tasks: rag-quality-score-unify
## 1. Score contract utilities
- [x] 1.1 Finalize `RetrievalScoreLabels` (`dense` / `hybrid` + canonicalize legacy aliases)
- [x] 1.2 Finalize `RetrievalScoreNormalizer.toQualityScore` (dense L2 formula; hybrid pure rank map with batchSize)
- [x] 1.3 Unit tests for normalizer: dense L2 edges; hybrid rank monotonicity; alias canonicalize
## 2. Store / search emission
- [x] 2.1 `searchDense`: emit `scoreLabel=dense`, L2 `score`, stable rank order
- [x] 2.2 `searchHybrid`: emit `scoreLabel=hybrid`; keep RRF order as `originalRank`; stop dense L2 overwrite and `bm25_only_*` labels; no parallel dense probe for score rewrite
- [x] 2.3 Update `VectorSearchService.SearchResult` / adapter comments so `score`+`scoreLabel` contract matches design
- [x] 2.4 Ensure `KnowledgeDocumentRetriever` / `VectorKnowledgeSearchAdapter` propagate `scoreLabel`, `score`, `rawScore`, `originalRank` unchanged
## 3. Post-process
- [x] 3.1 `KnowledgeEvidencePostProcessor`: compute quality via normalizer; sort by `originalRank` ASC (stable tie-break)
- [x] 3.2 Remove domain/entity/keyword/source_type additive boosts from ordering/`finalScore`
- [x] 3.3 Optional: L0 overlap only as explanatory `hitReasons` (no score delta)
- [x] 3.4 `relevance_level` / `isLowQuality` / `topSimilarity` use qualityScore only; drop `hasHintSupport` gate for PRECISE
- [x] 3.5 Keep evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate
## 4. Tests
- [x] 4.1 Update `KnowledgeEvidencePostProcessorTest` for rank order + caps under new scoring
- [x] 4.2 Update `LookupKnowledgeToolTest.rerankUsesHintMatchesAndContextPackPreservesMetadata` (no boost re-order; metadata/context pack still ok)
- [x] 4.3 Adjust any tests asserting `l2_distance` / boost reasons `:+0.xx` as needed
- [x] 4.4 Run targeted unit tests for touched classes
## 5. Docs
- [x] 5.1 Update `mvp/architecture/RAG知识检索架构.md` §6 score/post-process (replace L2-enrichment narrative)
- [x] 5.2 Align `application.yml` comments if still describing L2-only post-process for hybrid
## 6. Verify
- [x] 6.1 Confirm no production path still sets `bm25_only_no_dense` or overwrites hybrid scores with L2 for thresholds
- [x] 6.2 Note known limitation: hybrid quality is ordinal within batch; thresholds may need later calibration