feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
+1 -1
View File
@@ -65,8 +65,8 @@ uploads/
### Windows / Runtime Artifacts ### Windows / Runtime Artifacts
*.stackdump *.stackdump
NUL
### MVP Demo Generated Outputs ### MVP Demo Generated Outputs
mvp/demo/output/*.json mvp/demo/output/*.json
!mvp/demo/output/README.md !mvp/demo/output/README.md
.pi/extensions/emdash-hook.ts
+3 -1
View File
@@ -1,4 +1,4 @@
# devflow 索引 # devflow 索引
## Issue 生命周期 ## Issue 生命周期
@@ -12,6 +12,8 @@
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 | | 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|
| 2026-07-28 | rag-eval-hybrid-baseline | 离线 RAG eval 对齐 hybrid:search.mode 生成器、fixture meta、baseline 重刷;hybrid 质量闸门可用 denseDistance。 | RAG/eval/baseline | search.mode, fixture meta, kb_scope rag-eval, L0 filter fallback, denseDistance | openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline | archived |
| 2026-07-28 | rag-quality-score-unify | 统一 dense/hybrid scoreLabel 与 qualityScore;保检索序;去掉关键词 boost 改序与 hybrid L2 伪装。 | RAG/质量分/后处理 | qualityScore, scoreLabel dense/hybrid, originalRank, RetrievalScoreNormalizer, no boost rerank | openspec/changes/archive/2026-07-28-rag-quality-score-unify | archived |
| 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived | | 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived |
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived | | 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
| 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived | | 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived |
@@ -14,7 +14,7 @@ lookup_knowledge 同文档多 chunk 在后处理与投影阶段被 source 级去
## 范围 ## 范围
Delivery 1 only(见 `docs/milvus-hybrid-search-integration-checklist.md` §1.1)。 Delivery 1 only(见 `docs/Milvus-Hybrid接入清单.md` §1.1)。
## 非目标 ## 非目标
@@ -93,7 +93,7 @@ Reference files:
- `RagResultProjector.java` - `RagResultProjector.java`
- `LookupKnowledgeToolTest.java` - `LookupKnowledgeToolTest.java`
- `RagResultProjectorTest.java` - `RagResultProjectorTest.java`
- `docs/milvus-hybrid-search-integration-checklist.md` - `docs/Milvus-Hybrid接入清单.md`
Stack notes: Stack notes:
@@ -9,7 +9,7 @@
## 规格证据 ## 规格证据
- 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正) - 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正)
- `docs/milvus-hybrid-search-integration-checklist.md` §1.1 定义 Delivery 1 地基 - `docs/Milvus-Hybrid接入清单.md` §1.1 定义 Delivery 1 地基
## 验证证据 ## 验证证据
@@ -0,0 +1,39 @@
# Acceptance: rag-eval-hybrid-baseline
## Tasks
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
## 静态验证
- Snapshot generator path: no required `retrieval.vector-store.mode`.
- README documents hybrid generation and offline/live split.
## 脚本验证
```text
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
# exit 0
```
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
## 浏览器/人工
- 未做 UI 验证。
## 未验证 / 后续
- Dense vs hybrid dual-directory comparison report (knife-2).
- CI wiring of offline eval as required gate (optional process).
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
## Specs
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
@@ -0,0 +1,16 @@
# Brief: rag-eval-hybrid-baseline
## Background
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
## Goals (knife-1 only)
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
- Fixtures record `searchMode` / `kbScope`.
- README documents hybrid-era offline vs live loop.
- Best-effort live seed + regenerate fixtures + update baseline.
## Non-goals
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
@@ -0,0 +1,18 @@
# Decisions: rag-eval-hybrid-baseline(最终版)
## Process
sm-flow standard-lean: Discover → Commit → Apply → Archive.
## Key decisions
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
2. Fixture meta: `searchMode`, `kbScope` when set.
3. Knife-2 (dual fixtures / mustNot golden) deferred.
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
## Trade-offs
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
@@ -0,0 +1,17 @@
# Evidence: rag-eval-hybrid-baseline
## Pre-change
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
- Offline eval already supported Hit levels, recall@K, baseline diff.
## User decisions
- Scope: knife-1 only (no dual fixture dirs).
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
## Apply-discovered
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
@@ -0,0 +1,41 @@
# Acceptance: rag-quality-score-unify
## Tasks
OpenSpec `tasks.md` 全部 `[x]`(1.1–6.2)。
## 静态验证
- 生产路径 grep:无 `bm25_only_no_dense` 发射、无 hybrid L2 enrichment(仅 Labels canonicalize 兼容旧串)。
- 架构文档 §6 与 `application.yml` 注释已对齐 quality 契约。
## 脚本验证
```text
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
```
| 套件 | 结果 |
|---|---|
| RetrievalScoreNormalizerTest | 4 passed |
| KnowledgeEvidencePostProcessorTest | 6 passed |
| LookupKnowledgeToolTest | 7 passed |
| VectorSearchServiceTest | 2 passed |
| VectorKnowledgeSearchAdapterHybridTest | 1 passed |
(PowerShell 可能将 JVM warning 标为 exit 1;日志中为 BUILD SUCCESS / Failures: 0。)
## 浏览器 / 人工验证
- 未跑:live `lookup_knowledge` hybrid vs dense 对照、生产阈值标定。
## 未验证
| 项 | 风险 | 建议 |
|---|---|---|
| 真实 Milvus hybrid 联调 | 序/质量分布与单测 mock 有差 | 启动服务后固定 query 集切 mode 对比 |
| 阈值 0.75/0.5 在 hybrid rank 分下的标定 | retry/PRECISE 偏多或偏少 | 看 trace topSimilarity 再调 yml |
## Specs 同步
- 主规格新增:`openspec/specs/rag-retrieval-quality-score/spec.md`(archive 时从 delta 同步)。
@@ -0,0 +1,21 @@
# Brief: rag-quality-score-unify
## Background
真 BM25 hybrid(dense + BM25 + RRF)已上线,但后处理仍把 hybrid 结果伪装成 L2 做 `normalizeL2`,并用 L0 domain/entity/keyword contains 加分改序。排序权威与质量闸门分裂,词面信号被 BM25 与后处理双重计分。
## Goals
- 一级 `scoreLabel` 仅 `dense` | `hybrid`
- 唯一 `toQualityScore`;后处理 label-agnostic
- 排序主序 = 检索 `originalRank`;去掉关键词 boost 改序
- hybrid quality = 本轮 rank 纯映射(不做 max(rank, denseSim)、不为闸门回填 L2)
- 保留 `mode=dense` 作同库召回对照;线上默认 hybrid
## Scope
内部 RAG:store 发射、normalizer、evidence post-process、单测、架构文档 §6。
## Non-goals
精排 / query rewrite / 邻块、schema rebuild、改 Agent ACI 字段名、删除 dense 对照 mode。
@@ -0,0 +1,25 @@
# Decisions: rag-quality-score-unify(最终版)
## Scale / process
- sm-flow standard:Discover → Commit → Apply → Archive
- Committed OpenSpec:`openspec/changes/rag-quality-score-unify/`(归档后见 archive 目录)
- 废止:`rag-bm25-hybrid-drop-sdk` 中「dense L2 enrichment for threshold compatibility」
## Key decisions
1. **Label**:仅 `dense` | `hybrid`;旧别名 canonicalize。
2. **Normalizer**:唯一 `toQualityScore`;dense=L2 公式;hybrid=rank 线性映射(batchSize)。
3. **Store**:hybrid 不回填 L2、不发 `bm25_only_*`;返回序即 RRF 序。
4. **Post-process**:`originalRank` ASC;L0 重叠只写 hitReasons;relevance/low-quality 只看 qualityScore;PRECISE 不要求 hint support。
5. **Mode**:hybrid 主路径;dense 同库对照(架构 §6.0)。
## Trade-offs
- hybrid quality 为序数分,跨 query 绝对值不可比;阈值可能需后续标定。
- 去掉 boost 改序后,「词面热语义冷」不再被后处理抬升;词面交给 BM25+RRF。
## Risks accepted
- `relevance_level` / unfiltered retry 分布变化(产品已接受)。
- 未做 live E2E / 人工 hybrid 对照评测(见 acceptance 未验证项)。
@@ -0,0 +1,25 @@
# Evidence: rag-quality-score-unify
## Code (pre-change)
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后并行 dense 回填 L2;BM25-only → `bm25_only_no_dense` + maxL2。
- `KnowledgeEvidencePostProcessor`:一律 `normalizeL2(score)` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序;PRECISE 需 `hasHintSupport`。
## User decisions (grill)
| ID | 结论 |
|---|---|
| Q1 | 一级 label 仅 dense/hybrid;bm25_only 不作正式 label |
| Q2 | 后处理去掉 contains 加分改序,保 originalRank |
| Q3 | 唯一 toQualityScore;后处理统一 |
| Q4 | dense mode 保留作对照 |
| Q7 | 接受 relevance_level / retry 分布变化 |
| Q8 | hybrid quality = **纯 rank 映射** |
## Post-change anchors
- `RetrievalScoreLabels` / `RetrievalScoreNormalizer`
- `MilvusHybridKnowledgeStore`(无 L2 overwrite / 无 bm25_only 发射)
- `KnowledgeEvidencePostProcessor`(rank sort + explain-only L0 overlap)
- OpenSpec delta:`rag-retrieval-quality-score`
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
+12
View File
@@ -53,6 +53,18 @@ docs/
1. [分析笔记目录](analysis/) - 代码分析和问题分析 1. [分析笔记目录](analysis/) - 代码分析和问题分析
2. [临时报告目录](reports/) - 修复和验证报告 2. [临时报告目录](reports/) - 修复和验证报告
### RAG 设计讨论(docs 根目录)
1. [RAG 排序:多路召回与 RRF](RAG排序-多路召回与RRF.md) - K、融合、L0 边界
2. [Hybrid 之后的 qualityScore 与后处理](RAG-Hybrid质量分与后处理.md) - 上一代问题、L2 伪装、统一归一化
3. [Agent 如何读 relevance_level](RAG-Agent如何读relevance_level.md) - 粗相关度标签的含义与误读
4. [RAG 离线评测:讨论、设计与落地](RAG离线评测-基线设计.md) - Golden/Fixture、流程图、hybrid 对齐与闸门修复
5. [Milvus hybrid 接入清单](Milvus-Hybrid接入清单.md)
6. [RAG Trace / 审计(架构)](../mvp/architecture/RAG检索可观测性与审计.md) - 请求内 trace、tool_invocation、Trace API
### 诊断全流程(E2E 导读)
1. [一次诊断到底发生了什么](一次诊断全流程-E2E导读.md) - SUCCESS 全流程:阶段拆解、token、timeline、字段词典
2. [RAG 审计补丁 E2E:step_id + query](RAG审计补丁-stepid-query-E2E验收.md) - 审计字段 live 验收(含业务 FALLBACK 样本)
--- ---
## 📚 学习笔记 (learning/) ## 📚 学习笔记 (learning/)
@@ -4,7 +4,7 @@
**前提**:旧 Milvus SDK 直连检索路径后续废弃,不作为长期实现基础 **前提**:旧 Milvus SDK 直连检索路径后续废弃,不作为长期实现基础
**目标**:在现有 `lookup_knowledge` pipeline 上接入 dense + sparse/BM25 混合检索,融合优先走服务端 RRF **目标**:在现有 `lookup_knowledge` pipeline 上接入 dense + sparse/BM25 混合检索,融合优先走服务端 RRF
**关联文档**: **关联文档**:
- `docs/rag-ranking-multipath-retrieval-and-rrf.md`(排序与多路召回判断框架) - `docs/RAG排序-多路召回与RRF.md`(排序与多路召回判断框架)
- 本文后续实现讨论以本节 **「交付拆分:分块去重 + Hybrid 同规划」** 为基线 - 本文后续实现讨论以本节 **「交付拆分:分块去重 + Hybrid 同规划」** 为基线
--- ---
@@ -12,7 +12,7 @@
## 1. 结论先说 ## 1. 结论先说
可以接,而且和前面讨论的多路召回 / RRF 高度一致。 可以接,而且和前面讨论的多路召回 / RRF 高度一致。
但当前项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是: 但(写作当时)项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是:
```text ```text
1. schema 只有 dense,没有 sparse/BM25 字段 1. schema 只有 dense,没有 sparse/BM25 字段
@@ -32,7 +32,27 @@
- 分块去重与 hybrid 同规划、分里程碑交付(先共用地基,再开 hybrid) - 分块去重与 hybrid 同规划、分里程碑交付(先共用地基,再开 hybrid)
``` ```
## 实现状态(2026-07-27) ```mermaid
flowchart TB
subgraph gaps["写作时的缺口"]
G1[无 sparse/BM25 schema]
G2[写入只有 dense]
G3[单路 similaritySearch]
G4[后处理伪融合]
G5[source 级去重吞 chunk]
end
subgraph principles["接入原则"]
P1[端口抽象 · 不堆旧 SDK]
P2[融合下沉向量库 RRF]
P3[应用层:filter/dedup/return-n/投影]
P4[先地基后 hybrid 分里程碑]
end
gaps --> principles
```
## 实现状态(2026-07-27,后续已完成)
| 里程碑 | 状态 | 说明 | | 里程碑 | 状态 | 说明 |
|---|---|---| |---|---|---|
@@ -40,17 +60,42 @@
| 交付 2a 应用层 multi-path+RRF | **已完成并归档** | `2026-07-27-rag-hybrid-search-rrf`(已被 2b 取代为生产路径) | | 交付 2a 应用层 multi-path+RRF | **已完成并归档** | `2026-07-27-rag-hybrid-search-rrf`(已被 2b 取代为生产路径) |
| 交付 2b 真 BM25 hybrid + 废弃 SDK | **已完成并归档** | `2026-07-27-rag-bm25-hybrid-drop-sdk` | | 交付 2b 真 BM25 hybrid + 废弃 SDK | **已完成并归档** | `2026-07-27-rag-bm25-hybrid-drop-sdk` |
```mermaid
flowchart LR
D1[交付1<br/>chunk 身份/去重] --> D2a[交付2a<br/>app RRF]
D2a --> D2b[交付2b<br/>真 BM25 hybrid]
D2b --> NOW[生产:V2 store + hybrid mode]
```
**当前生产知识路径:** **当前生产知识路径:**
```text ```text
VectorIndexService / VectorSearchService VectorIndexService / VectorSearchService
-> MilvusHybridKnowledgeStore (MilvusClientV2 only) -> MilvusHybridKnowledgeStore (MilvusClientV2 only)
collection: milvus.collection (default biz_hybrid) collection: milvus.collection (default biz)
mode: retrieval.search.mode = dense | hybrid mode: retrieval.search.mode = dense | hybrid
hybrid: dense ANN + BM25 sparse ANN + RRFRanker hybrid: dense ANN + BM25 sparse ANN + RRFRanker
``` ```
**运维必做:** 全量重灌知识库到 `biz_hybrid`;旧 `biz` collection 不再被知识路径使用。 ```mermaid
flowchart TB
subgraph write["写入"]
UP[upload / init / rebuild] --> VIS[VectorIndexService]
VIS --> STORE[MilvusHybridKnowledgeStore]
end
subgraph read["检索"]
LK[lookup_knowledge] --> VSS[VectorSearchService]
VSS -->|dense| SD[searchDense]
VSS -->|hybrid| SH[searchHybrid + RRF]
SD --> STORE
SH --> STORE
end
STORE --> COL[(Milvus collection biz<br/>dense + BM25 schema)]
```
**运维必做:** 全量重灌知识库到 hybrid schema collection(配置名以 `milvus.collection` 为准,常见 `biz`);旧纯 dense collection 不能直接当 hybrid 用。
--- ---
@@ -83,6 +128,19 @@ VectorIndexService / VectorSearchService
-> return-n / Agent 投影 -> return-n / Agent 投影
``` ```
```mermaid
flowchart TB
HIT[检索命中] --> ID[候选身份<br/>docId / chunkIndex / evidenceKey]
ID --> DEDUP[chunk 去重 · 每文档上限]
DEDUP --> FUSE[排序融合 · 库内 RRF]
FUSE --> RET[return-n]
RET --> PROJ[Agent 投影]
ID -.->|交付1 地基| D1[chunk identity]
DEDUP -.-> D1
FUSE -.->|交付2| D2[hybrid]
```
若拆开且顺序错误: 若拆开且顺序错误:
| 只做一项 | 后果 | | 只做一项 | 后果 |
@@ -305,6 +363,27 @@ VectorIndexService / VectorSearchService
LookupResult / Projector LookupResult / Projector
``` ```
```mermaid
flowchart TB
subgraph write_port["写入边界"]
W[KnowledgeWritePort<br/>upsert / deleteByDocId]
end
subgraph search_port["检索边界"]
S[KnowledgeSearchPort<br/>mode DENSE / HYBRID]
end
W --> VDB[(Vector DB<br/>dense + BM25 sparse<br/>metadata filter)]
S --> VDB
UP[upload/init] --> W
LK[lookup_knowledge] --> S
S --> RET[DocumentRetriever]
RET --> POST[PostProcess<br/>dedup / cap / threshold / pack]
POST --> PROJ[LookupResult / Projector]
PROJ --> AG[Agent]
```
说明: 说明:
- **Port** 是应用边界,实现可换成 Spring AI、Milvus 新客户端、或其他封装。 - **Port** 是应用边界,实现可换成 Spring AI、Milvus 新客户端、或其他封装。
@@ -385,6 +464,19 @@ category/kb_scope/doc_id -> 标量过滤索引(如需要)
- 稳定 chunk id(doc_id + chunk_index 派生) - 稳定 chunk id(doc_id + chunk_index 派生)
``` ```
```mermaid
flowchart LR
CHUNK[DocumentChunk] --> EMB[dense embed]
CHUNK --> ST[search_text<br/>title+path+content]
CHUNK --> META[docId/chunkIndex<br/>category/kb_scope]
EMB --> ROW[upsert row]
ST --> ROW
META --> ROW
ROW --> FN[BM25 Function<br/>search_text → sparse]
ROW --> COL[(collection)]
FN --> COL
```
### 5.3 注意 ### 5.3 注意
- embedding 文本可以继续拼 `Title/Path/Content` - embedding 文本可以继续拼 `Title/Path/Content`
@@ -429,6 +521,16 @@ SearchHit {
} }
``` ```
```mermaid
flowchart TB
REQ[SearchRequest<br/>query · retrieveK · mode<br/>filter · rrfK] --> PORT[KnowledgeSearchPort]
PORT -->|DENSE| D[dense ANN only]
PORT -->|HYBRID| H[dense + BM25 + RRF]
D --> HIT[SearchHit 列表<br/>id/docId/chunk · content · ranks]
H --> HIT
HIT --> APP[后处理 / 投影]
```
### 6.2 `VectorSearchService` 怎么演进 ### 6.2 `VectorSearchService` 怎么演进
短期: 短期:
+396
View File
@@ -0,0 +1,396 @@
# Agent 如何读 `relevance_level`:它是什么、不是什么
**日期**:2026-07-28
**范围**:`lookup_knowledge` 投影给 Agent 的粗粒度相关度标签
**读者**:要在 Diagnosis Agent / 工具契约里正确使用知识库结果的工程与提示词同学
**关联**:
- ACI 契约:`RagToolResult.relevance_level` / `RagRelevanceLevel`
- 计算:`KnowledgeEvidencePostProcessor` → `qualityScore` 阈值
- 质量统一:`docs/RAG-Hybrid质量分与后处理.md`
- 多路与 RRF:`docs/RAG排序-多路召回与RRF.md`
- **运行时 Trace / 审计**:`mvp/architecture/RAG检索可观测性与审计.md`
---
## 1. 一句话定义
**`relevance_level` 是对「这一次 `lookup_knowledge` 调用整体有多相关」的粗档标签,不是某一条 evidence 的分数,也不是 0~1 的相似度。**
Agent 真正写诊断、做引用时,仍应以 `evidence[]` 里的 excerpt 为准;`relevance_level` 只帮助判断:**这批评据大概有多硬、还要不要再查。**
```mermaid
flowchart LR
subgraph tool["lookup_knowledge 结果"]
ES[evidence_status]
EV[evidence excerpt]
RL[relevance_level]
TR[truncated]
end
ES --> DEC{Agent 决策}
EV --> DEC
RL --> DEC
TR --> DEC
DEC --> A1[写结论 / 引用]
DEC --> A2[再查 / 换工具]
DEC --> A3[证据不足降级]
```
---
## 2. 它出现在哪里
成功(或有结果)的知识库工具投影里,典型形状:
```json
{
"evidence_status": "EVIDENCE_FOUND",
"tool_call_id": "call-…",
"query": "用户/Agent 的检索句",
"evidence": [
{
"document_id": "doc#chunk-0",
"source": "…",
"title": "…",
"breadcrumb": "…",
"excerpt": "…"
}
],
"returned_count": 3,
"relevance_level": "PRECISE",
"truncated": false
}
```
要点:
- JSON 字段名是 **`relevance_level`**(snake_case)
- Java 枚举:`RagRelevanceLevel`(`PRECISE` / `HIGHLY_RELEVANT` / `REFERENCE`)
- 无可用证据或质量不够时,字段常为 **null / 省略**(`NON_NULL`)
- Agent **看不到** raw L2、RRF 分、`retrievalTrace`、`rerankTrace`(ACI 有意裁掉)
```mermaid
flowchart TB
subgraph internal["内部 LookupResult(审计/调试)"]
QS[qualityScore / topSimilarity]
RT[retrievalTrace / rerankTrace]
CP[contextPack 全文]
SC[raw score / scoreLabel]
end
subgraph project["RagResultProjector"]
P[裁剪与规范化]
end
subgraph agent["Agent 可见 RagToolResult"]
A1[evidence_status]
A2[tool_call_id / query]
A3[evidence excerpt 列表]
A4[relevance_level 可选]
A5[returned_count / truncated]
end
internal --> P --> agent
```
---
## 3. 枚举值怎么理解
| 值 | 产品语义 | Agent 侧更合理的用法 |
|----|----------|----------------------|
| **PRECISE** | 整体很贴:当前 top 证据质量高,继续同题检索不太可能更准 | 优先依据 `evidence[]` 组织结论;避免无意义的重复 `lookup_knowledge` |
| **HIGHLY_RELEVANT** | 高度相关(枚举保留) | 与 PRECISE 类似,略保守表述即可 |
| **REFERENCE** | 可作参考,但不到「已精准命中」 | 可引用,但结论留余地;缺维度时换 query 再查或叠日志/指标 |
| **null / 不出现** | 没有可报的粗相关档(无证据或 top 质量偏低) | **不要**当成知识库已证实;按证据不足处理 |
### 和 `evidence_status` 的分工
| 字段 | 回答的问题 |
|------|------------|
| `evidence_status` | 这次有没有合法、可引用的证据(如 `EVIDENCE_FOUND` / `NO_EVIDENCE`) |
| `relevance_level` | **有证据时**,整体有多贴(粗档) |
| `evidence[]` | 具体可以引用哪些片段 |
没有证据时,不应指望靠 `relevance_level`「升级」出结论;契约上也不会用 level 把空结果扮成有证据。
```mermaid
flowchart TB
ES{evidence_status}
ES -->|NO_EVIDENCE| N1[不要当知识库已证实]
ES -->|EVIDENCE_FOUND| RL{relevance_level}
RL -->|PRECISE| U1[优先引用 excerpt · 少重复检索]
RL -->|REFERENCE| U2[可引用 · 结论留余地]
RL -->|null / 缺省| U3[有块但质量偏低 · 慎用强结论]
RL -->|HIGHLY_RELEVANT| U4[与 PRECISE 类似 · 略保守]
```
---
## 4. 它是怎么算出来的(实现口径)
### 4.1 只看「本轮第一名」的 qualityScore
后处理在完成排序、去重、截断之后:
```text
取 originalRank 最优(排序后第一条)的 qualityScore
≥ highly-relevant-threshold (默认 0.75) → PRECISE
≥ reference-threshold (默认 0.5) → REFERENCE
否则 / 无可用证据 → null
```
配置(`application.yml`):
```yaml
retrieval:
normalization:
max-l2-distance: 2.0
highly-relevant-threshold: 0.75
reference-threshold: 0.5
```
因此:
- level 描述的是 **整次调用的 top 质量**,不是每条 evidence 各打一档
- 列表里第 2、第 3 条即使偏弱,只要 top1 够高,整次仍可能是 `PRECISE`
```mermaid
flowchart TB
RET[检索有序候选] --> POST[后处理:保 originalRank · 去重 · 截断]
POST --> TOP[取排序后第一条 qualityScore]
TOP --> T1{≥ 0.75?}
T1 -->|是| PRECISE[PRECISE]
T1 -->|否| T2{≥ 0.5?}
T2 -->|是| REF[REFERENCE]
T2 -->|否| NULL[null / 不报档]
```
### 4.2 qualityScore 从哪来(和检索 mode 绑定)
统一经 `RetrievalScoreNormalizer.toQualityScore`:
| `retrieval.search.mode` | top qualityScore 含义 |
|-------------------------|------------------------|
| **dense** | top1 的 L2 归一化:约 `1 - L2 / maxL2Distance` |
| **hybrid** | **优先**同 id 的 `denseDistance`(绝对 L2 质量,供闸门/level);无 dense 邻域时 **rank 回退** |
```mermaid
flowchart LR
subgraph dense_mode["mode=dense"]
L2[score = L2] --> QD[quality = 1 - L2/max]
end
subgraph hybrid_mode["mode=hybrid"]
RRF[RRF 序 = originalRank] --> SORT[列表顺序]
DD[denseDistance 可选] --> QH{有 dense?}
QH -->|是| QL[quality = L2 归一化]
QH -->|否| QR[quality = rank 映射]
end
QD --> LV[relevance_level]
QL --> LV
QR --> LV
```
读 level 时注意:
> **dense 下的 PRECISE ≈「向量足够近」**
> **hybrid 下的 PRECISE ≈「top1 的绝对/回退 quality 跨过了 0.75」**;排序仍跟 RRF,不是「又变回只信 L2 排序」。
不要把 hybrid 的 level 读成与 dense **完全同一把尺子**,但也不要再假设「rank1 永远 PRECISE」(在附带 denseDistance 后,远邻 decoy 可以很低分并触发 filter fallback)。
### 4.3 关于 HIGHLY_RELEVANT
枚举和旧文档里仍有三档。历史上大致是:
```text
高分 + L0 hint 支撑 → PRECISE
高分但无 hint → HIGHLY_RELEVANT
中等分 → REFERENCE
```
质量分统一之后,当前实现是:**≥ 0.75 直接 PRECISE**,不再要求 L0 contains 才能精准。
因此运行时 **很少再单独产出 HIGHLY_RELEVANT**;读旧 trace / 旧快照时仍可能见到。
---
## 5. Agent 应该怎么读(建议协议)
### 5.1 推荐读法
```text
1. 先看 evidence_status
2. 再读 evidence[] 的 excerpt(唯一可引用正文)
3. 用 relevance_level 调节「敢多敢少」与「要不要再查」
```
```mermaid
flowchart TB
START[收到 RagToolResult] --> S1{evidence_status}
S1 -->|无证据| FAIL[不编造 · 换工具或安全降级]
S1 -->|有证据| S2[精读 evidence excerpt]
S2 --> S3{relevance_level}
S3 -->|PRECISE| C1[结论可较硬 · 少重复同 query 检索]
S3 -->|REFERENCE| C2[结论留余地 · 可换问法或叠日志指标]
S3 -->|缺省| C3[慎用强结论 · 优先补查]
C1 --> CITE[引用必须落在 excerpt]
C2 --> CITE
C3 --> CITE
```
| 组合 | 建议行为 |
|------|----------|
| FOUND + PRECISE | 以 excerpt 为主写结论;少重复同 query 检索 |
| FOUND + REFERENCE | 可引用,表述保守;缺关键事实则改写 query 或换工具 |
| FOUND 但 level 空 | 有块但质量闸门偏低:慎用强结论,优先补查 |
| NO_EVIDENCE | 不编造知识库依据;走其他证据工具或安全降级 |
### 5.2 明确不要这样读
1. **不要当逐条相关度**
没有 `evidence[i].relevance_level`;不能说「第 2 条是 REFERENCE」。
2. **不要当连续分数**
没有 0.83;只有粗档。不要在推理里假装有精确分。
3. **不要在 hybrid 下当成绝对语义相似度**
序数 quality 下 PRECISE 很常见,表示「本轮第一够格」,不等于「全局语义必近」。
4. **不要代替 excerpt 引用**
level 不能当证据正文;Gatekeeper / EvidenceGuard 认的是可核对片段与引用约束。
5. **不要和 Harness 验真混为一谈**
level 是检索侧粗标;工具生命周期、`evidence_status`、守卫校验是另一层。
6. **不要用它驱动跨 mode 对比**
同一 query 切 dense/hybrid 时,比命中集合与排名;别只比「是不是都 PRECISE」。
---
## 6. Agent 看不见、但会影响 level 的内部量
便于排查「为什么突然全是 PRECISE / 总是 null」:
| 内部量 | 作用 | Agent 是否可见 |
|--------|------|----------------|
| `qualityScore` / `topSimilarity` | 定 level、低质 retry | 否 |
| `originalRank` | 排序权威;hybrid quality 输入 | 否 |
| `score` + `scoreLabel` | dense=L2 / hybrid=融合侧 | 否 |
| L0 domain/keyword | 现仅 hitReasons 解释,不改序、不抬 level | 否(reasons 也可能被投影裁掉) |
| `completenessHint` | 内部完整度文案 | 通常否 |
| category filter + unfiltered retry | 低质时可能换一批 evidence 再定 level | 过程 trace 否 |
投影原则(ACI):模型只要能理解与引用结果;**不给 raw score、阈值、trace。**
```mermaid
flowchart TB
subgraph pipe["检索管道内部"]
STORE[Milvus hybrid/dense]
NORM[toQualityScore]
POST[PostProcessor]
STORE --> NORM --> POST
end
POST --> LV[relevance_level]
POST --> EB[evidenceBlocks]
POST --> TR[traces · 通常不投影]
EB --> PROJ[RagResultProjector]
LV --> PROJ
PROJ --> AGENT[Agent Observation]
```
---
## 7. 和相邻概念的边界
```text
evidence_status 有没有证据
relevance_level 有的话有多贴(粗)
evidence[].excerpt 贴在哪一段文字上(细、可引用)
truncated 列表是否被预算截断(可能还有更好的没展示)
information_gain 等 诊断环路里「这轮工具对任务有没有增益」(另一契约)
```
```mermaid
flowchart TB
subgraph fields["同一次工具结果里的分工"]
ES[evidence_status<br/>有没有]
RL[relevance_level<br/>有多贴·粗]
EX[excerpt<br/>说什么·细]
TC[truncated<br/>是否被截断]
end
ES --> USE[Agent 使用]
RL --> USE
EX --> USE
TC --> USE
USE --> NOTE[结论锚在 excerpt<br/>level 只调力度]
```
`truncated=true` 时:即使 `PRECISE`,也只说明 **已返回子集里的 top 很强**,不保证库内没有更相关却被截掉的块。
---
## 8. 提示词 / 产品文案可用的短说明
可直接给模型或文档的精简版:
```text
relevance_level 是本次知识库检索的整体相关度粗标:
- PRECISE:当前证据整体很贴,优先引用 evidence 写结论,避免无意义重复检索
- REFERENCE:仅供参考,结论需留余地,必要时换问法或改用其他工具
- 缺省:不要把本次结果当作高置信知识库证实
务必以 evidence 中的 excerpt 为唯一引用依据;不要编造未出现的文档内容。
在 hybrid 检索下,PRECISE 更多表示「本轮排序第一档」,不是精确相似度分数。
```
---
## 9. 常见误读示例
| 误读 | 更正 |
|------|------|
| 「PRECISE 所以三条 evidence 都精准」 | 只保证 top 质量跨线;其余条只是同批返回 |
| 「没有 relevance_level 就是工具失败」 | 更可能是无证据或质量偏低;看 `evidence_status` |
| 「hybrid 全是 PRECISE 说明召回完美」 | 可能只是 rank1→quality=1.0 的档位特性 |
| 「REFERENCE 的 excerpt 不能引用」 | 可以引用,但结论强度应下调 |
| 「level 高就可以跳过 excerpt」 | 不可;引用与验真仍看正文 |
---
## 10. 结语
`relevance_level` 是检索链路送给 Agent 的 **粗粒度驾驶辅助**:
- 告诉模型这批评据大概硬不硬
- **不**替代 excerpt,**不**暴露打分细节,**不**等于逐条标注
在 dense 模式下,它更接近「向量有多近」;
在 hybrid 主路径下,它更接近「本轮融合第一名是否跨过质量门槛」。
读的时候记住三句即可:
```text
1. 先 status,再 excerpt,最后才看 level
2. level 管「敢多敢少」,excerpt 管「说了什么」
3. hybrid 的 PRECISE ≠ 绝对语义满分
```
---
## 附录:代码与规格锚点
| 项 | 位置 |
|----|------|
| Agent 结果契约 | `RagToolResult` / `RagRelevanceLevel` |
| 投影 | `RagResultProjector` |
| 等级计算 | `KnowledgeEvidencePostProcessor.computeRelevance` |
| 质量分 | `RetrievalScoreNormalizer` |
| 规格 | `openspec/specs/aci-evidence-tool-contracts`、`rag-retrieval-quality-score` |
| 架构 | `mvp/architecture/RAG知识检索架构.md` §9 |
+592
View File
@@ -0,0 +1,592 @@
# 混合检索上线之后:为什么还要统一 qualityScore,以及上一代后处理错在哪
**日期**:2026-07-28
**范围**:hybrid 检索后的分数语义、后处理排序、相关度闸门、scoreLabel 约定
**读者**:已经(或准备)上 dense+BM25+RRF,却发现「召回变了、质量判断还拧着」的工程同学
**关联实现**:
- `lookup_knowledge` 模块化链路
- `MilvusHybridKnowledgeStore`(dense / hybrid)
- `RetrievalScoreNormalizer` / `KnowledgeEvidencePostProcessor`
- OpenSpec / devflow:`rag-quality-score-unify`
- 前置讨论:`docs/RAG排序-多路召回与RRF.md`
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
---
## 1. 引言:融合排好了序,不等于质量链路闭环了
上一篇文章(《诊断 Agent 场景下的 RAG 排序:从 K=3 规则加分,到多路召回与 RRF》)回答的是:
> 候选太少时别急着上精排;跨路不要硬加原始分;优先 RRF;L0 只做导航。
那一轮讨论之后,工程上陆续落地了:
1. **chunk 级证据身份与去重**(`docId#chunkIndex`,同文档多片段可并存)
2. **真 hybrid**:Milvus 服务端 dense ANN + BM25 sparse + `hybridSearch` + `RRFRanker`
3. **单一知识后端**(`MilvusClientV2`),去掉 sdk/spring 多路由主路径
4. **mode 开关**:`hybrid` 线上主路径,`dense` 同库对照评测
主缺口从「假 hybrid / 粗去重」变成了另一件事:
```text
库内:RRF 已经决定谁先谁后
应用:后处理仍假装每条 score 都是 L2
再用 L0 关键词 contains 加分改序
```
于是出现一种很拧的现象:
- 检索层已经是 **混合检索的世界**
- 质量层还活在 **单路 dense + 规则 boost 的世界**
本文记录的,就是这次对「拧」的拆解、拍板与落地口径:
**统一 qualityScore,废止 hybrid 的 L2 伪装,去掉关键词 boost 改序。**
目标不是再推一套更复杂的模型,而是回答:
> hybrid 上线之后,排序权威和质量闸门到底听谁的?后处理还该不该拿关键词打分?
```mermaid
flowchart TB
subgraph retrieval["检索层 · 已 hybrid"]
Q[query] --> D[dense ANN]
Q --> B[BM25 sparse]
D --> RRF[hybridSearch + RRF]
B --> RRF
RRF --> ORD[RRF 序]
end
subgraph post_old["后处理 · 仍 L2 世界"]
ORD --> FAKE[伪装 / 回填 L2]
FAKE --> BOOST[关键词 boost 改序]
BOOST --> GATE[阈值 / level / retry]
end
post_old --> PAIN[排序与闸门拧巴]
```
---
## 2. 上一代(hybrid 刚落地时)到底长什么样
### 2.1 检索侧:已经是真混合
`mode=hybrid` 时大致是:
```text
query
├─ dense ANN(query embedding) → vector / L2
└─ BM25 sparse(EmbeddedText) → sparse_vector / BM25
│
▼
Milvus hybridSearch + RRFRanker(k)
│
▼
融合后的 hit 列表(RRF 序)
```
```mermaid
flowchart LR
Q[query] --> EMB[embedding]
Q --> TXT[raw text]
EMB --> DA[dense ANN<br/>vector / L2]
TXT --> BA[BM25 ANN<br/>sparse]
DA --> HS[Milvus hybridSearch]
BA --> HS
HS --> RR[RRFRanker]
RR --> HITS[有序 hits]
```
这比应用层 sparse-lite / 伪 hybrid 前进了一大步:词面与语义在**库内**融合,chunk 身份也不会在后处理被文档级折叠吞掉。
### 2.2 分数侧:仍在「骗」后处理
后处理历史契约默认:
```text
score ≈ L2 距离(越小越好)
baseScore = 1 - clamp(L2) / maxL2Distance # 越大越好
再 + domain/entity/keyword boost
按 finalScore 重排
用 baseScore 定 relevance_level / 是否低质 retry
```
为了迁就这套契约,hybrid 路径做了补丁:
```text
hybrid 融合结果
-> 再跑一路 dense
-> 按 id 把 L2 回填到 score,label 改成 l2_distance
-> 仅 BM25 命中、dense 没命中:
score = maxL2Distance
label = bm25_only_no_dense
```
```mermaid
flowchart TB
H[hybrid RRF 结果] --> P[并行 dense 探测]
P --> M{同 id 有 L2?}
M -->|是| O1[score=L2 · label=l2_distance]
M -->|否| O2[score=maxL2 · label=bm25_only]
O1 --> N[normalizeL2 + boost 重排]
O2 --> N
N --> BAD[BM25-only 好证据被当成最差]
```
意图是好的:让 `normalizeL2` 和 0.75/0.5 阈值「还能用」。
副作用也很清楚:
| 现象 | 后果 |
|------|------|
| RRF 决定顺序,L2 决定「好不好」 | 两套真理,互相打架 |
| BM25-only 好证据被标成最远 L2 | quality≈0,像低质,甚至触发 unfiltered retry |
| `bm25_only_no_dense` 像第三种 label | 概念膨胀:mode 其实只有 dense/hybrid |
| 多打一路 dense 只为回填 | 延迟与复杂度,换来的是语义自洽的假象 |
一句话:
> **不是拿 RRF 分错误地套了 L2 公式,而是排序信 RRF,打分/闸门仍假装大家都是 L2。**
### 2.3 后处理侧:关键词 boost 改主序
典型逻辑:
```text
baseScore = normalizeL2(score)
finalScore = baseScore
+ domain_match (+0.15)
+ entity_match (+0.20)
+ keyword_match (+0.10)
+ source_type (+0.05)
按 finalScore 降序
```
在 **还没有库内 BM25** 时,这套东西多少能补一点词面。
在 **已经 hybrid** 之后,问题变成:
1. **双重计分**
BM25 已经在 RRF 里投过票;后处理再用 contains 加分,等于词面再抬一次。
2. **contains 比 BM25 更糙**
无 IDF、无文档长度、短词子串误命中——正好制造「词频/词面高、相关度低却排前面」。
3. **冲掉 RRF 序**
花了 hybrid 买到的融合序,被 L0 词表二次改写。
4. **PRECISE 还绑 hint**
高质量还要 `hasHintSupport`(同样是 contains),把导航层信号抬成等级门槛。
结合上一篇文章的判断——**L0 / 关键词适合做提示,不适合当最终裁判**——hybrid 上线后,后处理 boost 改序已经从「可接受的轻启发式」滑向「明确的设计债」。
---
## 3. 问题清单:chunk 去重 + hybrid 之后,还剩什么
可以分成四层(本次主要收口前两层):
### 3.1 正确性 / 契约(本次主战场)
1. hybrid **没有**独立的融合分归一化,只有 L2 兼容补丁
2. 后处理关键词打分不合理,会抬升词面热、语义冷的片段
3. `scoreLabel` 语义混乱:`l2_distance` / `rrf_fused` / `bm25_only_*` 混用
4. `mode=dense` 与 hybrid 内部 dense 子路概念易混(mode 是整次查询算法,不是「第三套库」)
### 3.2 质量上限(未在本 change 做完)
- 无固定 RAG 评测报表驱动阈值标定
- 无邻块扩展、真 query rewrite、cross-encoder 精排
- 中文 analyzer / 分词策略未产品化钉死
### 3.3 工程债(部分清理、部分保留)
- 应用层 `LexicalRanker` / 自研 `RrfFusion` 可能仍像「还有 app-layer hybrid」
- Spring AI starter 仍可作 sidecar,但不是知识主路径(starter 至 2.0.0 仍无 BM25 hybrid)
- 写入先删后插,非强 upsert;Milvus / MySQL / L0 三方一致性靠流程
### 3.4 运维边界
- hybrid 依赖 BM25 Function + sparse index
- 全量 rebuild 受 embedding 与写入延迟约束
- `totalVectors` 一类统计可能仍不可信
本次 change(`rag-quality-score-unify`)**有意只收口 3.1**:
让 hybrid 的排序权威与质量闸门重新对齐,而不是同时上精排模型。
---
## 4. 关键澄清:L2 是什么,它是不是「后处理」本身
讨论中容易把「L2」和「后处理打分流程」混成一个词。需要拆开:
### 4.1 L2 是度量
**L2 = 欧氏距离**,dense ANN 常用 metric:
- 越小越相似
- 单位向量场景下可用 `maxL2Distance≈2` 做上界
- 归一化相似度:`1 - clamp(L2) / maxL2`
### 4.2 后处理是流水线
后处理消费的是「约定好的 score」,历史上**假定**它是 L2,于是:
```text
score(L2) → normalizeL2 → baseScore → (+boost) → finalScore → 排序/等级
```
```mermaid
flowchart LR
subgraph metric["度量层"]
L2[L2 距离]
end
subgraph pipe["后处理流水线"]
N[normalize]
B[可选 boost]
S[排序 / 截断]
G[等级 / 闸门]
N --> B --> S --> G
end
L2 -.->|历史上假定输入是 L2| N
RRF[RRF 融合分] -.->|量纲不同 · 不能直接套| N
```
所以:
- L2 ≠ 后处理
- L2 = dense 路径的自然距离
- 后处理 = 把某种 score 变成 quality / 等级 / 截断结果的流程
hybrid 的问题是:**流程还在,输入契约已经不再总是 L2。**
### 4.3 category 降级也不是 L2 存在的唯一理由
filtered → unfiltered retry 用的是:
```text
isLowQuality = 无可用证据 或 topSimilarity < referenceThreshold
```
`topSimilarity` 来自归一化后的质量分。
unfiltered 只是**再检一次**,尺子本来就该是统一 quality,而不是「专为降级准备的 L2」。
```mermaid
flowchart TB
F[带 category 的检索] --> Q{isLowQuality?}
Q -->|是| U[unfiltered retry]
Q -->|否| K[采用本次结果]
U --> M[合并/替换为 retry 结果]
K --> OUT[后处理出口]
M --> OUT
```
---
## 5. 设计拍板:统一成什么
### 5.1 两层概念,不要混
| 层 | 只有什么 | 不是什么 |
|----|----------|----------|
| **检索 mode** | `dense` \| `hybrid` | 不是三套库 |
| **一级 scoreLabel** | `dense` \| `hybrid` | 不是 `bm25_only` 第三种模式 |
- `mode`:整次查询怎么跑(配置 `retrieval.search.mode`)
- `scoreLabel`:这条 hit 的 `score` 怎么解释
```mermaid
flowchart TB
CFG[retrieval.search.mode] --> M1[dense 整次只跑 ANN]
CFG --> M2[hybrid 整次 dense+BM25+RRF]
M1 --> L1[scoreLabel=dense]
M2 --> L2[scoreLabel=hybrid]
L2 -.->|不是| L3[bm25_only 第三种 mode]
```
`bm25_only_no_dense` **不是第三种检索**,只是旧链路里「这条 hybrid 命中没有 dense L2 可回填」的补丁标签。统一后应降级为历史别名(canonicalize → `hybrid`),不再一级发射。
### 5.2 检索回来带什么
建议最小契约:
```text
rank (originalRank) // 1 最好;hybrid = RRF 序;dense = ANN 序
score // 引擎主分;量纲由 label 解释
scoreLabel // dense | hybrid
rawScore? // 可选调试
denseDistance? // hybrid 可选:同 id 的 L2,仅供闸门
```
| label | score 含义 |
|-------|------------|
| `dense` | L2 距离(越小越好) |
| `hybrid` | 引擎融合分可放 raw/score;**排序不看其量纲** |
```mermaid
flowchart LR
subgraph emit["Store 发射"]
LAB[scoreLabel<br/>dense | hybrid]
SCR[score / rawScore]
RNK[列表序 → originalRank]
DD[denseDistance? 仅 hybrid]
end
LAB --> NORM
SCR --> NORM
RNK --> SORT
DD --> NORM
NORM[toQualityScore] --> QS[qualityScore]
SORT[按 rank 排序] --> LIST[evidence 顺序]
QS --> GATE[level / isLowQuality]
```
### 5.3 唯一归一化点
```text
qualityScore = toQualityScore(label, score, rank, batchSize, maxL2, denseDistance?)
// 输出统一:[0,1],越大越好
```
分支只允许出现在这里:
```text
dense → 1 - clamp(L2)/maxL2
hybrid → 优先 denseDistance 的 L2 归一化(绝对质量 / 闸门)
无 dense 时 rank 线性回退
```
```mermaid
flowchart TB
IN[label + score + rank + denseDistance?] --> C{canonicalize label}
C -->|dense| L2[l2ToQuality score]
C -->|hybrid| H{denseDistance?}
H -->|有| L2H[l2ToQuality denseDistance]
H -->|无| RK[rankToQuality]
L2 --> OUT[qualityScore 0..1]
L2H --> OUT
RK --> OUT
```
**演进说明:** 切片 1 曾用纯 rank 做 hybrid quality;eval 发现 rank1 恒高会杀死 L0 filter fallback。
现行约定:**排序仍纯 RRF;闸门可用 denseDistance 绝对质量**,且 **不得** 再把主分/label 伪装成 L2。
### 5.4 后处理:统一流程,不要按 label 再分叉业务
```text
candidates
→ 每条 toQualityScore(...) ← 唯一认 label 的地方
→ qualityScore + originalRank
→ 统一:按 rank 排序 / 去重 / 每文档 chunk 上限 / return-n
→ 统一:relevance_level、isLowQuality(只看 qualityScore)
```
```mermaid
flowchart TB
CAND[candidates] --> QS[toQualityScore 每条]
QS --> SORT[sort by originalRank ASC]
SORT --> DEDUP[evidenceKey 去重]
DEDUP --> CAP[max-chunks / return-n]
CAP --> REL[relevance_level]
CAP --> LOW[isLowQuality → filter retry]
CAP --> OUT[EvidenceBlocks]
```
可以记成:
> **Label 只活在进后处理之前的适配器里;后处理是 label-agnostic 的。**
> **排序听 rank;闸门听 qualityScore(hybrid 可含 denseDistance)。**
### 5.5 后处理还改不改?——要改,而且和归一化同一刀
后处理合理职责是 **裁剪与装配**,不是第二套检索:
| 保留 | 去掉或降级 |
|------|------------|
| evidenceKey 去重 | domain/entity/keyword **加分改序** |
| max-chunks-per-document | contains 当相关度代理 |
| return-n | PRECISE 强制 hint support |
| excerpt 截断、EvidenceBlock | |
| 统一 qualityScore 闸门 | |
L0 仍可:
- 导航:category filter(失败 unfiltered retry)
- 解释:`hitReasons` 记 `l0_keyword_overlap` 等(**零分值**)
词面该不该高:交给 **BM25 子路 + RRF**。
语义该不该近:交给 **dense 子路**(融合时已参与;dense-only mode 对照时单独看)。
### 5.6 mode=dense 还要不要
要,但定位清楚:
| 模式 | 定位 |
|------|------|
| **hybrid** | 线上主路径 / 默认 |
| **dense** | 同库对照、评测、排障——看「去掉 BM25+RRF 后差在哪」 |
注意:
- hybrid **入库**数据完全适用于 dense 查询(每条都写了 `vector`)
- hybrid **内部**仍有 dense 子路——那是融合的一部分,≠ `mode=dense`
- 对照时固定 `retrieve-k` / `return-n` / filter / query 集,只切 mode
- 优先比命中集合与排名;`relevance_level` 在 hybrid 下是序数 quality,慎作跨 mode 绝对值对比
---
## 6. 目标数据流(落地后)
```text
VectorSearchService (mode=dense|hybrid)
-> hits{ originalRank, score, scoreLabel=dense|hybrid, rawScore?, denseDistance? }
-> KnowledgeDocumentRetriever / SearchPort
-> KnowledgeEvidencePostProcessor
qualityScore = RetrievalScoreNormalizer.toQualityScore(...)
sort by originalRank ASC
evidenceKey dedup / max-chunks / return-n
relevance_level & topSimilarity from qualityScore
L0 overlap → hitReasons only
-> ContextPack / Assembler / Projector
```
```mermaid
flowchart TB
Q[query] --> VSS[VectorSearchService]
VSS -->|mode=dense| SD[searchDense]
VSS -->|mode=hybrid| SH[searchHybrid + 可选 denseDistance]
SD --> PORT[KnowledgeSearchPort / Adapter]
SH --> PORT
PORT --> POST[KnowledgeEvidencePostProcessor]
POST --> PACK[ContextPacker]
POST --> ASM[LookupResultAssembler]
ASM --> PROJ[RagResultProjector]
PROJ --> AGENT[Agent 可见契约]
```
与上一代对比:
| 环节 | 上一代 | 现在 |
|------|--------|------|
| hybrid score | 常被 L2 覆盖 | 保持融合侧;label=`hybrid` |
| BM25-only | maxL2 + `bm25_only_*` | 普通 hybrid hit,quality 看 rank |
| 归一化 | 一律当 L2 | 按 label 唯一转换 |
| 排序 | finalScore(含 boost) | originalRank |
| L0 关键词 | +分改序 | 仅解释 |
| 质量闸门 | baseScore(L2 兼容) | qualityScore |
---
## 7. 行为变化:必须说清楚的协议调整
这是**有意的行为变化**(对内检索质量语义;Agent ACI 字段名可不变):
1. hybrid 下证据顺序更贴近 **RRF**,不再被 contains 抬到前面
2. 「词面很准、dense 略远」的命中,不再被默认打成低质占位
3. `relevance_level` / category unfiltered retry 的触发分布可能变化
4. hybrid 的 quality 是**本轮序数分**,跨 query 绝对值不可比;阈值可能需后续标定
5. dense 对照模式:质量仍走 L2 归一化,行为更接近旧 dense 主路径
未改:
- Agent 可见字段结构(evidence 列表、relevance 枚举名等)
- Milvus hybrid schema / 不必为本次 rebuild
- `mode=dense` 开关本身
---
## 8. 和上一篇文章的衔接:阶段进度
对照 `RAG排序-多路召回与RRF.md` 的推进顺序:
| 阶段 | 内容 | 状态(截至 2026-07-28) |
|------|------|-------------------------|
| Phase 0 | chunk 去重、retrieve-k/return-n、身份 | **已落地** |
| Phase 1~2 | 多路 + RRF;真 BM25 hybrid | **已落地**(库内 hybrid,非 app-layer 伪融合) |
| 分数职责分离 | 排序 vs 可用性/质量闸门 | **本次收口**(qualityScore 统一) |
| 去掉 L0 当裁判 | 关键词不改主序 | **本次收口** |
| Phase 3 | 可插拔模型 Rerank | **未做**(候选池与评测闭环仍优先) |
| 邻块 / query rewrite | 上下文与问句改写 | **未做** |
因此,本次文章不是推翻上一篇,而是补上上一篇写到「融合之后」却还没写完的半截:
> 融合解决「谁进来、谁先排」;
> 归一化与后处理决定「算不算够好、会不会被规则再次打乱」。
---
## 9. 实现锚点(便于对照代码)
| 组件 | 职责 |
|------|------|
| `RetrievalScoreLabels` | `dense` / `hybrid` + 旧别名 canonicalize |
| `RetrievalScoreNormalizer` | 唯一 `toQualityScore` |
| `MilvusHybridKnowledgeStore` | 发射 label;hybrid 不 L2 覆盖 |
| `VectorSearchService` | mode 路由;SearchResult 契约注释 |
| `KnowledgeEvidencePostProcessor` | rank 保序、去 boost 改序、quality 闸门 |
| 架构 §6.0 | mode 用途:hybrid 主路径 / dense 对照 |
验证(单测,非 live E2E):
- normalizer:L2 边界、rank 单调、别名
- post-process:rank 不被 keyword 打乱;caps/return-n
- lookup tool:保序;context pack 元数据仍在
已知未验证:真实 Milvus 联调对照、阈值标定。
---
## 10. 实践清单:以后别再踩的坑
1. **不要**为了复用旧 `normalizeL2`,把 hybrid 结果伪装成 L2。
2. **不要**在已经 BM25 hybrid 之后,再用 L0 contains 大额加分改主序。
3. **不要**把 `bm25_only` 当成第三种检索模式。
4. **不要**把 hybrid 内部的 dense 子路,和 `mode=dense` 整次查询混为一谈。
5. **要**让 label 差异停在适配器;后处理只认 qualityScore + rank。
6. **要**用 dense mode 做召回对照,而不是第二套长期线上策略。
7. **要**接受:hybrid 序数 quality 与绝对阈值之间,需要观测后再调,而不是再发明一层伪装。
8. **下一步再考虑**精排模型——在契约掰直、有固定 query 回归集之后。
---
## 11. 结语
混合检索落地,解决的是「漏」和「跨路硬加分」里很大一块。
但若后处理仍活在 L2 + 关键词 boost 的旧世界,hybrid 买到的 RRF 序和质量信号会被悄悄改写,甚至惩罚「只在 BM25 路很强」的好证据。
这次收口的核心就三句:
```text
1. 一级 label 只有 dense / hybrid
2. 唯一 toQualityScore;后处理统一、保 rank
3. L0 关键词可以解释,不可以再当排序裁判
```
它不是 RAG 的终点,而是 hybrid 从「能跑」变成「分数语义自洽」的必要一步。
在此之后,评测闭环、阈值标定、邻块与精排,才有干净的基线可谈。
---
## 附录 A:术语
| 术语 | 含义 |
|------|------|
| L2 | 欧氏距离;dense ANN 常用;越小越相似 |
| RRF | Reciprocal Rank Fusion;用名次融合多路,不融合原始分 |
| scoreLabel | 一级分数语义:`dense` \| `hybrid` |
| qualityScore | 归一化后的 0~1 质量分(越大越好),供等级与低质闸门 |
| originalRank | 检索返回名次;后处理排序权威 |
| mode=dense | 整次只跑 dense ANN(对照) |
| mode=hybrid | dense+BM25+RRF(主路径) |
| L0 | query hint / 可选 category filter;不作事实证据、不改主序 |
## 附录 B:相关材料
| 材料 | 路径 |
|------|------|
| 多路与 RRF 讨论 | `docs/RAG排序-多路召回与RRF.md` |
| 当前架构 | `mvp/architecture/RAG知识检索架构.md` |
| OpenSpec 归档 | `openspec/changes/archive/2026-07-28-rag-quality-score-unify/` |
| 主规格 | `openspec/specs/rag-retrieval-quality-score/spec.md` |
| devflow | `devflow/projects/2026-07-28-rag-quality-score-unify/` |
@@ -0,0 +1,163 @@
# RAG 审计补丁 E2E:`step_id` + query(含业务 FALLBACK 样本)
**日期**:2026-07-28
**状态**:审计字段 live 验收记录
**关联主文档**:[一次诊断到底发生了什么(SUCCESS 全流程)](一次诊断全流程-E2E导读.md)
> 主文档只保留 **SUCCESS 完整诊断** 与 **现行审计能力说明**。
> 本页单独记录:改造后的一次 live 验收——**审计字段 PASS**,业务因 Milvus 空结果走了 **FALLBACK**。
---
## 1. 样本身份
| 项 | 值 |
|----|----|
| `session_id` | `e2e-audit-20260728171858` |
| `run_id` | `0be605f6-e036-40d3-b364-11b73672241f` |
| 入口 | `POST /api/chat`(与 SUCCESS 样例同构的 RAG-only 约束题) |
| 冷启动 | 含 `AgentStepAuditTracker` 等补丁后的 `mvn spring-boot:run` |
| 业务结局 | `release_outcome=FALLBACK`,`content_type=SAFE_FALLBACK` |
| 工具 | 1× `lookup_knowledge` |
本地产物(若仍在):`target/e2e-audit-sse.txt`、`target/e2e-audit-session.txt`、`target/e2e-audit-run.txt`。
---
## 2. 验收目标 vs 非目标
| 目标 | 是否本页重点 |
|------|----------------|
| `tool_invocation.step_id` = 发出 tool_call 的 `agent_step.id` | **是** |
| `input_params` 含安全 `query` 预览 | **是** |
| Trace / timeline 带回 `step_id` | **是** |
| 业务必须 SUCCESS | **否**(本 run 为 FALLBACK,归因环境) |
---
## 3. 审计结果:PASS
### 3.1 `agent_step`
| id | step_index | has_tool_call | 说明 |
|----|------------|---------------|------|
| **962** | 0 | 1 | 发出 `lookup_knowledge` |
| 963 | 1 | 0 | 无证据后的收尾轮 |
### 3.2 `tool_invocation`(id=860)
| 字段 | 值 |
|------|-----|
| `step_id` | **962**(= step0) |
| `tool_name` | `lookup_knowledge` |
| `success` | 1(工具跑完;无证据也算执行成功) |
| `search_mode` | hybrid |
| `evidence_status` | `NO_EVIDENCE` |
| `candidate_count` | 0 |
| `relevance_level` | null |
**`input_params` 实值:**
```json
{
"query": "MySQL connection pool exhausted HikariCP diagnosis",
"step_id": 962,
"tool_call_id": "call_00_LireJzbiFEsbZ9ZVvgYp8330",
"request_bytes": 62
}
```
### 3.3 Trace API
- `toolInvocations[0].stepId = 962`
- `inputParams.query` 有值
- timeline `TOOL_INVOCATION.details.step_id = 962`
### 3.4 对照表
| 检查项 | 预期 | 实际 | 判定 |
|--------|------|------|------|
| `tool_invocation.step_id` | = agent_step.id | 962 | **PASS** |
| `input_params.query` | 有预览 | 有 | **PASS** |
| `input_params.step_id` | 与列一致 | 962 | **PASS** |
| Trace `stepId` | 非空 | 962 | **PASS** |
| timeline `step_id` | 非空 | 962 | **PASS** |
| `search_mode` | hybrid | hybrid | **PASS** |
| 业务 outcome | (非本页 KPI) | FALLBACK | 见 §4 |
```mermaid
flowchart LR
S0[agent_step 962<br/>r1 tool_call] --> TI[tool_invocation 860<br/>step_id=962]
TI --> IP[input_params.query]
TI --> TR[Trace / timeline]
TI --> RAG[hybrid candidate_count=0]
RAG --> FB[FALLBACK<br/>NO_EVIDENCE]
```
---
## 4. 业务 FALLBACK 原因(与审计无关)
| 现象 | 说明 |
|------|------|
| 日志 | `found=false`,`evidenceBlocks=0` |
| audit | `attempts[].usable=false`,`candidate_count=0` |
| warm 检索 | 同期 `GET /api/search/similar` 亦失败 |
| 日志噪音 | Milvus channel 未正确 shutdown 等提示(环境/客户端生命周期) |
**结论**:hybrid **路径进了**,但当次 **0 候选** → 无证据可写报告 → `SAFE_FALLBACK` / `INSUFFICIENT_EVIDENCE`。
这不否定 `step_id` / `query` 落库;完整 SUCCESS 业务故事见主文档。
```mermaid
flowchart TB
Q[同一 RAG-only 题] --> LK[lookup_knowledge]
LK --> AUD[审计: step_id + query PASS]
LK --> HIT{有候选?}
HIT -->|SUCCESS 主文档 run| OK[PRECISE → DIAGNOSIS_REPORT]
HIT -->|本页 run| NO[0 候选 → FALLBACK]
```
---
## 5. 实现索引(补丁代码)
| 组件 | 职责 |
|------|------|
| `AgentStepAuditTracker` | run 级 bind/current/clear `agent_step.id` |
| `HarnessAgentAuditHook` | `beforeModel` 落 step 后 `bind` |
| `ToolBoundary.auditSafely` | 读 tracker,带上 `stepId` + `requestJson` |
| `JpaToolInvocationAuditSink` | 写 `step_id` 列;安全展开 `input_params` |
| `TraceAuditEvents.toolInvocation` | timeline details 的 `step_id` |
| `JpaChatRunStore.finish` | `clear(runId)` |
**`input_params` 安全规则摘要**:
- 始终:`tool_call_id`、`request_bytes`;有则:`step_id`
- 顶层 string/number/boolean/纯字符串数组;文本 ≤160 字符
- 键名含 password/token/secret/apikey 等 → 不落库
更完整的字段词典与 SUCCESS 阶段拆解见主文档 §3.4 / §3.4.1。
---
## 6. 复盘 SQL
```bash
python scripts/query_mysql.py "SELECT id, step_id, tool_name, relevance_level, CAST(input_params AS CHAR) AS inputp, LEFT(CAST(retrieval_details AS CHAR), 500) AS details FROM tool_invocation WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f'"
python scripts/query_mysql.py "SELECT id, step_index, agent_name, has_tool_call, token_count FROM agent_step WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f' ORDER BY step_index"
python scripts/query_mysql.py "SELECT run_id, status, release_outcome, tool_call_count, total_token_count FROM diagnosis_run WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f'"
```
```http
GET /api/diagnosis/e2e-audit-20260728171858/trace?runId=0be605f6-e036-40d3-b364-11b73672241f
```
---
## 7. 修订记录
| 日期 | 说明 |
|------|------|
| 2026-07-28 | 从主 walkthrough 拆出:专门记录 audit 验收 run(FALLBACK 业务 + step_id/query PASS) |
@@ -39,6 +39,18 @@
> 现在的 K、现有的 L0/L1、现有的规则加分,下一步到底该扩召回、该融合,还是该上精排? > 现在的 K、现有的 L0/L1、现有的规则加分,下一步到底该扩召回、该融合,还是该上精排?
```mermaid
flowchart TB
P[排序/质量问题] --> A{候选池 K 多大?}
A -->|K 很小 3~5| B[先扩召回 / 修去重 / 轻规则]
A -->|K 中等 15~30| C[多路 + RRF + 可选轻精排]
A -->|K 很大 50+| D[强 rerank 才划算]
P --> E{跨路分数?}
E -->|原始分硬加| F[尺度不同 · 易玄学]
E -->|RRF 名次投票| G[推荐]
```
--- ---
## 2. 现状解剖:有“重排”,不等于有“强 Rerank” ## 2. 现状解剖:有“重排”,不等于有“强 Rerank”
@@ -57,6 +69,17 @@ Agent query
-> RagResultProjector # 投影成 Agent 可见契约 -> RagResultProjector # 投影成 Agent 可见契约
``` ```
```mermaid
flowchart TB
AQ[Agent query] --> L0[KnowledgeQueryTransformer<br/>L0 hint / category filter]
L0 --> L1[KnowledgeDocumentRetriever<br/>L1 向量 topK]
L1 --> POST[KnowledgeEvidencePostProcessor<br/>归一化 + 规则 boost + 去重]
POST --> PACK[KnowledgeContextPacker]
PACK --> ASM[LookupResultAssembler]
ASM --> PROJ[RagResultProjector]
PROJ --> AGENT[Agent 可见结果]
```
其中“重排”发生在后处理阶段,名字也常叫 rerank,但实现通常是: 其中“重排”发生在后处理阶段,名字也常叫 rerank,但实现通常是:
```text ```text
@@ -69,6 +92,20 @@ finalScore = baseScore
再按 finalScore 降序 再按 finalScore 降序
``` ```
```mermaid
flowchart LR
BASE[baseScore<br/>向量相似度] --> SUM[finalScore]
D[+ domain]
E[+ entity]
K[+ keyword]
S[+ source_type]
D --> SUM
E --> SUM
K --> SUM
S --> SUM
SUM --> SORT[按 finalScore 降序]
```
同时会留下 `rerankTrace`(base/final score、boost reasons),便于内部审计。 同时会留下 `rerankTrace`(base/final score、boost reasons),便于内部审计。
### 2.2 这套做法解决了什么 ### 2.2 这套做法解决了什么
@@ -110,6 +147,18 @@ finalScore = baseScore
排序策略必须和 K 匹配。可以先用下面这张表做决策: 排序策略必须和 K 匹配。可以先用下面这张表做决策:
```mermaid
flowchart TB
K{retrieve-k / 候选规模}
K -->|3~5| S1[修去重 · 轻规则<br/>双路径 dense 融合]
K -->|15~30| S2[多路召回 + RRF<br/>可选轻精排]
K -->|50+| S3[强 Cross-Encoder / 托管 Rerank]
S1 -.->|先别上| X1[商业 Rerank]
S2 -.->|收益有限| X2[只继续调 keyword boost]
S3 -.->|避免| X3[无评测堆模型]
```
| 召回规模 K | 更适合做什么 | 不太值得先做什么 | | 召回规模 K | 更适合做什么 | 不太值得先做什么 |
|---|---|---| |---|---|---|
| 3 ~ 5 | 修去重、轻规则、双路径 dense 融合 | Cross-Encoder / 商业 Rerank | | 3 ~ 5 | 修去重、轻规则、双路径 dense 融合 | Cross-Encoder / 商业 Rerank |
@@ -157,6 +206,15 @@ topK = 3,召回、排序、返回都是 3
若结果差,再 unfiltered 重试 若结果差,再 unfiltered 重试
``` ```
```mermaid
flowchart TB
Q[query + L0 category?] --> F[filtered dense]
F --> BAD{结果差/空?}
BAD -->|是| U[unfiltered retry]
BAD -->|否| USE[采用 filtered]
U --> REP[常见:整锅替换 filtered]
```
这能工作,但常见实现是 **串行整锅替换**: 这能工作,但常见实现是 **串行整锅替换**:
- retry 成功后,直接丢掉第一次 filtered 的全部结果 - retry 成功后,直接丢掉第一次 filtered 的全部结果
@@ -170,6 +228,16 @@ topK = 3,召回、排序、返回都是 3
去重合并后一起排序 去重合并后一起排序
``` ```
```mermaid
flowchart LR
Q[query] --> A[路A filtered dense]
Q --> B[路B unfiltered dense]
A --> M[去重合并]
B --> M
M --> RRF[RRF / 统一排序]
RRF --> OUT[topN]
```
#### 为什么值得做? #### 为什么值得做?
因为两路解决的是不同失败模式: 因为两路解决的是不同失败模式:
@@ -333,6 +401,23 @@ RRF(d) = Σ 1 / (k + rank_i(d))
- 多路都靠前的候选,融合分自然更高 - 多路都靠前的候选,融合分自然更高
- 只在一路偶然靠前的候选,不会单靠绝对分尺度“爆掉” - 只在一路偶然靠前的候选,不会单靠绝对分尺度“爆掉”
```mermaid
flowchart TB
subgraph paths["各路有序结果"]
P1[dense ranks]
P2[BM25 ranks]
P3[filtered ranks · 可选]
end
P1 --> RRF["RRF(d) = Σ 1/(k + rank_i)"]
P2 --> RRF
P3 --> RRF
RRF --> OUT[融合序 · 不依赖原始分尺度]
L2[L2 原分] -.->|不直接相加| X[避免]
BM[BM25 原分] -.-> X
```
### 6.2 为什么适合 RAG 多路融合 ### 6.2 为什么适合 RAG 多路融合
RRF 特别适合下面这种现实约束: RRF 特别适合下面这种现实约束:
@@ -553,8 +638,29 @@ Query
Agent projection Agent projection
``` ```
```mermaid
flowchart TB
Q[Query] --> U[dense unfiltered top20]
Q --> F[dense filtered top10]
Q --> B[bm25/keyword top10]
U --> DEDUP[chunk 级去重<br/>docId#chunkIndex]
F --> DEDUP
B --> DEDUP
DEDUP --> RRF[RRF / 加权 RRF]
RRF --> RR[轻规则或模型精排 top5]
RR --> CAP[每文档 chunk 上限 + pack]
CAP --> AG[Agent projection]
```
### 8.2 分阶段推进 ### 8.2 分阶段推进
```mermaid
flowchart LR
P0[Phase0<br/>chunk 去重<br/>retrieve-k/return-n] --> P1[Phase1<br/>filtered+unfiltered RRF]
P1 --> P2[Phase2<br/>BM25 跨维度]
P2 --> P3[Phase3<br/>可插拔精排]
```
#### Phase 0:先修前提 #### Phase 0:先修前提
否则后面多路都会被吞: 否则后面多路都会被吞:
+583
View File
@@ -0,0 +1,583 @@
# RAG 离线评测:讨论、设计与落地
**日期**:2026-07-28
**范围**:`eval/rag-retrieval` 离线 baseline、fixture 生成、与 hybrid/quality 主路径对齐
**读者**:要维护或扩展知识库回归评测的工程同学
**关联实现 / 变更**:
- 目录:`eval/rag-retrieval/`
- 脚本:`scripts/eval_rag_retrieval.py`、`generate_rag_lookup_snapshots.ps1`、`prepare_rag_eval_seed.ps1`
- OpenSpec / devflow:`rag-eval-hybrid-baseline`(已归档)
- 前置:`docs/RAG-Hybrid质量分与后处理.md`、`docs/RAG-Agent如何读relevance_level.md`
---
## 1. 为什么要单独谈评测
hybrid、chunk 去重、qualityScore 统一之后,工程上仍缺一块:
> **改检索之后,用什么可重复的信号判断「变好了还是变坏了」?**
完整诊断 E2E(多工具 + 最终回答)太重、太噪。需要一层**只盯 `lookup_knowledge` 召回与管道行为**的回归。
本文汇总讨论中形成的:
1. 离线评测是什么、不是什么
2. Golden / Fixture 怎么设计、比什么
3. 项目现状是否符合定义
4. 改造复杂度与 sm-flow 落地(含 apply 中发现的闸门问题)
5. 指标、报告、纪律
---
## 2. 离线评测:概念边界
### 2.1 离线 vs 在线
| 说法 | 含义 |
|------|------|
| **在线** | 真跑检索:embedding、Milvus、完整 `lookup_knowledge` |
| **离线** | **不再访问检索栈**;用事先冻住的结果快照,和标准答案比对 |
```mermaid
flowchart TB
subgraph online["在线(贵、真、偶发)"]
S[Seed 语料] --> G[真 lookup / 检索管道]
G --> F[写入 Fixtures]
end
subgraph offline["离线(便宜、稳、可 CI)"]
C[Golden cases] --> E[比对脚本]
F2[已提交的 Fixtures] --> E
E --> R[报告 / baseline / diff]
end
F -.->|提交入库| F2
```
可以记成:
```text
在线:考试现场答题(环境会变)
离线:用标准答卷复印件批改(环境冻结)
```
### 2.2 离线适合 / 不适合回答的问题
**适合:**
- 契约有没有破(结构、关键字段、行为路径)
- 在「同一份检索结果」假设下,期望文档/关键词/attempt 是否仍满足
- 相对上一版 baseline 的回归 diff
**不适合单独承担:**
- hybrid 是否比 dense 更好 → 需要**同一时期**在线双跑
- 阈值 0.75 是否合适 → 需要在线统计 level 分布
- 换 embedding 后召回如何 → 必须重刷 fixture 或 live
- Agent 最终诊断对不对 → 诊断 E2E
```mermaid
flowchart LR
Q1[契约 / 期望回归] --> OFF[离线 baseline]
Q2[当前召回是否正确] --> ON[在线生成 fixture 或 live]
Q3[dense vs hybrid 增益] --> CMP[同期双 mode 对照]
Q4[诊断是否正确] --> E2E[诊断 harness]
```
---
## 3. 三块积木
```mermaid
flowchart TB
subgraph golden_box["① Golden(薄)"]
GQ[query]
GE[期望:doc / keyword / attempt / …]
end
subgraph fixture_box["② Fixture(冻)"]
FM[meta: caseId, searchMode, kbScope, time]
FL[lookupResult 结构化输出]
end
subgraph judge_box["③ 裁判(规则)"]
A[按 golden 字段断言]
M[衍生 Hit / rank / 通过率]
REP[json + md + 可选 diff]
end
golden_box -->|caseId 对齐| judge_box
fixture_box --> judge_box
```
| 积木 | 是什么 | 不是什么 |
|------|--------|----------|
| **Golden** | 问什么 + **应该**怎样 | 不是整包线上成功 JSON 原样当期望 |
| **Fixture** | 某次跑完**实际**怎样 | 不是每次离线评测都要重跑检索 |
| **裁判** | 关键字段比对 + 指标 + 报告 | 不是两个大 JSON deep equal |
时间线:
```text
① 设计 Golden
② (改检索 / 换库 / 换 mode 时)在线跑 → 写/更新 Fixtures
③ 日常:Fixtures × Golden → 报告(多数时候只做这一步)
```
---
## 4. 为什么不全量字段对比
讨论中的直觉「分数不固定」是对的,但原因不止这一条:
| 原因 | 说明 |
|------|------|
| 分数不稳定 | L2 / RRF / embedding 会漂 |
| 实现细节会变 | trace 结构、reason 文案、时间戳、tool_call_id |
| 截断与预算会变 | excerpt 长度、packedText |
| 目标是「对不对」 | 不是字节级一致 |
```mermaid
flowchart LR
F[Fixture JSON] --> K[只抽取关键字段]
G[Golden 期望] --> K
K --> P{断言}
P -->|通过| OK[Pass + 指标]
P -->|失败| FAIL[Fail + 原因列表]
F -.->|不做| FULL[全量 deep equal]
```
**全量对比**偶尔可用于「紧挨着两次生成器输出的工程 diff」,那不是 golden 质量标准。
---
## 5. Golden / Fixture 字段设计
### 5.1 Golden:一行长什么样(分层)
```mermaid
flowchart TB
ID[身份: caseId, scenario/tags, notes]
IN[输入: query]
P0[P0 命中: expectedDocIds/Sources, expectedKeywords]
P1[P1 行为/展示: attempt, fallback, status, breadcrumb]
P2[P2 细粒度: evidenceKey, minCount, firstRankMax]
P3[P3 观察: relevanceLevel — 慎作硬门禁]
NEG[负例: mustNotDocIds/Sources]
ID --> IN --> P0 --> P1
P1 --> P2
P1 --> NEG
P1 --> P3
```
| 优先级 | 比什么 | 作用 |
|--------|--------|------|
| P0 | docId / source、keywords | 召回对不对、段是否有用 |
| P1 | selectedAttempt、fallback、evidenceStatus | 路径有没有坏 |
| P1 | mustNot* | 硬负例 / decoy |
| P2 | chunk / evidenceKey、条数、首条相关 rank | 身份与排序 |
| P3 | relevance_level | 观察用;hybrid 下易松 |
**原则:** 期望对齐「用户/Agent 可感知的对错」,少锁实现细节。
### 5.2 Fixture:最少保留什么
```mermaid
flowchart TB
subgraph fix["fixture"]
META["meta<br/>caseId, query, retrievedAt<br/>searchMode, kbScope?"]
RES["result / lookupResult<br/>有序 evidence[]<br/>attempt / fallback / status<br/>contextPack? / traces?"]
DBG["debug 可选<br/>rawScore, denseDistance…"]
end
META --> RES
RES --> DBG
```
每条 evidence 最少:
```text
docId 或可对齐的 source
excerpt / content(关键词断言需要)
顺序 = rank(数组下标即可)
evidenceKey / chunkIndex(多 chunk case 需要)
title / breadcrumb(按需)
```
**故意不锁:** score 全文、完整 trace、tool_call_id、packedText 全文(除非单独立项)。
### 5.3 Golden → Fixture 取值对照
```mermaid
flowchart LR
subgraph g["Golden"]
g1[expectedDocIds]
g2[expectedKeywords]
g3[expectedSelectedAttempt]
g4[expectedFallbackReason]
g5[mustNotDocIds]
end
subgraph f["Fixture"]
f1[evidence[].docId/source]
f2[evidence[].excerpt 拼接]
f3[retrievalTrace.selectedAttempt]
f4[retrievalTrace.fallbackReason]
f5[evidence 全表扫描]
end
g1 --> f1
g2 --> f2
g3 --> f3
g4 --> f4
g5 --> f5
```
---
## 6. 指标与报告
### 6.1 两层指标
```mermaid
flowchart TB
subgraph gate["门禁主信号"]
PASS[逐 case Pass/Fail]
RATE[通过率 / 按 tag 通过率]
end
subgraph quality["质量刻度(报告展示)"]
HIT[Hit@n / hitLevel strong·medium·weak·miss]
RANK[first relevant rank / MRR]
KW[keyword coverage]
FB[fallback rate]
LV[level 直方图 — 观察]
end
PASS --> RATE
HIT --> RANK
```
| 指标 | 含义 |
|------|------|
| **Pass/Fail** | golden 声明的 expected* 是否全部满足 |
| **hitLevel** | strong / medium / weak / miss(项目已有) |
| **Recall@K** | strong+medium 算命中 |
| **firstExpectedRank** | 第一条期望文档的排名 |
| **fallback / attempt** | 行为路径 |
| **level 分布** | 宜观察,hybrid 下慎作硬门禁 |
### 6.2 报告长什么样
```mermaid
flowchart LR
EVAL[离线评测] --> J[baseline.json<br/>机器可读]
EVAL --> M[baseline.md<br/>人读表格]
EVAL --> D[baseline-diff.*<br/>相对上一版]
J --> CI[CI / 脚本解析]
M --> HUM[人看失败原因]
D --> REV[改代码还是改期望]
```
工程上的「得出结果」=:
1. **门禁**:must-pass 是否全绿
2. **诊断**:谁红、红在哪类断言
3. **趋势**:相对旧 baseline 变好还是变差
---
## 7. 项目现状审计(改造前)
讨论结论:**模型符合定义,内容偏旧(约 70%)**。
```mermaid
flowchart TB
subgraph ok["已符合"]
A1[golden × fixture × key-field]
A2[seed + kb_scope=rag-eval]
A3[Hit 分层 + recall + baseline diff]
A4[离线不连库]
end
subgraph gap["缺口"]
B1[生成器仍传 vector-store.mode=spring]
B2[fixture 无 searchMode/kbScope]
B3[无 dense/hybrid 双目录对照]
B4[无 mustNot / chunk 硬期望]
B5[快照停在 boost 改序时代]
end
ok --> gap
```
| 维度 | 符合度 |
|------|--------|
| 三件套架构 | 高 |
| 关键字段比对 | 高 |
| 报告 / diff | 高 |
| 与 hybrid 主路径同步 | 低(改造前) |
| chunk / 负例 / mode 矩阵 | 弱或无 |
---
## 8. 改造策略与复杂度
### 8.1 两刀切分
```mermaid
flowchart TB
K1["第一刀(已落地)<br/>search.mode 接线<br/>fixture meta<br/>README<br/>重刷 hybrid baseline"]
K2["第二刀(未做)<br/>fixtures/hybrid vs dense<br/>对照表<br/>golden tags/mustNot"]
K1 --> DONE[可门禁当前主路径]
K2 --> CMP[可回答 hybrid 增益]
```
**复杂度判断:中低。** 不必重写框架;成本在联调环境与 baseline 纪律,不在算法。
| 工作 | 复杂度 |
|------|--------|
| 改生成参数 / meta / README | 低 |
| seed + 重刷 fixture | 中低(看环境) |
| 双目录对照 | 中低(第二刀) |
| 换框架 / LLM judge | 高(不建议现在) |
### 8.2 sm-flow 落地范围(已确认)
- **仅第一刀**
- **接线必交**;fixture 刷新尽力(本次环境可用,已刷绿)
Change:`rag-eval-hybrid-baseline`(已归档)。
---
## 9. 落地后的主链路(当前)
### 9.1 日常离线
```mermaid
flowchart LR
GC[golden-cases.json] --> PY[eval_rag_retrieval.py]
FX[fixtures/*.json] --> PY
PY --> BR[reports/baseline.*]
```
```bash
python scripts/eval_rag_retrieval.py
```
### 9.2 改检索后的完整环
```mermaid
sequenceDiagram
participant Eng as 工程师
participant Seed as prepare_rag_eval_seed
participant Gen as generate_rag_lookup_snapshots
participant Tool as LookupKnowledgeTool
participant Off as eval_rag_retrieval.py
participant Git as 仓库 baseline
Eng->>Seed: 导入 seed-docs (kb_scope=rag-eval)
Seed-->>Eng: MySQL/L0/Milvus 就绪
Eng->>Gen: -SearchMode hybrid
Gen->>Tool: 每条 golden.query
Tool-->>Gen: LookupResult
Gen->>Gen: 写 fixture + searchMode/kbScope
Eng->>Off: fixtures × golden
Off-->>Eng: pass/fail + 指标
Eng->>Git: 意图变更则更新 baseline(带 diff 原因)
```
### 9.3 生成器配置(改造后)
| 参数 | 默认 | 含义 |
|------|------|------|
| `SearchMode` | `hybrid` | `retrieval.search.mode` |
| `KbScope` | `rag-eval` | 评测语料隔离 |
| (已删除) | — | `vector-store.mode=spring\|sdk` |
```powershell
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 # hybrid
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures ... -SkipEval # 对照用
```
---
## 10. Apply 中发现的关键问题:filter fallback 与 quality 闸门
### 10.1 现象
重刷 hybrid fixtures 后,`chat-l0-filter-fallback` 变红:
- 只命中 decoy(`overfilter-decoy`)
- `selectedAttempt=FILTERED_VECTOR`,**没有** unfiltered retry
- 根因:hybrid **纯 rank→quality** 时 rank1 恒为 ~1.0 → `isLowQuality` 永不成立
```mermaid
flowchart TB
subgraph before["纯 rank quality(有问题)"]
H1[hybrid RRF 序] --> R1[rank1 quality=1.0]
R1 --> N1[isLowQuality=false]
N1 --> X1[不 retry · decoy 留下]
end
subgraph after["denseDistance 闸门(已修)"]
H2[hybrid RRF 序 · 排序不变] --> D2[并行 dense 填 denseDistance]
D2 --> Q2[quality = L2 归一化]
Q2 --> L2{top quality < 0.5?}
L2 -->|是| RET[UNFILTERED_VECTOR_RETRY]
L2 -->|否| KEEP[保留 filtered 结果]
end
```
### 10.2 设计取舍(必须记清)
| 信号 | 用途 |
|------|------|
| **RRF / originalRank** | **排序权威**(谁在前) |
| **denseDistance → quality** | **绝对质量闸门**(要不要 retry、level 档) |
| **不**再:用 L2 覆盖 hybrid 主分 / label | 避免回到「伪装成 L2」 |
```mermaid
flowchart LR
subgraph sort["排序"]
RRF[RRF 返回序]
end
subgraph gate["质量闸门"]
L2[dense L2 若有]
RK[rank 回退若无 dense]
L2 --> QS[qualityScore]
RK --> QS
end
RRF --> LIST[evidence 列表顺序]
QS --> LV[relevance_level]
QS --> FB[isLowQuality → filter fallback]
```
这与 quality 统一文的精神一致:**排序与闸门分信号**;只是 hybrid 闸门不能**只**靠序数分。
验证(归档时):
- offline **7/7 pass**
- fallback case:`UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`
---
## 11. Baseline 纪律
```mermaid
flowchart TB
RED[离线变红] --> Q{实现退步还是预期变了?}
Q -->|退步| CODE[改代码]
Q -->|预期变了| GOLD[改 golden / 重刷 fixture]
GOLD --> NOTE[写清原因 · 更新 baseline]
Q -->|禁止| BLIND[不看 diff 整锅覆盖]
```
README 原话仍然成立:fixture 对不上,要么修链路,要么改期望——**二选一要显式**。
---
## 12. 与完整评测体系的位置
```mermaid
flowchart TB
subgraph L1["L1 离线契约 — 已有且已对齐 hybrid"]
OFF[fixtures × golden]
end
subgraph L2["L2 在线召回 — 生成器已接线"]
LIVE[seed → snapshot hybrid]
end
subgraph L3["L3 对照与标定 — 部分未做"]
DD[dense vs hybrid 双目录表]
CAL[level/阈值直方图标定]
end
subgraph L4["L4 诊断 E2E — 另一套"]
DIAG[多工具 · 最终回答]
end
L1 --> L2
L2 --> L3
L2 -.-> L4
```
| 层 | 状态 |
|----|------|
| L1 离线 | **已落地**,hybrid fixtures + baseline 绿 |
| L2 在线生成 | **已接线**,本机已成功重刷 |
| L3 双 mode 对照目录 | **未做**(第二刀) |
| L4 诊断 E2E | 独立 harness,非本文 |
---
## 13. 实践清单
1. **日常**:只跑离线 `eval_rag_retrieval.py`。
2. **改检索 / 索引 / mode / 闸门**:seed → hybrid 生成 → 离线 → 看 diff 再更新 baseline。
3. **对照 dense**:`-SearchMode dense` 指到另一 fixtures 目录(第二刀可产品化报表)。
4. **Golden** 锁业务真值;**score / 完整 trace** 默认不锁。
5. **relevance_level** 先观察,慎作硬门禁。
6. **排序听 RRF**;**retry/level 听绝对 quality(dense L2)**。
7. 评测语料固定 `kb_scope=rag-eval`,勿绑生产杂库。
---
## 14. 结语
离线评测不是「再造一个复杂平台」,而是:
```text
固定问题(Golden)
× 冻结答卷(Fixture)
× 关键字段裁判
→ 可 diff 的报告
```
项目原本骨架正确;本轮补上了 **hybrid 时代的生成接线、fixture meta、baseline 重刷**,并在真实跑通时修正了 **「序数 quality 杀死 filter fallback」** 的闸门设计。
下一有价值的增量是 **dense/hybrid 同期对照表(第二刀)** 与 **level 分布标定**,而不是换评测框架。
---
## 附录 A:目录与命令速查
| 路径 | 作用 |
|------|------|
| `eval/rag-retrieval/cases/golden-cases.json` | Golden |
| `eval/rag-retrieval/fixtures/*.json` | Fixtures(含 searchMode) |
| `eval/rag-retrieval/seed-docs/` | 评测语料 |
| `eval/rag-retrieval/reports/baseline.*` | 离线基线报告 |
| `scripts/eval_rag_retrieval.py` | 离线裁判 |
| `scripts/generate_rag_lookup_snapshots.ps1` | 在线生成 fixture |
| `scripts/prepare_rag_eval_seed.ps1` | 导入 seed |
```powershell
# 离线
python scripts\eval_rag_retrieval.py
# 完整刷新(需 embedding + Milvus 等)
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid
```
## 附录 B:相关文档
| 文档 | 内容 |
|------|------|
| `eval/rag-retrieval/README.md` | 操作说明(以仓库为准) |
| `docs/RAG-Hybrid质量分与后处理.md` | quality / 后处理 |
| `docs/RAG-Agent如何读relevance_level.md` | Agent 如何读 level |
| `docs/RAG排序-多路召回与RRF.md` | 多路与 RRF |
| `devflow/projects/2026-07-28-rag-eval-hybrid-baseline/` | 本 change 档案 |
| `openspec/specs/rag-eval-offline-baseline/spec.md` | 主规格 |
File diff suppressed because it is too large Load Diff
+89 -143
View File
@@ -1,38 +1,65 @@
# RAG Retrieval Baseline # RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor. Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
The baseline is intentionally narrower than full diagnosis evaluation. It checks It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
evidence keywords before changing L0 behavior, query augmentation, evidence Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
post-processing, or Spring AI VectorStore integration. `mode=dense` remains a same-collection baseline for recall comparison (not a second index).
Related design notes:
- `docs/RAG-Hybrid质量分与后处理.md`
- `docs/RAG-Agent如何读relevance_level.md`
- `mvp/architecture/RAG知识检索架构.md` §6
## Offline vs live
| Layer | What | Needs live stack? |
|-------|------|-------------------|
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
Daily CI / local quick check: **offline only**.
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
## Layout ## Layout
```text ```text
eval/rag-retrieval/ eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases cases/golden-cases.json Fixed queries + expectations
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
fixtures/*.json Saved retrieval fixtures for each case fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
reports/baseline.json Machine-readable baseline report reports/baseline.json|md Last accepted offline report
reports/baseline.md Human-readable baseline report reports/baseline-diff.* Optional diff vs previous report
reports/baseline-diff.* Optional diff reports
reports/live-post-reindex.* Optional live acceptance reports
``` ```
## Seed Docs + Import/Reindex ## Fixture shape (minimum)
The live-tool eval uses canonical seed documents so the real ```text
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of caseId
whatever ad hoc documents happen to exist in the local knowledge base. query
retrievedAt
searchMode # hybrid | dense (required on newly generated fixtures)
kbScope # e.g. rag-eval when generation used a scope
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
```
Seed documents live in: Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
Older fixtures may omit `searchMode`; regenerate to attach meta.
## Seed docs + import
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
```text ```text
eval/rag-retrieval/seed-docs/*.md eval/rag-retrieval/seed-docs/*.md
``` ```
Each seed doc uses frontmatter fields that are propagated into vector metadata: Frontmatter example:
```yaml ```yaml
source: mysql-connection-pool source: mysql-connection-pool
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval kb_scope: rag-eval
``` ```
Import or reindex the seed docs through the real upload pipeline: Import via real upload pipeline:
```powershell ```powershell
.\scripts\prepare_rag_eval_seed.ps1 .\scripts\prepare_rag_eval_seed.ps1
``` ```
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It Isolation:
deletes the existing document with the same `source`/`docId`, uploads the seed
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
Milvus chunks.
`kb_scope` isolates eval data: - App default may leave `retrieval.kb-scope` empty (all docs).
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
- default application config leaves `retrieval.kb-scope` empty, so legacy docs Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
without `kb_scope` remain searchable;
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
vector retrieval both use only the canonical eval seed docs;
- the fallback retry skips only the L0 category filter, not the `kb_scope`
boundary.
Frontmatter is not embedded as chunk content during upload. It feeds metadata, Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
L0, and document enrichment; only the Markdown body is chunked and embedded.
This keeps controlled L0 decoys from becoming semantically relevant just because
their frontmatter keywords matched the query.
## Run ## Offline run (no live stack)
From the repository root:
```bash ```bash
python scripts/eval_rag_retrieval.py python scripts/eval_rag_retrieval.py
``` ```
Custom paths are also supported: Custom paths:
```bash ```bash
python scripts/eval_rag_retrieval.py \ python scripts/eval_rag_retrieval.py \
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
--markdown-report eval/rag-retrieval/reports/baseline.md --markdown-report eval/rag-retrieval/reports/baseline.md
``` ```
## Generate Fixtures From LookupKnowledgeTool ## Generate fixtures (live stack)
Use the snapshot generator when fixtures should reflect the real
`LookupKnowledgeTool` pipeline:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
For the intended live loop, run seed import first:
```powershell ```powershell
.\scripts\prepare_rag_eval_seed.ps1 .\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 .\scripts\generate_rag_lookup_snapshots.ps1
python scripts\eval_rag_retrieval.py # default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
``` ```
The script runs a Spring test harness: Dense baseline snapshot (same seed, comparison only):
```text
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
```
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
bean, calls `lookupKnowledge(query)` for each case, writes
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
Custom paths are supported:
```powershell ```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 ` .\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
-Cases eval\rag-retrieval\cases\golden-cases.json `
-Fixtures eval\rag-retrieval\fixtures `
-RetrievedAt 2026-07-06T00:00:00Z
``` ```
The generator is disabled in normal test runs. It only executes when (Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
`rag.snapshot.enabled=true` is provided because it writes repository files and
depends on the configured runtime retrieval stack.
If generated fixtures fail the offline baseline, treat that as a real alignment Maven equivalent:
signal: either the golden expectations need to be adjusted to the current
knowledge base, or the knowledge base/indexing path needs to be fixed.
## Modular RAG Contract
Fixtures must use the current `lookupResult` shape, which mirrors the
`lookup_knowledge` output:
```text ```text
lookupResult.evidenceBlocks mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
lookupResult.contextPack -Drag.snapshot.enabled=true \
lookupResult.retrievalTrace -Dretrieval.kb-scope=rag-eval \
lookupResult.rerankTrace -Dretrieval.search.mode=hybrid \
test
``` ```
Golden cases can assert both retrieval quality and pipeline behavior: Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
## Golden assertions
Supported expectation fields include:
- `expectedSources` / `expectedDocIds` - `expectedSources` / `expectedDocIds`
- `expectedBreadcrumbs` - `expectedBreadcrumbs` / `expectedKeywords`
- `expectedKeywords`
- `expectedSelectedAttempt` - `expectedSelectedAttempt`
- `expectedFallbackReason` - `expectedFallbackReason` / `expectedFallbackReasons`
- `expectedFallbackReasons`
- `expectedEvidenceStatus` - `expectedEvidenceStatus`
- `expectedContextSources` - `expectedContextSources`
- `expectedRerankTopSource` - `expectedRerankTopSource`
This lets the baseline catch regressions such as losing the expected evidence Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
source, skipping context packing, changing the selected retrieval attempt, or
breaking the filtered-vector to unfiltered-retry fallback.
## Baseline Diff **Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
To compare a freshly generated report against an existing baseline: ## Baseline diff
```bash ```bash
python scripts/eval_rag_retrieval.py \ python scripts/eval_rag_retrieval.py \
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md --diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
``` ```
The diff reports aggregate regressions and case-level changes for: Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
Non-zero exit on case failure or regression in diff mode.
- pass rate, recall@K, strong hit rate, miss count ## Hit levels
- pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
The command exits non-zero when a case fails or the diff contains a regression. - `strong`: expected document found **and** breadcrumb or keyword coverage OK
- `medium`: expected document found, coverage incomplete
- `weak`: keyword hit without expected document
- `miss`: neither
## Hit Levels `Recall@K` counts `strong` + `medium`.
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied. ## Optional live smoke (post-reindex)
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
- `weak`: expected evidence keyword is found, but expected document is missing.
- `miss`: expected document and expected evidence are not found.
`Recall@K` counts `strong` and `medium` as retrieved. After reindex, with app up:
## Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
or the Spring Boot application. It is a regression harness for retrieval behavior,
not a claim that live production retrieval accuracy is complete.
## Live Post-Reindex Acceptance
When embedding input changes, existing vectors do not update by themselves. For
example, after adding `title` and `breadcrumb` to the embedding text, the live
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
semantic signal.
Use this optional live acceptance flow after the application is running and the
knowledge base has been reindexed:
```bash ```bash
python scripts/eval_rag_live_acceptance.py python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
``` ```
Custom service URL and output paths are supported: Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
```bash
python scripts/eval_rag_live_acceptance.py \
--base-url http://127.0.0.1:9900 \
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
```
The script calls:
```text
GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
results, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
@@ -1,81 +1,122 @@
{ {
"caseId": "aiops-payment-latency-alert", "caseId" : "aiops-payment-latency-alert",
"query": "Alert HighLatency on payment-service with p95 latency above threshold", "query" : "Alert HighLatency on payment-service with p95 latency above threshold",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "payment-service-latency", "evidenceBlocks" : [ {
"title": "Payment Service Latency Alert Playbook", "docId" : "payment-service-latency",
"breadcrumb": "AIOps > Service Alerts > Payment Latency", "chunkIndex" : 2,
"retrievalLayer": "L1", "evidenceKey" : "payment-service-latency#chunk-2",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.", "source" : "payment-service-latency",
"score": 0.84, "title" : "Payment Latency",
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"] "breadcrumb" : "AIOps > Service Alerts > Payment Latency",
"retrievalLayer" : "L1",
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "aiops-alert-scope-control",
"chunkIndex" : 1,
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
"source" : "aiops-alert-scope-control",
"title" : "Alert Scope Control",
"breadcrumb" : "AIOps > Alert Scope Control",
"retrievalLayer" : "L1",
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "payment-service-latency",
"chunkIndex" : 1,
"evidenceKey" : "payment-service-latency#chunk-1",
"source" : "payment-service-latency",
"title" : "Service Alerts",
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
"retrievalLayer" : "L1",
"content" : "## Service Alerts",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "aiops-alert-scope-control",
"chunkIndex" : 0,
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
"source" : "aiops-alert-scope-control",
"title" : "AIOps",
"breadcrumb" : "AIOps > Alert Scope Control",
"retrievalLayer" : "L1",
"content" : "# AIOps",
"score" : 0.015384615398943424,
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 2]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Service Alerts\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 2108,
"includedSources" : [ "payment-service-latency", "aiops-alert-scope-control", "payment-service-latency", "aiops-alert-scope-control" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "mysql-connection-pool", "originalQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
"title": "MySQL Connection Pool Troubleshooting", "rewrittenQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
"breadcrumb": "Database > MySQL > Connection Pool", "categoryFilter" : "aiops",
"retrievalLayer": "L1", "selectedAttempt" : "FILTERED_VECTOR",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.", "fallbackReason" : null,
"score": 0.68, "evidenceStatus" : "supported",
"hitReasons": ["keyword_match:+0.10"] "queryHints" : {
} "domains" : [ "aiops" ],
], "matched_keywords" : [ "HighLatency", "payment-service", "p95 latency" ],
"contextPack": { "entities" : [ "HighLatency", "payment-service", "p95 latency" ],
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.", "l0_titles" : [ "Payment Service Latency Alert" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 236,
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis", "query" : "Alert HighLatency on payment-service with p95 latency above threshold",
"categoryFilter": "AIOps", "categoryFilter" : "aiops",
"selectedAttempt": "FILTERED_VECTOR", "candidateCount" : 5,
"fallbackReason": null, "usable" : true,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 969,
"domains": ["AIOps"], "topScore" : 0.032786883413791656,
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"], "topSimilarity" : 0.7736010700464249
"entities": ["payment-service", "HighLatency"], } ]
"l0_titles": ["Payment Service Latency Alert Playbook"],
"l0_match_count": 1
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis", "source" : "payment-service-latency",
"categoryFilter": "AIOps", "baseScore" : 0.7736010700464249,
"candidateCount": 2, "finalScore" : 0.7736010700464249,
"usable": true, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 11, }, {
"topScore": 0.84, "finalRank" : 2,
"topSimilarity": 0.84 "source" : "aiops-alert-scope-control",
} "baseScore" : 0.4683566689491272,
] "finalScore" : 0.4683566689491272,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 3,
"source" : "payment-service-latency",
"baseScore" : 0.4981400966644287,
"finalScore" : 0.4981400966644287,
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"finalRank" : 4,
"source" : "aiops-alert-scope-control",
"baseScore" : 0.3814886808395386,
"finalScore" : 0.3814886808395386,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
}, },
"rerankTrace": { "evidenceCandidateCount" : 5,
"items": [ "evidenceBlockCount" : 4,
{ "relevanceLevel" : "PRECISE",
"finalRank": 1, "completenessHint" : "知识库中不存在比上述结果更精准的文档",
"source": "payment-service-latency", "retrievedDomainsThisSession" : null,
"baseScore": 0.84, "message" : null
"finalScore": 1.29,
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "mysql-connection-pool",
"baseScore": 0.68,
"finalScore": 0.78,
"boostReasons": ["keyword_match:+0.10"]
}
]
}
} }
} }
@@ -1,65 +1,122 @@
{ {
"caseId": "aiops-prometheus-alert-scope", "caseId" : "aiops-prometheus-alert-scope",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?", "query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "aiops-alert-scope-control", "evidenceBlocks" : [ {
"title": "AIOps Alert Scope Control", "docId" : "aiops-alert-scope-control",
"breadcrumb": "AIOps > Alert Scope Control", "chunkIndex" : 1,
"retrievalLayer": "L1", "evidenceKey" : "aiops-alert-scope-control#chunk-1",
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.", "source" : "aiops-alert-scope-control",
"score": 0.88, "title" : "Alert Scope Control",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "AIOps > Alert Scope Control",
} "retrievalLayer" : "L1",
], "content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
"contextPack": { "score" : 0.032786883413791656,
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.", "hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"strategy": "top_evidence_blocks", }, {
"charBudget": 3500, "docId" : "payment-service-latency",
"usedChars": 188, "chunkIndex" : 1,
"includedSources": ["aiops-alert-scope-control"], "evidenceKey" : "payment-service-latency#chunk-1",
"omittedSources": [] "source" : "payment-service-latency",
"title" : "Service Alerts",
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
"retrievalLayer" : "L1",
"content" : "## Service Alerts",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "payment-service-latency",
"chunkIndex" : 2,
"evidenceKey" : "payment-service-latency#chunk-2",
"source" : "payment-service-latency",
"title" : "Payment Latency",
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
"retrievalLayer" : "L1",
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "aiops-alert-scope-control",
"chunkIndex" : 0,
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
"source" : "aiops-alert-scope-control",
"title" : "AIOps",
"breadcrumb" : "AIOps > Alert Scope Control",
"retrievalLayer" : "L1",
"content" : "# AIOps",
"score" : 0.03076923079788685,
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 2]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Service Alerts\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 2069,
"includedSources" : [ "aiops-alert-scope-control", "payment-service-latency", "payment-service-latency", "aiops-alert-scope-control" ],
"omittedSources" : [ ]
}, },
"retrievalTrace": { "retrievalTrace" : {
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?", "originalQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis", "rewrittenQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"categoryFilter": "AIOps", "categoryFilter" : "aiops",
"selectedAttempt": "FILTERED_VECTOR", "selectedAttempt" : "FILTERED_VECTOR",
"fallbackReason": null, "fallbackReason" : null,
"evidenceStatus": "supported", "evidenceStatus" : "supported",
"queryHints": { "queryHints" : {
"domains": ["AIOps"], "domains" : [ "aiops" ],
"matched_keywords": ["payload", "unrelated active alerts", "scope"], "matched_keywords" : [ "alert payload", "unrelated active alerts" ],
"entities": ["alert payload"], "entities" : [ "alert payload", "unrelated active alerts" ],
"l0_titles": ["AIOps Alert Scope Control"], "l0_titles" : [ "AIOps Alert Scope Control" ],
"l0_match_count": 1 "l0_match_count" : 1
}, },
"attempts": [ "attempts" : [ {
{ "name" : "FILTERED_VECTOR",
"name": "FILTERED_VECTOR", "query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"query": "AIOps alert payload scope unrelated active alerts diagnosis", "categoryFilter" : "aiops",
"categoryFilter": "AIOps", "candidateCount" : 5,
"candidateCount": 1, "usable" : true,
"usable": true, "errorMessage" : null,
"durationMs": 8, "durationMs" : 1623,
"topScore": 0.88, "topScore" : 0.032786883413791656,
"topSimilarity": 0.88 "topSimilarity" : 0.7561411112546921
} } ]
]
}, },
"rerankTrace": { "rerankTrace" : {
"items": [ "items" : [ {
{ "finalRank" : 1,
"finalRank": 1, "source" : "aiops-alert-scope-control",
"source": "aiops-alert-scope-control", "baseScore" : 0.7561411112546921,
"baseScore": 0.88, "finalScore" : 0.7561411112546921,
"finalScore": 1.13, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"] }, {
} "finalRank" : 2,
] "source" : "payment-service-latency",
} "baseScore" : 0.503810703754425,
"finalScore" : 0.503810703754425,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 3,
"source" : "payment-service-latency",
"baseScore" : 0.5772626996040344,
"finalScore" : 0.5772626996040344,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 4,
"source" : "aiops-alert-scope-control",
"baseScore" : 0.36770421266555786,
"finalScore" : 0.36770421266555786,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
},
"evidenceCandidateCount" : 5,
"evidenceBlockCount" : 4,
"relevanceLevel" : "PRECISE",
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
"retrievedDomainsThisSession" : null,
"message" : null
} }
} }
@@ -1,81 +1,88 @@
{ {
"caseId": "chat-diagnosis-flow", "caseId" : "chat-diagnosis-flow",
"query": "What is the standard troubleshooting flow for an application incident?", "query" : "What is the standard troubleshooting flow for an application incident?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "incident-diagnosis-flow", "evidenceBlocks" : [ {
"title": "Incident Diagnosis Flow", "docId" : "incident-diagnosis-flow",
"breadcrumb": "AIOps > Diagnosis Flow", "chunkIndex" : 1,
"retrievalLayer": "L1", "evidenceKey" : "incident-diagnosis-flow#chunk-1",
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.", "source" : "incident-diagnosis-flow",
"score": 0.82, "title" : "Diagnosis Flow",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "AIOps > Diagnosis Flow",
"retrievalLayer" : "L1",
"content" : "## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "incident-diagnosis-flow",
"chunkIndex" : 0,
"evidenceKey" : "incident-diagnosis-flow#chunk-0",
"source" : "incident-diagnosis-flow",
"title" : "AIOps",
"breadcrumb" : "AIOps > Diagnosis Flow",
"retrievalLayer" : "L1",
"content" : "# AIOps",
"score" : 0.016129031777381897,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: incident-diagnosis-flow\ntitle: Diagnosis Flow\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.\n\n[Evidence 2]\nsource: incident-diagnosis-flow\ntitle: AIOps\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 1032,
"includedSources" : [ "incident-diagnosis-flow", "incident-diagnosis-flow" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "rag-chunk-context-reconstruction", "originalQuery" : "What is the standard troubleshooting flow for an application incident?",
"title": "RAG Chunk Context Reconstruction", "rewrittenQuery" : "What is the standard troubleshooting flow for an application incident?",
"breadcrumb": "RAG > Chunking > Context Reconstruction", "categoryFilter" : "ops",
"retrievalLayer": "L1", "selectedAttempt" : "FILTERED_VECTOR",
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.", "fallbackReason" : null,
"score": 0.55, "evidenceStatus" : "supported",
"hitReasons": [] "queryHints" : {
} "domains" : [ "ops" ],
], "matched_keywords" : [ "standard troubleshooting flow", "application incident" ],
"contextPack": { "entities" : [ "standard troubleshooting flow", "application incident" ],
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.", "l0_titles" : [ "Incident Diagnosis Flow" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 192,
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "What is the standard troubleshooting flow for an application incident?", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation", "query" : "What is the standard troubleshooting flow for an application incident?",
"categoryFilter": "AIOps", "categoryFilter" : "ops",
"selectedAttempt": "FILTERED_VECTOR", "candidateCount" : 2,
"fallbackReason": null, "usable" : true,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 1540,
"domains": ["AIOps"], "topScore" : 0.032786883413791656,
"matched_keywords": ["collect evidence", "verify", "remediation"], "topSimilarity" : 0.6828859150409698
"entities": ["application incident"], } ]
"l0_titles": ["Incident Diagnosis Flow"],
"l0_match_count": 1
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "standard application incident troubleshooting flow collect evidence verify remediation", "source" : "incident-diagnosis-flow",
"categoryFilter": "AIOps", "baseScore" : 0.6828859150409698,
"candidateCount": 2, "finalScore" : 0.6828859150409698,
"usable": true, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 10, }, {
"topScore": 0.82, "finalRank" : 2,
"topSimilarity": 0.82 "source" : "incident-diagnosis-flow",
} "baseScore" : 0.3166210651397705,
] "finalScore" : 0.3166210651397705,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
}, },
"rerankTrace": { "evidenceCandidateCount" : 2,
"items": [ "evidenceBlockCount" : 2,
{ "relevanceLevel" : "REFERENCE",
"finalRank": 1, "completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
"source": "incident-diagnosis-flow", "retrievedDomainsThisSession" : null,
"baseScore": 0.82, "message" : null
"finalScore": 1.07,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "rag-chunk-context-reconstruction",
"baseScore": 0.55,
"finalScore": 0.55,
"boostReasons": []
}
]
}
} }
} }
@@ -1,81 +1,122 @@
{ {
"caseId": "chat-l0-domain-hint", "caseId" : "chat-l0-domain-hint",
"query": "Should L0 keyword matching decide the final retrieval result?", "query" : "Should L0 keyword matching decide the final retrieval result?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "rag-l0-domain-entity-hint", "evidenceBlocks" : [ {
"title": "RAG L0 Domain Entity Hint", "docId" : "rag-l0-domain-entity-hint",
"breadcrumb": "RAG > L0 > Domain Entity Hint", "chunkIndex" : 2,
"retrievalLayer": "L1", "evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.", "source" : "rag-l0-domain-entity-hint",
"score": 0.88, "title" : "Domain Entity Hint",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "rag-chunk-context-reconstruction",
"chunkIndex" : 2,
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
"source" : "rag-chunk-context-reconstruction",
"title" : "Context Reconstruction",
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
"retrievalLayer" : "L1",
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
"score" : 0.032258063554763794,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "rag-l0-domain-entity-hint",
"chunkIndex" : 1,
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
"source" : "rag-l0-domain-entity-hint",
"title" : "L0",
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "## L0",
"score" : 0.0317460335791111,
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "rag-chunk-context-reconstruction",
"chunkIndex" : 1,
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
"source" : "rag-chunk-context-reconstruction",
"title" : "Chunking",
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
"retrievalLayer" : "L1",
"content" : "## Chunking",
"score" : 0.015625,
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 2]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## L0\n\n[Evidence 4]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 1965,
"includedSources" : [ "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "rag-l0-l1-fusion-ranking", "originalQuery" : "Should L0 keyword matching decide the final retrieval result?",
"title": "RAG L0 L1 Fusion Ranking", "rewrittenQuery" : "Should L0 keyword matching decide the final retrieval result?",
"breadcrumb": "RAG > Ranking > Fusion", "categoryFilter" : "rag",
"retrievalLayer": "L1", "selectedAttempt" : "FILTERED_VECTOR",
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.", "fallbackReason" : null,
"score": 0.75, "evidenceStatus" : "supported",
"hitReasons": ["domain_match:+0.15"] "queryHints" : {
} "domains" : [ "rag" ],
], "matched_keywords" : [ "L0 keyword matching", "final retrieval result" ],
"contextPack": { "entities" : [ "L0 keyword matching", "final retrieval result" ],
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.", "l0_titles" : [ "RAG L0 Domain Entity Hint" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 219,
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "Should L0 keyword matching decide the final retrieval result?", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision", "query" : "Should L0 keyword matching decide the final retrieval result?",
"categoryFilter": "RAG", "categoryFilter" : "rag",
"selectedAttempt": "FILTERED_VECTOR", "candidateCount" : 6,
"fallbackReason": null, "usable" : true,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 850,
"domains": ["RAG"], "topScore" : 0.032786883413791656,
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"], "topSimilarity" : 0.6438122987747192
"entities": ["L0"], } ]
"l0_titles": ["RAG L0 Domain Entity Hint"],
"l0_match_count": 1
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "RAG L0 keyword matching domain entity hint final retrieval decision", "source" : "rag-l0-domain-entity-hint",
"categoryFilter": "RAG", "baseScore" : 0.6438122987747192,
"candidateCount": 2, "finalScore" : 0.6438122987747192,
"usable": true, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 9, }, {
"topScore": 0.88, "finalRank" : 2,
"topSimilarity": 0.88 "source" : "rag-chunk-context-reconstruction",
} "baseScore" : 0.41426247358322144,
] "finalScore" : 0.41426247358322144,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 3,
"source" : "rag-l0-domain-entity-hint",
"baseScore" : 0.3964804410934448,
"finalScore" : 0.3964804410934448,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 4,
"source" : "rag-chunk-context-reconstruction",
"baseScore" : 0.28678786754608154,
"finalScore" : 0.28678786754608154,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
}, },
"rerankTrace": { "evidenceCandidateCount" : 6,
"items": [ "evidenceBlockCount" : 4,
{ "relevanceLevel" : "REFERENCE",
"finalRank": 1, "completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
"source": "rag-l0-domain-entity-hint", "retrievedDomainsThisSession" : null,
"baseScore": 0.88, "message" : null
"finalScore": 1.13,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "rag-l0-l1-fusion-ranking",
"baseScore": 0.75,
"finalScore": 0.9,
"boostReasons": ["domain_match:+0.15"]
}
]
}
} }
} }
@@ -1,91 +1,149 @@
{ {
"caseId": "chat-l0-filter-fallback", "caseId" : "chat-l0-filter-fallback",
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?", "query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "rag-l0-filter-fallback", "evidenceBlocks" : [ {
"title": "RAG L0 Filter Fallback", "docId" : "rag-l0-filter-fallback",
"breadcrumb": "RAG > Fallback > Unfiltered Retry", "chunkIndex" : 2,
"retrievalLayer": "L1", "evidenceKey" : "rag-l0-filter-fallback#chunk-2",
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.", "source" : "rag-l0-filter-fallback",
"score": 0.83, "title" : "Unfiltered Retry",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "RAG > Fallback > Unfiltered Retry",
"retrievalLayer" : "L1",
"content" : "### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:UNFILTERED_VECTOR_RETRY", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "rag-l0-filter-decoy",
"chunkIndex" : 1,
"evidenceKey" : "rag-l0-filter-decoy#chunk-1",
"source" : "rag-l0-filter-decoy",
"title" : "Approval Window",
"breadcrumb" : "RAG > Fallback > Decoy",
"retrievalLayer" : "L1",
"content" : "## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:2", "attempt:UNFILTERED_VECTOR_RETRY", "l0_domain_overlap" ]
}, {
"docId" : "rag-l0-domain-entity-hint",
"chunkIndex" : 2,
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
"source" : "rag-l0-domain-entity-hint",
"title" : "Domain Entity Hint",
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
"score" : 0.0320020467042923,
"hitReasons" : [ "semantic_rank:3", "attempt:UNFILTERED_VECTOR_RETRY" ]
}, {
"docId" : "rag-l0-domain-entity-hint",
"chunkIndex" : 1,
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
"source" : "rag-l0-domain-entity-hint",
"title" : "L0",
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "## L0",
"score" : 0.03125,
"hitReasons" : [ "semantic_rank:4", "attempt:UNFILTERED_VECTOR_RETRY" ]
}, {
"docId" : "rag-chunk-context-reconstruction",
"chunkIndex" : 2,
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
"source" : "rag-chunk-context-reconstruction",
"title" : "Context Reconstruction",
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
"retrievalLayer" : "L1",
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
"score" : 0.03053613007068634,
"hitReasons" : [ "semantic_rank:5", "attempt:UNFILTERED_VECTOR_RETRY" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: rag-l0-filter-fallback\ntitle: Unfiltered Retry\nbreadcrumb: RAG > Fallback > Unfiltered Retry\nlayer: L1\nreasons: semantic_rank:1, attempt:UNFILTERED_VECTOR_RETRY, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.\n\n[Evidence 2]\nsource: rag-l0-filter-decoy\ntitle: Approval Window\nbreadcrumb: RAG > Fallback > Decoy\nlayer: L1\nreasons: semantic_rank:2, attempt:UNFILTERED_VECTOR_RETRY, l0_domain_overlap\ncontent:\n## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n## L0\n\n[Evidence 5]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:5, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 2904,
"includedSources" : [ "rag-l0-filter-fallback", "rag-l0-filter-decoy", "rag-l0-domain-entity-hint", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "rag-l0-domain-entity-hint", "originalQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"title": "RAG L0 Domain Entity Hint", "rewrittenQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"breadcrumb": "RAG > L0 > Domain Entity Hint", "categoryFilter" : "overfilter-decoy",
"retrievalLayer": "L1", "selectedAttempt" : "UNFILTERED_VECTOR_RETRY",
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.", "fallbackReason" : "filtered_vector_low_quality",
"score": 0.66, "evidenceStatus" : "supported",
"hitReasons": ["domain_match:+0.15"] "queryHints" : {
} "domains" : [ "overfilter-decoy" ],
], "matched_keywords" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
"contextPack": { "entities" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.", "l0_titles" : [ "RAG L0 Filter Decoy" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 214,
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry", "query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"categoryFilter": "RAG", "categoryFilter" : "overfilter-decoy",
"selectedAttempt": "UNFILTERED_VECTOR_RETRY", "candidateCount" : 2,
"fallbackReason": "filtered_vector_low_quality", "usable" : false,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 722,
"domains": ["RAG"], "topScore" : 0.032786883413791656,
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"], "topSimilarity" : 0.47391992807388306
"entities": ["L0"], }, {
"l0_titles": ["RAG L0 Domain Entity Hint"], "name" : "UNFILTERED_VECTOR_RETRY",
"l0_match_count": 1 "query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"categoryFilter" : null,
"candidateCount" : 20,
"usable" : true,
"errorMessage" : null,
"durationMs" : 2606,
"topScore" : 0.032786883413791656,
"topSimilarity" : 0.7571567445993423
} ]
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "RAG L0 filtered vector low quality fallback unfiltered retry", "source" : "rag-l0-filter-fallback",
"categoryFilter": "RAG", "baseScore" : 0.7571567445993423,
"candidateCount": 1, "finalScore" : 0.7571567445993423,
"usable": false, "boostReasons" : [ "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 7, }, {
"topScore": 1.35, "finalRank" : 2,
"topSimilarity": 0.325 "source" : "rag-l0-filter-decoy",
"baseScore" : 0.47391992807388306,
"finalScore" : 0.47391992807388306,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 3,
"source" : "rag-l0-domain-entity-hint",
"baseScore" : 0.5912361443042755,
"finalScore" : 0.5912361443042755,
"boostReasons" : [ ]
}, {
"finalRank" : 4,
"source" : "rag-l0-domain-entity-hint",
"baseScore" : 0.473749577999115,
"finalScore" : 0.473749577999115,
"boostReasons" : [ ]
}, {
"finalRank" : 5,
"source" : "rag-chunk-context-reconstruction",
"baseScore" : 0.4492502808570862,
"finalScore" : 0.4492502808570862,
"boostReasons" : [ ]
} ]
}, },
{ "evidenceCandidateCount" : 20,
"name": "UNFILTERED_VECTOR_RETRY", "evidenceBlockCount" : 5,
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?", "relevanceLevel" : "PRECISE",
"categoryFilter": null, "completenessHint" : "知识库中不存在比上述结果更精准的文档",
"candidateCount": 2, "retrievedDomainsThisSession" : null,
"usable": true, "message" : null
"durationMs": 13,
"topScore": 0.83,
"topSimilarity": 0.83
}
]
},
"rerankTrace": {
"items": [
{
"finalRank": 1,
"source": "rag-l0-filter-fallback",
"baseScore": 0.83,
"finalScore": 1.08,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "rag-l0-domain-entity-hint",
"baseScore": 0.66,
"finalScore": 0.81,
"boostReasons": ["domain_match:+0.15"]
}
]
}
} }
} }
@@ -1,81 +1,88 @@
{ {
"caseId": "chat-mysql-connection-pool", "caseId" : "chat-mysql-connection-pool",
"query": "MySQL connection pool is exhausted. How should I diagnose it?", "query" : "MySQL connection pool is exhausted. How should I diagnose it?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "mysql-connection-pool", "evidenceBlocks" : [ {
"title": "MySQL Connection Pool Troubleshooting", "docId" : "mysql-connection-pool",
"breadcrumb": "Database > MySQL > Connection Pool", "chunkIndex" : 2,
"retrievalLayer": "L1", "evidenceKey" : "mysql-connection-pool#chunk-2",
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.", "source" : "mysql-connection-pool",
"score": 0.86, "title" : "Connection Pool",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "Database > MySQL > Connection Pool",
"retrievalLayer" : "L1",
"content" : "### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "mysql-connection-pool",
"chunkIndex" : 1,
"evidenceKey" : "mysql-connection-pool#chunk-1",
"source" : "mysql-connection-pool",
"title" : "MySQL",
"breadcrumb" : "Database > MySQL > Connection Pool",
"retrievalLayer" : "L1",
"content" : "## MySQL",
"score" : 0.032258063554763794,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: mysql-connection-pool\ntitle: Connection Pool\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.\n\n[Evidence 2]\nsource: mysql-connection-pool\ntitle: MySQL\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## MySQL",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 1135,
"includedSources" : [ "mysql-connection-pool", "mysql-connection-pool" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "incident-diagnosis-flow", "originalQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
"title": "Incident Diagnosis Flow", "rewrittenQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
"breadcrumb": "AIOps > Diagnosis Flow", "categoryFilter" : "database",
"retrievalLayer": "L1", "selectedAttempt" : "FILTERED_VECTOR",
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.", "fallbackReason" : null,
"score": 0.61, "evidenceStatus" : "supported",
"hitReasons": [] "queryHints" : {
} "domains" : [ "database" ],
], "matched_keywords" : [ "MySQL connection pool" ],
"contextPack": { "entities" : [ "MySQL connection pool" ],
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.", "l0_titles" : [ "MySQL Connection Pool Runbook" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 216,
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis", "query" : "MySQL connection pool is exhausted. How should I diagnose it?",
"categoryFilter": "Database", "categoryFilter" : "database",
"selectedAttempt": "FILTERED_VECTOR", "candidateCount" : 3,
"fallbackReason": null, "usable" : true,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 5267,
"domains": ["Database", "MySQL"], "topScore" : 0.032786883413791656,
"matched_keywords": ["connection pool", "HikariCP", "max_connections"], "topSimilarity" : 0.8114794194698334
"entities": ["MySQL", "HikariCP"], } ]
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
"l0_match_count": 1
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis", "source" : "mysql-connection-pool",
"categoryFilter": "Database", "baseScore" : 0.8114794194698334,
"candidateCount": 2, "finalScore" : 0.8114794194698334,
"usable": true, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 12, }, {
"topScore": 0.86, "finalRank" : 2,
"topSimilarity": 0.86 "source" : "mysql-connection-pool",
} "baseScore" : 0.49219560623168945,
] "finalScore" : 0.49219560623168945,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
}, },
"rerankTrace": { "evidenceCandidateCount" : 3,
"items": [ "evidenceBlockCount" : 2,
{ "relevanceLevel" : "PRECISE",
"finalRank": 1, "completenessHint" : "知识库中不存在比上述结果更精准的文档",
"source": "mysql-connection-pool", "retrievedDomainsThisSession" : null,
"baseScore": 0.86, "message" : null
"finalScore": 1.11,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "incident-diagnosis-flow",
"baseScore": 0.61,
"finalScore": 0.61,
"boostReasons": []
}
]
}
} }
} }
@@ -1,81 +1,122 @@
{ {
"caseId": "chat-rag-chunk-context", "caseId" : "chat-rag-chunk-context",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?", "query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
"retrievedAt": "2026-07-06T00:00:00Z", "retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"lookupResult": { "searchMode" : "hybrid",
"found": true, "kbScope" : "rag-eval",
"evidenceBlocks": [ "lookupResult" : {
{ "found" : true,
"source": "rag-chunk-context-reconstruction", "evidenceBlocks" : [ {
"title": "RAG Chunk Context Reconstruction", "docId" : "rag-chunk-context-reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction", "chunkIndex" : 2,
"retrievalLayer": "L1", "evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.", "source" : "rag-chunk-context-reconstruction",
"score": 0.79, "title" : "Context Reconstruction",
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"] "breadcrumb" : "RAG > Chunking > Context Reconstruction",
"retrievalLayer" : "L1",
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "rag-l0-domain-entity-hint",
"chunkIndex" : 2,
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
"source" : "rag-l0-domain-entity-hint",
"title" : "Domain Entity Hint",
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
"score" : 0.032258063554763794,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "rag-chunk-context-reconstruction",
"chunkIndex" : 1,
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
"source" : "rag-chunk-context-reconstruction",
"title" : "Chunking",
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
"retrievalLayer" : "L1",
"content" : "## Chunking",
"score" : 0.01587301678955555,
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
}, {
"docId" : "rag-l0-domain-entity-hint",
"chunkIndex" : 0,
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-0",
"source" : "rag-l0-domain-entity-hint",
"title" : "RAG",
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
"retrievalLayer" : "L1",
"content" : "# RAG",
"score" : 0.015625,
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 2]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 3]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: RAG\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# RAG",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 1966,
"includedSources" : [ "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint" ],
"omittedSources" : [ ]
}, },
{ "retrievalTrace" : {
"source": "rag-breadcrumb-embedding-gap", "originalQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
"title": "RAG Breadcrumb Embedding Gap", "rewrittenQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
"breadcrumb": "RAG > Embedding > Breadcrumb", "categoryFilter" : "rag",
"retrievalLayer": "L1", "selectedAttempt" : "FILTERED_VECTOR",
"content": "Embedding title and breadcrumb with content helps recover section semantics.", "fallbackReason" : null,
"score": 0.72, "evidenceStatus" : "supported",
"hitReasons": ["domain_match:+0.15"] "queryHints" : {
} "domains" : [ "rag" ],
], "matched_keywords" : [ "split into multiple chunks", "retrieval context" ],
"contextPack": { "entities" : [ "split into multiple chunks", "retrieval context" ],
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.", "l0_titles" : [ "RAG Chunk Context Reconstruction" ],
"strategy": "top_evidence_blocks", "l0_match_count" : 1
"charBudget": 3500,
"usedChars": 203,
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
"omittedSources": []
}, },
"retrievalTrace": { "attempts" : [ {
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?", "name" : "FILTERED_VECTOR",
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb", "query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
"categoryFilter": "RAG", "categoryFilter" : "rag",
"selectedAttempt": "FILTERED_VECTOR", "candidateCount" : 6,
"fallbackReason": null, "usable" : true,
"evidenceStatus": "supported", "errorMessage" : null,
"queryHints": { "durationMs" : 743,
"domains": ["RAG"], "topScore" : 0.032786883413791656,
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"], "topSimilarity" : 0.7487991750240326
"entities": ["chunk", "breadcrumb"], } ]
"l0_titles": ["RAG Chunk Context Reconstruction"],
"l0_match_count": 1
}, },
"attempts": [ "rerankTrace" : {
{ "items" : [ {
"name": "FILTERED_VECTOR", "finalRank" : 1,
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb", "source" : "rag-chunk-context-reconstruction",
"categoryFilter": "RAG", "baseScore" : 0.7487991750240326,
"candidateCount": 2, "finalScore" : 0.7487991750240326,
"usable": true, "boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
"durationMs": 9, }, {
"topScore": 0.79, "finalRank" : 2,
"topSimilarity": 0.79 "source" : "rag-l0-domain-entity-hint",
} "baseScore" : 0.4840593934059143,
] "finalScore" : 0.4840593934059143,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 3,
"source" : "rag-chunk-context-reconstruction",
"baseScore" : 0.42978107929229736,
"finalScore" : 0.42978107929229736,
"boostReasons" : [ "l0_domain_overlap" ]
}, {
"finalRank" : 4,
"source" : "rag-l0-domain-entity-hint",
"baseScore" : 0.3859822154045105,
"finalScore" : 0.3859822154045105,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
}, },
"rerankTrace": { "evidenceCandidateCount" : 6,
"items": [ "evidenceBlockCount" : 4,
{ "relevanceLevel" : "REFERENCE",
"finalRank": 1, "completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
"source": "rag-chunk-context-reconstruction", "retrievedDomainsThisSession" : null,
"baseScore": 0.79, "message" : null
"finalScore": 1.04,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "rag-breadcrumb-embedding-gap",
"baseScore": 0.72,
"finalScore": 0.87,
"boostReasons": ["domain_match:+0.15"]
}
]
}
} }
} }
@@ -0,0 +1,107 @@
{
"generatedAt": "2026-07-28T06:48:56.439554+00:00",
"baselineReport": "eval/rag-retrieval/cases/golden-cases.json",
"currentReport": "eval/rag-retrieval/cases/golden-cases.json",
"baselineCaseCount": 7,
"currentCaseCount": 7,
"baselinePassRate": 1.0,
"currentPassRate": 0.8571,
"baselineRecallAtK": 1.0,
"currentRecallAtK": 0.8571,
"regressionCount": 6,
"improvementCount": 0,
"changedCount": 3,
"hasRegression": true,
"items": [
{
"type": "REGRESSION",
"scope": "aggregate",
"caseId": null,
"metric": "passRate",
"baselineValue": "1.0",
"currentValue": "0.8571",
"delta": -0.14290000000000003,
"message": "aggregate passRate changed"
},
{
"type": "REGRESSION",
"scope": "aggregate",
"caseId": null,
"metric": "recallAtK",
"baselineValue": "1.0",
"currentValue": "0.8571",
"delta": -0.14290000000000003,
"message": "aggregate recallAtK changed"
},
{
"type": "REGRESSION",
"scope": "aggregate",
"caseId": null,
"metric": "strongHitRate",
"baselineValue": "1.0",
"currentValue": "0.8571",
"delta": -0.14290000000000003,
"message": "aggregate strongHitRate changed"
},
{
"type": "REGRESSION",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "passed",
"baselineValue": "True",
"currentValue": "False",
"delta": -1.0,
"message": "chat-l0-filter-fallback passed changed"
},
{
"type": "REGRESSION",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "hitLevel",
"baselineValue": "strong",
"currentValue": "weak",
"delta": -2.0,
"message": "chat-l0-filter-fallback hitLevel changed"
},
{
"type": "REGRESSION",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "firstExpectedRank",
"baselineValue": "1",
"currentValue": "-",
"delta": null,
"message": "chat-l0-filter-fallback firstExpectedRank changed"
},
{
"type": "CHANGED",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "selectedAttempt",
"baselineValue": "UNFILTERED_VECTOR_RETRY",
"currentValue": "FILTERED_VECTOR",
"delta": null,
"message": "chat-l0-filter-fallback selectedAttempt changed"
},
{
"type": "CHANGED",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "fallbackReason",
"baselineValue": "filtered_vector_low_quality",
"currentValue": "-",
"delta": null,
"message": "chat-l0-filter-fallback fallbackReason changed"
},
{
"type": "CHANGED",
"scope": "case",
"caseId": "chat-l0-filter-fallback",
"metric": "rerankTopSource",
"baselineValue": "rag-l0-filter-fallback",
"currentValue": "rag-l0-filter-decoy",
"delta": null,
"message": "chat-l0-filter-fallback rerankTopSource changed"
}
]
}
@@ -0,0 +1,31 @@
# RAG Retrieval Baseline Diff
Generated at: `2026-07-28T06:48:56.439554+00:00`
## Summary
| Metric | Value |
|---|---:|
| Baseline cases | 7 |
| Current cases | 7 |
| Baseline pass rate | 1.0 |
| Current pass rate | 0.8571 |
| Baseline recall@K | 1.0 |
| Current recall@K | 0.8571 |
| Regressions | 6 |
| Improvements | 0 |
| Changed | 3 |
## Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|---|---|---|---|---|---|---:|---|
| REGRESSION | aggregate | | passRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate passRate changed |
| REGRESSION | aggregate | | recallAtK | 1.0 | 0.8571 | -0.14290000000000003 | aggregate recallAtK changed |
| REGRESSION | aggregate | | strongHitRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate strongHitRate changed |
| REGRESSION | case | chat-l0-filter-fallback | passed | True | False | -1.0 | chat-l0-filter-fallback passed changed |
| REGRESSION | case | chat-l0-filter-fallback | hitLevel | strong | weak | -2.0 | chat-l0-filter-fallback hitLevel changed |
| REGRESSION | case | chat-l0-filter-fallback | firstExpectedRank | 1 | - | | chat-l0-filter-fallback firstExpectedRank changed |
| CHANGED | case | chat-l0-filter-fallback | selectedAttempt | UNFILTERED_VECTOR_RETRY | FILTERED_VECTOR | | chat-l0-filter-fallback selectedAttempt changed |
| CHANGED | case | chat-l0-filter-fallback | fallbackReason | filtered_vector_low_quality | - | | chat-l0-filter-fallback fallbackReason changed |
| CHANGED | case | chat-l0-filter-fallback | rerankTopSource | rag-l0-filter-fallback | rag-l0-filter-decoy | | chat-l0-filter-fallback rerankTopSource changed |
+38 -14
View File
@@ -1,5 +1,5 @@
{ {
"generatedAt": "2026-07-06T13:37:59.726351+00:00", "generatedAt": "2026-07-28T06:55:09.379164+00:00",
"caseFile": "eval/rag-retrieval/cases/golden-cases.json", "caseFile": "eval/rag-retrieval/cases/golden-cases.json",
"fixtureDir": "eval/rag-retrieval/fixtures", "fixtureDir": "eval/rag-retrieval/fixtures",
"aggregate": { "aggregate": {
@@ -28,7 +28,7 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:mysql-connection-pool", "1:mysql-connection-pool",
"2:incident-diagnosis-flow" "2:mysql-connection-pool"
], ],
"matchedKeywords": [ "matchedKeywords": [
"connection pool", "connection pool",
@@ -41,7 +41,7 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"mysql-connection-pool", "mysql-connection-pool",
"incident-diagnosis-flow" "mysql-connection-pool"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "mysql-connection-pool", "rerankTopSource": "mysql-connection-pool",
@@ -57,7 +57,7 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:incident-diagnosis-flow", "1:incident-diagnosis-flow",
"2:rag-chunk-context-reconstruction" "2:incident-diagnosis-flow"
], ],
"matchedKeywords": [ "matchedKeywords": [
"collect evidence", "collect evidence",
@@ -70,7 +70,7 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"incident-diagnosis-flow", "incident-diagnosis-flow",
"rag-chunk-context-reconstruction" "incident-diagnosis-flow"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "incident-diagnosis-flow", "rerankTopSource": "incident-diagnosis-flow",
@@ -86,7 +86,9 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:payment-service-latency", "1:payment-service-latency",
"2:mysql-connection-pool" "2:aiops-alert-scope-control",
"3:payment-service-latency",
"4:aiops-alert-scope-control"
], ],
"matchedKeywords": [ "matchedKeywords": [
"p95 latency", "p95 latency",
@@ -99,7 +101,9 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"payment-service-latency", "payment-service-latency",
"mysql-connection-pool" "aiops-alert-scope-control",
"payment-service-latency",
"aiops-alert-scope-control"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "payment-service-latency", "rerankTopSource": "payment-service-latency",
@@ -114,7 +118,10 @@
"passed": true, "passed": true,
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:aiops-alert-scope-control" "1:aiops-alert-scope-control",
"2:payment-service-latency",
"3:payment-service-latency",
"4:aiops-alert-scope-control"
], ],
"matchedKeywords": [ "matchedKeywords": [
"payload", "payload",
@@ -126,6 +133,9 @@
"fallbackReason": null, "fallbackReason": null,
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"aiops-alert-scope-control",
"payment-service-latency",
"payment-service-latency",
"aiops-alert-scope-control" "aiops-alert-scope-control"
], ],
"omittedSources": [], "omittedSources": [],
@@ -142,7 +152,9 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:rag-chunk-context-reconstruction", "1:rag-chunk-context-reconstruction",
"2:rag-breadcrumb-embedding-gap" "2:rag-l0-domain-entity-hint",
"3:rag-chunk-context-reconstruction",
"4:rag-l0-domain-entity-hint"
], ],
"matchedKeywords": [ "matchedKeywords": [
"neighbor chunk", "neighbor chunk",
@@ -155,7 +167,9 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"rag-chunk-context-reconstruction", "rag-chunk-context-reconstruction",
"rag-breadcrumb-embedding-gap" "rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction",
"rag-l0-domain-entity-hint"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "rag-chunk-context-reconstruction", "rerankTopSource": "rag-chunk-context-reconstruction",
@@ -171,7 +185,9 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:rag-l0-domain-entity-hint", "1:rag-l0-domain-entity-hint",
"2:rag-l0-l1-fusion-ranking" "2:rag-chunk-context-reconstruction",
"3:rag-l0-domain-entity-hint",
"4:rag-chunk-context-reconstruction"
], ],
"matchedKeywords": [ "matchedKeywords": [
"domain detector", "domain detector",
@@ -184,7 +200,9 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"rag-l0-domain-entity-hint", "rag-l0-domain-entity-hint",
"rag-l0-l1-fusion-ranking" "rag-chunk-context-reconstruction",
"rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "rag-l0-domain-entity-hint", "rerankTopSource": "rag-l0-domain-entity-hint",
@@ -200,7 +218,10 @@
"firstExpectedRank": 1, "firstExpectedRank": 1,
"topCandidates": [ "topCandidates": [
"1:rag-l0-filter-fallback", "1:rag-l0-filter-fallback",
"2:rag-l0-domain-entity-hint" "2:rag-l0-filter-decoy",
"3:rag-l0-domain-entity-hint",
"4:rag-l0-domain-entity-hint",
"5:rag-chunk-context-reconstruction"
], ],
"matchedKeywords": [ "matchedKeywords": [
"skip the l0 filter", "skip the l0 filter",
@@ -213,7 +234,10 @@
"evidenceStatus": "supported", "evidenceStatus": "supported",
"includedSources": [ "includedSources": [
"rag-l0-filter-fallback", "rag-l0-filter-fallback",
"rag-l0-domain-entity-hint" "rag-l0-filter-decoy",
"rag-l0-domain-entity-hint",
"rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction"
], ],
"omittedSources": [], "omittedSources": [],
"rerankTopSource": "rag-l0-filter-fallback", "rerankTopSource": "rag-l0-filter-fallback",
+8 -8
View File
@@ -1,6 +1,6 @@
# RAG Retrieval Baseline # RAG Retrieval Baseline
Generated at: `2026-07-06T13:37:59.726351+00:00` Generated at: `2026-07-28T06:55:09.379164+00:00`
## Aggregate ## Aggregate
@@ -24,10 +24,10 @@ Generated at: `2026-07-06T13:37:59.726351+00:00`
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks | | Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|---|---|---|---|---|---|---|---:|---|---| |---|---|---|---|---|---|---|---:|---|---|
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | | | chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | | | chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | | | aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | | | aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | | | chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | | | chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | | | chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-filter-decoy<br>3:rag-l0-domain-entity-hint<br>4:rag-l0-domain-entity-hint<br>5:rag-chunk-context-reconstruction | |
+245
View File
@@ -0,0 +1,245 @@
{
"generatedAt": "2026-07-28T06:48:56.421877+00:00",
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
"fixtureDir": "eval/rag-retrieval/fixtures",
"aggregate": {
"caseCount": 7,
"topK": 5,
"passedCount": 6,
"failedCount": 1,
"passRate": 0.8571,
"lookupResultCaseCount": 7,
"strongHitCount": 6,
"mediumHitCount": 0,
"weakHitCount": 1,
"missCount": 0,
"recallAtK": 0.8571,
"strongHitRate": 0.8571,
"averageFirstHitRank": 1.0
},
"results": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:mysql-connection-pool",
"2:mysql-connection-pool"
],
"matchedKeywords": [
"connection pool",
"max_connections",
"hikaricp"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"mysql-connection-pool",
"mysql-connection-pool"
],
"omittedSources": [],
"rerankTopSource": "mysql-connection-pool",
"failedChecks": []
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:incident-diagnosis-flow",
"2:incident-diagnosis-flow"
],
"matchedKeywords": [
"collect evidence",
"verify",
"remediation"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"incident-diagnosis-flow",
"incident-diagnosis-flow"
],
"omittedSources": [],
"rerankTopSource": "incident-diagnosis-flow",
"failedChecks": []
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:payment-service-latency",
"2:aiops-alert-scope-control",
"3:payment-service-latency",
"4:aiops-alert-scope-control"
],
"matchedKeywords": [
"p95 latency",
"payment-service",
"downstream dependency"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"payment-service-latency",
"aiops-alert-scope-control",
"payment-service-latency",
"aiops-alert-scope-control"
],
"omittedSources": [],
"rerankTopSource": "payment-service-latency",
"failedChecks": []
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:aiops-alert-scope-control",
"2:payment-service-latency",
"3:payment-service-latency",
"4:aiops-alert-scope-control"
],
"matchedKeywords": [
"payload",
"unrelated active alerts",
"scope"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"aiops-alert-scope-control",
"payment-service-latency",
"payment-service-latency",
"aiops-alert-scope-control"
],
"omittedSources": [],
"rerankTopSource": "aiops-alert-scope-control",
"failedChecks": []
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-chunk-context-reconstruction",
"2:rag-l0-domain-entity-hint",
"3:rag-chunk-context-reconstruction",
"4:rag-l0-domain-entity-hint"
],
"matchedKeywords": [
"neighbor chunk",
"same section",
"breadcrumb"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"rag-chunk-context-reconstruction",
"rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction",
"rag-l0-domain-entity-hint"
],
"omittedSources": [],
"rerankTopSource": "rag-chunk-context-reconstruction",
"failedChecks": []
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"dataShape": "lookupResult",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-l0-domain-entity-hint",
"2:rag-chunk-context-reconstruction",
"3:rag-l0-domain-entity-hint",
"4:rag-chunk-context-reconstruction"
],
"matchedKeywords": [
"domain detector",
"entity extractor",
"metadata filter"
],
"breadcrumbMatched": true,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction",
"rag-l0-domain-entity-hint",
"rag-chunk-context-reconstruction"
],
"omittedSources": [],
"rerankTopSource": "rag-l0-domain-entity-hint",
"failedChecks": []
},
{
"caseId": "chat-l0-filter-fallback",
"scenario": "chat",
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"dataShape": "lookupResult",
"hitLevel": "weak",
"passed": false,
"firstExpectedRank": null,
"topCandidates": [
"1:rag-l0-filter-decoy",
"2:rag-l0-filter-decoy"
],
"matchedKeywords": [
"low quality"
],
"breadcrumbMatched": false,
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"includedSources": [
"rag-l0-filter-decoy",
"rag-l0-filter-decoy"
],
"omittedSources": [],
"rerankTopSource": "rag-l0-filter-decoy",
"failedChecks": [
"expected document not found",
"selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR",
"fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none>",
"rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy",
"expected context sources missing: rag-l0-filter-fallback"
]
}
]
}
+33
View File
@@ -0,0 +1,33 @@
# RAG Retrieval Baseline
Generated at: `2026-07-28T06:48:56.421877+00:00`
## Aggregate
| Metric | Value |
|---|---:|
| Cases | 7 |
| Top K | 5 |
| Passed | 6 |
| Failed | 1 |
| Pass rate | 0.8571 |
| LookupResult fixtures | 7 |
| Recall@K | 0.8571 |
| Strong hit rate | 0.8571 |
| Strong hits | 6 |
| Medium hits | 0 |
| Weak hits | 1 |
| Misses | 0 |
| Average first hit rank | 1.0 |
## Cases
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|---|---|---|---|---|---|---|---:|---|---|
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
| chat-l0-filter-fallback | chat | false | weak | FILTERED_VECTOR | | supported | | 1:rag-l0-filter-decoy<br>2:rag-l0-filter-decoy | expected document not found<br>selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR<br>fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none><br>rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy<br>expected context sources missing: rag-l0-filter-fallback |
@@ -0,0 +1,394 @@
# RAG 检索可观测性、审计与 Trace(现行)
**更新日期**:2026-07-28
**状态**:当前可运行
**关联**:`lookup_knowledge`、Harness `ToolBoundary`、`tool_invocation`、`DiagnosisTraceService`、离线 eval
---
## 1. 三层边界
```mermaid
flowchart TB
subgraph A["A. 请求内 Trace"]
LR[LookupResult<br/>retrievalTrace / rerankTrace / relevanceLevel]
end
subgraph B["B. 持久化审计 + Trace API"]
TI[tool_invocation 表]
DT[diagnosis_trace 事件摘要]
API["GET /api/diagnosis/{sessionId}/trace"]
end
subgraph C["C. 质量回归"]
EV[eval/rag-retrieval offline baseline]
end
subgraph agent["Agent 可见(非审计)"]
RT[RagToolResult<br/>evidence + optional relevance_level]
end
LK[LookupKnowledgeTool] --> LR
LR --> PROJ[RagResultProjector]
PROJ --> RT
LR --> BOUND[ToolBoundary audit]
BOUND --> TI
BOUND --> DT
TI --> API
DT --> API
EV -.->|不替代运行时 Trace| LK
```
| 层 | 完善度 | 说明 |
|----|--------|------|
| A 请求内 | 高 | attempt / fallback / quality 齐全 |
| B 持久化 + Trace API | 中高 | RAG 富字段入 `tool_invocation`,经 Trace API 回放 |
| C 离线 eval | 高 | hybrid fixtures 回归 |
**Agent 看到的不是完整 Trace。** 完整检索轨迹在 A/B;Agent 只拿投影后的证据契约。
---
## 2. 端到端:从 lookup 到 Trace API
```mermaid
sequenceDiagram
participant Agent
participant Adapter as RagToolAdapter
participant Bound as ToolBoundary
participant Tool as LookupKnowledgeTool
participant Sink as JpaToolInvocationAuditSink
participant DB as tool_invocation
participant Trace as DiagnosisTraceService
participant API as GET .../trace
Agent->>Adapter: lookup_knowledge(query)
Adapter->>Bound: execute(legacy, projector)
Bound->>Tool: execute(query)
Tool-->>Bound: raw LookupResult JSON
Note over Tool: 内含 retrievalTrace / rerankTrace / evidenceBlocks
Bound->>Bound: project → RagToolResult
Bound->>Sink: AuditEvent + rawResultJson + agentResultJson
Sink->>Sink: RagLookupAuditEnricher
Sink->>DB: 富字段行
Bound-->>Agent: 投影后 agent_result(无完整 trace)
API->>Trace: sessionId + optional runId
Trace->>DB: find tool_invocation by run/session
Trace-->>API: DiagnosisTraceResponse.toolInvocations[]
```
---
## 3. A 层:请求内 Trace(`LookupResult`)
一次成功的 `lookup_knowledge` 内部出口是 **`LookupResult`**(比 Agent 契约更富)。
### 3.1 结构总览
```mermaid
flowchart TB
LR[LookupResult]
LR --> F[found]
LR --> EB[evidenceBlocks[]]
LR --> CP[contextPack]
LR --> RT[retrievalTrace]
LR --> RR[rerankTrace]
LR --> RL[relevanceLevel]
LR --> CH[completenessHint]
LR --> CNT[evidenceCandidateCount / evidenceBlockCount]
RT --> ATT[attempts[]]
RT --> SEL[selectedAttempt]
RT --> FB[fallbackReason]
RT --> HINT[queryHints L0]
```
| 字段 | 含义 |
|------|------|
| `found` | 是否有可用证据块 |
| `evidenceBlocks` | 后处理后的 chunk 级证据(含 evidenceKey、source、content…) |
| `contextPack` | 字符预算打包文本(内部/审计用) |
| `retrievalTrace` | **检索路径 Trace**(见下) |
| `rerankTrace` | 后处理排序/quality 痕迹(现多为保序后的 quality) |
| `relevanceLevel` | PRECISE / REFERENCE / null |
| `completenessHint` | 给模型的天花板提示文案 |
### 3.2 `retrievalTrace`(检索路径)
| 字段 | 含义 |
|------|------|
| `originalQuery` | 原始查询 |
| `rewrittenQuery` | L0/变换后用于检索的 query |
| `categoryFilter` | 首次过滤的 category(可 null) |
| `selectedAttempt` | 最终采用的 attempt 名 |
| `fallbackReason` | 如 `filtered_vector_low_quality`;未降级为 null |
| `evidenceStatus` | 内部:`supported` / `no_evidence` 等 |
| `queryHints` | L0:domains、keywords、entities、l0_match_count… |
| `attempts[]` | 每次检索尝试快照 |
**常见 `selectedAttempt`:**
| 值 | 含义 |
|----|------|
| `FILTERED_VECTOR` | 带 category 的首次检索即采用 |
| `UNFILTERED_VECTOR` | 无 category,直接全库检索 |
| `UNFILTERED_VECTOR_RETRY` | filtered 低质/无证据后去掉 category 重试 |
**单次 `attempts[]` 元素:**
| 字段 | 含义 |
|------|------|
| `name` | attempt 名 |
| `query` | 该次实际检索句 |
| `categoryFilter` | 该次 filter |
| `candidateCount` | 召回候选数 |
| `usable` | 后处理阈值后是否可用 |
| `topScore` / `topSimilarity` | 引擎分 / 归一化 quality(0~1) |
| `durationMs` | 耗时 |
| `errorMessage` | 失败时 |
### 3.3 一次典型路径(含 filter fallback)
```mermaid
flowchart TB
Q[query] --> L0[L0 hint → 可选 categoryFilter]
L0 --> A1[attempt FILTERED_VECTOR]
A1 --> PQ{isLowQuality?}
PQ -->|否| USE1[selectedAttempt = FILTERED_VECTOR]
PQ -->|是| A2[attempt UNFILTERED_VECTOR_RETRY]
A2 --> USE2[selectedAttempt = RETRY<br/>fallbackReason = low_quality / no_evidence]
USE1 --> POST[PostProcess · evidenceBlocks · relevanceLevel]
USE2 --> POST
POST --> LR[LookupResult 完整 Trace]
```
### 3.4 与 Agent 投影的关系
```mermaid
flowchart LR
LR[LookupResult 全量 Trace] --> PROJ[RagResultProjector]
PROJ --> AG[RagToolResult]
AG --> F1[evidence_status]
AG --> F2[evidence excerpt]
AG --> F3[relevance_level 可选]
AG --> F4[truncated / returned_count]
LR -.->|不投影| X1[retrievalTrace]
LR -.->|不投影| X2[rerankTrace]
LR -.->|不投影| X3[raw scores / contextPack 全文]
```
人/系统要「为什么这样检索」→ 看 **A 全量** 或 **B 落库摘要**,不要只看 Agent 字段。
---
## 4. B 层:持久化 + Trace API
### 4.1 写入路径
| 组件 | 职责 |
|------|------|
| `ToolBoundary` | 执行后发 `ToolInvocationAuditEvent`(含 raw LookupResult JSON + agent JSON) |
| `RagLookupAuditEnricher` | 从 LookupResult 抽有界 RAG 字段 |
| `JpaToolInvocationAuditSink` | 写入 `tool_invocation` |
| `TraceAuditEvents.toolInvocation` | 另写一条 diagnosis_trace 摘要事件(不含全文 LookupResult) |
### 4.2 `tool_invocation` 列(RAG)
| 列 | lookup_knowledge | 其它工具 |
|----|------------------|----------|
| `tool_name` | `lookup_knowledge` | 各自工具名 |
| `retrieval_layer` | 通常 `L1` | `HARNESS` |
| `relevance_level` | **PRECISE / REFERENCE / …** | **null**(不再写 evidence_status) |
| `l0_match_count` | queryHints | null |
| `l1_match_count` | evidence 块数等 | null |
| `is_truncated` | 投影 truncated | false |
| `retrieval_details` | JSON `rag_lookup_v1` | 通用 status 元数据 |
| `output_preview` | level/attempt 摘要 | status=… |
| `duration_ms` / `success` | 有 | 有 |
### 4.3 `retrieval_details`(rag_lookup_v1)示例
```json
{
"audit_schema": "rag_lookup_v1",
"search_mode": "hybrid",
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
"fallback_reason": "filtered_vector_low_quality",
"category_filter": "overfilter-decoy",
"evidence_keys": ["doc#chunk-0"],
"sources": ["doc"],
"evidence_candidate_count": 8,
"evidence_block_count": 2,
"l0_hints": { "domains": ["mysql"], "matched_keywords": ["pool"] },
"attempts": [
{
"name": "FILTERED_VECTOR",
"category_filter": "overfilter-decoy",
"candidate_count": 2,
"usable": false,
"top_similarity": 0.3,
"duration_ms": 12
},
{
"name": "UNFILTERED_VECTOR_RETRY",
"candidate_count": 5,
"usable": true,
"top_similarity": 0.9,
"duration_ms": 20
}
],
"truncated": false,
"returned_count": 2,
"evidence_status": "EVIDENCE_FOUND",
"invocation_status": "READY",
"tool_call_id": "call-…"
}
```
**默认不落库:** 原始 query 全文、chunk 正文 excerpt、完整 rerankTrace(体积与隐私)。
### 4.4 Trace API:人怎么读 RAG
**接口:**
```http
GET /api/diagnosis/{sessionId}/trace
GET /api/diagnosis/{sessionId}/trace?runId={runId}
```
**实现:** `DiagnosisTraceController` → `DiagnosisTraceService.getTrace`
按 `sessionId`(可选精确 `runId`)拉 run、steps、**toolInvocations**、摘要等。
**响应中与 RAG 相关的核心块:** `DiagnosisTraceResponse.toolInvocations[]`
| API 字段 | 来源列 | 读法 |
|----------|--------|------|
| `toolName` | `tool_name` | 是否为 `lookup_knowledge` |
| `retrievalLayer` | `retrieval_layer` | L1 / HARNESS |
| `relevanceLevel` | `relevance_level` | RAG 粗相关度(非 evidence_status) |
| `l0MatchCount` / `l1MatchCount` | 同名列 | L0/L1 规模提示 |
| `truncated` | `is_truncated` | 证据是否被投影截断 |
| `outputPreview` | `output_preview` | 一行摘要(level/attempt…) |
| `retrievalDetails` | 解析自 `retrieval_details` | **RAG Trace 主阵地** |
| `retrievalDetailsRaw` | 原始 JSON 字符串 | 调试 |
| `durationMs` / `success` / `errorMessage` | 同名列 | 耗时与成败 |
| `inputParams` | 通常仅 tool_call_id、request_bytes | **不含完整 query**(有意) |
```mermaid
flowchart TB
API["GET /api/diagnosis/{sessionId}/trace"] --> SVC[DiagnosisTraceService]
SVC --> ROW[tool_invocation 行]
ROW --> T1[列: relevanceLevel, L0/L1 count, layer…]
ROW --> T2[retrievalDetails Map]
T2 --> D1[search_mode]
T2 --> D2[selected_attempt / fallback_reason]
T2 --> D3[attempts[] / evidence_keys]
T2 --> D4[evidence_status 契约状态]
```
### 4.5 读 Trace 的推荐顺序(排查「这次知识库怎么检的」)
```mermaid
flowchart TB
S1[找到 toolName=lookup_knowledge 的 invocation] --> S2{success?}
S2 -->|否| E[看 errorMessage / evidence_status]
S2 -->|是| S3[看 retrievalDetails.search_mode]
S3 --> S4[看 selected_attempt + fallback_reason]
S4 --> S5[看 attempts[] 每次 candidate_count / top_similarity / usable]
S5 --> S6[看 evidence_keys / sources]
S6 --> S7[看 relevanceLevel 列]
S7 --> S8[需要原文?看 Agent 侧 evidence 或当时 canonical 存储 · 审计默认无 excerpt]
```
| 现象 | 优先看 |
|------|--------|
| 为何走了 retry | `fallback_reason` + 两次 `attempts` |
| 是否 hybrid | `search_mode` |
| 滤错域 | `category_filter` + L0 domains |
| 相关度档 | 列 `relevanceLevel`(PRECISE/REFERENCE) |
| 返回了哪些块 | `evidence_keys` / `sources`(无正文) |
| Agent 是否被截断 | `truncated` / `returned_count` |
### 4.6 diagnosis_trace 事件 vs tool_invocation 行
| 通道 | 内容 | 用途 |
|------|------|------|
| `tool_invocation` 行 | RAG 富字段完整摘要 | **主审计/回放** |
| `diagnosis_trace` 中 `TOOL_INVOCATION` | tool_call_id、status、字节数、`has_raw_result` 等薄摘要 | 时间线事件,**不含**完整 retrieval_details |
查 RAG 细节以 **`toolInvocations[].retrievalDetails`** 为准。
---
## 5. 与 Agent / Eval 的边界
```mermaid
flowchart LR
subgraph human["人 / 运维 / 评测"]
TRACE[Trace API]
EVAL[Offline eval]
end
subgraph model["模型"]
AGENT[RagToolResult only]
end
TI[(tool_invocation)] --> TRACE
FX[fixtures] --> EVAL
PROJ[Projector] --> AGENT
```
| 消费者 | 能看到 |
|--------|--------|
| Agent | evidence + 可选 relevance_level,无 attempt 细节 |
| Trace API | 落库摘要:mode/attempt/fallback/keys/level… |
| Offline eval | 冻结 fixture 全量 LookupResult(含 trace),与 golden 比对 |
---
## 6. 与旧文档差异
| 旧(archive `retrieval-observability`) | 现 |
|----------------------------------------|-----|
| `vector-store.mode` 多后端 | `search_mode` dense\|hybrid,单一 V2 store |
| sink 理想化未落地 | `RagLookupAuditEnricher` + 列回填 |
| `relevance_level` 混用 evidence_status | **列仅 RAG 等级**;契约状态在 details |
| 未写清 Trace API 读法 | 本文 §4.4–4.5 |
---
## 7. 代码锚点
| 职责 | 类 / 路径 |
|------|-----------|
| 内建 Trace | `LookupKnowledgeTool`、`RetrievalTrace`、`LookupResult` |
| 投影 | `RagResultProjector`、`RagToolResult` |
| 审计事件 | `ToolInvocationAuditEvent`、`ToolBoundary` |
| 富化 | `RagLookupAuditEnricher` |
| 落库 | `JpaToolInvocationAuditSink`、`ToolInvocation` |
| Trace API | `DiagnosisTraceController`、`DiagnosisTraceService`、`DiagnosisTraceResponse.ToolInvocationTrace` |
| 离线回归 | `eval/rag-retrieval/`、`scripts/eval_rag_retrieval.py` |
---
## 8. 已知限制
- 持久化 **不存** 完整 query/excerpt(有意);要正文需 Agent 侧证据或其它存储
- `dedup_reason` 列可能仍为空
- 非 `lookup_knowledge` 工具仍为薄审计
- **历史** `tool_invocation` 行可能仍把 evidence_status 写进 `relevance_level`(旧 sink)
- `diagnosis_trace` 时间线事件不替代 `retrieval_details`
---
## 9. 相关文档
| 文档 | 内容 |
|------|------|
| `mvp/architecture/RAG知识检索架构.md` | 检索主架构 |
| `docs/RAG-Agent如何读relevance_level.md` | Agent 如何读 level |
| `docs/RAG-Hybrid质量分与后处理.md` | quality / 排序闸门 |
| `docs/RAG离线评测-基线设计.md` | 离线评测(非运行时 Trace) |
| `mvp/architecture/session-trace-lifecycle.md` | 会话/run Trace 总览(若存在) |
@@ -168,7 +168,7 @@ sparse_vector -> SPARSE_INVERTED_INDEX + BM25
```yaml ```yaml
retrieval: retrieval:
search: search:
mode: hybrid # dense | hybrid mode: hybrid # dense | hybrid(见 6.0 用途约定)
hybrid: hybrid:
rrf-k: 60 rrf-k: 60
kb-scope: "" kb-scope: ""
@@ -178,7 +178,29 @@ rag:
max-chunks-per-document: 2 max-chunks-per-document: 2
``` ```
### 6.1 dense ### 6.0 模式用途约定(保留双 mode 的原因)
知识库 **只维护一套** dense + BM25 schema 数据(默认 collection `biz`)。
`retrieval.search.mode` 切换的是**同库上的查询算法**,不是两套互斥索引、也不是两套写入路径。
| 模式 | 定位 | 说明 |
|---|---|---|
| **hybrid** | **线上主路径 / 默认** | dense ANN + 服务端 BM25 + RRF;`lookup_knowledge` 正式召回只认此模式 |
| **dense** | **对照 / 评测 / 排障** | 仅 dense ANN,用于和 hybrid 对比召回效果(命中文档/chunk、排名差异等) |
约定:
1. 生产配置保持 `mode: hybrid`;不要把 dense 当成第二套长期并行的线上策略。
2. 需要看「去掉 BM25+RRF 后召回差在哪」时,临时切 `mode: dense`,其它参数(`retrieve-k`、`return-n`、category filter、query 集)尽量固定,再切回 hybrid。
3. hybrid 入库的数据 **完全适用于** dense-only 查询:每条 chunk 都写了 `vector`;dense 模式只是不使用 `sparse_vector` / BM25 子路。
4. 代码里 `@Value` 在配置缺失时的兜底仍可能是 `dense`(历史兼容);**以 `application.yml` 的 hybrid 为准**。若做回归,确认运行配置而不是只看注解默认值。
不建议的用法:
- 按请求/按租户在 dense 与 hybrid 之间当产品功能随意切换(当前也无稳定的 per-call mode 覆盖)。
- 把 dense 模式的相关度表现直接当成 hybrid 的最终质量结论(hybrid 排序信 RRF,后处理分数仍多 L2 兼容,见下节)。
### 6.1 dense(对照基线)
```text ```text
query query
@@ -187,7 +209,9 @@ query
-> topK -> topK
``` ```
### 6.2 hybrid(当前默认) 仅走 `vector` 字段的 L2 ANN。用于基线对比,不作为正式主路径。
### 6.2 hybrid(当前默认 / 主路径)
```text ```text
query query
@@ -197,11 +221,18 @@ query
-> topK fused hits -> topK fused hits
``` ```
阈值兼容: 说明:hybrid **内部**的 dense 子路是融合的一部分,与配置项 mode=dense(整次检索只跑单路 ANN)不是同一概念。
- 后处理仍按 dense 兼容 L2 距离做 `normalizeL2` 分数与后处理(quality 统一,2026-07-28):
- hybrid 命中若能从并行 dense 结果拿到同 id 的 L2,则回填该分数
- 仅 BM25 命中、无 dense 分时,按弱相关处理,避免虚高 `PRECISE` - 一级 scoreLabel 仅 **dense | hybrid**(旧别名 canonicalize)。
- **dense**:score = L2;qualityScore = 1 - clamp(L2)/maxL2Distance。
- **hybrid**:返回序 = RRF 序;qualityScore 由 **本轮 rank 线性映射**(不把 RRF 原分当 L2;不做 dense L2 回填覆盖主分;无 m25_only_* 一级 label)。
- 后处理:**统一**消费 qualityScore;排序主序 = originalRank;**不做** L0 关键词/domain contains 加分改序(重叠仅可写 hitReasons 解释)。
-
elevance_level / category 低质 unfiltered retry:只看 top qualityScore 与阈值。
- 实现:RetrievalScoreNormalizer、KnowledgeEvidencePostProcessor;详见 OpenSpec
ag-quality-score-unify。
### 6.3 category filter 与降级 ### 6.3 category filter 与降级
@@ -266,6 +297,18 @@ Agent 只看到有界 `RagToolResult`:
完整内部结果仍在 `LookupResult` 中,供审计与调试使用。 完整内部结果仍在 `LookupResult` 中,供审计与调试使用。
### 9.1 Trace 与审计(现行入口)
请求内 `retrievalTrace` / 落库 `tool_invocation` / Trace API 读法见:
**[RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)**
要点:
- Agent **看不到**完整 retrievalTrace;人通过 `GET /api/diagnosis/{sessionId}/trace` 的 `toolInvocations[].retrievalDetails` 回放。
- `relevance_level` 列存 RAG 等级(PRECISE/REFERENCE);`evidence_status` 在 details JSON。
- 默认审计不落原始 query 全文与 excerpt 正文。
## 10. 写入与重建 ## 10. 写入与重建
### 10.1 日常写入 ### 10.1 日常写入
@@ -326,3 +369,8 @@ POST /api/knowledge/rebuild-hybrid?confirm=REBUILD
- L0 关键词匹配仍较粗,只作 hint,不作主召回 - L0 关键词匹配仍较粗,只作 hint,不作主召回
- 尚未做邻块上下文自动扩展、cross-encoder rerank、真 query rewrite - 尚未做邻块上下文自动扩展、cross-encoder rerank、真 query rewrite
- `totalVectors` 统计接口仍可能返回 0,不代表 collection 为空;以 rebuild/init 结果与检索命中为准 - `totalVectors` 统计接口仍可能返回 0,不代表 collection 为空;以 rebuild/init 结果与检索命中为准
- `retrieval.search.mode=dense` 仅作召回对照,不是第二套主路径
- 同一 hybrid schema 数据可被 dense / hybrid 两种查询复用;从纯旧 dense-only collection 升级必须 rebuild
- hybrid 质量闸门优先用 `denseDistance` 绝对 L2;无 dense 时 rank 回退;排序仍跟 RRF
- 后处理不再用 L0 关键词 boost 改序;词面信号以库内 BM25+RRF 为准
- Trace/审计细节与限制见 [RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)
+5 -3
View File
@@ -12,8 +12,10 @@
| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 | | [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 | | [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 |
| [diagnosis-information-gain-stop-architecture.md](diagnosis-information-gain-stop-architecture.md) | 已实施的信息增益评价、Harness 饱和检测、Draft 合同失败降级与证据不足停止设计 | | [diagnosis-information-gain-stop-architecture.md](diagnosis-information-gain-stop-architecture.md) | 已实施的信息增益评价、Harness 饱和检测、Draft 合同失败降级与证据不足停止设计 |
| [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md) | 当前 `lookup_knowledge` 检索:MilvusClientV2 dense+BM25 hybrid、chunk 证据身份、重建运维 | | [RAG知识检索架构.md](RAG知识检索架构.md) | 当前 `lookup_knowledge` 检索:MilvusClientV2 dense+BM25 hybrid、chunk 证据身份、重建运维 |
| [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) | RAG Trace / 审计:请求内 retrievalTrace、tool_invocation 富字段、Trace API 读法 |
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。其中旧 RAG 描述(Spring AI VectorStore 主路径 + Milvus SDK fallback)已被当前 hybrid 实现取代,请以 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md) 为准。 2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。其中旧 RAG 描述(Spring AI VectorStore 主路径 + Milvus SDK fallback)已被当前 hybrid 实现取代,请以 [RAG知识检索架构.md](RAG知识检索架构.md) 为准。检索可观测与 Trace 以 [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) 为准(勿再依赖 archive 内旧 retrieval-observability)。
当前普通 Trace 与 Provider reasoning 审计使用独立存储和独立接口。Reasoning 访问控制、保留期限、加密要求以及真实 Provider/V015 验证仍由 ISS-015 跟踪,不能把“数据已分表”理解为“治理已经完成”。 当前普通 Trace 与 LLM 步骤审计(`agent_reasoning_audit`:`reasoning_content` + `assistant_text`)使用独立存储和独立接口。
DeepSeek thinking 捕获路径与 V015–V017 字段已 live 验证(2026-07-28)。Reasoning 访问控制、保留期限、加密要求仍由 ISS-015 跟踪,不能把“数据已分表 + 能抓到 thinking”理解为“治理已经完成”。
+30 -8
View File
@@ -56,14 +56,36 @@ sequenceDiagram
PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。 PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。
## 5. Provider Reasoning 审计 ## 5. Provider Reasoning 与 Assistant 正文审计
`HarnessAgentAuditHook` 在每次模型步骤结束后检查 `AssistantMessage` metadata。当前识别 `reasoning_content`、`reasoningContent`、`reasoning` 和 `thinking`,但只接受 Provider 实际返回的非空文本: `HarnessAgentAuditHook` 在每次模型步骤结束后写入独立表 `agent_reasoning_audit`,**同时**尝试捕获:
- 有内容时写入 `agent_reasoning_audit`,单条最多保留 32000 个字符,并记录 UTF-8 `content_bytes`。 | 列 | 含义 |
- 无内容时写入 `reasoning_available=false`、`reasoning_content=NULL`、`content_bytes=0`,不得根据最终回答反推或生成 reasoning。 |---|---|
- `agent_step.thought` 始终为空;步骤表只记录 message count、roles、是否有文本、Tool names、reasoning availability 和字节数等 metadata。 | `reasoning_content` | Provider thinking / CoT |
- 普通 `diagnosis_trace_event` 的 `AGENT_MODEL_STEP` 只记录 reasoning availability/bytes,不保存 reasoning 原文。 | `assistant_text` | 本步 assistant 可见正文,和/或 tool-call **计划**(不含 tool 结果) |
- Reasoning 只用于受限审计,不进入 Agent 后续上下文,不参与 EvidenceGuard、SemanticGuard 或 Release Policy 的事实判断。 | `content_source` | `PROVIDER_REASONING+ASSISTANT_TEXT` 等组合标记 |
| `content_bytes` | 两列截断后 UTF-8 字节合计 |
当前查询隔离已经实现,完整访问治理和真实 Provider 行为验证仍属于 ISS-015。 ### 捕获路径(DeepSeek 生产)
当前 Chat 为 Spring AI 原生 `DeepSeekChatModel`(`deepseek-v4-flash`)。API 返回的 `message.reasoning_content` 被映射到 **`DeepSeekAssistantMessage.getReasoningContent()`**,而不是普通 `AssistantMessage.metadata`。Hook 优先读该专用字段,再反射 `getReasoningContent()`,最后才回退 metadata 键(`reasoning_content` / `thinking` 等)。
只接受 Provider 实际返回的非空文本;不得根据最终回答反推或生成伪 reasoning。
### 边界
- 单字段最多保留 32000 字符;tool **结果**只在 `tool_invocation`。
- `agent_step.model_*` 仍为有界 metadata(含 `reasoning_available` / bytes / `content_source`)。
- `agent_step.thought` 为兼容镜像:优先 reasoning,否则 assistant 正文;完整双字段以 `agent_reasoning_audit` 为准。
- 普通 Timeline 的 `AGENT_MODEL_STEP` 不保存 reasoning/assistant 原文。
- Reasoning / assistant 审计原文只用于受限审计 API,不进入 Agent 后续上下文,不参与 EvidenceGuard、SemanticGuard 或 Release Policy 的事实判断。
### 运行级结论读出
Run 结束时 `JpaChatRunStore` 从安全发布 JSON 提取 `diagnosis_run.conclusion`(与 `query` 并列),便于 Trace/DB 直接读结论;它不是 Provider thinking。
### 验证与治理
- **已 live 验证(2026-07-28)**:DeepSeek thinking 模式下 `reasoning_available=true`,且 reasoning 与 assistant 可同时非空。
- 查询隔离已实现;访问控制、保留期限、加密等完整治理仍属 ISS-015。
+11 -11
View File
@@ -7,7 +7,7 @@
SuperBizAgent 是面向故障诊断的可追踪 Agent 应用。当前系统只保留一个拥有 Tool loop 的 `Diagnosis Agent`;Harness 负责确定性的预算、取消、工具边界、证据验真、语义审查和安全发布。 SuperBizAgent 是面向故障诊断的可追踪 Agent 应用。当前系统只保留一个拥有 Tool loop 的 `Diagnosis Agent`;Harness 负责确定性的预算、取消、工具边界、证据验真、语义审查和安全发布。
知识检索当前为显式 `lookup_knowledge` 工具 + 单一 MilvusClientV2 后端(dense / dense+BM25 hybrid)。详细链路见 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md)。 知识检索当前为显式 `lookup_knowledge` 工具 + 单一 MilvusClientV2 后端(dense / dense+BM25 hybrid)。详细链路见 [RAG知识检索架构.md](RAG知识检索架构.md)。
## 2. 分层 ## 2. 分层
@@ -86,7 +86,7 @@ Agent
- 证据按 chunk 级 `evidenceKey` 去重;Agent 侧 `document_id` 为 chunk 级身份。 - 证据按 chunk 级 `evidenceKey` 去重;Agent 侧 `document_id` 为 chunk 级身份。
- 知识库全量重建:`python scripts/rebuild_hybrid_knowledge.py --confirm REBUILD`,默认操作 collection `biz`。 - 知识库全量重建:`python scripts/rebuild_hybrid_knowledge.py --confirm REBUILD`,默认操作 collection `biz`。
完整 schema、模式、重建与历史差异见 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md)。 完整 schema、模式、重建与历史差异见 [RAG知识检索架构.md](RAG知识检索架构.md)。
## 5. Trace 与持久化 ## 5. Trace 与持久化
@@ -100,20 +100,20 @@ chat_session(sessionId)
``` ```
- `chat_session` 是 JPA Run 目录与多轮 metadata,不保存完整对话历史。 - `chat_session` 是 JPA Run 目录与多轮 metadata,不保存完整对话历史。
- `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果和预算汇总真理源。 - `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果、预算汇总真理源;另含与 `query` 并列的提取字段 `conclusion`(业务结论读出,非 thinking)。
- `agent_step` 只保存模型步骤 metadata,不保存 Prompt、消息正文、模型正文或 Thought。 - `agent_step` 保存模型步骤有界 metadata;`thought` 可为 reasoning/assistant 的兼容镜像,完整双字段不在此表。
- `tool_invocation` 只保存 Tool durable audit metadata;完整调用由 Redis canonical store 短期保存。 - `tool_invocation` 只保存 Tool durable audit metadata;完整调用由 Redis canonical store 短期保存。
- `diagnosis_trace_event` 是追加式统一 Timeline,记录 Run、Routing、Agent、Tool、Evidence、Semantic 和 Release 生命周期事件;`details` 只能保存有界安全 metadata。 - `diagnosis_trace_event` 是追加式统一 Timeline,记录 Run、Routing、Agent、Tool、Evidence、Semantic 和 Release 生命周期事件;`details` 只能保存有界安全 metadata。
- `agent_reasoning_audit` 与普通 Trace 分表,只保存 Provider 实际返回的 reasoning 或明确的 unavailable 记录;reasoning 不属于事实证据。 - `agent_reasoning_audit` 与普通 Trace 分表,按模型步骤保存 Provider `reasoning_content` 与 `assistant_text`(及 `content_source`),或明确的 unavailable;二者均不属于事实证据。
## 6. 公开 API ## 6. 公开 API
当前诊断执行入口只有 `POST /api/chat`。诊断审计读取分为: 当前诊断执行入口只有 `POST /api/chat`。诊断审计读取分为:
- `GET /api/diagnosis/{sessionId}/trace?runId={runId}`:普通 Trace,返回 Run、步骤、Tool metadata 和统一 Timeline,不返回 reasoning 原文。 - `GET /api/diagnosis/{sessionId}/trace?runId={runId}`:普通 Trace,返回 Run(含 `query`/`conclusion`)、步骤、Tool metadata 和统一 Timeline,**不**返回 reasoning/assistant 原文。
- `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`:独立 reasoning 审计读取,`runId` 必填并校验其属于 path `sessionId`。 - `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`:独立 LLM 步骤审计读取(`reasoningContent` + `assistantText`),`runId` 必填并校验其属于 path `sessionId`。
Reasoning endpoint 是敏感审计面,不属于普通业务 API。当前已完成数据和查询隔离;认证授权、保留期限、加密要求及真实 Provider 验证仍由 ISS-015 收敛。Feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。 Reasoning endpoint 是敏感审计面,不属于普通业务 API。数据和查询隔离、DeepSeek thinking 捕获路径已 live 验证;认证授权、保留期限、加密要求仍由 ISS-015 收敛。Feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。
与知识检索相关的独立 API: 与知识检索相关的独立 API:
@@ -124,9 +124,9 @@ Reasoning endpoint 是敏感审计面,不属于普通业务 API。当前已完
## 7. 安全边界 ## 7. 安全边界
- 普通 SSE、Trace、Evidence Snapshot、业务结果和应用日志不输出或保存 reasoning 原文。 - 普通 SSE、普通 Trace 的 steps/timeline、Evidence Snapshot、业务结果和应用日志不输出或保存 reasoning / assistant 审计原文。
- 仅当 Provider 在模型 metadata 中实际返回 reasoning 时,审计 Hook 才将其截断后写入独立表;Provider 未返回时不得伪造。 - 审计 Hook 仅当 Provider **实际返回** thinking(DeepSeek:`DeepSeekAssistantMessage.reasoningContent`;其它:metadata 键)时写入 `reasoning_content`;未返回时不得伪造。`assistant_text` 来自本步 assistant 可见输出与 tool-call 计划,不含 tool 结果。
- Reasoning 不能作为事实证据,也不能绕过 EvidenceGuard 或 SemanticGuard。 - Reasoning / assistant 审计原文不能作为事实证据,也不能绕过 EvidenceGuard 或 SemanticGuard。
- 不向 Agent 暴露 Redis、canonical key、完整 Tool 请求/响应或数据库凭据。 - 不向 Agent 暴露 Redis、canonical key、完整 Tool 请求/响应或数据库凭据。
- EvidenceGuard 只接受当前 Run 的 READY canonical invocation。 - EvidenceGuard 只接受当前 Run 的 READY canonical invocation。
- SemanticGuard 无 Tool、无记忆、无回调主 Agent 能力。 - SemanticGuard 无 Tool、无记忆、无回调主 Agent 能力。
+7 -6
View File
@@ -37,7 +37,8 @@ disconnect、timeout 与 send failure 通过同一个 `ChatRunControl` 请求取
Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当前 Run EvidenceGuard 验真。 Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当前 Run EvidenceGuard 验真。
普通 Trace 不读取 `agent_reasoning_audit`,也不返回 reasoning 原文。Agent 模型步骤和 Timeline 只暴露 `reasoning_available`、`reasoning_bytes` 等有界 metadata。 普通 Trace 不读取 `agent_reasoning_audit` 的 **原文**。Agent 模型步骤和 Timeline 只暴露 `reasoning_available`、`reasoning_bytes` / `assistant_bytes`、`content_source` 等有界 metadata。
普通 Trace 的 `run` 可返回 `query` 与提取后的 `conclusion`(业务结论读出,非 thinking)。
## 4. PreviousTurn ## 4. PreviousTurn
@@ -45,11 +46,11 @@ Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当
## 5. Reasoning 审计查询 ## 5. Reasoning 审计查询
`GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}` 提供独立 reasoning 审计读取: `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}` 提供独立 LLM 步骤审计读取:
- `runId` 必填,服务端先验证 Run 存在且属于 path `sessionId`,禁止跨 Session 串读。 - `runId` 必填,服务端先验证 Run 存在且属于 path `sessionId`,禁止跨 Session 串读。
- 结果按 `step_index` 返回 Agent、reasoning availability、受限原文、字节数和创建时间。 - 结果按 `step_index` 返回:`reasoningAvailable`、`reasoningContent`、`assistantText`、`contentSource`、`contentBytes`、创建时间等。
- Provider 未返回 reasoning 时仍保留 unavailable 记录,以区分“没有返回”与“审计遗漏”。 - Provider 未返回 reasoning 时仍保留记录(`reasoningAvailable=false`,可仍有 `assistantText`),以区分“没有 thinking”与“审计遗漏”。
- Reasoning 数据不回流到 PreviousTurn,不进入普通 Trace、SSE、Evidence Snapshot 或发布结果。 - Reasoning / assistant 审计原文不回流到 PreviousTurn,不进入普通 Trace 的 steps/timeline 原文、SSE、Evidence Snapshot 或发布结果。
该端点属于敏感审计面。当前完成了分表、独立查询和归属校验;身份认证、权限模型、保留期限、加密及真实 Provider/V015 验证尚未完成,由 ISS-015 阶段 3 收敛。 该端点属于敏感审计面。分表、独立查询、归属校验,以及 DeepSeek 真实 thinking 捕获路径(`DeepSeekAssistantMessage.reasoningContent`)与 V015–V017 迁移已验证;身份认证、权限模型、保留期限、加密仍由 ISS-015 阶段 3 收敛。
+8 -5
View File
@@ -2,11 +2,14 @@
- [ ] SSE metadata 的 `session_id`、`run_id` 非空且与数据库完全一致。 - [ ] SSE metadata 的 `session_id`、`run_id` 非空且与数据库完全一致。
- [ ] `diagnosis_run.intent=DIAGNOSIS`,status/release_outcome 与 done outcome 一致。 - [ ] `diagnosis_run.intent=DIAGNOSIS`,status/release_outcome 与 done outcome 一致。
- [ ] `diagnosis_run.conclusion` 与发布结论一致(SUCCESS 时非空短结论;FALLBACK 可为 type/message 摘要)。
- [ ] `agent_step.agent_name` 只出现 `diagnosis_agent`。 - [ ] `agent_step.agent_name` 只出现 `diagnosis_agent`。
- [ ] AgentStep model_input/model_output 只含 metadata,thought 为空。 - [ ] AgentStep `model_input`/`model_output` 只含有界 metadata(可含 `reasoning_available`/`content_source`/bytes);**不含** reasoning/assistant 原文。
- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内。 - [ ] `agent_step.thought` 若非空,应为 reasoning 或 assistant 的兼容镜像,不得替代 `/trace/reasoning` 双字段验收。
- [ ] ToolInvocation input/output/retrieval details 不含 SQL、日志 query、raw response 或 evidence body。 - [ ] `/trace/reasoning`:DeepSeek thinking 开启时应有 `reasoningAvailable=true` 且 `reasoningContent` 非空;`assistantText` 有正文和/或 tool-call 计划;`contentSource` 合理;**无** tool 结果体。
- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内;`step_id` 能挂到对应 `agent_step.id`。
- [ ] ToolInvocation input/output/retrieval details 不含 SQL 全文、raw response 或 evidence body 泄露(query 等允许的有界字段除外)。
- [ ] Run 的模型/Tool/Token/字节预算均未超过集中配置。 - [ ] Run 的模型/Tool/Token/字节预算均未超过集中配置。
- [ ] content 只出现一次并来自 Release Policy;failure 与 content 互斥。 - [ ] content 只出现一次并来自 Release Policy;failure 与 content 互斥。
- [ ] `logs/application.log` 不含 Prompt、Thought、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。 - [ ] `logs/application.log` 不含 Prompt、完整 reasoning 原文、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。
- [ ] query_logs 标记为 Mock;query_mysql 只使用隔离只读 datasource contract。 - [ ] query_logs 标记为 Mock;query_mysql 仅在配置了隔离 datasource 时注册。
@@ -40,8 +40,8 @@ Fallback 必须在不暴露 Prompt、原始 Draft、原始 Tool 载荷和内部
|---|---|---|---| |---|---|---|---|
| 阶段 1:Diagnosis Agent 硬停止策略 | 已完成 | ISS-016 已实现 Tool Scope 归一化与去重、`GAINED/NO_GAIN` 信息增益、连续 `NO_GAIN` 饱和停止、连续协议错误 `PROGRESS_PROTOCOL_VIOLATED` 受控停止、可修正协议 observation、受控 Fallback、回归测试和真实 E2E | 无;阶段 1 停止策略已收口 | | 阶段 1:Diagnosis Agent 硬停止策略 | 已完成 | ISS-016 已实现 Tool Scope 归一化与去重、`GAINED/NO_GAIN` 信息增益、连续 `NO_GAIN` 饱和停止、连续协议错误 `PROGRESS_PROTOCOL_VIOLATED` 受控停止、可修正协议 observation、受控 Fallback、回归测试和真实 E2E | 无;阶段 1 停止策略已收口 |
| 阶段 2:Evidence Repair Schema | 未完成 | Repair 仍保持无 Tool、有限重试、失败后安全降级;Diagnosis Agent 最终输出已使用真实 `DiagnosisDraft` Schema | Evidence Repair 自身尚未注入真实 `DiagnosisDraft` JSON Schema,仍主要依赖 Prompt 文本约束和严格解析 | | 阶段 2:Evidence Repair Schema | 未完成 | Repair 仍保持无 Tool、有限重试、失败后安全降级;Diagnosis Agent 最终输出已使用真实 `DiagnosisDraft` Schema | Evidence Repair 自身尚未注入真实 `DiagnosisDraft` JSON Schema,仍主要依赖 Prompt 文本约束和严格解析 |
| 阶段 3:Reasoning 审计验证与治理 | 部分完成 | V015、独立 `agent_reasoning_audit`、审计 Hook、`reasoning_available=false`、受限查询接口和 `sessionId + runId` 归属校验已实现 | 真实 Provider reasoning metadata 行为、V015 真实迁移验收、访问控制、保留期限和加密要求尚未收敛 | | 阶段 3:Reasoning 审计验证与治理 | 大部分完成 | V015–V017;`reasoning_content` + `assistant_text` + `content_source`;Hook 优先读 `DeepSeekAssistantMessage.getReasoningContent()`;受限查询返回双字段;`sessionId + runId` 归属校验;2026-07-28 DeepSeek live E2E 两步均 `reasoning_available=true` 且双字段非空;V016 `diagnosis_run.conclusion` | 访问控制、保留期限、加密;非 DeepSeek Provider 回归;reasoning unavailable 专项场景归档 |
| 阶段 4:Fallback 信息质量与最终 E2E | 大部分完成 | 已实现有界 `observed_facts`、`validation_issues`、证据范围/缺口和差异化 Fallback;已完成 Diagnosis SUCCESS、信息不足 FALLBACK、Trace/Tool/Token 对账 E2E | reasoning unavailable 场景及 Reasoning Audit 跨表精确核验尚未完成;全部失败路径仍需最终综合验收 | | 阶段 4:Fallback 信息质量与最终 E2E | 大部分完成 | 已实现有界 `observed_facts`、`validation_issues`、证据范围/缺口和差异化 Fallback;已完成 Diagnosis SUCCESS、信息不足 FALLBACK、Trace/Tool/Token 对账 E2E;SUCCESS 路径已跨表核验 reasoning 双字段与 run.conclusion | reasoning unavailable 场景归档;全部失败路径仍需最终综合验收 |
本表是当前进度事实,下面各阶段条目仍保留为完整目标。ISS-016 完成的是阶段 1 和阶段 4 的主体能力,并新增模型 Token 与 Tool 拒绝审计;它不替代阶段 2 的 Repair Schema,也不代表阶段 3 的 Reasoning 治理已经完成。 本表是当前进度事实,下面各阶段条目仍保留为完整目标。ISS-016 完成的是阶段 1 和阶段 4 的主体能力,并新增模型 Token 与 Tool 拒绝审计;它不替代阶段 2 的 Repair Schema,也不代表阶段 3 的 Reasoning 治理已经完成。
@@ -70,18 +70,20 @@ Fallback 必须在不暴露 Prompt、原始 Draft、原始 Tool 载荷和内部
- 保持 Repair 无 Tool、最多一次、失败即安全 Fallback 的现有边界。 - 保持 Repair 无 Tool、最多一次、失败即安全 Fallback 的现有边界。
- 覆盖合法修复、Schema 非法、解析失败和二次 EvidenceGuard 失败。 - 覆盖合法修复、Schema 非法、解析失败和二次 EvidenceGuard 失败。
### 阶段 3:Reasoning 审计验证与治理(部分完成) ### 阶段 3:Reasoning 审计验证与治理(大部分完成)
- 使用真实 Provider 验证 reasoning metadata 的键和返回行为。 - [x] 使用真实 DeepSeek Provider 验证 thinking 返回路径:专用字段 `DeepSeekAssistantMessage.reasoningContent`(非 metadata)。
- Provider 不返回 reasoning 时写入 `reasoning_available=false`,不得伪造内容。 - [x] 同表同时落库 `assistant_text`(正文 / tool-call 计划)与 `content_source`(V017)。
- 验证 V015 数据库迁移和 `sessionId + runId` 精确 reasoning 查询。 - [x] Provider 不返回 reasoning 时写入 `reasoning_available=false`,不得伪造内容(单测覆盖;live unavailable 场景可再归档)。
- 明确 reasoning 审计接口的访问控制、保留期限和加密要求。 - [x] 验证 V015–V017 迁移与 `sessionId + runId` 精确 reasoning 查询(含双字段 API)。
- [ ] 明确 reasoning 审计接口的访问控制、保留期限和加密要求。
### 阶段 4:Fallback 信息质量与最终 E2E(大部分完成) ### 阶段 4:Fallback 信息质量与最终 E2E(大部分完成)
- 工具成功但无法构造 verified snapshot 时,返回有界 `observed_facts` 和 `validation_issues`。 - 工具成功但无法构造 verified snapshot 时,返回有界 `observed_facts` 和 `validation_issues`。
- 确保普通 Trace 只记录 `reasoning_available/reasoning_bytes`,不返回 reasoning 原文。 - 确保普通 Trace 只记录 `reasoning_available/reasoning_bytes` 等有界 metadata,不返回 reasoning/assistant 原文;`run.conclusion` 可为业务结论读出。
- 运行至少一个 Diagnosis SUCCESS、一个信息化 FALLBACK 和一个 reasoning unavailable 的真实 E2E。 - [x] Diagnosis SUCCESS 真实 E2E:SSE + Trace + Reasoning 双字段 + ToolInvocation + `diagnosis_run.conclusion` 对账(2026-07-28)。
- [ ] 再归档一个 reasoning unavailable 与信息化 FALLBACK 的跨表精确核验样本。
- 按 exact `sessionId + runId` 核对 SSE、Trace、Reasoning Audit、ToolInvocation 和 Run 终态。 - 按 exact `sessionId + runId` 核对 SSE、Trace、Reasoning Audit、ToolInvocation 和 Run 终态。
## 6. 验收标准 ## 6. 验收标准
@@ -1,13 +1,18 @@
# Agent 推理审计表:agent_reasoning_audit # Agent 推理审计表:agent_reasoning_audit
**状态**:当前独立敏感审计表;治理待 ISS-015 收敛 **状态**:当前独立敏感审计表;治理待 ISS-015 收敛
**来源**:`V015__create_agent_reasoning_audit.sql`、`AgentReasoningAudit` **来源**:`V015__create_agent_reasoning_audit.sql`、`V017__add_content_source_to_agent_reasoning_audit.sql`、`AgentReasoningAudit`、`HarnessAgentAuditHook`
## 定位 ## 定位
`agent_reasoning_audit` 保存模型 Provider 在 Agent 步骤 metadata 中实际返回的 reasoning 内容。它与普通 Trace、AgentStep、Evidence Snapshot 和业务发布结果物理分离,不能作为事实证据或诊断结论来源。 `agent_reasoning_audit` 按 **Diagnosis Agent 每个模型步骤** 保存两类 LLM 面向文本,并与普通 Trace、AgentStep、Evidence Snapshot、业务发布结果物理分离:
Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没有返回”误判为“审计链路漏写”。系统不得从最终回答、Tool 调用或其他字段生成伪 reasoning。 1. **Provider reasoning / thinking**(`reasoning_content`)
2. **Assistant 可见正文**(`assistant_text`:最终 prose 和/或 tool-call **计划**)
它不能作为事实证据或诊断结论来源。Tool **结果**载荷不进本表(见 `tool_invocation`)。
Provider 未返回 reasoning 时仍写入记录(`reasoning_available=false`),防止把“没有返回”误判为“审计链路漏写”。系统不得从最终回答、Tool 调用或其他字段生成伪 reasoning。
## 字段 ## 字段
@@ -19,8 +24,10 @@ Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没
| `step_index` | INT | 是 | Agent 模型步骤序号 | | `step_index` | INT | 是 | Agent 模型步骤序号 |
| `agent_name` | VARCHAR(64) | 是 | Agent 身份,当前为 `diagnosis_agent` | | `agent_name` | VARCHAR(64) | 是 | Agent 身份,当前为 `diagnosis_agent` |
| `reasoning_available` | BOOLEAN | 是 | Provider 是否实际返回非空 reasoning | | `reasoning_available` | BOOLEAN | 是 | Provider 是否实际返回非空 reasoning |
| `reasoning_content` | LONGTEXT | 否 | Provider reasoning 原文;Hook 当前最多保留 32000 个字符 | | `reasoning_content` | LONGTEXT | 否 | Provider CoT / thinking 原文;Hook 单字段最多保留 32000 字符 |
| `content_bytes` | INT | 是 | 截断后 reasoning 的 UTF-8 字节数;unavailable 时为 0 | | `assistant_text` | LONGTEXT | 否 | 本步 assistant 可见正文;若有 tool_call,可附带 `tool_calls:` 计划预览(**不含** tool 结果) |
| `content_source` | VARCHAR(64) | 否 | 本步写入摘要:`PROVIDER_REASONING+ASSISTANT_TEXT` / `PROVIDER_REASONING` / `ASSISTANT_TEXT` / `TOOL_CALL_PLAN` / `NONE` |
| `content_bytes` | INT | 是 | 截断后 `reasoning_content` + `assistant_text` 的 UTF-8 字节合计 |
| `created_at` | DATETIME | 是 | 创建时间,默认当前时间 | | `created_at` | DATETIME | 是 | 创建时间,默认当前时间 |
## 索引 ## 索引
@@ -36,15 +43,36 @@ Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没
- `session_id` 逻辑关联 `chat_session.session_id`。 - `session_id` 逻辑关联 `chat_session.session_id`。
- 通过 `session_id + run_id + step_index` 与 `agent_step` 逻辑对应,不建立数据库外键。 - 通过 `session_id + run_id + step_index` 与 `agent_step` 逻辑对应,不建立数据库外键。
## 写入规则 ## 写入规则(HarnessAgentAuditHook)
- Provider 返回非空 reasoning:`reasoning_available=true`,保存截断后的原文及实际 UTF-8 字节数。 ### 捕获来源(按优先级)
- Provider 未返回 reasoning:`reasoning_available=false`,`reasoning_content=NULL`,`content_bytes=0`。
- `agent_step.thought` 继续保持为空;普通 Trace 仅保留 availability/bytes metadata。 1. **DeepSeek 生产主路径**:`DeepSeekAssistantMessage.getReasoningContent()`
- Reasoning 不进入 SSE、PreviousTurn、Evidence Snapshot、发布结果或应用日志。 (Spring AI 将 API 的 `message.reasoning_content` 映射到该专用字段,**不是** `AssistantMessage.metadata`)
2. 反射 `getReasoningContent()`(兼容序列化/子类)
3. metadata 键兜底:`reasoning_content` / `reasoningContent` / `reasoning` / `thinking` / `reasoning_text` / `reasoningText`(单测与其它 Provider)
`assistant_text` 来自 `AssistantMessage.getText()`;若本步有 tool_calls,追加有界的 `tool_calls:` 名称与 args 预览。
### 落库语义
- 有 Provider reasoning:`reasoning_available=true`,写入截断后的 `reasoning_content`。
- 无 Provider reasoning:`reasoning_available=false`,`reasoning_content=NULL`;仍可写入 `assistant_text`。
- `content_source` 反映本步实际写入组合;两者皆空则为 `NONE`。
- `content_bytes` 为两列截断后文本 UTF-8 字节之和。
- **禁止**把 tool 执行结果、raw Tool response、用户 Prompt 全文写入本表。
### 与其它表的边界
- 普通 Trace / Timeline 只暴露 `reasoning_available`、`reasoning_bytes` / `assistant_bytes`、`content_source` 等有界 metadata,不返回 reasoning/assistant 原文。
- `agent_step.model_input` / `model_output` 仍为有界 metadata。
- `agent_step.thought` 可作为兼容镜像:优先 Provider reasoning,否则 assistant 正文;完整双字段以本表为准。
- Reasoning / assistant 审计原文不进入 SSE、PreviousTurn、Evidence Snapshot、发布结果或应用日志。
## 查询与治理 ## 查询与治理
当前独立接口为 `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`。`runId` 必填,服务端校验其属于 path `sessionId`,并按 `step_index` 返回记录。 当前独立接口为 `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`。`runId` 必填,服务端校验其属于 path `sessionId`,并按 `step_index` 返回记录(含 `reasoningContent`、`assistantText`、`contentSource`)。
分表和归属校验已经实现,但不能等同于完整安全治理。身份认证、角色授权、保留/删除期限、静态与传输加密以及真实 Provider/V015 验证仍由 ISS-015 阶段 3 跟踪;治理完成前不应将该接口暴露给普通业务用户。 **已验证(2026-07-28 live E2E,DeepSeek `deepseek-v4-flash`)**:thinking 模式下两步模型调用均可得到 `reasoning_available=true` 且 `content_source=PROVIDER_REASONING+ASSISTANT_TEXT`。
分表和归属校验已经实现,但不能等同于完整安全治理。身份认证、角色授权、保留/删除期限、静态与传输加密仍由 ISS-015 跟踪;治理完成前不应将该接口暴露给普通业务用户。
+11 -7
View File
@@ -1,11 +1,13 @@
# Agent 步骤表:agent_step # Agent 步骤表:agent_step
**状态**:当前表 **状态**:当前表
**来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep` **来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep`、`HarnessAgentAuditHook`
## 定位 ## 定位
`agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。Provider reasoning 使用独立 `agent_reasoning_audit` 表,不复用历史 `thought` 字段。 `agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存完整 Prompt、完整消息列表、Tool 结果体。
**Provider reasoning 与 assistant 正文的完整双字段** 在独立表 `agent_reasoning_audit`;本表只保留步骤级摘要与兼容镜像。
## 字段 ## 字段
@@ -17,8 +19,8 @@
| `step_index` | INT | 是 | 步骤序号,从 0 开始 | | `step_index` | INT | 是 | 步骤序号,从 0 开始 |
| `agent_name` | VARCHAR(32) | 是 | 当前 Harness 写入固定为 `diagnosis_agent` | | `agent_name` | VARCHAR(32) | 是 | 当前 Harness 写入固定为 `diagnosis_agent` |
| `model_input` | TEXT | 否 | JSON metadata,仅包含 message count 与 roles | | `model_input` | TEXT | 否 | JSON metadata,仅包含 message count 与 roles |
| `model_output` | TEXT | 否 | JSON metadata,仅包含 text presence 与 Tool names | | `model_output` | TEXT | 否 | JSON metadata:`has_text`、`tool_names`、`reasoning_available`、`reasoning_bytes`、`assistant_bytes`、`content_source` |
| `thought` | TEXT | 否 | 当前 Harness 必须写空;字段仅保留历史兼容 | | `thought` | TEXT | 否 | 兼容镜像:优先 Provider reasoning,否则 assistant 正文;**完整双字段以 `agent_reasoning_audit` 为准** |
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 | | `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
| `duration_ms` | INT | 否 | 本步骤耗时 | | `duration_ms` | INT | 否 | 本步骤耗时 |
| `token_count` | INT | 否 | 本步骤 Token 消耗 | | `token_count` | INT | 否 | 本步骤 Token 消耗 |
@@ -36,12 +38,14 @@
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。 - `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。 - `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。 - `tool_invocation.step_id` 应关联当前模型步骤的 `agent_step.id`(由 `AgentStepAuditTracker` 在 beforeModel 绑定)。
- `agent_reasoning_audit` 通过相同的 `session_id + run_id + step_index` 逻辑定位模型步骤,不建立数据库外键。 - `agent_reasoning_audit` 通过相同的 `session_id + run_id + step_index` 逻辑定位模型步骤,不建立数据库外键。
## 注意点 ## 注意点
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。 - 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
- 新 Trace 和验收读路径必须按 exact `run_id` 取数,避免同一 `sessionId` 多次运行混入。 - 新 Trace 和验收读路径必须按 exact `run_id` 取数,避免同一 `sessionId` 多次运行混入。
- 当前 Run 若出现 `diagnosis_agent` 之外的新写入,或 `thought` 非空,视为审计边界违规。 - 当前 Run 若出现 `diagnosis_agent` 之外的新写入,视为审计边界违规。
- `model_output` 可保存 `has_text`、`tool_names`、`reasoning_available` 和 `reasoning_bytes` 等有界 metadata,不得保存 reasoning 原文。 - `model_output` **不得**保存 reasoning/assistant 原文;只允许有界 metadata。
- `thought` 非空 **不再**视为违规:它是受限镜像,便于旧读路径一眼看到“本步在想什么/说什么”;敏感完整审计仍走 `/trace/reasoning`。
- 需要同时查看 thinking 与 assistant 正文时,必须查 `agent_reasoning_audit` 或 reasoning Trace API。
+5 -3
View File
@@ -1,7 +1,7 @@
# 诊断运行表:diagnosis_run # 诊断运行表:diagnosis_run
**状态**:当前诊断运行主表 **状态**:当前诊断运行主表
**来源**:`V011__add_session_run_isolation.sql`、`DiagnosisRun` **来源**:`V011__add_session_run_isolation.sql`、`V012__add_chat_release_contract.sql`、`V016__add_conclusion_to_diagnosis_run.sql`、`DiagnosisRun`
## 定位 ## 定位
@@ -15,9 +15,10 @@
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID | | `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` | | `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
| `query` | TEXT | 是 | 本次 Chat 用户问题 | | `query` | TEXT | 是 | 本次 Chat 用户问题 |
| `conclusion` | TEXT | 否 | 从安全发布内容提取的短结论,与 `query` 并列便于读出;非 reasoning 原文 |
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` | | `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
| `agent_flow` | VARCHAR(32) | 否 | 历史兼容字段;当前公开执行统一来自 Chat Harness | | `agent_flow` | VARCHAR(32) | 否 | 历史兼容字段;当前公开执行统一来自 Chat Harness |
| `answer` | LONGTEXT | 否 | 安全发布的最终文本兼容字段 | | `answer` | LONGTEXT | 否 | 安全发布的最终内容 JSON/文本兼容字段 |
| `intent` | VARCHAR(32) | 否 | `SYSTEM_CHAT`、`KNOWLEDGE_QUERY` 或 `DIAGNOSIS` | | `intent` | VARCHAR(32) | 否 | `SYSTEM_CHAT`、`KNOWLEDGE_QUERY` 或 `DIAGNOSIS` |
| `release_outcome` | VARCHAR(16) | 否 | `SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` | | `release_outcome` | VARCHAR(16) | 否 | `SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
| `published_result` | JSON | 否 | Release Policy 允许发布的结构化安全结果 | | `published_result` | JSON | 否 | Release Policy 允许发布的结构化安全结果 |
@@ -55,4 +56,5 @@
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。 - latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。 - 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
- 当前诊断发布结果以 `release_outcome + published_result` 为准,不得从旧 self-evaluation 推断 Release Policy 结果。 - 当前诊断发布结果以 `release_outcome + published_result` 为准,不得从旧 self-evaluation 推断 Release Policy 结果。
- `diagnosis_run` 不保存 reasoning 原文;普通 Run/Trace 查询也不得通过聚合将其带出。 - `conclusion` 由 `RunConclusionExtractor` 在 `finish` 时从 `answer`(安全发布 JSON)提取:优先 `report.conclusion.text`,否则 fallback 的 `type + message` 等;最多约 4000 字符。它是 **业务结论读出字段**,不是 Provider thinking。
- `diagnosis_run` 不保存 reasoning 原文;普通 Run/Trace 查询也不得通过聚合 `agent_reasoning_audit` 将其带出。Trace 的 `run.conclusion` 可以返回上述短结论。
@@ -3,7 +3,7 @@
## Context ## Context
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector). - Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
- Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1. - Delivery baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1.
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary. - Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic. - Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
@@ -9,7 +9,7 @@
As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness. As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.
This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checklist.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived. This change is **Delivery 1** from `docs/Milvus-Hybrid接入清单.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
## What Changes ## What Changes
@@ -47,7 +47,7 @@ This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checkl
- Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests - Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
- Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant - Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
- Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids - Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids
- Docs baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1 - Docs baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1
## Risks ## Risks
@@ -0,0 +1,6 @@
Committed OpenSpec
change: rag-eval-hybrid-baseline
committed_at: 2026-07-28
scale: standard-lean
scope: knife-1-only
acceptance: wiring-required; fixture-refresh-best-effort
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-28
@@ -0,0 +1,13 @@
# Brief: rag-eval-hybrid-baseline
## Background
Offline RAG eval structure is correct but generator/docs/fixtures predate hybrid search mode.
## Goals
Knife 1 only: `search.mode` wiring, fixture meta, README, best-effort fixture refresh.
## Non-goals
Dual dense/hybrid fixture trees; golden mustNot/chunk/level hard gates; new frameworks.
@@ -0,0 +1,111 @@
# Decisions — rag-eval-hybrid-baseline
## sm-flow meta
- **Checkpoint**: Discover (in progress)
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
- **Slug**: `rag-eval-hybrid-baseline`
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
## Clarify summary
| Item | Content |
|------|---------|
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
| Non-goals | New framework, LLM judge, prod retrieval redesign |
## Context summary
| Source | Conclusion | Into OpenSpec |
|--------|------------|---------------|
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
## Question pool (grill)
| ID | Dim | Mode | Question | Status |
|----|-----|------|----------|--------|
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
### Q3–Q5 evidence
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
### Q1 用户确认
- **选择**: 仅第一刀(推荐)
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
### Q2 用户确认
- **选择**: 接线优先,刷新可记未验证
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
---
## Discover status
- [x] clarify
- [x] context
- [x] propose (`proposal.md`)
- [x] grill (Q1–Q5 closed)
**Discover checkpoint: 完成。**
---
## Commit checkpoint
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
- **Gate**: `.committed` written
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
---
## Apply checkpoint
- **Capability**: openspec-apply-change + Committed tasks
- **Authorization**: user「实现」
### Delivered (knife-1)
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
- `eval/rag-retrieval/README.md` hybrid-era docs
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
### Apply-discovered regression + fix
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
### Commands run
```text
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
```
**Apply checkpoint: 完成。** Ready for Archive when user requests.
@@ -0,0 +1,62 @@
# Design: rag-eval-hybrid-baseline
## Context
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
## Goals / Non-Goals
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
## Decisions
### D1 — Replace vector-store mode with search mode
| Before | After |
|--------|--------|
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
### D2 — Fixture meta minimum
```text
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
```
- `searchMode`: actual mode used for generation.
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
### D3 — LookupResult payload
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
### D4 — Acceptance if live refresh fails
Must deliver: ps1, test meta emission, README.
Should attempt: seed + generate + eval.
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
### D5 — Baseline update policy
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
## Risks
| Risk | Mitigation |
|------|------------|
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
## Interface impact
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
## Audit
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
@@ -0,0 +1,66 @@
# Change: Align RAG offline eval with hybrid + qualityScore era
## Why
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
- No first-class dense vs hybrid fixture split for recall comparison.
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
## What Changes
### Knife 1 (must) — make offline eval reflect current main path
1. **Generator wiring**
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
2. **Fixture meta**
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
3. **Refresh path**
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
4. **Docs**
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
### Out of scope this change (confirmed grill)
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
## Non-goals
- Rewriting eval into a new framework or LLM-as-judge.
- Full threshold calibration productization.
- Neighbor chunks / query rewrite / cross-encoder.
- Changing production retrieval code paths (except eval generator test harness props).
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
- Dense/hybrid dual fixture directories (later change).
## Context constraints
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
- Seed isolation via `kb_scope=rag-eval` stays.
## Impact
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
## Success
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
- Fixtures carry searchMode meta.
- README describes correct offline/live loop.
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
- (If knife 2) dual fixture roots documented and runnable.
@@ -0,0 +1,50 @@
# rag-eval-offline-baseline Specification
## Purpose
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
## ADDED Requirements
### Requirement: Snapshot generation SHALL use retrieval search mode
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
#### Scenario: Default hybrid generation
- **WHEN** the snapshot generator is invoked with default parameters
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
#### Scenario: Dense mode override for comparison runs
- **WHEN** the operator sets search mode to `dense`
- **THEN** fixture generation SHALL use dense retrieval for that run
### Requirement: Generated fixtures SHALL record search meta
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
#### Scenario: Meta fields present
- **WHEN** a fixture is written for a golden case
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
### Requirement: Offline evaluation SHALL remain dependency-free
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
#### Scenario: Offline eval without live stack
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
- **THEN** it SHALL produce pass/fail results using fixture contents only
### Requirement: Eval documentation SHALL describe the hybrid-era loop
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
#### Scenario: README main path
- **WHEN** an engineer follows the eval README happy path
- **THEN** the documented default generation mode SHALL be hybrid search mode
@@ -0,0 +1,34 @@
# Tasks: rag-eval-hybrid-baseline
## 1. Generator wiring
- [x] 1.1 Update `scripts/generate_rag_lookup_snapshots.ps1`: replace `VectorStoreMode` / `vector-store.mode` with `SearchMode` default `hybrid` and `-Dretrieval.search.mode=...`
- [x] 1.2 Keep `-Dretrieval.kb-scope` (default `rag-eval`); document `-SearchMode dense` override
- [x] 1.3 Ensure snapshot test picks up `retrieval.search.mode` (system property and/or `@SpringBootTest` properties)
## 2. Fixture meta
- [x] 2.1 `RagLookupSnapshotGeneratorTest` writes `searchMode` and `kbScope` (when set) on each fixture
- [x] 2.2 Confirm offline `eval_rag_retrieval.py` still loads fixtures (ignore extra meta)
## 3. Docs
- [x] 3.1 Rewrite `eval/rag-retrieval/README.md` hybrid-era loop; remove spring vector-store as default generation path
- [x] 3.2 Note offline vs live responsibilities; point to qualityScore/hybrid main path briefly
## 4. Refresh attempt (best-effort)
- [x] 4.1 Attempt `prepare_rag_eval_seed` + hybrid snapshot generate + offline eval when environment allows
- [x] 4.2 On success: update fixtures and `reports/baseline.*` if needed after diff review
- [x] 4.3 On failure: record exact commands, error summary, and “unverified refresh” in change decisions/acceptance notes — do not block wiring delivery
## 5. Verify
- [x] 5.1 Static: grep shows no required `vector-store.mode` in snapshot generator path
- [x] 5.2 Offline: `python scripts/eval_rag_retrieval.py` runs on committed fixtures (pass or documented baseline drift)
## 6. Apply-discovered fix (hybrid quality gate)
- [x] 6.1 Hybrid attaches optional `denseDistance` without overwriting RRF order/scoreLabel
- [x] 6.2 `toQualityScore(hybrid)` prefers dense L2 for absolute gates; rank fallback if no dense
- [x] 6.3 Regenerate fixtures; L0 filter fallback case green again; baseline updated
@@ -0,0 +1,6 @@
Committed OpenSpec
change: rag-quality-score-unify
committed_at: 2026-07-28
scale: standard
interface_impact: L2
gate: proposal+design+specs+tasks+cross-artifact+audit
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-28
@@ -0,0 +1,20 @@
# Brief: rag-quality-score-unify
## Background
True BM25 hybrid retrieval is live, but post-processing still normalizes as if every score were dense L2 and re-ranks with L0 keyword contains boosts. That splits ranking authority from quality gates and double-counts lexical signal.
## Goals
- Unify score labels to `dense` | `hybrid`.
- Single `toQualityScore`; label-agnostic post-process.
- Preserve retrieval rank; remove boost re-rank.
- Hybrid quality = pure rank mapping (slice 1).
## Scope
Internal RAG pipeline: store emission, normalizer, evidence post-process, tests, architecture docs.
## Non-goals
Fine re-rankers, query rewrite, neighbor chunks, schema rebuild, ACI field renames, removing dense comparison mode.
@@ -0,0 +1,136 @@
# Decisions — rag-quality-score-unify
## sm-flow meta
- **Checkpoint**: Discover(clarify + context + propose + grill)
- **Scale**: standard
- **Capability**: sm-flow 内置协议;openspec CLI `new change`;grill 使用内置协议(conversation-confirmed + evidence-driven),标注 fallback:未调用外部 `grill-with-docs` skill 文件执行器
- **Slug**: `rag-quality-score-unify`
- **OpenSpec path**: `openspec/changes/rag-quality-score-unify/`
## Clarify summary
| 项 | 内容 |
|---|---|
| 问题 | hybrid 已 RRF 融合,后处理仍 L2 伪装 + 关键词 boost 改序,质量信号不统一 |
| 期望 | label 仅 dense/hybrid;toQualityScore 唯一归一化;后处理保 rank、去 boost 改序 |
| 影响代码 | `MilvusHybridKnowledgeStore`, `VectorSearchService`, `KnowledgeEvidencePostProcessor`, retrieval 包新类, DTO 注释, 测试, 架构文档 |
| 非目标 | 精排/rewrite/邻块、删 dense mode、改 ACI 字段结构、改 schema |
## Context summary (devflow)
| 来源 | 结论 | 需进 OpenSpec |
|---|---|---|
| `devflow/index.md` | rag-chunk-identity / bm25-hybrid / hybrid-rrf 均 archived | 是:承接不回退 |
| `rag-bm25-hybrid-drop-sdk/decisions.md` | 曾要求 dense L2 enrichment 兼容阈值 | **是:本 change 废止该 decision** |
| glossary | lookup_knowledge 为证据工具;不在此改 ACI 主结构 | 是:非目标 |
| 架构文档 §6.0 | mode dense=对照,hybrid=主路径 | 是:保留 |
**index 使用状态**: 已命中相关 RAG 条目。
## Question pool (grill)
| ID | 维度 | 模式 | 问题 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | user-interview(对话已确认) | 一级 scoreLabel 是否只有 dense/hybrid,bm25_only 不作正式 label? | **已确认** |
| Q2 | 边界 | user-interview(对话已确认) | 后处理是否去掉关键词 contains 加分改序,仅保 originalRank? | **已确认** |
| Q3 | 边界 | user-interview(对话已确认) | 归一化是否唯一 toQualityScore;后处理 label-agnostic? | **已确认** |
| Q4 | 验收 | user-interview(对话已确认) | dense mode 保留作召回对照;主路径 hybrid? | **已确认** |
| Q5 | 技术 | evidence-driven | 当前代码是否仍 L2 回填 + boost 重排? | **已查证** |
| Q6 | 技术 | evidence-driven | Agent ACI 是否暴露 scoreLabel? | **已查证** |
| Q7 | 验收 | user-interview(对话已确认) | 接受 relevance_level / retry 分布变化? | **已确认** |
| Q8 | 边界 | user-interview | hybrid quality 切片 1 是否采用**纯 rank 映射**(不做 max(rank,denseSim))? | **已确认** |
### Q1–Q4, Q7 用户确认摘录(本会话)
- Label:「应该只有 hybrid 和 dense」「bm25_only 不是第三种」→ 同意收成两种。
- 归一化:「抽取抽象转换…后处理抽象统一」→ 同意。
- 后处理:「关键词打分不合理」「可以,就按照这个」(去 boost 改序 + 归一化一起做)。
- mode:保留 dense 作对照,写入架构 §6.0。
- 行为变化:讨论中已说明 hybrid 顺序/等级/retry 会变,用户要求按该方案实施(经 sm-flow)。
### Q5 evidence-driven
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后仍 dense 回填 L2 / `bm25_only_no_dense`。
- `KnowledgeEvidencePostProcessor.score`:`normalizeL2` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序。
- `LookupKnowledgeTool`:`isLowQuality` 看 `topSimilarity`(来自 baseScore)。
### Q6 evidence-driven
- Agent 主契约 `RagToolResult` / projector 暴露 evidence 列表与 relevance_level,不依赖 scoreLabel 字符串;改 label 为内部/L2 影响。
### Q8 用户确认(2026-07-28)
- **问题原文**: hybrid 的 qualityScore(切片 1)采用哪种映射?
- **用户选择**: 纯 rank 映射(推荐)
- **确认状态**: 已确认
- **实现约束**: `toQualityScore(hybrid)` = `rankToQuality(originalRank, batchSize)`;不看 RRF 原分量纲;不做 `max(rank, denseSim)`;不在 hybrid 路径为质量闸门再查/回填 dense L2。
---
## Discover status
- [x] clarify
- [x] context
- [x] propose (`proposal.md`)
- [x] grill 完成(Q1–Q8 均已关闭)
**Discover checkpoint: 完成。**
---
## Commit checkpoint
### Capability
- specify: sm-flow 内置 + openspec status/instructions(fallback:按 template 手写 design/specs/tasks)
- audit: sm-flow 内置协议(未调用外部 zoom-out)
- commit gate: 文件完整性 + 一致性检查后写入 `.committed`
### Cross-artifact 对齐
| 链路 | 状态 |
|---|---|
| brief/proposal 目标范围 → design | 已对齐 |
| design 决策(label/normalizer/保序/去 boost/纯 rank)→ specs | 已对齐 |
| specs 可观察行为 → tasks 可执行切片 | 已对齐 |
| decisions Q1–Q8 → proposal/design/specs | 已对齐 |
### Audit(≤5 句)
1. 链路仍是 Tool→Retriever→Store→Post→Pack→Project,无新外部系统。
2. 分数所有权上收 store 发射 + normalizer;后处理只裁剪与质量闸门。
3. 废止 bm25-hybrid 的「dense L2 enrichment」决策,属有意行为变化(L2 接口影响)。
4. 风险主要是 hybrid 序数 quality 与阈值标定,已记入 design Risks。
5. 不触及 Agent ACI 字段名与 Milvus schema。
### Commit gate checklist
- [x] proposal / design / specs / tasks / brief 存在
- [x] 核心概念在 design 有对应
- [x] design 关键决策在 tasks 有任务
- [x] tasks 可验证(checkbox 纵向切片)
- [x] 无未确认 user-interview
- [x] `.committed` 已创建
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
**下一步**: 等待用户明确授权 **Apply**(例如「开始 apply / 实现」)。未授权前不改业务接线代码。
---
## Apply checkpoint
- **Capability**: openspec-apply-change + Committed OpenSpec tasks
- **授权**: 用户「实现」
- **完成**: tasks.md 全部勾选
- **验证**:
- `RetrievalScoreNormalizerTest` 4 passed
- `KnowledgeEvidencePostProcessorTest` 6 passed
- `LookupKnowledgeToolTest` 7 passed
- `VectorSearchServiceTest` 2 passed
- `VectorKnowledgeSearchAdapterHybridTest` 1 passed
- **已知限制**: hybrid quality 为本轮 rank 序数映射,跨 query 绝对值不可比;阈值可能需后续标定
- **行为变化**: 已落地(去 L2 回填、去 boost 改序、label dense/hybrid)
**Apply checkpoint: 完成。** 可进入 Archive(需用户确认是否 archive OpenSpec)。
@@ -0,0 +1,134 @@
# Design: rag-quality-score-unify
## Context
- Knowledge path already uses single `MilvusHybridKnowledgeStore` (MilvusClientV2) with dense + BM25 + RRF.
- Chunk-level `evidenceKey` dedup and `retrieve-k` / `return-n` are landed.
- Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
- Prior design in `rag-bm25-hybrid-drop-sdk` required dense L2 enrichment for threshold compatibility — **this change supersedes that decision**.
Stakeholders: `lookup_knowledge` internal pipeline; Agent ACI field *names* unchanged; operators comparing `retrieval.search.mode=dense|hybrid`.
## Goals / Non-Goals
**Goals:**
1. First-class `scoreLabel` values: only `dense` | `hybrid` (aliases canonicalize).
2. Single `toQualityScore` adapter; post-process is label-agnostic.
3. Preserve retrieval `originalRank` as sort authority; remove keyword/domain boost re-ranking.
4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
5. Keep `mode=dense` for offline recall comparison; production default remains hybrid.
**Non-Goals:**
- Cross-encoder / query rewrite / neighbor chunks.
- Schema rebuild or collection rename.
- Changing Agent-facing ACI JSON field names.
- Configurable `max(rank, denseSim)` quality (future).
## Decisions
### D1 — Two labels only
| label | `score` meaning | quality mapping |
|---|---|---|
| `dense` | L2 distance (smaller better) | `1 - clamp(l2)/maxL2Distance` |
| `hybrid` | engine fused score optional in `rawScore`; **not** used as L2 | `rankToQuality(originalRank, batchSize)` |
Canonicalize legacy strings: `l2_distance`→dense; `rrf_fused` / `bm25_only_*`→hybrid.
**Why not keep `bm25_only`:** it is not a search mode; it was a L2-fake patch. Hybrid path hits are all `hybrid`.
### D2 — Stop dense L2 overwrite on hybrid hits
`searchHybrid` SHALL:
1. Run `hybridSearch` + RRFRanker.
2. Emit hits in RRF order with `scoreLabel=hybrid`, `originalRank=1..n`.
3. Set `rawScore` from engine when present; `score` MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption.
4. MUST NOT set `bm25_only_no_dense` or force `score=maxL2Distance` for threshold faking.
5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).
`searchDense` SHALL emit `scoreLabel=dense` and L2 in `score`.
### D3 — `RetrievalScoreNormalizer` is the only label branch
```text
qualityScore = RetrievalScoreNormalizer.toQualityScore(
scoreLabel, score, originalRank, batchSize, maxL2Distance)
```
- Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
- Dense: existing L2 formula (behavior parity for dense mode).
Post-processor, `isLowQuality`, and `relevance_level` consume **only** `qualityScore` (exposed today as `baseScore` / `topSimilarity` fields for minimal DTO churn).
### D4 — Post-process sort and boosts
```text
order = originalRank ASC, then stable evidenceKey
// NO finalScore = base + 0.15 domain + ...
```
- Remove additive boosts from sort key and from `finalScore` used for ordering.
- Optional: if L0 hint string matches, append explanatory `hitReasons` only (e.g. `l0_keyword_overlap`) — zero score delta.
- Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
- `relevance_level`: compare top `qualityScore` to existing thresholds; **remove** `hasHintSupport` gate for PRECISE.
- `isLowQuality` / category unfiltered retry: unchanged control flow, new score semantics.
### D5 — Trace fields
- `RerankTrace` may keep `baseScore`/`finalScore` names but both equal `qualityScore` when boosts are zero; `boostReasons` empty or explanation-only reasons without `:+0.xx` score deltas.
- Prefer renaming comments to “quality trace”; no Agent contract change required.
### D6 — WIP files
Workspace may contain draft `RetrievalScoreLabels` / `RetrievalScoreNormalizer`. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.
## Interface impact
- **Level: L2** — internal DTO semantics (`score`, `scoreLabel`), post-process ordering and relevance distribution.
- Agent ACI: field names stable; `relevance_level` *values distribution* may change (accepted behavioral change).
## Data flow (target)
```text
VectorSearchService (mode dense|hybrid)
-> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
-> KnowledgeDocumentRetriever candidates
-> KnowledgeEvidencePostProcessor
for each: qualityScore = toQualityScore(...)
sort by originalRank
dedup / caps / return-n
relevance + topSimilarity from qualityScore
-> pack / assemble / project
```
## Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute |
| Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” |
| PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains |
| Tests assert boost re-order | Update `LookupKnowledgeToolTest.rerankUsesHintMatches...` |
| Old label strings in sidecar/eval | canonicalize in normalizer |
## Migration Plan
1. Deploy code; no Milvus schema migration.
2. Default `retrieval.search.mode=hybrid` unchanged.
3. Rollback: revert change; old L2-enrichment behavior returns.
4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).
## Open Questions
- None for slice-1 (Q8 confirmed: pure rank).
- Follow-up: threshold calibration after live traces; optional denseSim blend.
## Audit notes (inline)
Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
No new cross-module lifecycle. Couples only internal RAG pipeline.
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.
@@ -0,0 +1,80 @@
# Change: Unify RAG quality score (dense/hybrid labels) and stop keyword boost re-rank
## Why
BM25 hybrid 已在库内完成 dense + BM25 + RRF 融合,但后处理仍:
1. 把 hybrid 结果**伪装成 L2** 再 `normalizeL2`(含 `bm25_only_no_dense` 弱分占位);
2. 用 L0 domain/entity/keyword **contains 加分改主序**。
这导致:排序信号与质量闸门分裂;词面信号被 BM25 与后处理**双重计分**;「词面热、语义冷」的片段可能被抬到前面;hybrid 的 RRF 序被冲掉。
需要统一:**检索负责序,后处理只做 quality 归一化 + 裁剪装配**。
## What Changes
### 检索层(`MilvusHybridKnowledgeStore` / `VectorSearchService`)
- 一级 `scoreLabel` 仅两种:`dense` | `hybrid`(与 `retrieval.search.mode` 对齐)。
- **废弃**正式一级 label:`l2_distance` / `rrf_fused` / `bm25_only_no_dense`(可读兼容映射到 dense/hybrid)。
- `mode=dense`:`score` = L2 距离,`label=dense`,`originalRank` = ANN 序。
- `mode=hybrid`:`label=hybrid`;**不再**用 dense L2 覆盖主 `score`;**不再**对 BM25-only 伪造 maxL2;`originalRank` = RRF 返回序;`rawScore` 可保留引擎融合分。
- `mode=dense|hybrid` **保留**:hybrid 为线上主路径;dense 为同库对照/评测(已写入架构 §6.0)。
### 归一化(新)
- 新增唯一转换点 `RetrievalScoreNormalizer.toQualityScore(label, score, rank, batchSize, maxL2)` → `qualityScore ∈ [0,1]`(越大越好)。
- `dense`:`1 - clamp(L2)/maxL2Distance`
- `hybrid`:按 **rank** 映射(本轮 batch 线性),不把 RRF 原分当 L2 套公式。
- Label 差异**只**在此消化。
### 后处理(`KnowledgeEvidencePostProcessor`)
- **统一流程**,只消费 `qualityScore` + `originalRank`(label-agnostic)。
- **排序主序 = `originalRank` 升序**(保检索序);去掉 domain/entity/keyword/source_type **加分改序**。
- L0 contains 匹配若保留,仅写入 `hitReasons` / trace 解释,**不参与 sort key、不加 finalScore**。
- `relevance_level` / `isLowQuality` / category unfiltered retry:只看 top `qualityScore` 与既有阈值;**不再**要求 `hasHintSupport` 才能 PRECISE。
- 保留:evidenceKey 去重、`max-chunks-per-document`、`return-n`、excerpt 截断、EvidenceBlock 装配。
### 文档 / 测试
- 更新 `mvp/architecture/RAG知识检索架构.md` §6 分数与后处理约定。
- 单测:dense 路径 quality 与现 L2 归一化一致;hybrid 保序且不被 keyword 打乱;无 `bm25_only` 一级 label;旧 label 别名可 canonicalize。
## Non-goals
- Cross-encoder / listwise 精排、query rewrite、邻块扩展。
- 删除 `mode=dense` 对照开关。
- 改变 Agent 可见 ACI 字段结构(`evidence[]` / `relevance_level` 枚举名可不变;**分布会变**)。
- 修改 Milvus schema / 强制全量 rebuild(本 change 不改 collection 结构)。
- 上线可配 `max(rank, denseSim)` 混合 quality(可后续迭代;本 change 切片 1 用纯 rank 映射 hybrid)。
## Context constraints (from devflow)
- 承接 archived:`rag-chunk-evidence-identity-dedup`、`rag-bm25-hybrid-drop-sdk`、`rag-hybrid-search-rrf`、`modular-rag-pipeline`。
- 单一知识后端仍为 `MilvusHybridKnowledgeStore`(V2);不得恢复 sdk/spring 主路径路由。
- Agent 投影仍不暴露 raw fused score / 完整 contextPack 作为主契约(内部 LookupResult/trace 可保留调试字段)。
- 历史 decision「Dense L2 enrichment for threshold compatibility」**本 change 有意废止**,改为 qualityScore 统一闸门。
## Impact
- **行为变化(对内检索质量语义)**:
- hybrid 下证据顺序更贴近 RRF;
- 词面命中不再被后处理 contains 二次抬序;
- `relevance_level` 与 unfiltered retry 触发分布可能变化;
- BM25-only 命中不再被标成 quality≈0。
- **接口影响**:L2(内部 DTO/注释/scoreLabel 字符串约定);Agent ACI 字段名不变。
- **风险**:hybrid rank→quality 为序数映射,绝对值不跨 query 可比;阈值 0.75/0.5 可能需后续观测再调(本 change 先沿用配置项)。
## Scale
- **standard**(多文件、有意行为变化、需 design + specs + tasks + 测试)。
## Depends on
- 已落地 hybrid schema + chunk evidenceKey(archived changes 如上)。
- 对话已确认的设计口径(见 change `decisions.md`)。
## WIP note
- 工作区可能已有未接线的 `RetrievalScoreLabels` / `RetrievalScoreNormalizer` 草稿文件;apply 阶段以 **Committed OpenSpec** 为准接入或改写,不视为已完成实现。
@@ -0,0 +1,94 @@
# rag-retrieval-quality-score Specification
## Purpose
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
## ADDED Requirements
### Requirement: Primary score labels SHALL be only dense or hybrid
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
#### Scenario: Dense mode labels hits as dense
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
- **AND** `score` SHALL be the dense L2 distance
#### Scenario: Hybrid mode labels hits as hybrid
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
#### Scenario: Legacy aliases canonicalize
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
#### Scenario: No L2 enrichment overwrite
- **WHEN** hybrid search completes
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
### Requirement: Quality score SHALL be produced by a single normalizer
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
#### Scenario: Dense quality from L2
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
#### Scenario: Hybrid quality from rank
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
### Requirement: Post-process SHALL preserve retrieval rank order
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
#### Scenario: Keyword overlap does not promote lower rank
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
- **AND** B matches more L0 keywords via string contains than A
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
#### Scenario: Structural caps still apply after rank order
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
- **AND** `rag.return-n` SHALL still bound total blocks
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
#### Scenario: Low quality uses top qualityScore
- **WHEN** post-process finishes with at least one evidence block
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
#### Scenario: PRECISE does not require hint support
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
### Requirement: Dense mode remains available for recall comparison
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
#### Scenario: Mode dense still callable
- **WHEN** mode is `dense`
- **THEN** search SHALL call dense ANN only and label hits `dense`
@@ -0,0 +1,39 @@
# Tasks: rag-quality-score-unify
## 1. Score contract utilities
- [x] 1.1 Finalize `RetrievalScoreLabels` (`dense` / `hybrid` + canonicalize legacy aliases)
- [x] 1.2 Finalize `RetrievalScoreNormalizer.toQualityScore` (dense L2 formula; hybrid pure rank map with batchSize)
- [x] 1.3 Unit tests for normalizer: dense L2 edges; hybrid rank monotonicity; alias canonicalize
## 2. Store / search emission
- [x] 2.1 `searchDense`: emit `scoreLabel=dense`, L2 `score`, stable rank order
- [x] 2.2 `searchHybrid`: emit `scoreLabel=hybrid`; keep RRF order as `originalRank`; stop dense L2 overwrite and `bm25_only_*` labels; no parallel dense probe for score rewrite
- [x] 2.3 Update `VectorSearchService.SearchResult` / adapter comments so `score`+`scoreLabel` contract matches design
- [x] 2.4 Ensure `KnowledgeDocumentRetriever` / `VectorKnowledgeSearchAdapter` propagate `scoreLabel`, `score`, `rawScore`, `originalRank` unchanged
## 3. Post-process
- [x] 3.1 `KnowledgeEvidencePostProcessor`: compute quality via normalizer; sort by `originalRank` ASC (stable tie-break)
- [x] 3.2 Remove domain/entity/keyword/source_type additive boosts from ordering/`finalScore`
- [x] 3.3 Optional: L0 overlap only as explanatory `hitReasons` (no score delta)
- [x] 3.4 `relevance_level` / `isLowQuality` / `topSimilarity` use qualityScore only; drop `hasHintSupport` gate for PRECISE
- [x] 3.5 Keep evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate
## 4. Tests
- [x] 4.1 Update `KnowledgeEvidencePostProcessorTest` for rank order + caps under new scoring
- [x] 4.2 Update `LookupKnowledgeToolTest.rerankUsesHintMatchesAndContextPackPreservesMetadata` (no boost re-order; metadata/context pack still ok)
- [x] 4.3 Adjust any tests asserting `l2_distance` / boost reasons `:+0.xx` as needed
- [x] 4.4 Run targeted unit tests for touched classes
## 5. Docs
- [x] 5.1 Update `mvp/architecture/RAG知识检索架构.md` §6 score/post-process (replace L2-enrichment narrative)
- [x] 5.2 Align `application.yml` comments if still describing L2-only post-process for hybrid
## 6. Verify
- [x] 6.1 Confirm no production path still sets `bm25_only_no_dense` or overwrites hybrid scores with L2 for thresholds
- [x] 6.2 Note known limitation: hybrid quality is ordinal within batch; thresholds may need later calibration
@@ -0,0 +1,50 @@
# rag-eval-offline-baseline Specification
## Purpose
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
## ADDED Requirements
### Requirement: Snapshot generation SHALL use retrieval search mode
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
#### Scenario: Default hybrid generation
- **WHEN** the snapshot generator is invoked with default parameters
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
#### Scenario: Dense mode override for comparison runs
- **WHEN** the operator sets search mode to `dense`
- **THEN** fixture generation SHALL use dense retrieval for that run
### Requirement: Generated fixtures SHALL record search meta
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
#### Scenario: Meta fields present
- **WHEN** a fixture is written for a golden case
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
### Requirement: Offline evaluation SHALL remain dependency-free
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
#### Scenario: Offline eval without live stack
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
- **THEN** it SHALL produce pass/fail results using fixture contents only
### Requirement: Eval documentation SHALL describe the hybrid-era loop
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
#### Scenario: README main path
- **WHEN** an engineer follows the eval README happy path
- **THEN** the documented default generation mode SHALL be hybrid search mode
@@ -0,0 +1,104 @@
# rag-retrieval-quality-score Specification
## Purpose
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
## Requirements
### Requirement: Primary score labels SHALL be only dense or hybrid
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
#### Scenario: Dense mode labels hits as dense
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
- **AND** `score` SHALL be the dense L2 distance
#### Scenario: Hybrid mode labels hits as hybrid
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
#### Scenario: Legacy aliases canonicalize
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
#### Scenario: No L2 enrichment overwrite
- **WHEN** hybrid search completes
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
### Requirement: Quality score SHALL be produced by a single normalizer
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
#### Scenario: Dense quality from L2
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
#### Scenario: Hybrid quality from rank
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
### Requirement: Post-process SHALL preserve retrieval rank order
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
#### Scenario: Keyword overlap does not promote lower rank
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
- **AND** B matches more L0 keywords via string contains than A
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
#### Scenario: Structural caps still apply after rank order
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
- **AND** `rag.return-n` SHALL still bound total blocks
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
#### Scenario: Low quality uses top qualityScore
- **WHEN** post-process finishes with at least one evidence block
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
#### Scenario: PRECISE does not require hint support
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
### Requirement: Dense mode remains available for recall comparison
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
#### Scenario: Mode dense still callable
- **WHEN** mode is `dense`
- **THEN** search SHALL call dense ANN only and label hits `dense`
### Requirement: Hybrid quality gates MAY use optional dense distance without reordering
When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.
#### Scenario: Dense distance does not replace hybrid label
- **WHEN** a hybrid hit includes denseDistance
- **THEN** scoreLabel SHALL still canonicalize to hybrid
- **AND** post-process sort order SHALL still follow originalRank from hybrid results
+6 -6
View File
@@ -90,12 +90,12 @@
<groupId>org.springframework.boot</groupId> <groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId> <artifactId>spring-boot-starter-web</artifactId>
</dependency> </dependency>
<dependency> <!--
<groupId>org.springframework.boot</groupId> spring-boot-devtools removed on purpose.
<artifactId>spring-boot-devtools</artifactId> Classpath restart (restartedMain) recreates beans without reliably closing
<scope>runtime</scope> MilvusClientV2 gRPC channels, causing orphan channels and long hybrid RPC retries.
<optional>true</optional> Prefer full process restart: stop then `mvn spring-boot:run`.
</dependency> -->
<dependency> <dependency>
<groupId>io.milvus</groupId> <groupId>io.milvus</groupId>
<artifactId>milvus-sdk-java</artifactId> <artifactId>milvus-sdk-java</artifactId>
+11 -3
View File
@@ -3,7 +3,9 @@ param(
[string]$Fixtures = "eval\rag-retrieval\fixtures", [string]$Fixtures = "eval\rag-retrieval\fixtures",
[string]$RetrievedAt = "", [string]$RetrievedAt = "",
[string]$KbScope = "rag-eval", [string]$KbScope = "rag-eval",
[string]$VectorStoreMode = "spring", # hybrid = production main path; dense = same-collection recall baseline
[ValidateSet("hybrid", "dense")]
[string]$SearchMode = "hybrid",
[switch]$SkipEval [switch]$SkipEval
) )
@@ -16,7 +18,7 @@ $mavenArgs = @(
"-Drag.snapshot.cases=$Cases", "-Drag.snapshot.cases=$Cases",
"-Drag.snapshot.fixtures=$Fixtures", "-Drag.snapshot.fixtures=$Fixtures",
"-Dretrieval.kb-scope=$KbScope", "-Dretrieval.kb-scope=$KbScope",
"-Dretrieval.vector-store.mode=$VectorStoreMode" "-Dretrieval.search.mode=$SearchMode"
) )
if ($RetrievedAt -ne "") { if ($RetrievedAt -ne "") {
@@ -25,10 +27,16 @@ if ($RetrievedAt -ne "") {
$mavenArgs += "test" $mavenArgs += "test"
Write-Host "Generating RAG lookupResult fixtures from real LookupKnowledgeTool..." Write-Host "Generating RAG lookupResult fixtures (search.mode=$SearchMode, kb-scope=$KbScope)..."
& mvn @mavenArgs & mvn @mavenArgs
if ($LASTEXITCODE -ne 0) {
throw "Fixture generation failed with exit code $LASTEXITCODE"
}
if (-not $SkipEval) { if (-not $SkipEval) {
Write-Host "Running offline RAG retrieval baseline..." Write-Host "Running offline RAG retrieval baseline..."
& python scripts\eval_rag_retrieval.py --cases $Cases --fixtures $Fixtures & python scripts\eval_rag_retrieval.py --cases $Cases --fixtures $Fixtures
if ($LASTEXITCODE -ne 0) {
throw "Offline eval failed with exit code $LASTEXITCODE"
}
} }
@@ -4,6 +4,7 @@ import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.config.MysqlToolProperties.DataSourceProperties; import com.superbiz.agent.config.MysqlToolProperties.DataSourceProperties;
import com.superbiz.agent.harness.agent.HarnessEvidenceTools; import com.superbiz.agent.harness.agent.HarnessEvidenceTools;
import com.superbiz.agent.harness.agent.DiagnosisAgentFactory; import com.superbiz.agent.harness.agent.DiagnosisAgentFactory;
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
import com.superbiz.agent.harness.agent.DiagnosisAgentLimits; import com.superbiz.agent.harness.agent.DiagnosisAgentLimits;
import com.superbiz.agent.harness.agent.DiagnosisAgentUseCase; import com.superbiz.agent.harness.agent.DiagnosisAgentUseCase;
import com.superbiz.agent.harness.application.ChatApplicationUseCase; import com.superbiz.agent.harness.application.ChatApplicationUseCase;
@@ -153,8 +154,9 @@ public class HarnessChatConfiguration {
@Bean @Bean
public ToolBoundary toolBoundary(DiagnosisHarnessCore core, ToolCallKeyFactory keyFactory, public ToolBoundary toolBoundary(DiagnosisHarnessCore core, ToolCallKeyFactory keyFactory,
CanonicalInvocationStore store, ObjectMapper objectMapper, Clock clock, CanonicalInvocationStore store, ObjectMapper objectMapper, Clock clock,
ToolInvocationAuditSink auditSink) { ToolInvocationAuditSink auditSink,
return new ToolBoundary(core, keyFactory, store, objectMapper, clock, auditSink); AgentStepAuditTracker stepTracker) {
return new ToolBoundary(core, keyFactory, store, objectMapper, clock, auditSink, stepTracker);
} }
@Bean @Bean
@@ -240,8 +242,15 @@ public class HarnessChatConfiguration {
@Bean @Bean
public HarnessEvidenceTools harnessEvidenceTools(RagToolAdapter rag, QueryLogsToolAdapter logs, public HarnessEvidenceTools harnessEvidenceTools(RagToolAdapter rag, QueryLogsToolAdapter logs,
MysqlToolAdapter mysql) { MysqlToolAdapter mysql,
return HarnessEvidenceTools.fromAdapters(rag, logs, mysql); MysqlToolProperties mysqlToolProperties) {
// query_mysql is only exposed when at least one logical datasource is configured.
// Default application.yml has harness.mysql-tools.data-sources: {} — do not inject a dead tool.
boolean mysqlEnabled = mysqlToolProperties != null
&& mysqlToolProperties.getDataSources() != null
&& mysqlToolProperties.getDataSources().values().stream()
.anyMatch(ds -> ds != null && ds.getJdbcUrl() != null && !ds.getJdbcUrl().isBlank());
return HarnessEvidenceTools.fromAdapters(rag, logs, mysqlEnabled ? mysql : null);
} }
@Bean @Bean
@@ -250,10 +259,12 @@ public class HarnessChatConfiguration {
AgentStepRepository steps, AgentStepRepository steps,
AgentReasoningAuditRepository reasoningAudits, AgentReasoningAuditRepository reasoningAudits,
DiagnosisTraceRecorder traceRecorder, DiagnosisTraceRecorder traceRecorder,
ModelCallAuditor modelCallAuditor) { ModelCallAuditor modelCallAuditor,
AgentStepAuditTracker stepTracker) {
return new DiagnosisAgentFactory(chatModel, core, tools, mapper, return new DiagnosisAgentFactory(chatModel, core, tools, mapper,
List.of(new HarnessAgentAuditHook( List.of(new HarnessAgentAuditHook(
steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits)), steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits,
stepTracker)),
traceRecorder, modelCallAuditor); traceRecorder, modelCallAuditor);
} }
@@ -6,10 +6,15 @@ import org.slf4j.LoggerFactory;
import org.springframework.context.annotation.Configuration; import org.springframework.context.annotation.Configuration;
/** /**
* Milvus configuration notes. * Milvus 知识路径配置说明(无额外 Bean 装配)。
* *
* <p>Knowledge RAG uses {@link MilvusHybridKnowledgeStore} (MilvusClientV2) exclusively. * <p>知识库 RAG 唯一实现:{@link MilvusHybridKnowledgeStore}({@code MilvusClientV2})。</p>
* Legacy {@code MilvusServiceClient} bean is no longer created for the knowledge path.</p> * <ul>
* <li>支持 dense 与 dense+BM25 {@code hybridSearch}+RRF。</li>
* <li>不再为知识路径创建 legacy {@code MilvusServiceClient} Bean。</li>
* <li>Spring AI {@code VectorStore} starter 仍可存在于 classpath,但只作 sidecar,
* 不作 lookup_knowledge 主路径(starter 无 BM25 hybrid API)。</li>
* </ul>
*/ */
@Configuration @Configuration
public class MilvusConfig { public class MilvusConfig {
@@ -42,12 +42,32 @@ public class AgentReasoningAudit {
@Column(name = "agent_name", nullable = false, length = 64) @Column(name = "agent_name", nullable = false, length = 64)
private String agentName; private String agentName;
/** True when provider returned a non-blank reasoning/thinking chain. */
@Column(name = "reasoning_available", nullable = false) @Column(name = "reasoning_available", nullable = false)
private Boolean reasoningAvailable; private Boolean reasoningAvailable;
/**
* Provider chain-of-thought / reasoning_content when available.
* Never stores tool result payloads.
*/
@Column(name = "reasoning_content", columnDefinition = "LONGTEXT") @Column(name = "reasoning_content", columnDefinition = "LONGTEXT")
private String reasoningContent; private String reasoningContent;
/**
* Assistant-visible text for this model step (final prose and/or tool-call plan).
* Tool result message bodies are not stored here.
*/
@Column(name = "assistant_text", columnDefinition = "LONGTEXT")
private String assistantText;
/**
* Summary of what was stored:
* PROVIDER_REASONING+ASSISTANT_TEXT | PROVIDER_REASONING | ASSISTANT_TEXT | TOOL_CALL_PLAN | NONE
*/
@Column(name = "content_source", length = 64)
private String contentSource;
/** UTF-8 byte length of reasoning_content + assistant_text combined. */
@Column(name = "content_bytes", nullable = false) @Column(name = "content_bytes", nullable = false)
private Integer contentBytes; private Integer contentBytes;
@@ -41,6 +41,10 @@ public class DiagnosisRun {
@Column(name = "query", nullable = false, columnDefinition = "TEXT") @Column(name = "query", nullable = false, columnDefinition = "TEXT")
private String query; private String query;
/** Extracted final conclusion text, parallel to query for simple readout. */
@Column(name = "conclusion", columnDefinition = "TEXT")
private String conclusion;
@Column(name = "status", length = 16) @Column(name = "status", length = 16)
private String status = "PENDING"; private String status = "PENDING";
@@ -71,6 +71,8 @@ public class DiagnosisTraceResponse {
private String runId; private String runId;
private String sessionId; private String sessionId;
private String query; private String query;
/** Extracted conclusion text, parallel to query. */
private String conclusion;
private String status; private String status;
private String agentFlow; private String agentFlow;
private String intent; private String intent;
@@ -164,7 +166,11 @@ public class DiagnosisTraceResponse {
private Integer stepIndex; private Integer stepIndex;
private String agentName; private String agentName;
private Boolean reasoningAvailable; private Boolean reasoningAvailable;
/** Provider thinking / CoT when available. */
private String reasoningContent; private String reasoningContent;
/** Assistant visible text and/or tool-call plan (no tool results). */
private String assistantText;
private String contentSource;
private Integer contentBytes; private Integer contentBytes;
private LocalDateTime createdAt; private LocalDateTime createdAt;
} }
@@ -31,7 +31,7 @@ public class EvidencePostprocessResult {
private RerankTrace rerankTrace; private RerankTrace rerankTrace;
/** /**
* 排序第一名的 baseScore(0~1 相似度,不含规则 boost)。 * 排序第一名(originalRank 最优)的 qualityScore(0~1,越大越好)。
* 用于 isLowQuality 与 attempt.topSimilarity。 * 用于 isLowQuality 与 attempt.topSimilarity。
*/ */
private Double topSimilarity; private Double topSimilarity;
@@ -9,70 +9,59 @@ import java.util.Map;
/** /**
* L1 向量命中后、后处理前的统一候选结构。 * L1 向量命中后、后处理前的统一候选结构。
* *
* <p>由检索适配器从 {@code KnowledgeSearchHit} / 向量结果映射而来。 * <p>后处理:{@code RetrievalScoreNormalizer} → qualityScore;按 {@link #originalRank} 保序;
* 后处理会基于它做归一化、规则 boost、chunk 级去重并生成 {@link EvidenceBlock}。</p> * evidenceKey 去重 / 截断;不再用关键词 boost 改序。</p>
*/ */
@Data @Data
@Builder @Builder
public class RetrievedEvidenceCandidate { public class RetrievedEvidenceCandidate {
/** 向量库记录 id。 */
private String id; private String id;
/**
* 文档级 id(metadata.docId 等)。
* 用于每文档 chunk 上限;不等于 evidenceKey。
*/
private String docId; private String docId;
/** 文档内切片序号;可能为空(老数据)。 */
private Integer chunkIndex; private Integer chunkIndex;
/**
* 片段级去重/投影主键。
* 通常为 docId#chunk-N,fallback 为 vector:{id}。
*/
private String evidenceKey; private String evidenceKey;
/**
* 来源标识(_source / source / filePath / docId 等)。
* 可与同文档其他 chunk 重复;不再作为唯一去重键。
*/
private String source; private String source;
private String title; private String title;
private String breadcrumb; private String breadcrumb;
/** chunk 正文原文(后处理前未截断或仅底层原样)。 */
private String content; private String content;
/** 固定为 L1(向量层);预留多路召回标记。 */
private String retrievalLayer; private String retrievalLayer;
/** 所属 attempt 名,如 FILTERED_VECTOR。 */
private String retrievalAttempt; private String retrievalAttempt;
/** /**
* 兼容 L2 距离分(越小越相似),后处理会 normalize 成 baseScore。 * 引擎主分:dense=L2;hybrid=融合分。量纲由 {@link #scoreLabel} 解释。
*/ */
private Double score; private Double score;
/** 底层原始分。 */
private Double rawScore; private Double rawScore;
/** rawScore 语义标签:l2_distance / similarity。 */ /**
* {@code dense} | {@code hybrid}(及可被 canonicalize 的历史别名)。
*/
private String scoreLabel; private String scoreLabel;
/** 向量召回顺序(从 1 起),规则 rerank 前的名次。 */ /**
* 检索返回名次(从 1 起)——后处理排序权威。
*/
private Integer originalRank; private Integer originalRank;
/** /**
* 扁平化 metadata(string map)。 * hybrid 命中可选的 dense L2,仅供质量闸门;不参与排序。
* 可能含 docId、chunkIndex、category、kb_scope 等。
*/ */
private Double denseDistance;
private Map<String, String> metadata; private Map<String, String> metadata;
/** 初步命中原因,后处理会追加 boost reasons。 */ /**
* 初步命中原因;后处理可追加 L0 重叠解释(无分值)。
*/
private List<String> hitReasons; private List<String> hitReasons;
} }
@@ -31,6 +31,10 @@ public final class HarnessEvidenceTools {
private final List<ToolCallback> callbacks; private final List<ToolCallback> callbacks;
private final Map<String, EvidenceToolInvoker> invokers; private final Map<String, EvidenceToolInvoker> invokers;
/**
* @param mysqlInvoker optional; when null, {@code query_mysql} is not registered
* (no logical datasource configured / tool unavailable).
*/
public HarnessEvidenceTools(EvidenceToolInvoker ragInvoker, public HarnessEvidenceTools(EvidenceToolInvoker ragInvoker,
EvidenceToolInvoker logsInvoker, EvidenceToolInvoker logsInvoker,
EvidenceToolInvoker mysqlInvoker) { EvidenceToolInvoker mysqlInvoker) {
@@ -39,28 +43,38 @@ public final class HarnessEvidenceTools {
Objects.requireNonNull(ragInvoker, "ragInvoker must not be null")); Objects.requireNonNull(ragInvoker, "ragInvoker must not be null"));
registered.put(AgentToolContracts.QUERY_LOGS, registered.put(AgentToolContracts.QUERY_LOGS,
Objects.requireNonNull(logsInvoker, "logsInvoker must not be null")); Objects.requireNonNull(logsInvoker, "logsInvoker must not be null"));
registered.put(AgentToolContracts.QUERY_MYSQL, if (mysqlInvoker != null) {
Objects.requireNonNull(mysqlInvoker, "mysqlInvoker must not be null")); registered.put(AgentToolContracts.QUERY_MYSQL, mysqlInvoker);
}
this.invokers = Map.copyOf(registered); this.invokers = Map.copyOf(registered);
this.callbacks = List.of(
definition(AgentToolContracts.LOOKUP_KNOWLEDGE, List<ToolCallback> built = new java.util.ArrayList<>();
AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolCall.class), built.add(definition(AgentToolContracts.LOOKUP_KNOWLEDGE,
definition(AgentToolContracts.QUERY_LOGS, AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolCall.class));
AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsToolCall.class), built.add(definition(AgentToolContracts.QUERY_LOGS,
definition(AgentToolContracts.QUERY_MYSQL, AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsToolCall.class));
if (mysqlInvoker != null) {
built.add(definition(AgentToolContracts.QUERY_MYSQL,
AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolCall.class)); AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolCall.class));
} }
this.callbacks = List.copyOf(built);
}
/**
* @param mysqlAdapter optional; omit registration when null or when no datasources are wired
*/
public static HarnessEvidenceTools fromAdapters(RagToolAdapter ragAdapter, public static HarnessEvidenceTools fromAdapters(RagToolAdapter ragAdapter,
QueryLogsToolAdapter logsAdapter, QueryLogsToolAdapter logsAdapter,
MysqlToolAdapter mysqlAdapter) { MysqlToolAdapter mysqlAdapter) {
Objects.requireNonNull(ragAdapter, "ragAdapter must not be null"); Objects.requireNonNull(ragAdapter, "ragAdapter must not be null");
Objects.requireNonNull(logsAdapter, "logsAdapter must not be null"); Objects.requireNonNull(logsAdapter, "logsAdapter must not be null");
Objects.requireNonNull(mysqlAdapter, "mysqlAdapter must not be null"); EvidenceToolInvoker mysql = mysqlAdapter == null
? null
: bridge(AgentToolContracts.QUERY_MYSQL, mysqlAdapter::execute);
return new HarnessEvidenceTools( return new HarnessEvidenceTools(
bridge(AgentToolContracts.LOOKUP_KNOWLEDGE, ragAdapter::execute), bridge(AgentToolContracts.LOOKUP_KNOWLEDGE, ragAdapter::execute),
bridge(AgentToolContracts.QUERY_LOGS, logsAdapter::execute), bridge(AgentToolContracts.QUERY_LOGS, logsAdapter::execute),
bridge(AgentToolContracts.QUERY_MYSQL, mysqlAdapter::execute)); mysql);
} }
public List<ToolCallback> callbacks() { public List<ToolCallback> callbacks() {
@@ -6,18 +6,21 @@ import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.ObjectReader; import com.fasterxml.jackson.databind.ObjectReader;
import com.superbiz.agent.domain.entity.ChatSession; import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun; import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.harness.audit.RunConclusionExtractor;
import com.superbiz.agent.harness.contract.IntentType; import com.superbiz.agent.harness.contract.IntentType;
import com.superbiz.agent.harness.contract.PreviousTurn; import com.superbiz.agent.harness.contract.PreviousTurn;
import com.superbiz.agent.harness.contract.PublishedResult; import com.superbiz.agent.harness.contract.PublishedResult;
import com.superbiz.agent.harness.contract.ReleaseOutcome; import com.superbiz.agent.harness.contract.ReleaseOutcome;
import com.superbiz.agent.harness.core.RunBudgetUsage; import com.superbiz.agent.harness.core.RunBudgetUsage;
import com.superbiz.agent.harness.core.RunContext; import com.superbiz.agent.harness.core.RunContext;
import com.superbiz.agent.harness.audit.ModelCallComponent;
import com.superbiz.agent.repository.ChatSessionRepository; import com.superbiz.agent.repository.ChatSessionRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository; import com.superbiz.agent.repository.DiagnosisRunRepository;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.lang.Nullable;
import org.springframework.stereotype.Component; import org.springframework.stereotype.Component;
import org.springframework.transaction.annotation.Transactional; import org.springframework.transaction.annotation.Transactional;
import org.springframework.beans.factory.annotation.Autowired;
import java.time.LocalDateTime; import java.time.LocalDateTime;
import java.util.List; import java.util.List;
@@ -32,19 +35,28 @@ public class JpaChatRunStore implements ChatRunStore {
private final ObjectMapper objectMapper; private final ObjectMapper objectMapper;
private final ObjectReader publishedReader; private final ObjectReader publishedReader;
private final PublishedResultPolicy publishedPolicy; private final PublishedResultPolicy publishedPolicy;
private final AgentStepAuditTracker stepTracker;
public JpaChatRunStore(ChatSessionRepository chatSessions, public JpaChatRunStore(ChatSessionRepository chatSessions,
DiagnosisRunRepository runs, DiagnosisRunRepository runs,
ObjectMapper objectMapper) { ObjectMapper objectMapper) {
this(chatSessions, runs, objectMapper, this(chatSessions, runs, objectMapper,
new PublishedResultPolicy(PreviousTurnLimits.defaults())); new PublishedResultPolicy(PreviousTurnLimits.defaults()), null);
}
public JpaChatRunStore(ChatSessionRepository chatSessions,
DiagnosisRunRepository runs,
ObjectMapper objectMapper,
PublishedResultPolicy publishedPolicy) {
this(chatSessions, runs, objectMapper, publishedPolicy, null);
} }
@Autowired @Autowired
public JpaChatRunStore(ChatSessionRepository chatSessions, public JpaChatRunStore(ChatSessionRepository chatSessions,
DiagnosisRunRepository runs, DiagnosisRunRepository runs,
ObjectMapper objectMapper, ObjectMapper objectMapper,
PublishedResultPolicy publishedPolicy) { PublishedResultPolicy publishedPolicy,
@Nullable AgentStepAuditTracker stepTracker) {
this.chatSessions = Objects.requireNonNull(chatSessions, "chatSessions must not be null"); this.chatSessions = Objects.requireNonNull(chatSessions, "chatSessions must not be null");
this.runs = Objects.requireNonNull(runs, "runs must not be null"); this.runs = Objects.requireNonNull(runs, "runs must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
@@ -52,6 +64,7 @@ public class JpaChatRunStore implements ChatRunStore {
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES) .with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS); .with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
this.publishedPolicy = Objects.requireNonNull(publishedPolicy, "publishedPolicy must not be null"); this.publishedPolicy = Objects.requireNonNull(publishedPolicy, "publishedPolicy must not be null");
this.stepTracker = stepTracker;
} }
@Override @Override
@@ -110,11 +123,13 @@ public class JpaChatRunStore implements ChatRunStore {
String safeContentJson, PublishedResult publishedResult, int durationMs) { String safeContentJson, PublishedResult publishedResult, int durationMs) {
Objects.requireNonNull(context, "context must not be null"); Objects.requireNonNull(context, "context must not be null");
Objects.requireNonNull(outcome, "outcome must not be null"); Objects.requireNonNull(outcome, "outcome must not be null");
try {
DiagnosisRun run = requiredRun(context.runId()); DiagnosisRun run = requiredRun(context.runId());
run.setIntent(intent); run.setIntent(intent);
run.setReleaseOutcome(outcome); run.setReleaseOutcome(outcome);
run.setStatus(status(outcome)); run.setStatus(status(outcome));
run.setAnswer(safeContentJson); run.setAnswer(safeContentJson);
run.setConclusion(RunConclusionExtractor.extract(objectMapper, safeContentJson));
run.setPublishedResult(outcome == ReleaseOutcome.SUCCESS && intent == IntentType.DIAGNOSIS run.setPublishedResult(outcome == ReleaseOutcome.SUCCESS && intent == IntentType.DIAGNOSIS
&& publishedResult != null ? write(publishedResult) : null); && publishedResult != null ? write(publishedResult) : null);
run.setTotalDurationMs(Math.max(0, durationMs)); run.setTotalDurationMs(Math.max(0, durationMs));
@@ -132,6 +147,11 @@ public class JpaChatRunStore implements ChatRunStore {
chatSessions.save(session); chatSessions.save(session);
}); });
} }
} finally {
if (stepTracker != null) {
stepTracker.clear(context.runId());
}
}
} }
private Optional<PreviousTurn> readPreviousTurn(DiagnosisRun run) { private Optional<PreviousTurn> readPreviousTurn(DiagnosisRun run) {
@@ -0,0 +1,39 @@
package com.superbiz.agent.harness.audit;
import org.springframework.stereotype.Component;
import java.util.concurrent.ConcurrentHashMap;
/**
* Binds the in-flight {@code agent_step.id} for a run so tool audits can set {@code step_id}.
*
* <p>Lifecycle: {@link #bind} on Agent {@code beforeModel} (step row created). Binding stays
* until the next {@code beforeModel} for the same run, covering tool execution that happens
* after {@code afterModel} emits tool_calls. {@link #clear} on run finish is optional cleanup.</p>
*/
@Component
public final class AgentStepAuditTracker {
private final ConcurrentHashMap<String, Long> currentStepIdByRun = new ConcurrentHashMap<>();
public void bind(String runId, Long stepId) {
if (runId == null || runId.isBlank() || stepId == null) {
return;
}
currentStepIdByRun.put(runId, stepId);
}
public Long currentStepId(String runId) {
if (runId == null || runId.isBlank()) {
return null;
}
return currentStepIdByRun.get(runId);
}
public void clear(String runId) {
if (runId == null || runId.isBlank()) {
return;
}
currentStepIdByRun.remove(runId);
}
}
@@ -7,32 +7,53 @@ import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook; import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
import com.fasterxml.jackson.core.JsonProcessingException; import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.AgentReasoningAudit; import com.superbiz.agent.domain.entity.AgentReasoningAudit;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.repository.AgentReasoningAuditRepository; import com.superbiz.agent.repository.AgentReasoningAuditRepository;
import com.superbiz.agent.repository.AgentStepRepository; import com.superbiz.agent.repository.AgentStepRepository;
import org.slf4j.Logger; import org.slf4j.Logger;
import org.slf4j.LoggerFactory; import org.slf4j.LoggerFactory;
import org.springframework.ai.chat.messages.AssistantMessage; import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.messages.Message; import org.springframework.ai.chat.messages.Message;
import org.springframework.ai.deepseek.DeepSeekAssistantMessage;
import java.nio.charset.StandardCharsets;
import java.lang.reflect.Method;
import java.util.ArrayList;
import java.util.LinkedHashMap; import java.util.LinkedHashMap;
import java.util.List; import java.util.List;
import java.util.Map; import java.util.Map;
import java.nio.charset.StandardCharsets;
import java.util.Objects; import java.util.Objects;
import java.util.concurrent.ConcurrentHashMap; import java.util.concurrent.ConcurrentHashMap;
/**
* Persists per-model-step audit for the diagnosis agent.
*
* <p>LLM-facing content policy for {@link AgentReasoningAudit}:</p>
* <ul>
* <li>{@code reasoning_content}: provider thinking/CoT when present</li>
* <li>{@code assistant_text}: assistant visible text and/or tool-call plan</li>
* <li>Tool <em>result</em> payloads are never stored here (use {@code tool_invocation})</li>
* </ul>
*/
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL}) @HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
public final class HarnessAgentAuditHook extends MessagesModelHook { public final class HarnessAgentAuditHook extends MessagesModelHook {
private static final Logger log = LoggerFactory.getLogger(HarnessAgentAuditHook.class); private static final Logger log = LoggerFactory.getLogger(HarnessAgentAuditHook.class);
private static final int MAX_TEXT_CHARS = 32_000;
public static final String SOURCE_BOTH = "PROVIDER_REASONING+ASSISTANT_TEXT";
public static final String SOURCE_PROVIDER = "PROVIDER_REASONING";
public static final String SOURCE_ASSISTANT = "ASSISTANT_TEXT";
public static final String SOURCE_TOOL_PLAN = "TOOL_CALL_PLAN";
public static final String SOURCE_NONE = "NONE";
private final AgentStepRepository repository; private final AgentStepRepository repository;
private final ObjectMapper objectMapper; private final ObjectMapper objectMapper;
private final String agentName; private final String agentName;
private final DiagnosisTraceRecorder traceRecorder; private final DiagnosisTraceRecorder traceRecorder;
private final AgentReasoningAuditRepository reasoningRepository; private final AgentReasoningAuditRepository reasoningRepository;
private final AgentStepAuditTracker stepTracker;
private final ConcurrentHashMap<String, Integer> stepCounters = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String, Integer> stepCounters = new ConcurrentHashMap<>();
private final ConcurrentHashMap<String, PendingStep> pendingSteps = new ConcurrentHashMap<>(); private final ConcurrentHashMap<String, PendingStep> pendingSteps = new ConcurrentHashMap<>();
@@ -42,12 +63,19 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper, public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
String agentName, DiagnosisTraceRecorder traceRecorder) { String agentName, DiagnosisTraceRecorder traceRecorder) {
this(repository, objectMapper, agentName, traceRecorder, null); this(repository, objectMapper, agentName, traceRecorder, null, null);
} }
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper, public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
String agentName, DiagnosisTraceRecorder traceRecorder, String agentName, DiagnosisTraceRecorder traceRecorder,
AgentReasoningAuditRepository reasoningRepository) { AgentReasoningAuditRepository reasoningRepository) {
this(repository, objectMapper, agentName, traceRecorder, reasoningRepository, null);
}
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
String agentName, DiagnosisTraceRecorder traceRecorder,
AgentReasoningAuditRepository reasoningRepository,
AgentStepAuditTracker stepTracker) {
this.repository = Objects.requireNonNull(repository, "repository must not be null"); this.repository = Objects.requireNonNull(repository, "repository must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
if (agentName == null || agentName.isBlank()) { if (agentName == null || agentName.isBlank()) {
@@ -56,6 +84,7 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
this.agentName = agentName; this.agentName = agentName;
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null"); this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
this.reasoningRepository = reasoningRepository; this.reasoningRepository = reasoningRepository;
this.stepTracker = stepTracker;
} }
@Override @Override
@@ -82,6 +111,9 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
.hasToolCall(false) .hasToolCall(false)
.build()); .build());
pending = new PendingStep(saved.getId(), pending.startedNanos(), pending.input()); pending = new PendingStep(saved.getId(), pending.startedNanos(), pending.input());
if (stepTracker != null && saved.getId() != null) {
stepTracker.bind(identity.runId(), saved.getId());
}
} catch (RuntimeException exception) { } catch (RuntimeException exception) {
log.warn("Failed to persist AgentStep audit before model: agent={}", agentName); log.warn("Failed to persist AgentStep audit before model: agent={}", agentName);
} }
@@ -106,14 +138,15 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
? List.of() ? List.of()
: assistant.getToolCalls().stream().map(AssistantMessage.ToolCall::name) : assistant.getToolCalls().stream().map(AssistantMessage.ToolCall::name)
.distinct().sorted().toList(); .distinct().sorted().toList();
ReasoningContent reasoning = reasoningContent(assistant); LlmTurnContent turn = llmTurnContent(assistant, toolNames);
Map<String, Object> output = outputMetadata(assistant, toolNames, reasoning); Map<String, Object> output = outputMetadata(assistant, toolNames, turn);
int durationMs = durationMillis(pending.startedNanos()); int durationMs = durationMillis(pending.startedNanos());
try { try {
AgentStep step = pending.id() == null ? null : repository.findById(pending.id()).orElse(null); AgentStep step = pending.id() == null ? null : repository.findById(pending.id()).orElse(null);
if (step != null) { if (step != null) {
step.setModelOutput(write(output)); step.setModelOutput(write(output));
step.setThought(null); // Keep thought aligned with provider reasoning when present; else assistant text.
step.setThought(firstNonBlank(turn.reasoningContent(), turn.assistantText()));
step.setHasToolCall(!toolNames.isEmpty()); step.setHasToolCall(!toolNames.isEmpty());
step.setDurationMs(durationMs); step.setDurationMs(durationMs);
repository.save(step); repository.save(step);
@@ -121,7 +154,7 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
} catch (RuntimeException exception) { } catch (RuntimeException exception) {
log.warn("Failed to complete AgentStep audit: agent={}", agentName); log.warn("Failed to complete AgentStep audit: agent={}", agentName);
} }
persistReasoning(identity, stepIndex, reasoning); persistReasoning(identity, stepIndex, turn);
traceRecorder.record(TraceAuditEvents.agentModelStep( traceRecorder.record(TraceAuditEvents.agentModelStep(
identity.sessionId(), identity.runId(), agentName, stepIndex, identity.sessionId(), identity.runId(), agentName, stepIndex,
durationMs, pending.input(), output)); durationMs, pending.input(), output));
@@ -139,32 +172,148 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
} }
private Map<String, Object> outputMetadata(AssistantMessage assistant, List<String> toolNames, private Map<String, Object> outputMetadata(AssistantMessage assistant, List<String> toolNames,
ReasoningContent reasoning) { LlmTurnContent turn) {
Map<String, Object> metadata = new LinkedHashMap<>(); Map<String, Object> metadata = new LinkedHashMap<>();
metadata.put("has_text", assistant != null metadata.put("has_text", assistant != null
&& assistant.getText() != null && !assistant.getText().isBlank()); && assistant.getText() != null && !assistant.getText().isBlank());
metadata.put("tool_names", toolNames); metadata.put("tool_names", toolNames);
metadata.put("reasoning_available", reasoning.available()); metadata.put("reasoning_available", turn.reasoningAvailable());
metadata.put("reasoning_bytes", reasoning.bytes()); metadata.put("reasoning_bytes", turn.reasoningBytes());
metadata.put("assistant_bytes", turn.assistantBytes());
metadata.put("content_source", turn.contentSource());
return metadata; return metadata;
} }
private ReasoningContent reasoningContent(AssistantMessage assistant) { /**
if (assistant == null || assistant.getMetadata() == null) { * Captures both provider reasoning and assistant-visible text.
return ReasoningContent.empty(); * Tool <em>results</em> are never included.
} */
for (String key : List.of("reasoning_content", "reasoningContent", "reasoning", "thinking")) { static LlmTurnContent llmTurnContent(AssistantMessage assistant, List<String> toolNames) {
Object value = assistant.getMetadata().get(key); String reasoning = providerReasoning(assistant);
if (value instanceof CharSequence text && !text.toString().isBlank()) { String assistantBody = assistantText(assistant);
String content = bounded(text.toString()); String toolPlan = toolCallPlan(assistant, toolNames);
return new ReasoningContent(content, true,
content.getBytes(StandardCharsets.UTF_8).length); String assistantCombined = joinNonBlank("\n", assistantBody, toolPlan);
} boolean hasReasoning = hasText(reasoning);
} boolean hasAssistant = hasText(assistantCombined);
return ReasoningContent.empty();
String source;
if (hasReasoning && hasAssistant) {
source = SOURCE_BOTH;
} else if (hasReasoning) {
source = SOURCE_PROVIDER;
} else if (hasText(assistantBody)) {
source = SOURCE_ASSISTANT;
} else if (hasText(toolPlan)) {
source = SOURCE_TOOL_PLAN;
} else {
source = SOURCE_NONE;
} }
private void persistReasoning(AuditIdentity identity, int stepIndex, ReasoningContent reasoning) { String reasoningBound = bound(reasoning);
String assistantBound = bound(assistantCombined);
int bytes = utf8Bytes(reasoningBound) + utf8Bytes(assistantBound);
return new LlmTurnContent(
hasReasoning,
reasoningBound,
assistantBound,
source,
utf8Bytes(reasoningBound),
utf8Bytes(assistantBound),
bytes);
}
/**
* DeepSeek puts CoT on {@link DeepSeekAssistantMessage#getReasoningContent()},
* <em>not</em> on {@link AssistantMessage#getMetadata()}. Older docs/tests used metadata keys;
* keep those as fallback for mocks and non-DeepSeek providers.
*/
static String providerReasoning(AssistantMessage assistant) {
if (assistant == null) {
return null;
}
// 1) Native DeepSeek message field (primary path in production)
if (assistant instanceof DeepSeekAssistantMessage deepSeek) {
String nativeReasoning = blankToNull(deepSeek.getReasoningContent());
if (nativeReasoning != null) {
return nativeReasoning;
}
}
// 2) Reflective getReasoningContent() for subclasses / reloaded types
String reflective = invokeReasoningGetter(assistant);
if (reflective != null) {
return reflective;
}
// 3) Metadata keys (tests / other providers)
Map<String, Object> metadata = assistant.getMetadata();
if (metadata != null) {
for (String key : List.of(
"reasoning_content", "reasoningContent", "reasoning", "thinking",
"reasoning_text", "reasoningText")) {
Object value = metadata.get(key);
if (value instanceof CharSequence text) {
String trimmed = blankToNull(text.toString());
if (trimmed != null) {
return trimmed;
}
}
}
}
return null;
}
private static String invokeReasoningGetter(AssistantMessage assistant) {
try {
Method method = assistant.getClass().getMethod("getReasoningContent");
Object value = method.invoke(assistant);
return value instanceof CharSequence text ? blankToNull(text.toString()) : null;
} catch (ReflectiveOperationException ignored) {
return null;
}
}
private static String blankToNull(String value) {
if (value == null) {
return null;
}
String trimmed = value.trim();
return trimmed.isEmpty() ? null : trimmed;
}
private static String assistantText(AssistantMessage assistant) {
if (assistant == null || assistant.getText() == null || assistant.getText().isBlank()) {
return null;
}
return assistant.getText().trim();
}
/**
* Records which tools the model decided to call (names + arg preview), not tool outputs.
*/
private static String toolCallPlan(AssistantMessage assistant, List<String> toolNames) {
if (assistant == null || assistant.getToolCalls() == null || assistant.getToolCalls().isEmpty()) {
return null;
}
List<String> lines = new ArrayList<>();
lines.add("tool_calls:");
for (AssistantMessage.ToolCall call : assistant.getToolCalls()) {
if (call == null) {
continue;
}
String name = call.name() == null ? "?" : call.name();
String args = call.arguments() == null ? "" : call.arguments().trim();
if (args.length() > 500) {
args = args.substring(0, 500) + "...";
}
lines.add("- " + name + (args.isEmpty() ? "" : " args=" + args));
}
if (lines.size() == 1 && toolNames != null && !toolNames.isEmpty()) {
lines.add("- " + String.join(", ", toolNames));
}
return lines.size() <= 1 ? null : String.join("\n", lines);
}
private void persistReasoning(AuditIdentity identity, int stepIndex, LlmTurnContent turn) {
if (reasoningRepository == null) { if (reasoningRepository == null) {
return; return;
} }
@@ -174,18 +323,51 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
.runId(identity.runId()) .runId(identity.runId())
.stepIndex(stepIndex) .stepIndex(stepIndex)
.agentName(agentName) .agentName(agentName)
.reasoningAvailable(reasoning.available()) .reasoningAvailable(turn.reasoningAvailable())
.reasoningContent(reasoning.content()) .reasoningContent(turn.reasoningContent())
.contentBytes(reasoning.bytes()) .assistantText(turn.assistantText())
.contentSource(turn.contentSource())
.contentBytes(turn.totalBytes())
.build()); .build());
} catch (RuntimeException exception) { } catch (RuntimeException exception) {
log.warn("Failed to persist reasoning audit: agent={}, step={}", agentName, stepIndex); log.warn("Failed to persist reasoning audit: agent={}, step={}", agentName, stepIndex);
} }
} }
private static String bounded(String value) { private static String bound(String value) {
int maxChars = 32_000; if (value == null) {
return value.length() <= maxChars ? value : value.substring(0, maxChars); return null;
}
return value.length() <= MAX_TEXT_CHARS ? value : value.substring(0, MAX_TEXT_CHARS);
}
private static int utf8Bytes(String value) {
return value == null ? 0 : value.getBytes(StandardCharsets.UTF_8).length;
}
private static String joinNonBlank(String sep, String a, String b) {
boolean ha = hasText(a);
boolean hb = hasText(b);
if (ha && hb) {
return a + sep + b;
}
if (ha) {
return a;
}
if (hb) {
return b;
}
return null;
}
private static String firstNonBlank(String a, String b) {
if (hasText(a)) {
return a;
}
if (hasText(b)) {
return b;
}
return null;
} }
private AuditIdentity identity(RunnableConfig config) { private AuditIdentity identity(RunnableConfig config) {
@@ -233,10 +415,14 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
private record AuditIdentity(String sessionId, String runId) { private record AuditIdentity(String sessionId, String runId) {
} }
private record ReasoningContent(String content, boolean available, int bytes) { record LlmTurnContent(
private static ReasoningContent empty() { boolean reasoningAvailable,
return new ReasoningContent(null, false, 0); String reasoningContent,
} String assistantText,
String contentSource,
int reasoningBytes,
int assistantBytes,
int totalBytes) {
} }
private record PendingStep(Long id, long startedNanos, Map<String, Object> input) { private record PendingStep(Long id, long startedNanos, Map<String, Object> input) {
@@ -1,66 +1,247 @@
package com.superbiz.agent.harness.audit; package com.superbiz.agent.harness.audit;
import com.fasterxml.jackson.core.JsonProcessingException; import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation; import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.harness.contract.InvocationStatus; import com.superbiz.agent.harness.contract.InvocationStatus;
import com.superbiz.agent.repository.ToolInvocationRepository; import com.superbiz.agent.repository.ToolInvocationRepository;
import org.springframework.beans.factory.ObjectProvider;
import org.springframework.beans.factory.annotation.Autowired; import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.stereotype.Component; import org.springframework.stereotype.Component;
import java.util.ArrayList;
import java.util.Iterator;
import java.util.LinkedHashMap; import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map; import java.util.Map;
import java.util.Objects; import java.util.Objects;
import java.util.Set;
/**
* Persists tool invocation audit rows.
*
* <p>For {@code lookup_knowledge}, enriches RAG-specific columns from internal LookupResult
* JSON when present. {@code relevance_level} stores RAG PRECISE/REFERENCE (not evidence_status).
* {@code evidence_status} remains in {@code retrieval_details} for all tools.</p>
*/
@Component @Component
public final class JpaToolInvocationAuditSink implements ToolInvocationAuditSink { public final class JpaToolInvocationAuditSink implements ToolInvocationAuditSink {
/** Max chars stored for free-text query / sql / topic style fields. */
private static final int MAX_TEXT_PREVIEW = 160;
private static final int MAX_REQUEST_KEYS = 12;
private static final Set<String> SENSITIVE_KEY_FRAGMENTS = Set.of(
"password", "passwd", "secret", "token", "apikey", "api_key",
"authorization", "credential", "private_key", "access_key");
private final ToolInvocationRepository repository; private final ToolInvocationRepository repository;
private final ObjectMapper objectMapper; private final ObjectMapper objectMapper;
private final DiagnosisTraceRecorder traceRecorder; private final DiagnosisTraceRecorder traceRecorder;
private final RagLookupAuditEnricher ragEnricher;
public JpaToolInvocationAuditSink(ToolInvocationRepository repository, ObjectMapper objectMapper) {
this(repository, objectMapper, DiagnosisTraceRecorder.noop());
}
@Autowired @Autowired
public JpaToolInvocationAuditSink(ToolInvocationRepository repository, ObjectMapper objectMapper, public JpaToolInvocationAuditSink(ToolInvocationRepository repository,
DiagnosisTraceRecorder traceRecorder) { ObjectMapper objectMapper,
DiagnosisTraceRecorder traceRecorder,
ObjectProvider<RagLookupAuditEnricher> ragEnricherProvider) {
this.repository = Objects.requireNonNull(repository, "repository must not be null"); this.repository = Objects.requireNonNull(repository, "repository must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null"); this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
this.ragEnricher = ragEnricherProvider == null ? null : ragEnricherProvider.getIfAvailable();
}
/** Test helper with explicit enricher (may be null). */
static JpaToolInvocationAuditSink forTest(ToolInvocationRepository repository,
ObjectMapper objectMapper,
DiagnosisTraceRecorder traceRecorder,
RagLookupAuditEnricher ragEnricher) {
return new JpaToolInvocationAuditSink(repository, objectMapper, traceRecorder, ragEnricher, true);
}
private JpaToolInvocationAuditSink(ToolInvocationRepository repository,
ObjectMapper objectMapper,
DiagnosisTraceRecorder traceRecorder,
RagLookupAuditEnricher ragEnricher,
boolean testMarker) {
this.repository = Objects.requireNonNull(repository, "repository must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
this.ragEnricher = ragEnricher;
} }
@Override @Override
public void record(ToolInvocationAuditEvent event) { public void record(ToolInvocationAuditEvent event) {
Objects.requireNonNull(event, "event must not be null"); Objects.requireNonNull(event, "event must not be null");
traceRecorder.record(TraceAuditEvents.toolInvocation(event)); traceRecorder.record(TraceAuditEvents.toolInvocation(event));
Map<String, Object> details = baseResultMetadata(event);
String retrievalLayer = "HARNESS";
Integer l0 = null;
Integer l1 = null;
String relevanceLevel = null;
Boolean truncated = false;
String dedupReason = null;
String outputPreview = "status=%s,evidence_status=%s".formatted(
event.status(), event.evidenceStatus());
if (ragEnricher != null && ragEnricher.supports(event.toolName())) {
RagLookupAuditEnricher.Enrichment enrichment =
ragEnricher.enrich(event.rawResultJson(), event.agentResultJson());
if (enrichment.retrievalDetails() != null) {
details.putAll(enrichment.retrievalDetails());
}
if (enrichment.retrievalLayer() != null && !enrichment.retrievalLayer().isBlank()) {
retrievalLayer = enrichment.retrievalLayer();
}
l0 = enrichment.l0MatchCount();
l1 = enrichment.l1MatchCount();
// RAG semantic level only — never overwrite with evidence_status enum names
relevanceLevel = enrichment.relevanceLevel();
if (enrichment.truncated() != null) {
truncated = enrichment.truncated();
}
dedupReason = enrichment.dedupReason();
if (enrichment.outputPreview() != null && !enrichment.outputPreview().isBlank()) {
outputPreview = enrichment.outputPreview()
+ ",status=" + event.status()
+ ",evidence_status=" + event.evidenceStatus();
}
}
// Always keep harness evidence_status in details (column relevance_level is RAG-only when set)
details.put("evidence_status", event.evidenceStatus().name());
details.put("invocation_status", event.status().name());
repository.save(ToolInvocation.builder() repository.save(ToolInvocation.builder()
.sessionId(event.sessionId()) .sessionId(event.sessionId())
.runId(event.runId()) .runId(event.runId())
.stepId(event.stepId())
.toolName(event.toolName()) .toolName(event.toolName())
.inputParams(write(inputMetadata(event))) .inputParams(write(inputMetadata(event)))
.outputPreview("status=%s,evidence_status=%s".formatted( .outputPreview(outputPreview)
event.status(), event.evidenceStatus()))
.outputLength(event.agentResultBytes()) .outputLength(event.agentResultBytes())
.retrievalLayer("HARNESS") .retrievalLayer(retrievalLayer)
.isTruncated(false) .l0MatchCount(l0)
.relevanceLevel(event.evidenceStatus().name()) .l1MatchCount(l1)
.retrievalDetails(write(resultMetadata(event))) .isTruncated(Boolean.TRUE.equals(truncated))
.relevanceLevel(relevanceLevel)
.dedupReason(dedupReason)
.retrievalDetails(write(details))
.durationMs(event.durationMs()) .durationMs(event.durationMs())
.success(event.status() == InvocationStatus.READY) .success(event.status() == InvocationStatus.READY)
.errorMessage(event.errorCode()) .errorMessage(event.errorCode())
.build()); .build());
} }
private Map<String, Object> inputMetadata(ToolInvocationAuditEvent event) { /**
* Bounded request audit: always tool_call_id + request_bytes; optionally step_id and
* safe scalar fields from request JSON (query/topic/region/limit/...). Never stores
* password/token-like keys or nested blobs wholesale.
*/
Map<String, Object> inputMetadata(ToolInvocationAuditEvent event) {
Map<String, Object> metadata = new LinkedHashMap<>(); Map<String, Object> metadata = new LinkedHashMap<>();
metadata.put("tool_call_id", event.toolCallId()); metadata.put("tool_call_id", event.toolCallId());
metadata.put("request_bytes", event.requestBytes()); metadata.put("request_bytes", event.requestBytes());
if (event.stepId() != null) {
metadata.put("step_id", event.stepId());
}
appendSafeRequestFields(metadata, event.requestJson());
return metadata; return metadata;
} }
private Map<String, Object> resultMetadata(ToolInvocationAuditEvent event) { private void appendSafeRequestFields(Map<String, Object> metadata, String requestJson) {
if (requestJson == null || requestJson.isBlank()) {
return;
}
try {
JsonNode root = objectMapper.readTree(requestJson);
if (root == null || !root.isObject()) {
return;
}
int added = 0;
Iterator<Map.Entry<String, JsonNode>> fields = root.fields();
while (fields.hasNext() && added < MAX_REQUEST_KEYS) {
Map.Entry<String, JsonNode> entry = fields.next();
String key = entry.getKey();
if (key == null || key.isBlank() || isSensitiveKey(key)) {
continue;
}
JsonNode value = entry.getValue();
if (value == null || value.isNull()) {
continue;
}
if (value.isTextual()) {
String text = value.asText();
if (text == null || text.isBlank()) {
continue;
}
metadata.put(key, truncate(text, MAX_TEXT_PREVIEW));
if (text.length() > MAX_TEXT_PREVIEW) {
metadata.put(key + "_truncated", true);
metadata.put(key + "_chars", text.length());
}
added++;
} else if (value.isNumber()) {
metadata.put(key, value.numberValue());
added++;
} else if (value.isBoolean()) {
metadata.put(key, value.booleanValue());
added++;
} else if (value.isArray() && isStringArray(value)) {
List<String> items = new ArrayList<>();
for (int i = 0; i < value.size() && items.size() < 8; i++) {
JsonNode item = value.get(i);
if (item != null && item.isTextual() && !item.asText().isBlank()) {
items.add(truncate(item.asText(), 64));
}
}
if (!items.isEmpty()) {
metadata.put(key, items);
added++;
}
}
// objects / mixed arrays intentionally omitted
}
} catch (Exception ignored) {
metadata.put("request_parse", "failed");
}
}
private static boolean isStringArray(JsonNode value) {
if (value == null || !value.isArray() || value.isEmpty()) {
return false;
}
for (JsonNode n : value) {
if (n == null || !n.isTextual()) {
return false;
}
}
return true;
}
private static boolean isSensitiveKey(String key) {
String normalized = key.toLowerCase(Locale.ROOT);
for (String fragment : SENSITIVE_KEY_FRAGMENTS) {
if (normalized.contains(fragment)) {
return true;
}
}
return false;
}
private static String truncate(String value, int maxChars) {
if (value == null) {
return null;
}
if (value.length() <= maxChars) {
return value;
}
return value.substring(0, maxChars);
}
private Map<String, Object> baseResultMetadata(ToolInvocationAuditEvent event) {
Map<String, Object> metadata = new LinkedHashMap<>(); Map<String, Object> metadata = new LinkedHashMap<>();
metadata.put("tool_call_id", event.toolCallId()); metadata.put("tool_call_id", event.toolCallId());
metadata.put("status", event.status().name()); metadata.put("status", event.status().name());
@@ -0,0 +1,335 @@
package com.superbiz.agent.harness.audit;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.springframework.beans.factory.annotation.Value;
import org.springframework.stereotype.Component;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Objects;
/**
* Builds durable RAG audit fields from internal {@code LookupResult} JSON
* (and optional projected agent result) without dumping full excerpts/traces.
*/
@Component
public class RagLookupAuditEnricher {
public static final String TOOL_LOOKUP_KNOWLEDGE = "lookup_knowledge";
private static final int MAX_EVIDENCE_KEYS = 12;
private static final int MAX_HINT_VALUES = 8;
private static final int PREVIEW_CHARS = 160;
private final ObjectMapper objectMapper;
private final String searchMode;
public RagLookupAuditEnricher(
ObjectMapper objectMapper,
@Value("${retrieval.search.mode:hybrid}") String searchMode) {
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper");
this.searchMode = searchMode == null || searchMode.isBlank()
? "hybrid"
: searchMode.trim().toLowerCase(Locale.ROOT);
}
public boolean supports(String toolName) {
return TOOL_LOOKUP_KNOWLEDGE.equals(toolName);
}
/**
* @param rawLookupJson internal LookupResult JSON (may be null on hard failures)
* @param agentResultJson projected RagToolResult JSON (may be null)
*/
public Enrichment enrich(String rawLookupJson, String agentResultJson) {
Map<String, Object> details = new LinkedHashMap<>();
details.put("audit_schema", "rag_lookup_v1");
details.put("search_mode", searchMode);
String retrievalLayer = "L1";
Integer l0MatchCount = null;
Integer l1MatchCount = null;
String relevanceLevel = null;
Boolean truncated = null;
String dedupReason = null;
String outputPreview = null;
try {
if (rawLookupJson != null && !rawLookupJson.isBlank()) {
JsonNode root = objectMapper.readTree(rawLookupJson);
if (root != null && root.isObject()) {
relevanceLevel = text(root, "relevanceLevel");
if (relevanceLevel == null) {
relevanceLevel = text(root, "relevance_level");
}
JsonNode blocks = root.get("evidenceBlocks");
if (blocks != null && blocks.isArray()) {
l1MatchCount = blocks.size();
List<String> keys = new ArrayList<>();
List<String> sources = new ArrayList<>();
String layer = null;
for (int i = 0; i < blocks.size() && keys.size() < MAX_EVIDENCE_KEYS; i++) {
JsonNode b = blocks.get(i);
if (b == null || !b.isObject()) {
continue;
}
String key = firstText(b, "evidenceKey", "evidence_key");
if (key != null) {
keys.add(key);
}
String source = text(b, "source");
if (source != null && sources.size() < MAX_EVIDENCE_KEYS) {
sources.add(source);
}
if (layer == null) {
layer = text(b, "retrievalLayer");
}
}
if (!keys.isEmpty()) {
details.put("evidence_keys", keys);
}
if (!sources.isEmpty()) {
details.put("sources", sources);
}
if (layer != null && !layer.isBlank()) {
retrievalLayer = layer;
}
}
Integer candidateCount = intVal(root, "evidenceCandidateCount");
if (candidateCount != null) {
details.put("evidence_candidate_count", candidateCount);
if (l1MatchCount == null) {
l1MatchCount = candidateCount;
}
}
Integer blockCount = intVal(root, "evidenceBlockCount");
if (blockCount != null) {
details.put("evidence_block_count", blockCount);
}
String completenessHint = text(root, "completenessHint");
if (completenessHint != null) {
details.put("completeness_hint", truncate(completenessHint, PREVIEW_CHARS));
}
JsonNode trace = root.get("retrievalTrace");
if (trace != null && trace.isObject()) {
putText(details, "selected_attempt", text(trace, "selectedAttempt"));
putText(details, "fallback_reason", text(trace, "fallbackReason"));
putText(details, "retrieval_evidence_status", text(trace, "evidenceStatus"));
putText(details, "category_filter", text(trace, "categoryFilter"));
// query texts intentionally omitted from durable audit by default (PII/size)
JsonNode hints = trace.get("queryHints");
if (hints != null && hints.isObject()) {
Integer l0 = intVal(hints, "l0_match_count");
if (l0 == null) {
l0 = arraySize(hints.get("domains"));
}
l0MatchCount = l0;
Map<String, Object> hintSnap = new LinkedHashMap<>();
putLimitedList(hintSnap, "domains", hints.get("domains"));
putLimitedList(hintSnap, "matched_keywords", hints.get("matched_keywords"));
if (hintSnap.isEmpty()) {
putLimitedList(hintSnap, "matched_keywords", hints.get("matchedKeywords"));
}
if (!hintSnap.isEmpty()) {
details.put("l0_hints", hintSnap);
}
}
JsonNode attempts = trace.get("attempts");
if (attempts != null && attempts.isArray()) {
List<Map<String, Object>> attemptSnap = new ArrayList<>();
for (JsonNode a : attempts) {
if (a == null || !a.isObject()) {
continue;
}
Map<String, Object> row = new LinkedHashMap<>();
putText(row, "name", text(a, "name"));
putText(row, "category_filter", text(a, "categoryFilter"));
if (a.has("candidateCount") && a.get("candidateCount").canConvertToInt()) {
row.put("candidate_count", a.get("candidateCount").asInt());
}
if (a.has("usable") && a.get("usable").isBoolean()) {
row.put("usable", a.get("usable").asBoolean());
}
if (a.has("topSimilarity") && a.get("topSimilarity").isNumber()) {
row.put("top_similarity", a.get("topSimilarity").asDouble());
}
if (a.has("durationMs") && a.get("durationMs").canConvertToInt()) {
row.put("duration_ms", a.get("durationMs").asInt());
}
if (!row.isEmpty()) {
attemptSnap.add(row);
}
}
if (!attemptSnap.isEmpty()) {
details.put("attempts", attemptSnap);
}
}
}
JsonNode pack = root.get("contextPack");
if (pack != null && pack.isObject()) {
if (pack.has("usedChars") && pack.get("usedChars").canConvertToInt()) {
details.put("context_used_chars", pack.get("usedChars").asInt());
}
if (pack.has("charBudget") && pack.get("charBudget").canConvertToInt()) {
details.put("context_char_budget", pack.get("charBudget").asInt());
}
}
// compact preview for list UIs
outputPreview = buildPreview(relevanceLevel, details);
}
}
} catch (Exception ignored) {
details.put("enrich_error", "lookup_result_parse_failed");
}
try {
if (agentResultJson != null && !agentResultJson.isBlank()) {
JsonNode agent = objectMapper.readTree(agentResultJson);
if (agent != null && agent.isObject()) {
if (agent.has("truncated") && agent.get("truncated").isBoolean()) {
truncated = agent.get("truncated").asBoolean();
details.put("truncated", truncated);
}
if (relevanceLevel == null) {
relevanceLevel = text(agent, "relevance_level");
if (relevanceLevel == null) {
relevanceLevel = text(agent, "relevanceLevel");
}
}
if (agent.has("returned_count") && agent.get("returned_count").canConvertToInt()) {
details.put("returned_count", agent.get("returned_count").asInt());
}
}
}
} catch (Exception ignored) {
details.put("agent_result_parse", "failed");
}
if (outputPreview == null) {
outputPreview = buildPreview(relevanceLevel, details);
}
return new Enrichment(
retrievalLayer,
l0MatchCount,
l1MatchCount,
relevanceLevel,
truncated,
dedupReason,
details,
outputPreview
);
}
private static String buildPreview(String relevanceLevel, Map<String, Object> details) {
String attempt = details.get("selected_attempt") == null ? null : String.valueOf(details.get("selected_attempt"));
String fallback = details.get("fallback_reason") == null ? null : String.valueOf(details.get("fallback_reason"));
StringBuilder sb = new StringBuilder("lookup_knowledge");
if (relevanceLevel != null) {
sb.append(" level=").append(relevanceLevel);
}
if (attempt != null) {
sb.append(" attempt=").append(attempt);
}
if (fallback != null) {
sb.append(" fallback=").append(fallback);
}
Object keys = details.get("evidence_keys");
if (keys instanceof List<?> list) {
sb.append(" keys=").append(list.size());
}
return truncate(sb.toString(), PREVIEW_CHARS);
}
private static void putLimitedList(Map<String, Object> target, String key, JsonNode node) {
if (node == null || !node.isArray() || node.isEmpty()) {
return;
}
List<String> values = new ArrayList<>();
for (int i = 0; i < node.size() && values.size() < MAX_HINT_VALUES; i++) {
JsonNode n = node.get(i);
if (n != null && n.isTextual() && !n.asText().isBlank()) {
values.add(n.asText());
}
}
if (!values.isEmpty()) {
target.put(key, values);
}
}
private static void putText(Map<String, Object> map, String key, String value) {
if (value != null && !value.isBlank()) {
map.put(key, value);
}
}
private static String firstText(JsonNode node, String... fields) {
for (String f : fields) {
String v = text(node, f);
if (v != null) {
return v;
}
}
return null;
}
private static String text(JsonNode node, String field) {
if (node == null || field == null || !node.has(field) || node.get(field).isNull()) {
return null;
}
String v = node.get(field).asText(null);
return v == null || v.isBlank() ? null : v;
}
private static Integer intVal(JsonNode node, String field) {
if (node == null || !node.has(field) || node.get(field).isNull()) {
return null;
}
JsonNode n = node.get(field);
if (n.isIntegralNumber() || n.canConvertToInt()) {
return n.asInt();
}
return null;
}
private static Integer arraySize(JsonNode node) {
if (node != null && node.isArray()) {
return node.size();
}
return null;
}
private static String truncate(String value, int max) {
if (value == null) {
return null;
}
if (value.length() <= max) {
return value;
}
return value.substring(0, max) + "...";
}
public record Enrichment(
String retrievalLayer,
Integer l0MatchCount,
Integer l1MatchCount,
String relevanceLevel,
Boolean truncated,
String dedupReason,
Map<String, Object> retrievalDetails,
String outputPreview
) {
}
}
@@ -0,0 +1,114 @@
package com.superbiz.agent.harness.audit;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
/**
* Extracts a short human-readable conclusion string from released safe content JSON
* for {@code diagnosis_run.conclusion}.
*/
public final class RunConclusionExtractor {
private static final int MAX_CHARS = 4_000;
private RunConclusionExtractor() {
}
public static String extract(ObjectMapper objectMapper, String safeContentJson) {
if (safeContentJson == null || safeContentJson.isBlank()) {
return null;
}
try {
JsonNode root = objectMapper.readTree(safeContentJson);
if (root == null || !root.isObject()) {
return bound(safeContentJson.trim());
}
// DIAGNOSIS_REPORT payload shape stored in answer
String fromReport = text(root.path("report").path("conclusion").path("text"));
if (fromReport != null) {
return bound(fromReport);
}
// nested content envelope (defensive)
String nested = text(root.path("payload").path("report").path("conclusion").path("text"));
if (nested != null) {
return bound(nested);
}
// SAFE_FALLBACK envelope: {"fallback":{type,message,...}}
String fromFallback = fallbackConclusion(root.path("fallback"));
if (fromFallback != null) {
return bound(fromFallback);
}
String nestedFb = fallbackConclusion(root.path("payload").path("fallback"));
if (nestedFb != null) {
return bound(nestedFb);
}
// bare SafeFallback object (defensive)
if (root.hasNonNull("type") || root.has("message") || root.has("conclusion")) {
String bare = fallbackConclusion(root);
if (bare != null) {
return bound(bare);
}
}
// knowledge / plain answer shapes
String plain = text(root.path("answer"));
if (plain != null) {
return bound(plain);
}
return null;
} catch (Exception ignored) {
return bound(safeContentJson.trim());
}
}
/**
* Builds a short conclusion from a SafeFallback-shaped node.
* Prefer explicit conclusion text, else "type: message", else either alone.
*/
private static String fallbackConclusion(JsonNode node) {
if (node == null || node.isMissingNode() || node.isNull() || !node.isObject()) {
return null;
}
String conclusion = text(node.path("conclusion"));
if (conclusion != null) {
return conclusion;
}
String type = enumOrText(node.path("type"));
String message = text(node.path("message"));
if (type != null && message != null) {
return type + ": " + message;
}
if (message != null) {
return message;
}
return type;
}
private static String enumOrText(JsonNode node) {
if (node == null || node.isMissingNode() || node.isNull()) {
return null;
}
if (node.isTextual() || node.isNumber() || node.isBoolean()) {
String value = node.asText();
return value == null || value.isBlank() ? null : value.trim();
}
return null;
}
private static String text(JsonNode node) {
if (node == null || node.isMissingNode() || node.isNull()) {
return null;
}
if (!node.isTextual()) {
return null;
}
String value = node.asText();
return value == null || value.isBlank() ? null : value.trim();
}
private static String bound(String value) {
if (value == null) {
return null;
}
return value.length() <= MAX_CHARS ? value : value.substring(0, MAX_CHARS);
}
}
@@ -5,6 +5,17 @@ import com.superbiz.agent.harness.contract.InvocationStatus;
import java.util.Objects; import java.util.Objects;
/**
* Durable tool-invocation audit event.
*
* <p>{@code rawResultJson} is optional internal executor output (e.g. full {@code LookupResult}
* before Agent projection). It is used only to enrich durable RAG fields and must not be
* echoed wholesale into agent-facing views.</p>
*
* <p>{@code stepId} links to {@code agent_step.id} when the in-flight Agent step is known.
* {@code requestJson} is the tool request envelope body used only to derive bounded
* {@code input_params} (e.g. query preview) — never dump secrets wholesale.</p>
*/
public record ToolInvocationAuditEvent( public record ToolInvocationAuditEvent(
String sessionId, String sessionId,
String runId, String runId,
@@ -15,7 +26,11 @@ public record ToolInvocationAuditEvent(
String errorCode, String errorCode,
int durationMs, int durationMs,
int requestBytes, int requestBytes,
int agentResultBytes) { int agentResultBytes,
String rawResultJson,
String agentResultJson,
Long stepId,
String requestJson) {
public ToolInvocationAuditEvent { public ToolInvocationAuditEvent {
requireText(sessionId, "sessionId"); requireText(sessionId, "sessionId");
@@ -35,6 +50,38 @@ public record ToolInvocationAuditEvent(
} }
} }
/** Backward-compatible constructor without raw/agent JSON / step / request body. */
public ToolInvocationAuditEvent(String sessionId,
String runId,
String toolCallId,
String toolName,
InvocationStatus status,
EvidenceStatus evidenceStatus,
String errorCode,
int durationMs,
int requestBytes,
int agentResultBytes) {
this(sessionId, runId, toolCallId, toolName, status, evidenceStatus, errorCode,
durationMs, requestBytes, agentResultBytes, null, null, null, null);
}
/** Backward-compatible constructor with raw/agent JSON only. */
public ToolInvocationAuditEvent(String sessionId,
String runId,
String toolCallId,
String toolName,
InvocationStatus status,
EvidenceStatus evidenceStatus,
String errorCode,
int durationMs,
int requestBytes,
int agentResultBytes,
String rawResultJson,
String agentResultJson) {
this(sessionId, runId, toolCallId, toolName, status, evidenceStatus, errorCode,
durationMs, requestBytes, agentResultBytes, rawResultJson, agentResultJson, null, null);
}
private static void requireText(String value, String name) { private static void requireText(String value, String name) {
if (value == null || value.isBlank()) { if (value == null || value.isBlank()) {
throw new IllegalArgumentException(name + " must not be blank"); throw new IllegalArgumentException(name + " must not be blank");
@@ -129,9 +129,14 @@ public final class TraceAuditEvents {
details.put("evidence_status", tool.evidenceStatus().name()); details.put("evidence_status", tool.evidenceStatus().name());
details.put("request_bytes", tool.requestBytes()); details.put("request_bytes", tool.requestBytes());
details.put("agent_result_bytes", tool.agentResultBytes()); details.put("agent_result_bytes", tool.agentResultBytes());
details.put("has_raw_result", tool.rawResultJson() != null && !tool.rawResultJson().isBlank());
if (tool.stepId() != null) {
details.put("step_id", tool.stepId());
}
if (tool.errorCode() != null) { if (tool.errorCode() != null) {
details.put("error_code", tool.errorCode()); details.put("error_code", tool.errorCode());
} }
// Do not embed raw LookupResult / agent JSON here — durable RAG fields go to tool_invocation.
return new DiagnosisTraceAuditEvent( return new DiagnosisTraceAuditEvent(
tool.sessionId(), tool.runId(), TracePhase.TOOL, TraceEventType.TOOL_INVOCATION, tool.sessionId(), tool.runId(), TracePhase.TOOL, TraceEventType.TOOL_INVOCATION,
tool.status() == com.superbiz.agent.harness.contract.InvocationStatus.READY tool.status() == com.superbiz.agent.harness.contract.InvocationStatus.READY
@@ -3,9 +3,9 @@ package com.superbiz.agent.harness.tool.boundary;
import com.fasterxml.jackson.core.JsonProcessingException; import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode; import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
import com.superbiz.agent.harness.audit.ToolInvocationAuditEvent; import com.superbiz.agent.harness.audit.ToolInvocationAuditEvent;
import com.superbiz.agent.harness.audit.ToolInvocationAuditSink; import com.superbiz.agent.harness.audit.ToolInvocationAuditSink;
import com.superbiz.agent.harness.contract.EvidenceStatus;
import com.superbiz.agent.harness.core.BudgetExceededException; import com.superbiz.agent.harness.core.BudgetExceededException;
import com.superbiz.agent.harness.core.DiagnosisHarnessCore; import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
import com.superbiz.agent.harness.core.RunAbortedException; import com.superbiz.agent.harness.core.RunAbortedException;
@@ -19,10 +19,10 @@ import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
import org.slf4j.Logger; import org.slf4j.Logger;
import org.slf4j.LoggerFactory; import org.slf4j.LoggerFactory;
import java.nio.charset.StandardCharsets;
import java.time.Clock; import java.time.Clock;
import java.time.Duration; import java.time.Duration;
import java.time.Instant; import java.time.Instant;
import java.nio.charset.StandardCharsets;
import java.util.Objects; import java.util.Objects;
public final class ToolBoundary { public final class ToolBoundary {
@@ -35,13 +35,14 @@ public final class ToolBoundary {
private final ObjectMapper objectMapper; private final ObjectMapper objectMapper;
private final Clock clock; private final Clock clock;
private final ToolInvocationAuditSink auditSink; private final ToolInvocationAuditSink auditSink;
private final AgentStepAuditTracker stepTracker;
public ToolBoundary(DiagnosisHarnessCore core, public ToolBoundary(DiagnosisHarnessCore core,
ToolCallKeyFactory keyFactory, ToolCallKeyFactory keyFactory,
CanonicalInvocationStore store, CanonicalInvocationStore store,
ObjectMapper objectMapper, ObjectMapper objectMapper,
Clock clock) { Clock clock) {
this(core, keyFactory, store, objectMapper, clock, ToolInvocationAuditSink.noop()); this(core, keyFactory, store, objectMapper, clock, ToolInvocationAuditSink.noop(), null);
} }
public ToolBoundary(DiagnosisHarnessCore core, public ToolBoundary(DiagnosisHarnessCore core,
@@ -50,12 +51,23 @@ public final class ToolBoundary {
ObjectMapper objectMapper, ObjectMapper objectMapper,
Clock clock, Clock clock,
ToolInvocationAuditSink auditSink) { ToolInvocationAuditSink auditSink) {
this(core, keyFactory, store, objectMapper, clock, auditSink, null);
}
public ToolBoundary(DiagnosisHarnessCore core,
ToolCallKeyFactory keyFactory,
CanonicalInvocationStore store,
ObjectMapper objectMapper,
Clock clock,
ToolInvocationAuditSink auditSink,
AgentStepAuditTracker stepTracker) {
this.core = Objects.requireNonNull(core, "core must not be null"); this.core = Objects.requireNonNull(core, "core must not be null");
this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null"); this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null");
this.store = Objects.requireNonNull(store, "store must not be null"); this.store = Objects.requireNonNull(store, "store must not be null");
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null"); this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
this.clock = Objects.requireNonNull(clock, "clock must not be null"); this.clock = Objects.requireNonNull(clock, "clock must not be null");
this.auditSink = Objects.requireNonNull(auditSink, "auditSink must not be null"); this.auditSink = Objects.requireNonNull(auditSink, "auditSink must not be null");
this.stepTracker = stepTracker;
} }
public ToolBoundaryResult execute(RunContext context, public ToolBoundaryResult execute(RunContext context,
@@ -63,12 +75,12 @@ public final class ToolBoundary {
ToolExecutor executor, ToolExecutor executor,
ToolResultProjector projector) { ToolResultProjector projector) {
Instant startedAt = clock.instant(); Instant startedAt = clock.instant();
ToolBoundaryResult result = executeCanonical(context, request, executor, projector); ExecutionOutcome outcome = executeCanonical(context, request, executor, projector);
auditSafely(context, request, result, startedAt); auditSafely(context, request, outcome.result(), outcome.rawResponse(), startedAt);
return result; return outcome.result();
} }
private ToolBoundaryResult executeCanonical(RunContext context, private ExecutionOutcome executeCanonical(RunContext context,
ToolCallRequestEnvelope request, ToolCallRequestEnvelope request,
ToolExecutor executor, ToolExecutor executor,
ToolResultProjector projector) { ToolResultProjector projector) {
@@ -79,23 +91,23 @@ public final class ToolBoundary {
core.beforeToolCall(context, request.toolName()); core.beforeToolCall(context, request.toolName());
long requestBytes = store.limits().utf8Bytes(request.requestJson()); long requestBytes = store.limits().utf8Bytes(request.requestJson());
if (requestBytes > store.limits().maxRecordBytes()) { if (requestBytes > store.limits().maxRecordBytes()) {
return errorAndNoRecord(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE); return ExecutionOutcome.of(errorAndNoRecord(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE));
} }
core.reserveRunBytes(context, requestBytes); core.reserveRunBytes(context, requestBytes);
store.begin(key, CanonicalToolInvocation.projecting( store.begin(key, CanonicalToolInvocation.projecting(
request.toolCallId(), request.runId(), request.toolName(), request.toolCallId(), request.runId(), request.toolName(),
request.requestJson(), clock.instant())); request.requestJson(), clock.instant()));
} catch (DuplicateInvocationException e) { } catch (DuplicateInvocationException e) {
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.DUPLICATE_TOOL_CALL); return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.DUPLICATE_TOOL_CALL));
} catch (RunAbortedException | BudgetExceededException e) { } catch (RunAbortedException | BudgetExceededException e) {
return ToolBoundaryResult.error(toolCallId, return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId,
e instanceof BudgetExceededException e instanceof BudgetExceededException
? ToolBoundaryErrorCode.BUDGET_EXHAUSTED ? ToolBoundaryErrorCode.BUDGET_EXHAUSTED
: ToolBoundaryErrorCode.RUN_INACTIVE); : ToolBoundaryErrorCode.RUN_INACTIVE));
} catch (IllegalArgumentException e) { } catch (IllegalArgumentException e) {
return ToolBoundaryResult.error(toolCallId, classifyPreflightError(e)); return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, classifyPreflightError(e)));
} catch (CanonicalStoreException e) { } catch (CanonicalStoreException e) {
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.STORE_ERROR); return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.STORE_ERROR));
} }
String rawResponse; String rawResponse;
@@ -107,7 +119,7 @@ public final class ToolBoundary {
} }
} catch (Exception e) { } catch (Exception e) {
markErrorSafely(key, null, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR); markErrorSafely(key, null, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR); return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR));
} }
try { try {
@@ -115,13 +127,16 @@ public final class ToolBoundary {
core.reserveRunBytes(context, store.limits().utf8Bytes(rawResponse)); core.reserveRunBytes(context, store.limits().utf8Bytes(rawResponse));
} catch (ResultTooLargeException e) { } catch (ResultTooLargeException e) {
markErrorSafely(key, null, ToolBoundaryErrorCode.RESULT_TOO_LARGE); markErrorSafely(key, null, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE), rawResponse);
} catch (BudgetExceededException e) { } catch (BudgetExceededException e) {
markErrorSafely(key, null, ToolBoundaryErrorCode.BUDGET_EXHAUSTED); markErrorSafely(key, null, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED), rawResponse);
} catch (RunAbortedException e) { } catch (RunAbortedException e) {
markErrorSafely(key, null, ToolBoundaryErrorCode.RUN_INACTIVE); markErrorSafely(key, null, ToolBoundaryErrorCode.RUN_INACTIVE);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE), rawResponse);
} }
ProjectedToolResult projected; ProjectedToolResult projected;
@@ -133,29 +148,35 @@ public final class ToolBoundary {
} }
} catch (Exception e) { } catch (Exception e) {
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.PROJECTION_ERROR); markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.PROJECTION_ERROR);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.PROJECTION_ERROR); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.PROJECTION_ERROR), rawResponse);
} }
try { try {
store.limits().validateAgentResult(projected.agentResult()); store.limits().validateAgentResult(projected.agentResult());
core.reserveRunBytes(context, store.limits().utf8Bytes(projected.agentResult())); core.reserveRunBytes(context, store.limits().utf8Bytes(projected.agentResult()));
store.markReady(key, rawResponse, projected.agentResult(), projected.evidenceStatus(), clock.instant()); store.markReady(key, rawResponse, projected.agentResult(), projected.evidenceStatus(), clock.instant());
return ToolBoundaryResult.ready(toolCallId, projected.agentResult(), projected.evidenceStatus()); return new ExecutionOutcome(
ToolBoundaryResult.ready(toolCallId, projected.agentResult(), projected.evidenceStatus()),
rawResponse);
} catch (ResultTooLargeException e) { } catch (ResultTooLargeException e) {
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RESULT_TOO_LARGE); markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE), rawResponse);
} catch (BudgetExceededException e) { } catch (BudgetExceededException e) {
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.BUDGET_EXHAUSTED); markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED), rawResponse);
} catch (RunAbortedException e) { } catch (RunAbortedException e) {
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RUN_INACTIVE); markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RUN_INACTIVE);
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE); return new ExecutionOutcome(
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE), rawResponse);
} catch (CanonicalStoreException e) { } catch (CanonicalStoreException e) {
ToolBoundaryErrorCode code = e instanceof ResultTooLargeException ToolBoundaryErrorCode code = e instanceof ResultTooLargeException
? ToolBoundaryErrorCode.RESULT_TOO_LARGE ? ToolBoundaryErrorCode.RESULT_TOO_LARGE
: ToolBoundaryErrorCode.PROJECTION_ERROR; : ToolBoundaryErrorCode.PROJECTION_ERROR;
markErrorSafely(key, rawResponse, code); markErrorSafely(key, rawResponse, code);
return ToolBoundaryResult.error(toolCallId, code); return new ExecutionOutcome(ToolBoundaryResult.error(toolCallId, code), rawResponse);
} }
} }
@@ -222,6 +243,7 @@ public final class ToolBoundary {
private void auditSafely(RunContext context, private void auditSafely(RunContext context,
ToolCallRequestEnvelope request, ToolCallRequestEnvelope request,
ToolBoundaryResult result, ToolBoundaryResult result,
String rawResponse,
Instant startedAt) { Instant startedAt) {
if (context == null || request == null || result == null if (context == null || request == null || result == null
|| !context.runId().equals(request.runId()) || !context.runId().equals(request.runId())
@@ -230,11 +252,16 @@ public final class ToolBoundary {
return; return;
} }
try { try {
Long stepId = stepTracker == null ? null : stepTracker.currentStepId(context.runId());
auditSink.record(new ToolInvocationAuditEvent( auditSink.record(new ToolInvocationAuditEvent(
context.sessionId(), context.runId(), request.toolCallId(), request.toolName(), context.sessionId(), context.runId(), request.toolCallId(), request.toolName(),
result.status(), result.evidenceStatus(), result.errorCode(), result.status(), result.evidenceStatus(), result.errorCode(),
saturatingInt(Math.max(0L, Duration.between(startedAt, clock.instant()).toMillis())), saturatingInt(Math.max(0L, Duration.between(startedAt, clock.instant()).toMillis())),
utf8Bytes(request.requestJson()), utf8Bytes(result.agentResult()))); utf8Bytes(request.requestJson()), utf8Bytes(result.agentResult()),
rawResponse,
result.agentResult(),
stepId,
request.requestJson()));
} catch (RuntimeException exception) { } catch (RuntimeException exception) {
log.warn("Failed to persist Tool durable audit: tool={}, status={}", log.warn("Failed to persist Tool durable audit: tool={}, status={}",
request.toolName(), result.status()); request.toolName(), result.status());
@@ -249,6 +276,12 @@ public final class ToolBoundary {
return value >= Integer.MAX_VALUE ? Integer.MAX_VALUE : (int) value; return value >= Integer.MAX_VALUE ? Integer.MAX_VALUE : (int) value;
} }
private record ExecutionOutcome(ToolBoundaryResult result, String rawResponse) {
static ExecutionOutcome of(ToolBoundaryResult result) {
return new ExecutionOutcome(result, null);
}
}
private static final class InvalidToolCallIdException extends IllegalArgumentException { private static final class InvalidToolCallIdException extends IllegalArgumentException {
private InvalidToolCallIdException(Throwable cause) { private InvalidToolCallIdException(Throwable cause) {
super("Invalid tool call ID", cause); super("Invalid tool call ID", cause);
@@ -184,6 +184,7 @@ public class DiagnosisTraceService {
.runId(run.getRunId()) .runId(run.getRunId())
.sessionId(run.getSessionId()) .sessionId(run.getSessionId())
.query(run.getQuery()) .query(run.getQuery())
.conclusion(run.getConclusion())
.status(run.getStatus()) .status(run.getStatus())
.agentFlow(run.getAgentFlow()) .agentFlow(run.getAgentFlow())
.intent(run.getIntent() == null ? null : run.getIntent().name()) .intent(run.getIntent() == null ? null : run.getIntent().name())
@@ -239,6 +240,8 @@ public class DiagnosisTraceService {
.agentName(audit.getAgentName()) .agentName(audit.getAgentName())
.reasoningAvailable(audit.getReasoningAvailable()) .reasoningAvailable(audit.getReasoningAvailable())
.reasoningContent(audit.getReasoningContent()) .reasoningContent(audit.getReasoningContent())
.assistantText(audit.getAssistantText())
.contentSource(audit.getContentSource())
.contentBytes(audit.getContentBytes()) .contentBytes(audit.getContentBytes())
.createdAt(audit.getCreatedAt()) .createdAt(audit.getCreatedAt())
.build(); .build();
@@ -86,6 +86,7 @@ public class KnowledgeDocumentRetriever {
.rawScore(hit.rawScore()) .rawScore(hit.rawScore())
.scoreLabel(hit.scoreLabel()) .scoreLabel(hit.scoreLabel())
.originalRank(hit.originalRank()) .originalRank(hit.originalRank())
.denseDistance(hit.denseDistance())
.metadata(hit.metadata() == null ? java.util.Map.of() : hit.metadata()) .metadata(hit.metadata() == null ? java.util.Map.of() : hit.metadata())
.hitReasons(List.of("semantic_rank:" + hit.originalRank(), "attempt:" + attemptName)) .hitReasons(List.of("semantic_rank:" + hit.originalRank(), "attempt:" + attemptName))
.build()); .build());
@@ -6,6 +6,7 @@ import com.superbiz.agent.dto.KnowledgeQuery;
import com.superbiz.agent.dto.RerankTrace; import com.superbiz.agent.dto.RerankTrace;
import com.superbiz.agent.dto.RetrievedEvidenceCandidate; import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
import com.superbiz.agent.service.retrieval.EvidenceIdentity; import com.superbiz.agent.service.retrieval.EvidenceIdentity;
import com.superbiz.agent.service.retrieval.RetrievalScoreNormalizer;
import org.springframework.beans.factory.annotation.Value; import org.springframework.beans.factory.annotation.Value;
import org.springframework.stereotype.Service; import org.springframework.stereotype.Service;
@@ -20,28 +21,25 @@ import java.util.Map;
import java.util.Set; import java.util.Set;
/** /**
* 检索后处理:分数归一化、规则 rerank、chunk 级证据组装、相关等级判定。 * 检索后处理:统一 qualityScore、保检索序、chunk 级装配、相关等级判定。
* *
* <h3>处理步骤</h3> * <h3>处理步骤</h3>
* <ol> * <ol>
* <li>把候选 L2 距离归一成 0~1 的 baseScore</li> * <li>{@link RetrievalScoreNormalizer#toQualityScore} → qualityScore(唯一 label 分支)</li>
* <li>用 L0 hint 做规则加分,得到 finalScore</li> * <li>按 {@code originalRank} 升序(检索权威序;不做关键词 boost 改序)</li>
* <li>按 finalScore 降序排序</li> * <li>evidenceKey 去重 / maxChunksPerDocument / return-n</li>
* <li>按 evidenceKey 去重</li> * <li>用 top qualityScore 定 relevanceLevel / 低质闸门字段</li>
* <li>按 maxChunksPerDocument 截断同文档 chunk</li>
* <li>按 return-n 截断最终 evidence 条数</li>
* <li>根据 top baseScore + hint 支撑计算 relevanceLevel</li>
* </ol> * </ol>
*
* <p>L0 domain/keyword 命中仅写入解释性 {@code hitReasons},不改变分数与主序。</p>
*/ */
@Service @Service
public class KnowledgeEvidencePostProcessor { public class KnowledgeEvidencePostProcessor {
private static final String LEVEL_PRECISE = "PRECISE"; private static final String LEVEL_PRECISE = "PRECISE";
private static final String LEVEL_HIGHLY_RELEVANT = "HIGHLY_RELEVANT";
private static final String LEVEL_REFERENCE = "REFERENCE"; private static final String LEVEL_REFERENCE = "REFERENCE";
private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档"; private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档";
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度"; private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
@Value("${retrieval.normalization.max-l2-distance:2.0}") @Value("${retrieval.normalization.max-l2-distance:2.0}")
@@ -53,22 +51,21 @@ public class KnowledgeEvidencePostProcessor {
@Value("${retrieval.normalization.reference-threshold:0.5}") @Value("${retrieval.normalization.reference-threshold:0.5}")
private double referenceThreshold = 0.5; private double referenceThreshold = 0.5;
/** 同一 docId 最多保留的 chunk 数。 */
@Value("${rag.max-chunks-per-document:2}") @Value("${rag.max-chunks-per-document:2}")
private int maxChunksPerDocument = 2; private int maxChunksPerDocument = 2;
/**
* 后处理后最多返回的 evidence 条数。
* 0 或负数表示不在此层截断(仍可能被 projector 预算截断)。
*/
@Value("${rag.return-n:5}") @Value("${rag.return-n:5}")
private int returnN = 5; private int returnN = 5;
public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) { public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) {
List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates; List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates;
int batchSize = safeCandidates.size();
List<ScoredCandidate> ranked = safeCandidates.stream() List<ScoredCandidate> ranked = safeCandidates.stream()
.map(candidate -> score(query, candidate)) .map(candidate -> score(query, candidate, batchSize))
.sorted(Comparator.comparingDouble(ScoredCandidate::finalScore).reversed()) .sorted(Comparator
.comparingInt((ScoredCandidate s) -> rankOrMax(s.candidate().getOriginalRank()))
.thenComparing(s -> resolveEvidenceKey(s.candidate()), Comparator.nullsLast(String::compareTo)))
.toList(); .toList();
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>(); Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
@@ -104,15 +101,15 @@ public class KnowledgeEvidencePostProcessor {
traceItems.add(RerankTrace.Item.builder() traceItems.add(RerankTrace.Item.builder()
.finalRank(finalRank++) .finalRank(finalRank++)
.source(candidate.getSource()) .source(candidate.getSource())
.baseScore(scored.baseScore()) .baseScore(scored.qualityScore())
.finalScore(scored.finalScore()) .finalScore(scored.qualityScore())
.boostReasons(scored.boostReasons()) .boostReasons(scored.explainReasons())
.build()); .build());
} }
List<EvidenceBlock> blocks = new ArrayList<>(deduped.values()); List<EvidenceBlock> blocks = new ArrayList<>(deduped.values());
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).baseScore(); Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).qualityScore();
RelevanceAssessment assessment = computeRelevance(query, ranked); RelevanceAssessment assessment = computeRelevance(ranked);
return EvidencePostprocessResult.builder() return EvidencePostprocessResult.builder()
.candidateCount(safeCandidates.size()) .candidateCount(safeCandidates.size())
.evidenceBlockCount(blocks.size()) .evidenceBlockCount(blocks.size())
@@ -132,12 +129,11 @@ public class KnowledgeEvidencePostProcessor {
return topSimilarity == null || topSimilarity < referenceThreshold; return topSimilarity == null || topSimilarity < referenceThreshold;
} }
/**
* @deprecated 保留给旧测试/调用;新路径请用 {@link RetrievalScoreNormalizer#l2ToQuality}。
*/
public double normalizeL2(Double l2Score) { public double normalizeL2(Double l2Score) {
if (l2Score == null) { return RetrievalScoreNormalizer.l2ToQuality(l2Score, maxL2Distance);
return 0.0;
}
double clamped = Math.min(l2Score, maxL2Distance);
return Math.max(0.0, 1.0 - clamped / maxL2Distance);
} }
public double getReferenceThreshold() { public double getReferenceThreshold() {
@@ -157,7 +153,7 @@ public class KnowledgeEvidencePostProcessor {
.retrievalLayer(candidate.getRetrievalLayer()) .retrievalLayer(candidate.getRetrievalLayer())
.content(truncate(candidate.getContent(), 800)) .content(truncate(candidate.getContent(), 800))
.score(candidate.getScore()) .score(candidate.getScore())
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons())) .hitReasons(mergeReasons(candidate.getHitReasons(), scored.explainReasons()))
.build(); .build();
} }
@@ -177,33 +173,29 @@ public class KnowledgeEvidencePostProcessor {
if (docId != null) { if (docId != null) {
return docId; return docId;
} }
// No docId: do not collapse unrelated fallback keys under one bucket.
return evidenceKey; return evidenceKey;
} }
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate) { private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate, int batchSize) {
double baseScore = normalizeL2(candidate.getScore()); double quality = RetrievalScoreNormalizer.toQualityScore(
double finalScore = baseScore; candidate.getScoreLabel(),
List<String> boosts = new ArrayList<>(); candidate.getScore(),
candidate.getOriginalRank(),
batchSize,
maxL2Distance,
candidate.getDenseDistance());
List<String> explain = new ArrayList<>();
// L0 重叠仅解释,不改变 quality / 排序
if (matchesAny(candidate, query.getDomainHints())) { if (matchesAny(candidate, query.getDomainHints())) {
finalScore += 0.15; explain.add("l0_domain_overlap");
boosts.add("domain_match:+0.15");
} }
if (matchesAny(candidate, query.getEntities())) { if (matchesAny(candidate, query.getEntities())) {
finalScore += 0.20; explain.add("l0_entity_overlap");
boosts.add("entity_match:+0.20");
} }
if (matchesAny(candidate, query.getMatchedKeywords())) { if (matchesAny(candidate, query.getMatchedKeywords())) {
finalScore += 0.10; explain.add("l0_keyword_overlap");
boosts.add("keyword_match:+0.10");
} }
if (isPreferredSourceType(candidate)) { return new ScoredCandidate(candidate, quality, explain);
finalScore += 0.05;
boosts.add("source_type:+0.05");
}
return new ScoredCandidate(candidate, baseScore, finalScore, boosts);
} }
private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) { private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) {
@@ -225,42 +217,21 @@ public class KnowledgeEvidencePostProcessor {
return false; return false;
} }
private boolean isPreferredSourceType(RetrievedEvidenceCandidate candidate) { private RelevanceAssessment computeRelevance(List<ScoredCandidate> ranked) {
Map<String, String> metadata = candidate.getMetadata();
if (metadata == null || metadata.isEmpty()) {
return false;
}
String type = firstNonBlank(metadata.get("source_type"), metadata.get("documentType"), metadata.get("type"));
if (type == null) {
return false;
}
String normalized = type.toLowerCase(Locale.ROOT);
return normalized.contains("runbook") || normalized.contains("guide") || normalized.contains("case");
}
private RelevanceAssessment computeRelevance(KnowledgeQuery query, List<ScoredCandidate> ranked) {
if (ranked.isEmpty()) { if (ranked.isEmpty()) {
return new RelevanceAssessment(null, null); return new RelevanceAssessment(null, null);
} }
ScoredCandidate top = ranked.get(0); double top = ranked.get(0).qualityScore();
if (top.baseScore() >= highlyRelevantThreshold && hasHintSupport(query, top)) { if (top >= highlyRelevantThreshold) {
// 不再要求 hasHintSupport;rank/L2 quality 足够即高相关/精准
return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE); return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE);
} }
if (top.baseScore() >= highlyRelevantThreshold) { if (top >= referenceThreshold) {
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
}
if (top.baseScore() >= referenceThreshold) {
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE); return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
} }
return new RelevanceAssessment(null, null); return new RelevanceAssessment(null, null);
} }
private boolean hasHintSupport(KnowledgeQuery query, ScoredCandidate top) {
return matchesAny(top.candidate(), query.getDomainHints())
|| matchesAny(top.candidate(), query.getEntities())
|| matchesAny(top.candidate(), query.getMatchedKeywords());
}
private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) { private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) {
Set<String> reasons = new LinkedHashSet<>(); Set<String> reasons = new LinkedHashSet<>();
if (existing.getHitReasons() != null) { if (existing.getHitReasons() != null) {
@@ -276,13 +247,13 @@ public class KnowledgeEvidencePostProcessor {
} }
} }
private List<String> mergeReasons(List<String> base, List<String> boosts) { private List<String> mergeReasons(List<String> base, List<String> extra) {
Set<String> merged = new LinkedHashSet<>(); Set<String> merged = new LinkedHashSet<>();
if (base != null) { if (base != null) {
merged.addAll(base); merged.addAll(base);
} }
if (boosts != null) { if (extra != null) {
merged.addAll(boosts); merged.addAll(extra);
} }
return new ArrayList<>(merged); return new ArrayList<>(merged);
} }
@@ -294,8 +265,8 @@ public class KnowledgeEvidencePostProcessor {
return text.substring(0, maxLength) + "..."; return text.substring(0, maxLength) + "...";
} }
private String firstNonBlank(String... values) { private static int rankOrMax(Integer rank) {
return EvidenceIdentity.firstNonBlank(values); return rank == null || rank < 1 ? Integer.MAX_VALUE : rank;
} }
private String nullToEmpty(String value) { private String nullToEmpty(String value) {
@@ -303,9 +274,8 @@ public class KnowledgeEvidencePostProcessor {
} }
private record ScoredCandidate(RetrievedEvidenceCandidate candidate, private record ScoredCandidate(RetrievedEvidenceCandidate candidate,
double baseScore, double qualityScore,
double finalScore, List<String> explainReasons) {
List<String> boostReasons) {
} }
private record RelevanceAssessment(String level, String hint) { private record RelevanceAssessment(String level, String hint) {
@@ -23,8 +23,14 @@ import java.util.Map;
/** /**
* 向量索引写入服务(RAG 入库侧)。 * 向量索引写入服务(RAG 入库侧)。
* *
* <p>写入单一后端 {@link MilvusHybridKnowledgeStore}(dense + BM25 search_text)。 * <p>唯一后端 {@link MilvusHybridKnowledgeStore}(Milvus SDK v2):</p>
* 不再使用 legacy {@code MilvusServiceClient} insert/delete。</p> * <ul>
* <li>dense:应用侧 embedding → 字段 {@code vector}</li>
* <li>BM25:{@link #buildSearchText} → 字段 {@code search_text};
* sparse 由 collection 上 BM25 Function 自动生成,本类不写 sparse</li>
* </ul>
* <p>不再使用 legacy {@code MilvusServiceClient} insert/delete,
* 也不走 Spring AI {@code VectorStore#add}(starter 无 hybrid schema/BM25 Function)。</p>
*/ */
@Service @Service
public class VectorIndexService { public class VectorIndexService {
@@ -118,12 +124,13 @@ public class VectorIndexService {
for (int i = 0; i < chunks.size(); i++) { for (int i = 0; i < chunks.size(); i++) {
DocumentChunk chunk = chunks.get(i); DocumentChunk chunk = chunks.get(i);
try { try {
// dense embedding 与 BM25 search_text 同源(title/path 增强)
List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk)); List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk));
Map<String, Object> metadata = buildMetadata(path.toString(), chunk, chunks.size()); Map<String, Object> metadata = buildMetadata(path.toString(), chunk, chunks.size());
knowledgeStore.upsertChunk( knowledgeStore.upsertChunk(
chunk.getContent(), chunk.getContent(), // 返回原文
buildSearchText(chunk), buildSearchText(chunk), // BM25 语料;sparse 由 Milvus Function 生成
vector, vector, // dense 向量
metadata, metadata,
chunk.getChunkIndex()); chunk.getChunkIndex());
logger.info("分片 {}/{} 索引成功", i + 1, chunks.size()); logger.info("分片 {}/{} 索引成功", i + 1, chunks.size());
@@ -153,12 +160,13 @@ public class VectorIndexService {
for (int i = 0; i < chunks.size(); i++) { for (int i = 0; i < chunks.size(); i++) {
DocumentChunk chunk = chunks.get(i); DocumentChunk chunk = chunks.get(i);
try { try {
// dense embedding 与 BM25 search_text 同源(title/path 增强)
List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk)); List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk));
Map<String, Object> metadata = buildDocumentMetadata(docId, chunk, chunks.size(), category, frontmatter); Map<String, Object> metadata = buildDocumentMetadata(docId, chunk, chunks.size(), category, frontmatter);
knowledgeStore.upsertChunk( knowledgeStore.upsertChunk(
chunk.getContent(), chunk.getContent(), // 返回原文
buildSearchText(chunk), buildSearchText(chunk), // BM25 语料;sparse 由 Milvus Function 生成
vector, vector, // dense 向量
metadata, metadata,
chunk.getChunkIndex()); chunk.getChunkIndex());
logger.info("文档分块 {}/{} 索引成功,docId: {}", i + 1, chunks.size(), docId); logger.info("文档分块 {}/{} 索引成功,docId: {}", i + 1, chunks.size(), docId);
@@ -212,12 +220,18 @@ public class VectorIndexService {
return metadata; return metadata;
} }
/**
* Dense embedding 输入。与 {@link #buildSearchText} 同源,保证 dense/BM25 看到同一增强文本。
*/
static String buildEmbeddingText(DocumentChunk chunk) { static String buildEmbeddingText(DocumentChunk chunk) {
return buildSearchText(chunk); return buildSearchText(chunk);
} }
/** /**
* Text used for BM25 {@code search_text} and dense embedding. * 构造写入 Milvus 的检索文本(BM25 {@code search_text},并复用为 dense embedding 输入)。
*
* <p>在正文前拼接 title / breadcrumb,提高「按标题或路径关键词」的 BM25 命中率,
* 同时让 dense 向量也编码结构信息。无标题路径时退回纯 content。</p>
*/ */
static String buildSearchText(DocumentChunk chunk) { static String buildSearchText(DocumentChunk chunk) {
String content = trimToEmpty(chunk.getContent()); String content = trimToEmpty(chunk.getContent());
@@ -1,6 +1,7 @@
package com.superbiz.agent.service; package com.superbiz.agent.service;
import com.superbiz.agent.service.milvus.MilvusHybridKnowledgeStore; import com.superbiz.agent.service.milvus.MilvusHybridKnowledgeStore;
import com.superbiz.agent.service.retrieval.RetrievalScoreLabels;
import lombok.Getter; import lombok.Getter;
import lombok.Setter; import lombok.Setter;
import org.slf4j.Logger; import org.slf4j.Logger;
@@ -13,11 +14,18 @@ import java.util.List;
import java.util.Locale; import java.util.Locale;
/** /**
* Knowledge vector retrieval facade. * 知识库向量检索门面(lookup_knowledge / RAG 召回入口)。
* *
* <p><b>Single backend:</b> {@link MilvusHybridKnowledgeStore} (Milvus Java SDK v2). * <p><b>唯一后端:</b>{@link MilvusHybridKnowledgeStore}(Milvus Java SDK v2)。</p>
* Legacy {@code MilvusServiceClient} search and Spring AI VectorStore routing for *
* {@code lookup_knowledge} have been removed.</p> * <h3>模式切换</h3>
* <p>{@code retrieval.search.mode}(同库查询算法,非两套写入):</p>
* <ul>
* <li>{@code hybrid} —— 线上主路径:dense + 服务端 BM25 + RRF</li>
* <li>{@code dense} —— 对照/评测:仅 dense ANN</li>
* </ul>
* <p>命中 {@link SearchResult#scoreLabel} 仅为 {@link RetrievalScoreLabels#DENSE} /
* {@link RetrievalScoreLabels#HYBRID}。质量分由后处理 {@code RetrievalScoreNormalizer} 统一计算。</p>
*/ */
@Service @Service
public class VectorSearchService { public class VectorSearchService {
@@ -31,7 +39,7 @@ public class VectorSearchService {
private VectorEmbeddingService embeddingService; private VectorEmbeddingService embeddingService;
/** /**
* dense | hybrid * 检索模式:{@code hybrid}(主路径)| {@code dense}(召回对照)。
*/ */
@Value("${retrieval.search.mode:dense}") @Value("${retrieval.search.mode:dense}")
private String searchMode = "dense"; private String searchMode = "dense";
@@ -53,18 +61,34 @@ public class VectorSearchService {
return knowledgeStore.searchDense(query, queryVector, topK, category); return knowledgeStore.searchDense(query, queryVector, topK, category);
} }
/**
* 单条召回结果。列表顺序即检索权威序(adapter 赋 originalRank=1..n)。
*
* <ul>
* <li>{@code scoreLabel=dense}:{@link #score} = L2 距离(越小越好)</li>
* <li>{@code scoreLabel=hybrid}:{@link #score}/{@link #rawScore} = 引擎融合分;
* 后处理 quality 主要按 rank 映射,不把 score 当 L2</li>
* </ul>
*/
@Setter @Setter
@Getter @Getter
public static class SearchResult { public static class SearchResult {
private String id; private String id;
private String content; private String content;
/** /**
* Compatibility score for post-process normalizeL2. * 引擎主分:dense=L2;hybrid=融合分(量纲由 scoreLabel 解释)。
* Dense path: L2 distance. Hybrid path: dense L2 when available.
*/ */
private float score; private float score;
/** 引擎原始分(与 score 同源或更细,便于调试)。 */
private Double rawScore; private Double rawScore;
/** {@link RetrievalScoreLabels#DENSE} 或 {@link RetrievalScoreLabels#HYBRID}。 */
private String scoreLabel; private String scoreLabel;
/**
* Optional dense L2 for the same id (hybrid path only).
* Used for absolute quality / low-quality gates; does <b>not</b> replace sort order.
*/
private Double denseDistance;
/** metadata JSON 字符串(docId、source、title…)。 */
private String metadata; private String metadata;
} }
} }
@@ -5,6 +5,7 @@ import com.google.gson.JsonObject;
import com.superbiz.agent.config.MilvusProperties; import com.superbiz.agent.config.MilvusProperties;
import com.superbiz.agent.constant.MilvusConstants; import com.superbiz.agent.constant.MilvusConstants;
import com.superbiz.agent.service.VectorSearchService; import com.superbiz.agent.service.VectorSearchService;
import com.superbiz.agent.service.retrieval.RetrievalScoreLabels;
import io.milvus.common.clientenum.FunctionType; import io.milvus.common.clientenum.FunctionType;
import io.milvus.v2.client.ConnectConfig; import io.milvus.v2.client.ConnectConfig;
import io.milvus.v2.client.MilvusClientV2; import io.milvus.v2.client.MilvusClientV2;
@@ -41,10 +42,35 @@ import java.util.Map;
import java.util.UUID; import java.util.UUID;
/** /**
* Single knowledge vector backend (Milvus Java SDK v2). * 知识库向量后端(Milvus Java SDK v2)—— dense + BM25 混合检索的唯一实现。
* *
* <p>Supports dense ANN and dense+BM25 hybrid search via {@code hybridSearch} + {@link RRFRanker}. * <h3>为什么不用 Spring AI {@code spring-ai-starter-vector-store-milvus}</h3>
* Legacy {@code MilvusServiceClient} search is not used.</p> * <ul>
* <li>Spring AI Milvus starter(截至 2.0.0 / 1.1.8)只封装 dense {@code similaritySearch}。</li>
* <li>底层仍是 V1 {@code MilvusServiceClient} + 单路 {@code SearchParam},无 {@code hybridSearch} /
* BM25 Function / {@link RRFRanker}。</li>
* <li>真混合检索(dense ANN + 服务端 BM25 sparse,再 RRF 融合)必须走 Milvus SDK v2,
* 见 {@link #searchHybrid}。</li>
* </ul>
*
* <h3>Collection schema(默认名 {@code biz})</h3>
* <pre>
* id VarChar PK
* content VarChar —— 原文,返回给上层
* search_text VarChar+analyzer —— BM25 输入文本(可含 title/path 增强)
* sparse_vector SparseFloatVector —— 由 BM25 Function 从 search_text 自动生成,写入时不必填
* vector FloatVector —— dense 向量(应用侧 embedding)
* metadata JSON —— docId / source / category / kb_scope 等
* </pre>
*
* <h3>检索模式</h3>
* <ul>
* <li>{@link #searchDense}:单路 L2 ANN;{@code scoreLabel=dense}。</li>
* <li>{@link #searchHybrid}:dense + BM25 + 服务端 {@link RRFRanker};{@code scoreLabel=hybrid};
* 返回序即 RRF 序,不再用 dense L2 覆盖主分。</li>
* </ul>
*
* <p>配置入口:{@code milvus.collection}、{@code retrieval.search.mode}、{@code retrieval.hybrid.rrf-k}。</p>
*/ */
@Service @Service
public class MilvusHybridKnowledgeStore { public class MilvusHybridKnowledgeStore {
@@ -52,11 +78,20 @@ public class MilvusHybridKnowledgeStore {
private static final Logger log = LoggerFactory.getLogger(MilvusHybridKnowledgeStore.class); private static final Logger log = LoggerFactory.getLogger(MilvusHybridKnowledgeStore.class);
private static final Gson GSON = new Gson(); private static final Gson GSON = new Gson();
/** 主键(稳定 UUID,由 source + chunkIndex 派生,便于幂等重写)。 */
public static final String FIELD_ID = "id"; public static final String FIELD_ID = "id";
/** 返回给 LLM / 上层的原文 chunk。 */
public static final String FIELD_CONTENT = "content"; public static final String FIELD_CONTENT = "content";
/**
* BM25 输入字段。写入明文;Milvus 侧 analyzer + BM25 Function 生成 {@link #FIELD_SPARSE}。
* 通常比 content 多带 title/path 等检索增强词。
*/
public static final String FIELD_SEARCH_TEXT = "search_text"; public static final String FIELD_SEARCH_TEXT = "search_text";
/** 稀疏向量字段;由 BM25 Function 自动产出,insert 时不要手动填。 */
public static final String FIELD_SPARSE = "sparse_vector"; public static final String FIELD_SPARSE = "sparse_vector";
/** Dense 向量字段(应用侧 EmbeddingModel 生成)。 */
public static final String FIELD_DENSE = "vector"; public static final String FIELD_DENSE = "vector";
/** 业务元数据 JSON(过滤、证据身份、展示用)。 */
public static final String FIELD_METADATA = "metadata"; public static final String FIELD_METADATA = "metadata";
private final MilvusProperties milvusProperties; private final MilvusProperties milvusProperties;
@@ -64,12 +99,14 @@ public class MilvusHybridKnowledgeStore {
@Value("${milvus.collection:biz}") @Value("${milvus.collection:biz}")
private String collectionName = "biz"; private String collectionName = "biz";
/**
* RRF 平滑参数 k:score(d) = Σ 1/(k + rank_i(d))。
* k 越大,各路排名差异被压得越平;默认 60 与常见 RRF 设定一致。
*/
@Value("${retrieval.hybrid.rrf-k:60}") @Value("${retrieval.hybrid.rrf-k:60}")
private int rrfK = 60; private int rrfK = 60;
@Value("${retrieval.normalization.max-l2-distance:2.0}") /** 非空时追加 {@code metadata.kb_scope} 过滤,实现多知识域隔离。 */
private double maxL2Distance = 2.0;
@Value("${retrieval.kb-scope:}") @Value("${retrieval.kb-scope:}")
private String kbScope = ""; private String kbScope = "";
@@ -79,6 +116,10 @@ public class MilvusHybridKnowledgeStore {
this.milvusProperties = milvusProperties; this.milvusProperties = milvusProperties;
} }
/**
* 懒连接:首次调用时建连、确保 collection schema 存在并 load。
* 线程安全;后续检索/写入复用同一 {@link MilvusClientV2}。
*/
public synchronized MilvusClientV2 client() { public synchronized MilvusClientV2 client() {
if (client == null) { if (client == null) {
client = connect(); client = connect();
@@ -92,6 +133,21 @@ public class MilvusHybridKnowledgeStore {
return collectionName; return collectionName;
} }
/**
* 写入单个 chunk(dense + BM25 所需明文)。
*
* <p>只插入 {@code content / search_text / vector / metadata};
* {@code sparse_vector} 由 collection 上的 BM25 Function 在服务端从 {@code search_text} 生成。</p>
*
* <p>id 由 {@code source|docId + chunkIndex} 的 nameUUID 派生,同一 chunk 重复写入会得到相同 id
*(配合先 delete 再 insert 的上层逻辑实现覆盖)。</p>
*
* @param content 原文(返回字段)
* @param searchText BM25 / 可与 dense embedding 同源的检索文本
* @param denseVector 应用侧 embedding
* @param metadata 须尽量带 {@code _source} 或 {@code docId},供 id 与过滤使用
* @param chunkIndex 分片序号
*/
public void upsertChunk(String content, public void upsertChunk(String content,
String searchText, String searchText,
List<Float> denseVector, List<Float> denseVector,
@@ -110,6 +166,7 @@ public class MilvusHybridKnowledgeStore {
JsonObject row = new JsonObject(); JsonObject row = new JsonObject();
row.addProperty(FIELD_ID, id); row.addProperty(FIELD_ID, id);
row.addProperty(FIELD_CONTENT, content == null ? "" : content); row.addProperty(FIELD_CONTENT, content == null ? "" : content);
// 仅写明文;sparse 由 BM25 Function(search_text -> sparse_vector) 自动生成
row.addProperty(FIELD_SEARCH_TEXT, searchText == null ? "" : searchText); row.addProperty(FIELD_SEARCH_TEXT, searchText == null ? "" : searchText);
row.add(FIELD_DENSE, GSON.toJsonTree(denseVector)); row.add(FIELD_DENSE, GSON.toJsonTree(denseVector));
row.add(FIELD_METADATA, GSON.toJsonTree(metadata == null ? Map.of() : metadata)); row.add(FIELD_METADATA, GSON.toJsonTree(metadata == null ? Map.of() : metadata));
@@ -120,6 +177,7 @@ public class MilvusHybridKnowledgeStore {
.build()); .build());
} }
/** 按 metadata.docId 删除该文档全部 chunk(重建/覆盖前调用)。 */
public void deleteByDocId(String docId) { public void deleteByDocId(String docId) {
if (docId == null || docId.isBlank()) { if (docId == null || docId.isBlank()) {
return; return;
@@ -131,6 +189,7 @@ public class MilvusHybridKnowledgeStore {
.build()); .build());
} }
/** 按 metadata._source(规范化路径)删除,用于按文件路径重索引。 */
public void deleteBySource(String sourcePath) { public void deleteBySource(String sourcePath) {
if (sourcePath == null || sourcePath.isBlank()) { if (sourcePath == null || sourcePath.isBlank()) {
return; return;
@@ -144,8 +203,8 @@ public class MilvusHybridKnowledgeStore {
} }
/** /**
* Drop the configured knowledge collection (if present) and recreate empty dense+BM25 schema. * 删除并重建当前知识 collection(空的 dense+BM25 schema)。
* Used by knowledge rebuild scripts. Existing vectors in this collection are destroyed. * 供 {@code /api/knowledge/rebuild-hybrid} 与重建脚本使用;会销毁该 collection 全部向量。
*/ */
public synchronized Map<String, Object> dropAndRecreateCollection() { public synchronized Map<String, Object> dropAndRecreateCollection() {
Map<String, Object> result = new LinkedHashMap<>(); Map<String, Object> result = new LinkedHashMap<>();
@@ -178,6 +237,10 @@ public class MilvusHybridKnowledgeStore {
return result; return result;
} }
/**
* 单路 dense ANN(L2)。
* {@code score} = L2 距离(越小越好);{@code scoreLabel} = {@link RetrievalScoreLabels#DENSE}。
*/
public List<VectorSearchService.SearchResult> searchDense(String queryEmbeddingText, public List<VectorSearchService.SearchResult> searchDense(String queryEmbeddingText,
List<Float> queryVector, List<Float> queryVector,
int topK, int topK,
@@ -194,12 +257,22 @@ public class MilvusHybridKnowledgeStore {
builder.filter(filter); builder.filter(filter);
} }
SearchResp resp = client().search(builder.build()); SearchResp resp = client().search(builder.build());
return toSearchResults(resp, "l2_distance", false); return toSearchResults(resp, RetrievalScoreLabels.DENSE);
} }
/** /**
* Dense + BM25 hybrid fused by RRF. Dense L2 scores are attached when the same id * Dense + BM25 真混合检索(Milvus 服务端融合)。
* appears in a parallel dense search so quality thresholds stay meaningful. *
* <ol>
* <li>dense 子路:{@code vector},L2</li>
* <li>BM25 子路:{@code sparse_vector} + {@link EmbeddedText}</li>
* <li>{@link HybridSearchReq} + {@link RRFRanker} → 返回序即权威序</li>
* </ol>
*
* <p>{@code scoreLabel=hybrid};{@code score}/{@code rawScore} 保留引擎融合分,
* <b>不</b>用 dense L2 覆盖主分或改 label。可选并行 dense 探测仅填充
* {@link VectorSearchService.SearchResult#setDenseDistance},供后处理绝对质量闸门
* (如 L0 filter low-quality → unfiltered retry),排序仍以 RRF 返回序为准。</p>
*/ */
public List<VectorSearchService.SearchResult> searchHybrid(String queryText, public List<VectorSearchService.SearchResult> searchHybrid(String queryText,
List<Float> queryVector, List<Float> queryVector,
@@ -236,37 +309,45 @@ public class MilvusHybridKnowledgeStore {
.build(); .build();
SearchResp hybridResp = client().hybridSearch(hybridReq); SearchResp hybridResp = client().hybridSearch(hybridReq);
List<VectorSearchService.SearchResult> fused = toSearchResults(hybridResp, "rrf_fused", true); List<VectorSearchService.SearchResult> fused = toSearchResults(hybridResp, RetrievalScoreLabels.HYBRID);
attachDenseDistances(fused, queryText, queryVector, pathTopK, category);
// Attach dense-compatible L2 when available.
Map<String, Float> denseScores = new HashMap<>();
try {
for (VectorSearchService.SearchResult denseHit :
searchDense(queryText, queryVector, pathTopK, category)) {
if (denseHit.getId() != null) {
denseScores.put(denseHit.getId(), denseHit.getScore());
}
}
} catch (Exception e) {
log.warn("Dense score enrichment failed: {}", e.getMessage());
}
for (VectorSearchService.SearchResult hit : fused) {
Float dense = denseScores.get(hit.getId());
if (dense != null) {
hit.setScore(dense);
hit.setScoreLabel("l2_distance");
} else {
// BM25-only hit: treat as weak for legacy thresholds
hit.setScore((float) maxL2Distance);
hit.setScoreLabel("bm25_only_no_dense");
}
}
return fused; return fused;
} }
private List<VectorSearchService.SearchResult> toSearchResults(SearchResp resp, /**
String scoreLabel, * Attach dense L2 by id for quality gates only — never overwrites hybrid score/label/order.
boolean fused) { */
private void attachDenseDistances(List<VectorSearchService.SearchResult> fused,
String queryText,
List<Float> queryVector,
int pathTopK,
String category) {
if (fused == null || fused.isEmpty()) {
return;
}
try {
Map<String, Float> denseById = new HashMap<>();
for (VectorSearchService.SearchResult denseHit :
searchDense(queryText, queryVector, pathTopK, category)) {
if (denseHit.getId() != null) {
denseById.put(denseHit.getId(), denseHit.getScore());
}
}
for (VectorSearchService.SearchResult hit : fused) {
Float l2 = denseById.get(hit.getId());
if (l2 != null) {
hit.setDenseDistance(l2.doubleValue());
}
}
} catch (Exception e) {
log.warn("Dense distance attach for hybrid quality gate failed: {}", e.getMessage());
}
}
/**
* 将 Milvus {@link SearchResp} 映射为上层结果;列表顺序即检索权威序(adapter 赋 originalRank)。
*/
private List<VectorSearchService.SearchResult> toSearchResults(SearchResp resp, String scoreLabel) {
List<VectorSearchService.SearchResult> out = new ArrayList<>(); List<VectorSearchService.SearchResult> out = new ArrayList<>();
if (resp == null || resp.getSearchResults() == null || resp.getSearchResults().isEmpty()) { if (resp == null || resp.getSearchResults() == null || resp.getSearchResults().isEmpty()) {
return out; return out;
@@ -293,27 +374,16 @@ public class MilvusHybridKnowledgeStore {
Float score = row.getScore(); Float score = row.getScore();
mapped.setRawScore(score == null ? null : score.doubleValue()); mapped.setRawScore(score == null ? null : score.doubleValue());
mapped.setScoreLabel(scoreLabel); mapped.setScoreLabel(scoreLabel);
if (fused) { // dense: L2;hybrid: 引擎融合分(后处理 quality 主要看 rank,不依赖此量纲)
// temporary; may be overwritten with dense L2 mapped.setScore(score == null ? 0f : score);
mapped.setScore(score == null ? (float) maxL2Distance : invertUnknownScore(score));
} else {
mapped.setScore(score == null ? (float) maxL2Distance : score);
}
out.add(mapped); out.add(mapped);
} }
return out; return out;
} }
private float invertUnknownScore(float score) { /**
// RRF-like small scores: map higher better -> small L2-like distance * 组装标量过滤表达式:category、kb_scope(配置级)可叠加,用 {@code &&} 连接。
double bounded = Math.max(0.0, Math.min(1.0, score)); */
if (score > 1.0f) {
// already distance-like
return score;
}
return (float) ((1.0 - bounded) * maxL2Distance);
}
private String buildFilter(String category) { private String buildFilter(String category) {
List<String> parts = new ArrayList<>(); List<String> parts = new ArrayList<>();
String categoryFilter = trimToNull(category); String categoryFilter = trimToNull(category);
@@ -352,6 +422,17 @@ public class MilvusHybridKnowledgeStore {
return new MilvusClientV2(builder.build()); return new MilvusClientV2(builder.build());
} }
/**
* 若不存在则创建 dense+BM25 hybrid collection。
*
* <p>关键点:</p>
* <ul>
* <li>{@code search_text} 开启 analyzer,作为 BM25 语料。</li>
* <li>{@link FunctionType#BM25}:input={@code search_text} → output={@code sparse_vector}。</li>
* <li>dense:IVF_FLAT + L2;sparse:SPARSE_INVERTED_INDEX + BM25。</li>
* </ul>
* <p>已存在的 collection 不会改 schema;schema 变更需走 {@link #dropAndRecreateCollection()}。</p>
*/
private void ensureCollection(MilvusClientV2 milvusClient) { private void ensureCollection(MilvusClientV2 milvusClient) {
Boolean exists = milvusClient.hasCollection(HasCollectionReq.builder() Boolean exists = milvusClient.hasCollection(HasCollectionReq.builder()
.collectionName(collectionName) .collectionName(collectionName)
@@ -376,6 +457,7 @@ public class MilvusHybridKnowledgeStore {
.dataType(DataType.VarChar) .dataType(DataType.VarChar)
.maxLength(MilvusConstants.CONTENT_MAX_LENGTH) .maxLength(MilvusConstants.CONTENT_MAX_LENGTH)
.build()); .build());
// BM25 语料字段:必须 enableAnalyzer,Function 才能从文本生成 sparse
schema.addField(AddFieldReq.builder() schema.addField(AddFieldReq.builder()
.fieldName(FIELD_SEARCH_TEXT) .fieldName(FIELD_SEARCH_TEXT)
.dataType(DataType.VarChar) .dataType(DataType.VarChar)
@@ -395,6 +477,7 @@ public class MilvusHybridKnowledgeStore {
.fieldName(FIELD_METADATA) .fieldName(FIELD_METADATA)
.dataType(DataType.JSON) .dataType(DataType.JSON)
.build()); .build());
// 写入 search_text 时,Milvus 自动维护 sparse_vector(应用层 insert 不填 sparse)
schema.addFunction(CreateCollectionReq.Function.builder() schema.addFunction(CreateCollectionReq.Function.builder()
.functionType(FunctionType.BM25) .functionType(FunctionType.BM25)
.name("bm25_fn") .name("bm25_fn")
@@ -446,6 +529,7 @@ public class MilvusHybridKnowledgeStore {
} }
} }
/** 过滤表达式字符串转义,防止引号打断 expr。 */
private static String escapeFilter(String value) { private static String escapeFilter(String value) {
return value.replace("\\", "\\\\").replace("\"", "\\\""); return value.replace("\\", "\\\\").replace("\"", "\\\"");
} }
@@ -19,6 +19,8 @@ public record KnowledgeSearchHit(
String source, String source,
String title, String title,
String breadcrumb, String breadcrumb,
int originalRank int originalRank,
/** Optional dense L2 for hybrid quality gates; null on dense-only hits. */
Double denseDistance
) { ) {
} }
@@ -1,10 +1,20 @@
package com.superbiz.agent.service.retrieval; package com.superbiz.agent.service.retrieval;
/** /**
* Retrieval mode for {@link KnowledgeSearchPort}. * {@link KnowledgeSearchPort} 检索模式。
* Delivery 1 only requires {@link #DENSE}; hybrid arrives in a later change. *
* <ul>
* <li>{@link #DENSE} —— 单路向量 ANN(L2)</li>
* <li>{@link #HYBRID} —— dense + Milvus 服务端 BM25 + RRF 融合</li>
* </ul>
*
* <p>当前实际生效模式由全局配置 {@code retrieval.search.mode} 决定
*(见 {@link com.superbiz.agent.service.VectorSearchService});
* 请求里的 mode 预留作将来 per-call 覆盖,adapter 暂未按请求切换。</p>
*/ */
public enum KnowledgeSearchMode { public enum KnowledgeSearchMode {
/** 仅 dense 向量检索。 */
DENSE, DENSE,
/** dense + BM25 hybrid(Milvus {@code hybridSearch} + RRF)。 */
HYBRID HYBRID
} }
@@ -3,8 +3,11 @@ package com.superbiz.agent.service.retrieval;
import java.util.List; import java.util.List;
/** /**
* Application boundary for knowledge semantic search. * 知识语义检索的应用边界端口。
* Implementations may wrap VectorStore, hybrid engines, etc. without leaking SDK details upward. *
* <p>实现可对接 dense / hybrid 等引擎,但不得向上层泄漏 SDK 类型。
* 当前实现:{@link VectorKnowledgeSearchAdapter} → {@code VectorSearchService}
* → {@code MilvusHybridKnowledgeStore}(Milvus SDK v2 dense 或 dense+BM25 RRF)。</p>
*/ */
public interface KnowledgeSearchPort { public interface KnowledgeSearchPort {
@@ -0,0 +1,45 @@
package com.superbiz.agent.service.retrieval;
/**
* 检索结果一级 {@code scoreLabel} 约定。
*
* <p>只区分两种检索形态(与 {@code retrieval.search.mode} 对齐),
* 不再使用 {@code bm25_only_*} 等作为正式一级 label。</p>
*/
public final class RetrievalScoreLabels {
/** dense-only ANN:{@code score} 为 L2 距离(越小越好)。 */
public static final String DENSE = "dense";
/** hybrid(dense+BM25+RRF):{@code score}/raw 为融合侧信号;质量分主要看 rank。 */
public static final String HYBRID = "hybrid";
private RetrievalScoreLabels() {
}
/**
* 将历史/别名 label 归一到 {@link #DENSE} 或 {@link #HYBRID}。
* 未知或空 → dense(保守,按 L2 解释失败时 quality 偏低)。
*/
public static String canonicalize(String scoreLabel) {
if (scoreLabel == null || scoreLabel.isBlank()) {
return DENSE;
}
String label = scoreLabel.trim().toLowerCase();
return switch (label) {
case DENSE, "l2_distance", "l2" -> DENSE;
case HYBRID, "rrf_fused", "rrf", "bm25_only_no_dense", "bm25_only" -> HYBRID;
default -> label.contains("hybrid") || label.contains("rrf") || label.contains("bm25")
? HYBRID
: DENSE;
};
}
public static boolean isHybrid(String scoreLabel) {
return HYBRID.equals(canonicalize(scoreLabel));
}
public static boolean isDense(String scoreLabel) {
return DENSE.equals(canonicalize(scoreLabel));
}
}
@@ -0,0 +1,71 @@
package com.superbiz.agent.service.retrieval;
/**
* 检索分 → 统一 {@code qualityScore ∈ [0,1]}(越大越好)的唯一转换点。
*
* <p>后处理排序仍按 {@code originalRank};本类只负责质量闸门 / relevance 用分。</p>
*
* <ul>
* <li>{@link RetrievalScoreLabels#DENSE}:{@code score} = L2 → {@code 1 - clamp(l2)/maxL2}</li>
* <li>{@link RetrievalScoreLabels#HYBRID}:优先用可选 {@code denseDistance} 做绝对质量
* (恢复 L0 filter low-quality 等闸门);无 dense 时回退 rank 映射</li>
* </ul>
*/
public final class RetrievalScoreNormalizer {
private RetrievalScoreNormalizer() {
}
/**
* @param scoreLabel {@link RetrievalScoreLabels#DENSE} / {@link RetrievalScoreLabels#HYBRID}
* @param score 引擎主分:dense=L2;hybrid=融合分(hybrid 质量不依赖其量纲)
* @param originalRank 检索名次(1-based)
* @param batchSize 本轮候选数(rank 回退映射用)
* @param maxL2Distance L2 上界
* @param denseDistance hybrid 命中上可选的 dense L2;dense 模式可传 null
*/
public static double toQualityScore(String scoreLabel,
Double score,
Integer originalRank,
int batchSize,
double maxL2Distance,
Double denseDistance) {
String label = RetrievalScoreLabels.canonicalize(scoreLabel);
if (RetrievalScoreLabels.HYBRID.equals(label)) {
if (denseDistance != null) {
return l2ToQuality(denseDistance, maxL2Distance);
}
// BM25-only hybrid hit (no dense neighbor): conservative mid quality via rank
return rankToQuality(originalRank, batchSize);
}
return l2ToQuality(score, maxL2Distance);
}
/** Backward-compatible overload without denseDistance. */
public static double toQualityScore(String scoreLabel,
Double score,
Integer originalRank,
int batchSize,
double maxL2Distance) {
return toQualityScore(scoreLabel, score, originalRank, batchSize, maxL2Distance, null);
}
public static double l2ToQuality(Double l2Score, double maxL2Distance) {
if (l2Score == null) {
return 0.0;
}
double max = maxL2Distance > 0 ? maxL2Distance : 2.0;
double clamped = Math.min(Math.max(l2Score, 0.0), max);
return Math.max(0.0, 1.0 - clamped / max);
}
public static double rankToQuality(Integer originalRank, int batchSize) {
int rank = originalRank == null || originalRank < 1 ? 1 : originalRank;
int n = batchSize > 0 ? Math.max(batchSize, rank) : Math.max(rank, 1);
if (n <= 1) {
return 1.0;
}
double quality = 1.0 - (rank - 1) / (double) n;
return Math.max(1.0 / n, Math.min(1.0, quality));
}
}
@@ -10,11 +10,11 @@ import java.util.List;
import java.util.Map; import java.util.Map;
/** /**
* {@link KnowledgeSearchPort} adapter. * {@link KnowledgeSearchPort} 适配器:把向量检索结果映射为带 evidenceKey 的命中结构。
* *
* <p>Delegates to {@link VectorSearchService}, which is backed solely by * <p>委托 {@link VectorSearchService}(背后仅 {@code MilvusHybridKnowledgeStore}):
* Milvus V2 dense / dense+BM25 hybrid store. Mode selection lives in * dense 或 dense+BM25 hybrid 由配置 {@code retrieval.search.mode} 选择。
* {@code retrieval.search.mode}.</p> * 本类负责 metadata 解析、docId/chunk 身份与 evidenceKey,不碰 SDK。</p>
*/ */
@Component @Component
public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort { public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort {
@@ -76,7 +76,8 @@ public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort {
source, source,
EvidenceIdentity.metadataValue(metadata, "title"), EvidenceIdentity.metadataValue(metadata, "title"),
EvidenceIdentity.metadataValue(metadata, "breadcrumb"), EvidenceIdentity.metadataValue(metadata, "breadcrumb"),
originalRank originalRank,
result.getDenseDistance()
); );
} }
+13 -7
View File
@@ -167,17 +167,21 @@ rag:
enabled: false enabled: false
content-preview-limit: 300 content-preview-limit: 300
# 检索配置(单一 Milvus V2 后端;已移除 sdk/spring/auto 路由) # 检索配置
# 知识主路径:Milvus Java SDK v2(MilvusHybridKnowledgeStore),非 Spring AI VectorStore starter。
# 原因:starter(含 2.0.0)仅 dense similarity,无 hybridSearch / BM25 Function / RRFRanker。
# 已移除 legacy sdk/spring/auto 多后端路由。
retrieval: retrieval:
kb-scope: "" # empty means search all documents in hybrid collection kb-scope: "" # 非空则过滤 metadata.kb_scope;空=不过滤
search: search:
mode: hybrid # dense | hybrid (dense + BM25 RRF) # hybrid=线上主路径;dense=同库对照/评测/排障(非第二套线上策略)。见 mvp/architecture/rag-knowledge-retrieval-architecture.md §6.0
mode: hybrid # dense=单路L2对照 | hybrid=dense+服务端BM25+RRF
hybrid: hybrid:
rrf-k: 60 rrf-k: 60 # RRF 平滑参数 k,score=Σ 1/(k+rank)
normalization: normalization:
max-l2-distance: 2.0 # L2 距离上界(BGE-M3 单位向量 = 2.0) max-l2-distance: 2.0 # dense quality:L2 上界(单位向量 ≈ 2.0)
highly-relevant-threshold: 0.75 # similarity >= 0.75 → HIGHLY_RELEVANT highly-relevant-threshold: 0.75 # qualityScore >= 0.75 → PRECISE(hybrid 为序数分,见架构 §6)
reference-threshold: 0.5 # similarity >= 0.5 → REFERENCE reference-threshold: 0.5 # qualityScore >= 0.5 → REFERENCE;低于则低质/可 unfiltered retry
# Prometheus 配置 # Prometheus 配置
prometheus: prometheus:
@@ -213,6 +217,8 @@ logging:
# Agent-facing MySQL Tool uses independent logical datasources only. # Agent-facing MySQL Tool uses independent logical datasources only.
# Production entries are supplied by a dedicated profile and Secret injection. # Production entries are supplied by a dedicated profile and Secret injection.
# When data-sources is empty (default), query_mysql is NOT registered on the Diagnosis Agent
# (avoids a permanently broken tool that the model can still call).
harness: harness:
chat: chat:
worker-core-pool-size: 2 worker-core-pool-size: 2

Some files were not shown because too many files have changed in this diff Show More