feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
+1
-1
@@ -65,8 +65,8 @@ uploads/
|
||||
|
||||
### Windows / Runtime Artifacts
|
||||
*.stackdump
|
||||
NUL
|
||||
|
||||
### MVP Demo Generated Outputs
|
||||
mvp/demo/output/*.json
|
||||
!mvp/demo/output/README.md
|
||||
.pi/extensions/emdash-hook.ts
|
||||
|
||||
+3
-1
@@ -1,4 +1,4 @@
|
||||
# devflow 索引
|
||||
# devflow 索引
|
||||
|
||||
## Issue 生命周期
|
||||
|
||||
@@ -12,6 +12,8 @@
|
||||
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-28 | rag-eval-hybrid-baseline | 离线 RAG eval 对齐 hybrid:search.mode 生成器、fixture meta、baseline 重刷;hybrid 质量闸门可用 denseDistance。 | RAG/eval/baseline | search.mode, fixture meta, kb_scope rag-eval, L0 filter fallback, denseDistance | openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline | archived |
|
||||
| 2026-07-28 | rag-quality-score-unify | 统一 dense/hybrid scoreLabel 与 qualityScore;保检索序;去掉关键词 boost 改序与 hybrid L2 伪装。 | RAG/质量分/后处理 | qualityScore, scoreLabel dense/hybrid, originalRank, RetrievalScoreNormalizer, no boost rerank | openspec/changes/archive/2026-07-28-rag-quality-score-unify | archived |
|
||||
| 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived |
|
||||
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
|
||||
| 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived |
|
||||
|
||||
@@ -14,7 +14,7 @@ lookup_knowledge 同文档多 chunk 在后处理与投影阶段被 source 级去
|
||||
|
||||
## 范围
|
||||
|
||||
Delivery 1 only(见 `docs/milvus-hybrid-search-integration-checklist.md` §1.1)。
|
||||
Delivery 1 only(见 `docs/Milvus-Hybrid接入清单.md` §1.1)。
|
||||
|
||||
## 非目标
|
||||
|
||||
|
||||
@@ -93,7 +93,7 @@ Reference files:
|
||||
- `RagResultProjector.java`
|
||||
- `LookupKnowledgeToolTest.java`
|
||||
- `RagResultProjectorTest.java`
|
||||
- `docs/milvus-hybrid-search-integration-checklist.md`
|
||||
- `docs/Milvus-Hybrid接入清单.md`
|
||||
|
||||
Stack notes:
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
## 规格证据
|
||||
|
||||
- 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正)
|
||||
- `docs/milvus-hybrid-search-integration-checklist.md` §1.1 定义 Delivery 1 地基
|
||||
- `docs/Milvus-Hybrid接入清单.md` §1.1 定义 Delivery 1 地基
|
||||
|
||||
## 验证证据
|
||||
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
# Acceptance: rag-eval-hybrid-baseline
|
||||
|
||||
## Tasks
|
||||
|
||||
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
|
||||
|
||||
## 静态验证
|
||||
|
||||
- Snapshot generator path: no required `retrieval.vector-store.mode`.
|
||||
- README documents hybrid generation and offline/live split.
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
# exit 0
|
||||
```
|
||||
|
||||
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
|
||||
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
|
||||
|
||||
## 浏览器/人工
|
||||
|
||||
- 未做 UI 验证。
|
||||
|
||||
## 未验证 / 后续
|
||||
|
||||
- Dense vs hybrid dual-directory comparison report (knife-2).
|
||||
- CI wiring of offline eval as required gate (optional process).
|
||||
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
|
||||
|
||||
## Specs
|
||||
|
||||
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
|
||||
@@ -0,0 +1,16 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
|
||||
|
||||
## Goals (knife-1 only)
|
||||
|
||||
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
|
||||
- Fixtures record `searchMode` / `kbScope`.
|
||||
- README documents hybrid-era offline vs live loop.
|
||||
- Best-effort live seed + regenerate fixtures + update baseline.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
|
||||
@@ -0,0 +1,18 @@
|
||||
# Decisions: rag-eval-hybrid-baseline(最终版)
|
||||
|
||||
## Process
|
||||
|
||||
sm-flow standard-lean: Discover → Commit → Apply → Archive.
|
||||
|
||||
## Key decisions
|
||||
|
||||
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
|
||||
2. Fixture meta: `searchMode`, `kbScope` when set.
|
||||
3. Knife-2 (dual fixtures / mustNot golden) deferred.
|
||||
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
|
||||
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
|
||||
|
||||
## Trade-offs
|
||||
|
||||
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
|
||||
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
|
||||
@@ -0,0 +1,17 @@
|
||||
# Evidence: rag-eval-hybrid-baseline
|
||||
|
||||
## Pre-change
|
||||
|
||||
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
|
||||
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
|
||||
- Offline eval already supported Hit levels, recall@K, baseline diff.
|
||||
|
||||
## User decisions
|
||||
|
||||
- Scope: knife-1 only (no dual fixture dirs).
|
||||
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
|
||||
|
||||
## Apply-discovered
|
||||
|
||||
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
|
||||
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
|
||||
@@ -0,0 +1,41 @@
|
||||
# Acceptance: rag-quality-score-unify
|
||||
|
||||
## Tasks
|
||||
|
||||
OpenSpec `tasks.md` 全部 `[x]`(1.1–6.2)。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- 生产路径 grep:无 `bm25_only_no_dense` 发射、无 hybrid L2 enrichment(仅 Labels canonicalize 兼容旧串)。
|
||||
- 架构文档 §6 与 `application.yml` 注释已对齐 quality 契约。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```text
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
```
|
||||
|
||||
| 套件 | 结果 |
|
||||
|---|---|
|
||||
| RetrievalScoreNormalizerTest | 4 passed |
|
||||
| KnowledgeEvidencePostProcessorTest | 6 passed |
|
||||
| LookupKnowledgeToolTest | 7 passed |
|
||||
| VectorSearchServiceTest | 2 passed |
|
||||
| VectorKnowledgeSearchAdapterHybridTest | 1 passed |
|
||||
|
||||
(PowerShell 可能将 JVM warning 标为 exit 1;日志中为 BUILD SUCCESS / Failures: 0。)
|
||||
|
||||
## 浏览器 / 人工验证
|
||||
|
||||
- 未跑:live `lookup_knowledge` hybrid vs dense 对照、生产阈值标定。
|
||||
|
||||
## 未验证
|
||||
|
||||
| 项 | 风险 | 建议 |
|
||||
|---|---|---|
|
||||
| 真实 Milvus hybrid 联调 | 序/质量分布与单测 mock 有差 | 启动服务后固定 query 集切 mode 对比 |
|
||||
| 阈值 0.75/0.5 在 hybrid rank 分下的标定 | retry/PRECISE 偏多或偏少 | 看 trace topSimilarity 再调 yml |
|
||||
|
||||
## Specs 同步
|
||||
|
||||
- 主规格新增:`openspec/specs/rag-retrieval-quality-score/spec.md`(archive 时从 delta 同步)。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Brief: rag-quality-score-unify
|
||||
|
||||
## Background
|
||||
|
||||
真 BM25 hybrid(dense + BM25 + RRF)已上线,但后处理仍把 hybrid 结果伪装成 L2 做 `normalizeL2`,并用 L0 domain/entity/keyword contains 加分改序。排序权威与质量闸门分裂,词面信号被 BM25 与后处理双重计分。
|
||||
|
||||
## Goals
|
||||
|
||||
- 一级 `scoreLabel` 仅 `dense` | `hybrid`
|
||||
- 唯一 `toQualityScore`;后处理 label-agnostic
|
||||
- 排序主序 = 检索 `originalRank`;去掉关键词 boost 改序
|
||||
- hybrid quality = 本轮 rank 纯映射(不做 max(rank, denseSim)、不为闸门回填 L2)
|
||||
- 保留 `mode=dense` 作同库召回对照;线上默认 hybrid
|
||||
|
||||
## Scope
|
||||
|
||||
内部 RAG:store 发射、normalizer、evidence post-process、单测、架构文档 §6。
|
||||
|
||||
## Non-goals
|
||||
|
||||
精排 / query rewrite / 邻块、schema rebuild、改 Agent ACI 字段名、删除 dense 对照 mode。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Decisions: rag-quality-score-unify(最终版)
|
||||
|
||||
## Scale / process
|
||||
|
||||
- sm-flow standard:Discover → Commit → Apply → Archive
|
||||
- Committed OpenSpec:`openspec/changes/rag-quality-score-unify/`(归档后见 archive 目录)
|
||||
- 废止:`rag-bm25-hybrid-drop-sdk` 中「dense L2 enrichment for threshold compatibility」
|
||||
|
||||
## Key decisions
|
||||
|
||||
1. **Label**:仅 `dense` | `hybrid`;旧别名 canonicalize。
|
||||
2. **Normalizer**:唯一 `toQualityScore`;dense=L2 公式;hybrid=rank 线性映射(batchSize)。
|
||||
3. **Store**:hybrid 不回填 L2、不发 `bm25_only_*`;返回序即 RRF 序。
|
||||
4. **Post-process**:`originalRank` ASC;L0 重叠只写 hitReasons;relevance/low-quality 只看 qualityScore;PRECISE 不要求 hint support。
|
||||
5. **Mode**:hybrid 主路径;dense 同库对照(架构 §6.0)。
|
||||
|
||||
## Trade-offs
|
||||
|
||||
- hybrid quality 为序数分,跨 query 绝对值不可比;阈值可能需后续标定。
|
||||
- 去掉 boost 改序后,「词面热语义冷」不再被后处理抬升;词面交给 BM25+RRF。
|
||||
|
||||
## Risks accepted
|
||||
|
||||
- `relevance_level` / unfiltered retry 分布变化(产品已接受)。
|
||||
- 未做 live E2E / 人工 hybrid 对照评测(见 acceptance 未验证项)。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Evidence: rag-quality-score-unify
|
||||
|
||||
## Code (pre-change)
|
||||
|
||||
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后并行 dense 回填 L2;BM25-only → `bm25_only_no_dense` + maxL2。
|
||||
- `KnowledgeEvidencePostProcessor`:一律 `normalizeL2(score)` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序;PRECISE 需 `hasHintSupport`。
|
||||
|
||||
## User decisions (grill)
|
||||
|
||||
| ID | 结论 |
|
||||
|---|---|
|
||||
| Q1 | 一级 label 仅 dense/hybrid;bm25_only 不作正式 label |
|
||||
| Q2 | 后处理去掉 contains 加分改序,保 originalRank |
|
||||
| Q3 | 唯一 toQualityScore;后处理统一 |
|
||||
| Q4 | dense mode 保留作对照 |
|
||||
| Q7 | 接受 relevance_level / retry 分布变化 |
|
||||
| Q8 | hybrid quality = **纯 rank 映射** |
|
||||
|
||||
## Post-change anchors
|
||||
|
||||
- `RetrievalScoreLabels` / `RetrievalScoreNormalizer`
|
||||
- `MilvusHybridKnowledgeStore`(无 L2 overwrite / 无 bm25_only 发射)
|
||||
- `KnowledgeEvidencePostProcessor`(rank sort + explain-only L0 overlap)
|
||||
- OpenSpec delta:`rag-retrieval-quality-score`
|
||||
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
|
||||
@@ -53,6 +53,18 @@ docs/
|
||||
1. [分析笔记目录](analysis/) - 代码分析和问题分析
|
||||
2. [临时报告目录](reports/) - 修复和验证报告
|
||||
|
||||
### RAG 设计讨论(docs 根目录)
|
||||
1. [RAG 排序:多路召回与 RRF](RAG排序-多路召回与RRF.md) - K、融合、L0 边界
|
||||
2. [Hybrid 之后的 qualityScore 与后处理](RAG-Hybrid质量分与后处理.md) - 上一代问题、L2 伪装、统一归一化
|
||||
3. [Agent 如何读 relevance_level](RAG-Agent如何读relevance_level.md) - 粗相关度标签的含义与误读
|
||||
4. [RAG 离线评测:讨论、设计与落地](RAG离线评测-基线设计.md) - Golden/Fixture、流程图、hybrid 对齐与闸门修复
|
||||
5. [Milvus hybrid 接入清单](Milvus-Hybrid接入清单.md)
|
||||
6. [RAG Trace / 审计(架构)](../mvp/architecture/RAG检索可观测性与审计.md) - 请求内 trace、tool_invocation、Trace API
|
||||
|
||||
### 诊断全流程(E2E 导读)
|
||||
1. [一次诊断到底发生了什么](一次诊断全流程-E2E导读.md) - SUCCESS 全流程:阶段拆解、token、timeline、字段词典
|
||||
2. [RAG 审计补丁 E2E:step_id + query](RAG审计补丁-stepid-query-E2E验收.md) - 审计字段 live 验收(含业务 FALLBACK 样本)
|
||||
|
||||
---
|
||||
|
||||
## 📚 学习笔记 (learning/)
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
**前提**:旧 Milvus SDK 直连检索路径后续废弃,不作为长期实现基础
|
||||
**目标**:在现有 `lookup_knowledge` pipeline 上接入 dense + sparse/BM25 混合检索,融合优先走服务端 RRF
|
||||
**关联文档**:
|
||||
- `docs/rag-ranking-multipath-retrieval-and-rrf.md`(排序与多路召回判断框架)
|
||||
- `docs/RAG排序-多路召回与RRF.md`(排序与多路召回判断框架)
|
||||
- 本文后续实现讨论以本节 **「交付拆分:分块去重 + Hybrid 同规划」** 为基线
|
||||
|
||||
---
|
||||
@@ -12,7 +12,7 @@
|
||||
## 1. 结论先说
|
||||
|
||||
可以接,而且和前面讨论的多路召回 / RRF 高度一致。
|
||||
但当前项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是:
|
||||
但(写作当时)项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是:
|
||||
|
||||
```text
|
||||
1. schema 只有 dense,没有 sparse/BM25 字段
|
||||
@@ -32,7 +32,27 @@
|
||||
- 分块去重与 hybrid 同规划、分里程碑交付(先共用地基,再开 hybrid)
|
||||
```
|
||||
|
||||
## 实现状态(2026-07-27)
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph gaps["写作时的缺口"]
|
||||
G1[无 sparse/BM25 schema]
|
||||
G2[写入只有 dense]
|
||||
G3[单路 similaritySearch]
|
||||
G4[后处理伪融合]
|
||||
G5[source 级去重吞 chunk]
|
||||
end
|
||||
|
||||
subgraph principles["接入原则"]
|
||||
P1[端口抽象 · 不堆旧 SDK]
|
||||
P2[融合下沉向量库 RRF]
|
||||
P3[应用层:filter/dedup/return-n/投影]
|
||||
P4[先地基后 hybrid 分里程碑]
|
||||
end
|
||||
|
||||
gaps --> principles
|
||||
```
|
||||
|
||||
## 实现状态(2026-07-27,后续已完成)
|
||||
|
||||
| 里程碑 | 状态 | 说明 |
|
||||
|---|---|---|
|
||||
@@ -40,17 +60,42 @@
|
||||
| 交付 2a 应用层 multi-path+RRF | **已完成并归档** | `2026-07-27-rag-hybrid-search-rrf`(已被 2b 取代为生产路径) |
|
||||
| 交付 2b 真 BM25 hybrid + 废弃 SDK | **已完成并归档** | `2026-07-27-rag-bm25-hybrid-drop-sdk` |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
D1[交付1<br/>chunk 身份/去重] --> D2a[交付2a<br/>app RRF]
|
||||
D2a --> D2b[交付2b<br/>真 BM25 hybrid]
|
||||
D2b --> NOW[生产:V2 store + hybrid mode]
|
||||
```
|
||||
|
||||
**当前生产知识路径:**
|
||||
|
||||
```text
|
||||
VectorIndexService / VectorSearchService
|
||||
-> MilvusHybridKnowledgeStore (MilvusClientV2 only)
|
||||
collection: milvus.collection (default biz_hybrid)
|
||||
collection: milvus.collection (default biz)
|
||||
mode: retrieval.search.mode = dense | hybrid
|
||||
hybrid: dense ANN + BM25 sparse ANN + RRFRanker
|
||||
```
|
||||
|
||||
**运维必做:** 全量重灌知识库到 `biz_hybrid`;旧 `biz` collection 不再被知识路径使用。
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph write["写入"]
|
||||
UP[upload / init / rebuild] --> VIS[VectorIndexService]
|
||||
VIS --> STORE[MilvusHybridKnowledgeStore]
|
||||
end
|
||||
|
||||
subgraph read["检索"]
|
||||
LK[lookup_knowledge] --> VSS[VectorSearchService]
|
||||
VSS -->|dense| SD[searchDense]
|
||||
VSS -->|hybrid| SH[searchHybrid + RRF]
|
||||
SD --> STORE
|
||||
SH --> STORE
|
||||
end
|
||||
|
||||
STORE --> COL[(Milvus collection biz<br/>dense + BM25 schema)]
|
||||
```
|
||||
|
||||
**运维必做:** 全量重灌知识库到 hybrid schema collection(配置名以 `milvus.collection` 为准,常见 `biz`);旧纯 dense collection 不能直接当 hybrid 用。
|
||||
|
||||
---
|
||||
|
||||
@@ -83,6 +128,19 @@ VectorIndexService / VectorSearchService
|
||||
-> return-n / Agent 投影
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
HIT[检索命中] --> ID[候选身份<br/>docId / chunkIndex / evidenceKey]
|
||||
ID --> DEDUP[chunk 去重 · 每文档上限]
|
||||
DEDUP --> FUSE[排序融合 · 库内 RRF]
|
||||
FUSE --> RET[return-n]
|
||||
RET --> PROJ[Agent 投影]
|
||||
|
||||
ID -.->|交付1 地基| D1[chunk identity]
|
||||
DEDUP -.-> D1
|
||||
FUSE -.->|交付2| D2[hybrid]
|
||||
```
|
||||
|
||||
若拆开且顺序错误:
|
||||
|
||||
| 只做一项 | 后果 |
|
||||
@@ -305,6 +363,27 @@ VectorIndexService / VectorSearchService
|
||||
LookupResult / Projector
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph write_port["写入边界"]
|
||||
W[KnowledgeWritePort<br/>upsert / deleteByDocId]
|
||||
end
|
||||
|
||||
subgraph search_port["检索边界"]
|
||||
S[KnowledgeSearchPort<br/>mode DENSE / HYBRID]
|
||||
end
|
||||
|
||||
W --> VDB[(Vector DB<br/>dense + BM25 sparse<br/>metadata filter)]
|
||||
S --> VDB
|
||||
|
||||
UP[upload/init] --> W
|
||||
LK[lookup_knowledge] --> S
|
||||
S --> RET[DocumentRetriever]
|
||||
RET --> POST[PostProcess<br/>dedup / cap / threshold / pack]
|
||||
POST --> PROJ[LookupResult / Projector]
|
||||
PROJ --> AG[Agent]
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- **Port** 是应用边界,实现可换成 Spring AI、Milvus 新客户端、或其他封装。
|
||||
@@ -385,6 +464,19 @@ category/kb_scope/doc_id -> 标量过滤索引(如需要)
|
||||
- 稳定 chunk id(doc_id + chunk_index 派生)
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
CHUNK[DocumentChunk] --> EMB[dense embed]
|
||||
CHUNK --> ST[search_text<br/>title+path+content]
|
||||
CHUNK --> META[docId/chunkIndex<br/>category/kb_scope]
|
||||
EMB --> ROW[upsert row]
|
||||
ST --> ROW
|
||||
META --> ROW
|
||||
ROW --> FN[BM25 Function<br/>search_text → sparse]
|
||||
ROW --> COL[(collection)]
|
||||
FN --> COL
|
||||
```
|
||||
|
||||
### 5.3 注意
|
||||
|
||||
- embedding 文本可以继续拼 `Title/Path/Content`
|
||||
@@ -429,6 +521,16 @@ SearchHit {
|
||||
}
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
REQ[SearchRequest<br/>query · retrieveK · mode<br/>filter · rrfK] --> PORT[KnowledgeSearchPort]
|
||||
PORT -->|DENSE| D[dense ANN only]
|
||||
PORT -->|HYBRID| H[dense + BM25 + RRF]
|
||||
D --> HIT[SearchHit 列表<br/>id/docId/chunk · content · ranks]
|
||||
H --> HIT
|
||||
HIT --> APP[后处理 / 投影]
|
||||
```
|
||||
|
||||
### 6.2 `VectorSearchService` 怎么演进
|
||||
|
||||
短期:
|
||||
@@ -0,0 +1,396 @@
|
||||
# Agent 如何读 `relevance_level`:它是什么、不是什么
|
||||
|
||||
**日期**:2026-07-28
|
||||
**范围**:`lookup_knowledge` 投影给 Agent 的粗粒度相关度标签
|
||||
**读者**:要在 Diagnosis Agent / 工具契约里正确使用知识库结果的工程与提示词同学
|
||||
**关联**:
|
||||
|
||||
- ACI 契约:`RagToolResult.relevance_level` / `RagRelevanceLevel`
|
||||
- 计算:`KnowledgeEvidencePostProcessor` → `qualityScore` 阈值
|
||||
- 质量统一:`docs/RAG-Hybrid质量分与后处理.md`
|
||||
- 多路与 RRF:`docs/RAG排序-多路召回与RRF.md`
|
||||
- **运行时 Trace / 审计**:`mvp/architecture/RAG检索可观测性与审计.md`
|
||||
|
||||
---
|
||||
|
||||
## 1. 一句话定义
|
||||
|
||||
**`relevance_level` 是对「这一次 `lookup_knowledge` 调用整体有多相关」的粗档标签,不是某一条 evidence 的分数,也不是 0~1 的相似度。**
|
||||
|
||||
Agent 真正写诊断、做引用时,仍应以 `evidence[]` 里的 excerpt 为准;`relevance_level` 只帮助判断:**这批评据大概有多硬、还要不要再查。**
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph tool["lookup_knowledge 结果"]
|
||||
ES[evidence_status]
|
||||
EV[evidence excerpt]
|
||||
RL[relevance_level]
|
||||
TR[truncated]
|
||||
end
|
||||
|
||||
ES --> DEC{Agent 决策}
|
||||
EV --> DEC
|
||||
RL --> DEC
|
||||
TR --> DEC
|
||||
|
||||
DEC --> A1[写结论 / 引用]
|
||||
DEC --> A2[再查 / 换工具]
|
||||
DEC --> A3[证据不足降级]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 它出现在哪里
|
||||
|
||||
成功(或有结果)的知识库工具投影里,典型形状:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "EVIDENCE_FOUND",
|
||||
"tool_call_id": "call-…",
|
||||
"query": "用户/Agent 的检索句",
|
||||
"evidence": [
|
||||
{
|
||||
"document_id": "doc#chunk-0",
|
||||
"source": "…",
|
||||
"title": "…",
|
||||
"breadcrumb": "…",
|
||||
"excerpt": "…"
|
||||
}
|
||||
],
|
||||
"returned_count": 3,
|
||||
"relevance_level": "PRECISE",
|
||||
"truncated": false
|
||||
}
|
||||
```
|
||||
|
||||
要点:
|
||||
|
||||
- JSON 字段名是 **`relevance_level`**(snake_case)
|
||||
- Java 枚举:`RagRelevanceLevel`(`PRECISE` / `HIGHLY_RELEVANT` / `REFERENCE`)
|
||||
- 无可用证据或质量不够时,字段常为 **null / 省略**(`NON_NULL`)
|
||||
- Agent **看不到** raw L2、RRF 分、`retrievalTrace`、`rerankTrace`(ACI 有意裁掉)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph internal["内部 LookupResult(审计/调试)"]
|
||||
QS[qualityScore / topSimilarity]
|
||||
RT[retrievalTrace / rerankTrace]
|
||||
CP[contextPack 全文]
|
||||
SC[raw score / scoreLabel]
|
||||
end
|
||||
|
||||
subgraph project["RagResultProjector"]
|
||||
P[裁剪与规范化]
|
||||
end
|
||||
|
||||
subgraph agent["Agent 可见 RagToolResult"]
|
||||
A1[evidence_status]
|
||||
A2[tool_call_id / query]
|
||||
A3[evidence excerpt 列表]
|
||||
A4[relevance_level 可选]
|
||||
A5[returned_count / truncated]
|
||||
end
|
||||
|
||||
internal --> P --> agent
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 枚举值怎么理解
|
||||
|
||||
| 值 | 产品语义 | Agent 侧更合理的用法 |
|
||||
|----|----------|----------------------|
|
||||
| **PRECISE** | 整体很贴:当前 top 证据质量高,继续同题检索不太可能更准 | 优先依据 `evidence[]` 组织结论;避免无意义的重复 `lookup_knowledge` |
|
||||
| **HIGHLY_RELEVANT** | 高度相关(枚举保留) | 与 PRECISE 类似,略保守表述即可 |
|
||||
| **REFERENCE** | 可作参考,但不到「已精准命中」 | 可引用,但结论留余地;缺维度时换 query 再查或叠日志/指标 |
|
||||
| **null / 不出现** | 没有可报的粗相关档(无证据或 top 质量偏低) | **不要**当成知识库已证实;按证据不足处理 |
|
||||
|
||||
### 和 `evidence_status` 的分工
|
||||
|
||||
| 字段 | 回答的问题 |
|
||||
|------|------------|
|
||||
| `evidence_status` | 这次有没有合法、可引用的证据(如 `EVIDENCE_FOUND` / `NO_EVIDENCE`) |
|
||||
| `relevance_level` | **有证据时**,整体有多贴(粗档) |
|
||||
| `evidence[]` | 具体可以引用哪些片段 |
|
||||
|
||||
没有证据时,不应指望靠 `relevance_level`「升级」出结论;契约上也不会用 level 把空结果扮成有证据。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
ES{evidence_status}
|
||||
ES -->|NO_EVIDENCE| N1[不要当知识库已证实]
|
||||
ES -->|EVIDENCE_FOUND| RL{relevance_level}
|
||||
|
||||
RL -->|PRECISE| U1[优先引用 excerpt · 少重复检索]
|
||||
RL -->|REFERENCE| U2[可引用 · 结论留余地]
|
||||
RL -->|null / 缺省| U3[有块但质量偏低 · 慎用强结论]
|
||||
RL -->|HIGHLY_RELEVANT| U4[与 PRECISE 类似 · 略保守]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 它是怎么算出来的(实现口径)
|
||||
|
||||
### 4.1 只看「本轮第一名」的 qualityScore
|
||||
|
||||
后处理在完成排序、去重、截断之后:
|
||||
|
||||
```text
|
||||
取 originalRank 最优(排序后第一条)的 qualityScore
|
||||
≥ highly-relevant-threshold (默认 0.75) → PRECISE
|
||||
≥ reference-threshold (默认 0.5) → REFERENCE
|
||||
否则 / 无可用证据 → null
|
||||
```
|
||||
|
||||
配置(`application.yml`):
|
||||
|
||||
```yaml
|
||||
retrieval:
|
||||
normalization:
|
||||
max-l2-distance: 2.0
|
||||
highly-relevant-threshold: 0.75
|
||||
reference-threshold: 0.5
|
||||
```
|
||||
|
||||
因此:
|
||||
|
||||
- level 描述的是 **整次调用的 top 质量**,不是每条 evidence 各打一档
|
||||
- 列表里第 2、第 3 条即使偏弱,只要 top1 够高,整次仍可能是 `PRECISE`
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
RET[检索有序候选] --> POST[后处理:保 originalRank · 去重 · 截断]
|
||||
POST --> TOP[取排序后第一条 qualityScore]
|
||||
TOP --> T1{≥ 0.75?}
|
||||
T1 -->|是| PRECISE[PRECISE]
|
||||
T1 -->|否| T2{≥ 0.5?}
|
||||
T2 -->|是| REF[REFERENCE]
|
||||
T2 -->|否| NULL[null / 不报档]
|
||||
```
|
||||
|
||||
### 4.2 qualityScore 从哪来(和检索 mode 绑定)
|
||||
|
||||
统一经 `RetrievalScoreNormalizer.toQualityScore`:
|
||||
|
||||
| `retrieval.search.mode` | top qualityScore 含义 |
|
||||
|-------------------------|------------------------|
|
||||
| **dense** | top1 的 L2 归一化:约 `1 - L2 / maxL2Distance` |
|
||||
| **hybrid** | **优先**同 id 的 `denseDistance`(绝对 L2 质量,供闸门/level);无 dense 邻域时 **rank 回退** |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph dense_mode["mode=dense"]
|
||||
L2[score = L2] --> QD[quality = 1 - L2/max]
|
||||
end
|
||||
|
||||
subgraph hybrid_mode["mode=hybrid"]
|
||||
RRF[RRF 序 = originalRank] --> SORT[列表顺序]
|
||||
DD[denseDistance 可选] --> QH{有 dense?}
|
||||
QH -->|是| QL[quality = L2 归一化]
|
||||
QH -->|否| QR[quality = rank 映射]
|
||||
end
|
||||
|
||||
QD --> LV[relevance_level]
|
||||
QL --> LV
|
||||
QR --> LV
|
||||
```
|
||||
|
||||
读 level 时注意:
|
||||
|
||||
> **dense 下的 PRECISE ≈「向量足够近」**
|
||||
> **hybrid 下的 PRECISE ≈「top1 的绝对/回退 quality 跨过了 0.75」**;排序仍跟 RRF,不是「又变回只信 L2 排序」。
|
||||
|
||||
不要把 hybrid 的 level 读成与 dense **完全同一把尺子**,但也不要再假设「rank1 永远 PRECISE」(在附带 denseDistance 后,远邻 decoy 可以很低分并触发 filter fallback)。
|
||||
### 4.3 关于 HIGHLY_RELEVANT
|
||||
|
||||
枚举和旧文档里仍有三档。历史上大致是:
|
||||
|
||||
```text
|
||||
高分 + L0 hint 支撑 → PRECISE
|
||||
高分但无 hint → HIGHLY_RELEVANT
|
||||
中等分 → REFERENCE
|
||||
```
|
||||
|
||||
质量分统一之后,当前实现是:**≥ 0.75 直接 PRECISE**,不再要求 L0 contains 才能精准。
|
||||
因此运行时 **很少再单独产出 HIGHLY_RELEVANT**;读旧 trace / 旧快照时仍可能见到。
|
||||
|
||||
---
|
||||
|
||||
## 5. Agent 应该怎么读(建议协议)
|
||||
|
||||
### 5.1 推荐读法
|
||||
|
||||
```text
|
||||
1. 先看 evidence_status
|
||||
2. 再读 evidence[] 的 excerpt(唯一可引用正文)
|
||||
3. 用 relevance_level 调节「敢多敢少」与「要不要再查」
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
START[收到 RagToolResult] --> S1{evidence_status}
|
||||
S1 -->|无证据| FAIL[不编造 · 换工具或安全降级]
|
||||
S1 -->|有证据| S2[精读 evidence excerpt]
|
||||
S2 --> S3{relevance_level}
|
||||
S3 -->|PRECISE| C1[结论可较硬 · 少重复同 query 检索]
|
||||
S3 -->|REFERENCE| C2[结论留余地 · 可换问法或叠日志指标]
|
||||
S3 -->|缺省| C3[慎用强结论 · 优先补查]
|
||||
C1 --> CITE[引用必须落在 excerpt]
|
||||
C2 --> CITE
|
||||
C3 --> CITE
|
||||
```
|
||||
|
||||
| 组合 | 建议行为 |
|
||||
|------|----------|
|
||||
| FOUND + PRECISE | 以 excerpt 为主写结论;少重复同 query 检索 |
|
||||
| FOUND + REFERENCE | 可引用,表述保守;缺关键事实则改写 query 或换工具 |
|
||||
| FOUND 但 level 空 | 有块但质量闸门偏低:慎用强结论,优先补查 |
|
||||
| NO_EVIDENCE | 不编造知识库依据;走其他证据工具或安全降级 |
|
||||
|
||||
### 5.2 明确不要这样读
|
||||
|
||||
1. **不要当逐条相关度**
|
||||
没有 `evidence[i].relevance_level`;不能说「第 2 条是 REFERENCE」。
|
||||
|
||||
2. **不要当连续分数**
|
||||
没有 0.83;只有粗档。不要在推理里假装有精确分。
|
||||
|
||||
3. **不要在 hybrid 下当成绝对语义相似度**
|
||||
序数 quality 下 PRECISE 很常见,表示「本轮第一够格」,不等于「全局语义必近」。
|
||||
|
||||
4. **不要代替 excerpt 引用**
|
||||
level 不能当证据正文;Gatekeeper / EvidenceGuard 认的是可核对片段与引用约束。
|
||||
|
||||
5. **不要和 Harness 验真混为一谈**
|
||||
level 是检索侧粗标;工具生命周期、`evidence_status`、守卫校验是另一层。
|
||||
|
||||
6. **不要用它驱动跨 mode 对比**
|
||||
同一 query 切 dense/hybrid 时,比命中集合与排名;别只比「是不是都 PRECISE」。
|
||||
|
||||
---
|
||||
|
||||
## 6. Agent 看不见、但会影响 level 的内部量
|
||||
|
||||
便于排查「为什么突然全是 PRECISE / 总是 null」:
|
||||
|
||||
| 内部量 | 作用 | Agent 是否可见 |
|
||||
|--------|------|----------------|
|
||||
| `qualityScore` / `topSimilarity` | 定 level、低质 retry | 否 |
|
||||
| `originalRank` | 排序权威;hybrid quality 输入 | 否 |
|
||||
| `score` + `scoreLabel` | dense=L2 / hybrid=融合侧 | 否 |
|
||||
| L0 domain/keyword | 现仅 hitReasons 解释,不改序、不抬 level | 否(reasons 也可能被投影裁掉) |
|
||||
| `completenessHint` | 内部完整度文案 | 通常否 |
|
||||
| category filter + unfiltered retry | 低质时可能换一批 evidence 再定 level | 过程 trace 否 |
|
||||
|
||||
投影原则(ACI):模型只要能理解与引用结果;**不给 raw score、阈值、trace。**
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph pipe["检索管道内部"]
|
||||
STORE[Milvus hybrid/dense]
|
||||
NORM[toQualityScore]
|
||||
POST[PostProcessor]
|
||||
STORE --> NORM --> POST
|
||||
end
|
||||
|
||||
POST --> LV[relevance_level]
|
||||
POST --> EB[evidenceBlocks]
|
||||
POST --> TR[traces · 通常不投影]
|
||||
|
||||
EB --> PROJ[RagResultProjector]
|
||||
LV --> PROJ
|
||||
PROJ --> AGENT[Agent Observation]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. 和相邻概念的边界
|
||||
|
||||
```text
|
||||
evidence_status 有没有证据
|
||||
relevance_level 有的话有多贴(粗)
|
||||
evidence[].excerpt 贴在哪一段文字上(细、可引用)
|
||||
truncated 列表是否被预算截断(可能还有更好的没展示)
|
||||
information_gain 等 诊断环路里「这轮工具对任务有没有增益」(另一契约)
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph fields["同一次工具结果里的分工"]
|
||||
ES[evidence_status<br/>有没有]
|
||||
RL[relevance_level<br/>有多贴·粗]
|
||||
EX[excerpt<br/>说什么·细]
|
||||
TC[truncated<br/>是否被截断]
|
||||
end
|
||||
|
||||
ES --> USE[Agent 使用]
|
||||
RL --> USE
|
||||
EX --> USE
|
||||
TC --> USE
|
||||
|
||||
USE --> NOTE[结论锚在 excerpt<br/>level 只调力度]
|
||||
```
|
||||
|
||||
`truncated=true` 时:即使 `PRECISE`,也只说明 **已返回子集里的 top 很强**,不保证库内没有更相关却被截掉的块。
|
||||
|
||||
---
|
||||
|
||||
## 8. 提示词 / 产品文案可用的短说明
|
||||
|
||||
可直接给模型或文档的精简版:
|
||||
|
||||
```text
|
||||
relevance_level 是本次知识库检索的整体相关度粗标:
|
||||
- PRECISE:当前证据整体很贴,优先引用 evidence 写结论,避免无意义重复检索
|
||||
- REFERENCE:仅供参考,结论需留余地,必要时换问法或改用其他工具
|
||||
- 缺省:不要把本次结果当作高置信知识库证实
|
||||
|
||||
务必以 evidence 中的 excerpt 为唯一引用依据;不要编造未出现的文档内容。
|
||||
在 hybrid 检索下,PRECISE 更多表示「本轮排序第一档」,不是精确相似度分数。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. 常见误读示例
|
||||
|
||||
| 误读 | 更正 |
|
||||
|------|------|
|
||||
| 「PRECISE 所以三条 evidence 都精准」 | 只保证 top 质量跨线;其余条只是同批返回 |
|
||||
| 「没有 relevance_level 就是工具失败」 | 更可能是无证据或质量偏低;看 `evidence_status` |
|
||||
| 「hybrid 全是 PRECISE 说明召回完美」 | 可能只是 rank1→quality=1.0 的档位特性 |
|
||||
| 「REFERENCE 的 excerpt 不能引用」 | 可以引用,但结论强度应下调 |
|
||||
| 「level 高就可以跳过 excerpt」 | 不可;引用与验真仍看正文 |
|
||||
|
||||
---
|
||||
|
||||
## 10. 结语
|
||||
|
||||
`relevance_level` 是检索链路送给 Agent 的 **粗粒度驾驶辅助**:
|
||||
|
||||
- 告诉模型这批评据大概硬不硬
|
||||
- **不**替代 excerpt,**不**暴露打分细节,**不**等于逐条标注
|
||||
|
||||
在 dense 模式下,它更接近「向量有多近」;
|
||||
在 hybrid 主路径下,它更接近「本轮融合第一名是否跨过质量门槛」。
|
||||
|
||||
读的时候记住三句即可:
|
||||
|
||||
```text
|
||||
1. 先 status,再 excerpt,最后才看 level
|
||||
2. level 管「敢多敢少」,excerpt 管「说了什么」
|
||||
3. hybrid 的 PRECISE ≠ 绝对语义满分
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 附录:代码与规格锚点
|
||||
|
||||
| 项 | 位置 |
|
||||
|----|------|
|
||||
| Agent 结果契约 | `RagToolResult` / `RagRelevanceLevel` |
|
||||
| 投影 | `RagResultProjector` |
|
||||
| 等级计算 | `KnowledgeEvidencePostProcessor.computeRelevance` |
|
||||
| 质量分 | `RetrievalScoreNormalizer` |
|
||||
| 规格 | `openspec/specs/aci-evidence-tool-contracts`、`rag-retrieval-quality-score` |
|
||||
| 架构 | `mvp/architecture/RAG知识检索架构.md` §9 |
|
||||
@@ -0,0 +1,592 @@
|
||||
# 混合检索上线之后:为什么还要统一 qualityScore,以及上一代后处理错在哪
|
||||
|
||||
**日期**:2026-07-28
|
||||
**范围**:hybrid 检索后的分数语义、后处理排序、相关度闸门、scoreLabel 约定
|
||||
**读者**:已经(或准备)上 dense+BM25+RRF,却发现「召回变了、质量判断还拧着」的工程同学
|
||||
**关联实现**:
|
||||
|
||||
- `lookup_knowledge` 模块化链路
|
||||
- `MilvusHybridKnowledgeStore`(dense / hybrid)
|
||||
- `RetrievalScoreNormalizer` / `KnowledgeEvidencePostProcessor`
|
||||
- OpenSpec / devflow:`rag-quality-score-unify`
|
||||
- 前置讨论:`docs/RAG排序-多路召回与RRF.md`
|
||||
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
|
||||
|
||||
---
|
||||
|
||||
## 1. 引言:融合排好了序,不等于质量链路闭环了
|
||||
|
||||
上一篇文章(《诊断 Agent 场景下的 RAG 排序:从 K=3 规则加分,到多路召回与 RRF》)回答的是:
|
||||
|
||||
> 候选太少时别急着上精排;跨路不要硬加原始分;优先 RRF;L0 只做导航。
|
||||
|
||||
那一轮讨论之后,工程上陆续落地了:
|
||||
|
||||
1. **chunk 级证据身份与去重**(`docId#chunkIndex`,同文档多片段可并存)
|
||||
2. **真 hybrid**:Milvus 服务端 dense ANN + BM25 sparse + `hybridSearch` + `RRFRanker`
|
||||
3. **单一知识后端**(`MilvusClientV2`),去掉 sdk/spring 多路由主路径
|
||||
4. **mode 开关**:`hybrid` 线上主路径,`dense` 同库对照评测
|
||||
|
||||
主缺口从「假 hybrid / 粗去重」变成了另一件事:
|
||||
|
||||
```text
|
||||
库内:RRF 已经决定谁先谁后
|
||||
应用:后处理仍假装每条 score 都是 L2
|
||||
再用 L0 关键词 contains 加分改序
|
||||
```
|
||||
|
||||
于是出现一种很拧的现象:
|
||||
|
||||
- 检索层已经是 **混合检索的世界**
|
||||
- 质量层还活在 **单路 dense + 规则 boost 的世界**
|
||||
|
||||
本文记录的,就是这次对「拧」的拆解、拍板与落地口径:
|
||||
**统一 qualityScore,废止 hybrid 的 L2 伪装,去掉关键词 boost 改序。**
|
||||
|
||||
目标不是再推一套更复杂的模型,而是回答:
|
||||
|
||||
> hybrid 上线之后,排序权威和质量闸门到底听谁的?后处理还该不该拿关键词打分?
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph retrieval["检索层 · 已 hybrid"]
|
||||
Q[query] --> D[dense ANN]
|
||||
Q --> B[BM25 sparse]
|
||||
D --> RRF[hybridSearch + RRF]
|
||||
B --> RRF
|
||||
RRF --> ORD[RRF 序]
|
||||
end
|
||||
|
||||
subgraph post_old["后处理 · 仍 L2 世界"]
|
||||
ORD --> FAKE[伪装 / 回填 L2]
|
||||
FAKE --> BOOST[关键词 boost 改序]
|
||||
BOOST --> GATE[阈值 / level / retry]
|
||||
end
|
||||
|
||||
post_old --> PAIN[排序与闸门拧巴]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 上一代(hybrid 刚落地时)到底长什么样
|
||||
|
||||
### 2.1 检索侧:已经是真混合
|
||||
|
||||
`mode=hybrid` 时大致是:
|
||||
|
||||
```text
|
||||
query
|
||||
├─ dense ANN(query embedding) → vector / L2
|
||||
└─ BM25 sparse(EmbeddedText) → sparse_vector / BM25
|
||||
│
|
||||
▼
|
||||
Milvus hybridSearch + RRFRanker(k)
|
||||
│
|
||||
▼
|
||||
融合后的 hit 列表(RRF 序)
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Q[query] --> EMB[embedding]
|
||||
Q --> TXT[raw text]
|
||||
EMB --> DA[dense ANN<br/>vector / L2]
|
||||
TXT --> BA[BM25 ANN<br/>sparse]
|
||||
DA --> HS[Milvus hybridSearch]
|
||||
BA --> HS
|
||||
HS --> RR[RRFRanker]
|
||||
RR --> HITS[有序 hits]
|
||||
```
|
||||
|
||||
这比应用层 sparse-lite / 伪 hybrid 前进了一大步:词面与语义在**库内**融合,chunk 身份也不会在后处理被文档级折叠吞掉。
|
||||
|
||||
### 2.2 分数侧:仍在「骗」后处理
|
||||
|
||||
后处理历史契约默认:
|
||||
|
||||
```text
|
||||
score ≈ L2 距离(越小越好)
|
||||
baseScore = 1 - clamp(L2) / maxL2Distance # 越大越好
|
||||
再 + domain/entity/keyword boost
|
||||
按 finalScore 重排
|
||||
用 baseScore 定 relevance_level / 是否低质 retry
|
||||
```
|
||||
|
||||
为了迁就这套契约,hybrid 路径做了补丁:
|
||||
|
||||
```text
|
||||
hybrid 融合结果
|
||||
-> 再跑一路 dense
|
||||
-> 按 id 把 L2 回填到 score,label 改成 l2_distance
|
||||
-> 仅 BM25 命中、dense 没命中:
|
||||
score = maxL2Distance
|
||||
label = bm25_only_no_dense
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
H[hybrid RRF 结果] --> P[并行 dense 探测]
|
||||
P --> M{同 id 有 L2?}
|
||||
M -->|是| O1[score=L2 · label=l2_distance]
|
||||
M -->|否| O2[score=maxL2 · label=bm25_only]
|
||||
O1 --> N[normalizeL2 + boost 重排]
|
||||
O2 --> N
|
||||
N --> BAD[BM25-only 好证据被当成最差]
|
||||
```
|
||||
|
||||
意图是好的:让 `normalizeL2` 和 0.75/0.5 阈值「还能用」。
|
||||
副作用也很清楚:
|
||||
|
||||
| 现象 | 后果 |
|
||||
|------|------|
|
||||
| RRF 决定顺序,L2 决定「好不好」 | 两套真理,互相打架 |
|
||||
| BM25-only 好证据被标成最远 L2 | quality≈0,像低质,甚至触发 unfiltered retry |
|
||||
| `bm25_only_no_dense` 像第三种 label | 概念膨胀:mode 其实只有 dense/hybrid |
|
||||
| 多打一路 dense 只为回填 | 延迟与复杂度,换来的是语义自洽的假象 |
|
||||
|
||||
一句话:
|
||||
|
||||
> **不是拿 RRF 分错误地套了 L2 公式,而是排序信 RRF,打分/闸门仍假装大家都是 L2。**
|
||||
|
||||
### 2.3 后处理侧:关键词 boost 改主序
|
||||
|
||||
典型逻辑:
|
||||
|
||||
```text
|
||||
baseScore = normalizeL2(score)
|
||||
finalScore = baseScore
|
||||
+ domain_match (+0.15)
|
||||
+ entity_match (+0.20)
|
||||
+ keyword_match (+0.10)
|
||||
+ source_type (+0.05)
|
||||
按 finalScore 降序
|
||||
```
|
||||
|
||||
在 **还没有库内 BM25** 时,这套东西多少能补一点词面。
|
||||
在 **已经 hybrid** 之后,问题变成:
|
||||
|
||||
1. **双重计分**
|
||||
BM25 已经在 RRF 里投过票;后处理再用 contains 加分,等于词面再抬一次。
|
||||
|
||||
2. **contains 比 BM25 更糙**
|
||||
无 IDF、无文档长度、短词子串误命中——正好制造「词频/词面高、相关度低却排前面」。
|
||||
|
||||
3. **冲掉 RRF 序**
|
||||
花了 hybrid 买到的融合序,被 L0 词表二次改写。
|
||||
|
||||
4. **PRECISE 还绑 hint**
|
||||
高质量还要 `hasHintSupport`(同样是 contains),把导航层信号抬成等级门槛。
|
||||
|
||||
结合上一篇文章的判断——**L0 / 关键词适合做提示,不适合当最终裁判**——hybrid 上线后,后处理 boost 改序已经从「可接受的轻启发式」滑向「明确的设计债」。
|
||||
|
||||
---
|
||||
|
||||
## 3. 问题清单:chunk 去重 + hybrid 之后,还剩什么
|
||||
|
||||
可以分成四层(本次主要收口前两层):
|
||||
|
||||
### 3.1 正确性 / 契约(本次主战场)
|
||||
|
||||
1. hybrid **没有**独立的融合分归一化,只有 L2 兼容补丁
|
||||
2. 后处理关键词打分不合理,会抬升词面热、语义冷的片段
|
||||
3. `scoreLabel` 语义混乱:`l2_distance` / `rrf_fused` / `bm25_only_*` 混用
|
||||
4. `mode=dense` 与 hybrid 内部 dense 子路概念易混(mode 是整次查询算法,不是「第三套库」)
|
||||
|
||||
### 3.2 质量上限(未在本 change 做完)
|
||||
|
||||
- 无固定 RAG 评测报表驱动阈值标定
|
||||
- 无邻块扩展、真 query rewrite、cross-encoder 精排
|
||||
- 中文 analyzer / 分词策略未产品化钉死
|
||||
|
||||
### 3.3 工程债(部分清理、部分保留)
|
||||
|
||||
- 应用层 `LexicalRanker` / 自研 `RrfFusion` 可能仍像「还有 app-layer hybrid」
|
||||
- Spring AI starter 仍可作 sidecar,但不是知识主路径(starter 至 2.0.0 仍无 BM25 hybrid)
|
||||
- 写入先删后插,非强 upsert;Milvus / MySQL / L0 三方一致性靠流程
|
||||
|
||||
### 3.4 运维边界
|
||||
|
||||
- hybrid 依赖 BM25 Function + sparse index
|
||||
- 全量 rebuild 受 embedding 与写入延迟约束
|
||||
- `totalVectors` 一类统计可能仍不可信
|
||||
|
||||
本次 change(`rag-quality-score-unify`)**有意只收口 3.1**:
|
||||
让 hybrid 的排序权威与质量闸门重新对齐,而不是同时上精排模型。
|
||||
|
||||
---
|
||||
|
||||
## 4. 关键澄清:L2 是什么,它是不是「后处理」本身
|
||||
|
||||
讨论中容易把「L2」和「后处理打分流程」混成一个词。需要拆开:
|
||||
|
||||
### 4.1 L2 是度量
|
||||
|
||||
**L2 = 欧氏距离**,dense ANN 常用 metric:
|
||||
|
||||
- 越小越相似
|
||||
- 单位向量场景下可用 `maxL2Distance≈2` 做上界
|
||||
- 归一化相似度:`1 - clamp(L2) / maxL2`
|
||||
|
||||
### 4.2 后处理是流水线
|
||||
|
||||
后处理消费的是「约定好的 score」,历史上**假定**它是 L2,于是:
|
||||
|
||||
```text
|
||||
score(L2) → normalizeL2 → baseScore → (+boost) → finalScore → 排序/等级
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph metric["度量层"]
|
||||
L2[L2 距离]
|
||||
end
|
||||
|
||||
subgraph pipe["后处理流水线"]
|
||||
N[normalize]
|
||||
B[可选 boost]
|
||||
S[排序 / 截断]
|
||||
G[等级 / 闸门]
|
||||
N --> B --> S --> G
|
||||
end
|
||||
|
||||
L2 -.->|历史上假定输入是 L2| N
|
||||
RRF[RRF 融合分] -.->|量纲不同 · 不能直接套| N
|
||||
```
|
||||
|
||||
所以:
|
||||
|
||||
- L2 ≠ 后处理
|
||||
- L2 = dense 路径的自然距离
|
||||
- 后处理 = 把某种 score 变成 quality / 等级 / 截断结果的流程
|
||||
|
||||
hybrid 的问题是:**流程还在,输入契约已经不再总是 L2。**
|
||||
|
||||
### 4.3 category 降级也不是 L2 存在的唯一理由
|
||||
|
||||
filtered → unfiltered retry 用的是:
|
||||
|
||||
```text
|
||||
isLowQuality = 无可用证据 或 topSimilarity < referenceThreshold
|
||||
```
|
||||
|
||||
`topSimilarity` 来自归一化后的质量分。
|
||||
unfiltered 只是**再检一次**,尺子本来就该是统一 quality,而不是「专为降级准备的 L2」。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
F[带 category 的检索] --> Q{isLowQuality?}
|
||||
Q -->|是| U[unfiltered retry]
|
||||
Q -->|否| K[采用本次结果]
|
||||
U --> M[合并/替换为 retry 结果]
|
||||
K --> OUT[后处理出口]
|
||||
M --> OUT
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. 设计拍板:统一成什么
|
||||
|
||||
### 5.1 两层概念,不要混
|
||||
|
||||
| 层 | 只有什么 | 不是什么 |
|
||||
|----|----------|----------|
|
||||
| **检索 mode** | `dense` \| `hybrid` | 不是三套库 |
|
||||
| **一级 scoreLabel** | `dense` \| `hybrid` | 不是 `bm25_only` 第三种模式 |
|
||||
|
||||
- `mode`:整次查询怎么跑(配置 `retrieval.search.mode`)
|
||||
- `scoreLabel`:这条 hit 的 `score` 怎么解释
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
CFG[retrieval.search.mode] --> M1[dense 整次只跑 ANN]
|
||||
CFG --> M2[hybrid 整次 dense+BM25+RRF]
|
||||
|
||||
M1 --> L1[scoreLabel=dense]
|
||||
M2 --> L2[scoreLabel=hybrid]
|
||||
|
||||
L2 -.->|不是| L3[bm25_only 第三种 mode]
|
||||
```
|
||||
|
||||
`bm25_only_no_dense` **不是第三种检索**,只是旧链路里「这条 hybrid 命中没有 dense L2 可回填」的补丁标签。统一后应降级为历史别名(canonicalize → `hybrid`),不再一级发射。
|
||||
### 5.2 检索回来带什么
|
||||
|
||||
建议最小契约:
|
||||
|
||||
```text
|
||||
rank (originalRank) // 1 最好;hybrid = RRF 序;dense = ANN 序
|
||||
score // 引擎主分;量纲由 label 解释
|
||||
scoreLabel // dense | hybrid
|
||||
rawScore? // 可选调试
|
||||
denseDistance? // hybrid 可选:同 id 的 L2,仅供闸门
|
||||
```
|
||||
|
||||
| label | score 含义 |
|
||||
|-------|------------|
|
||||
| `dense` | L2 距离(越小越好) |
|
||||
| `hybrid` | 引擎融合分可放 raw/score;**排序不看其量纲** |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph emit["Store 发射"]
|
||||
LAB[scoreLabel<br/>dense | hybrid]
|
||||
SCR[score / rawScore]
|
||||
RNK[列表序 → originalRank]
|
||||
DD[denseDistance? 仅 hybrid]
|
||||
end
|
||||
|
||||
LAB --> NORM
|
||||
SCR --> NORM
|
||||
RNK --> SORT
|
||||
DD --> NORM
|
||||
NORM[toQualityScore] --> QS[qualityScore]
|
||||
SORT[按 rank 排序] --> LIST[evidence 顺序]
|
||||
QS --> GATE[level / isLowQuality]
|
||||
```
|
||||
|
||||
### 5.3 唯一归一化点
|
||||
|
||||
```text
|
||||
qualityScore = toQualityScore(label, score, rank, batchSize, maxL2, denseDistance?)
|
||||
// 输出统一:[0,1],越大越好
|
||||
```
|
||||
|
||||
分支只允许出现在这里:
|
||||
|
||||
```text
|
||||
dense → 1 - clamp(L2)/maxL2
|
||||
hybrid → 优先 denseDistance 的 L2 归一化(绝对质量 / 闸门)
|
||||
无 dense 时 rank 线性回退
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
IN[label + score + rank + denseDistance?] --> C{canonicalize label}
|
||||
C -->|dense| L2[l2ToQuality score]
|
||||
C -->|hybrid| H{denseDistance?}
|
||||
H -->|有| L2H[l2ToQuality denseDistance]
|
||||
H -->|无| RK[rankToQuality]
|
||||
L2 --> OUT[qualityScore 0..1]
|
||||
L2H --> OUT
|
||||
RK --> OUT
|
||||
```
|
||||
|
||||
**演进说明:** 切片 1 曾用纯 rank 做 hybrid quality;eval 发现 rank1 恒高会杀死 L0 filter fallback。
|
||||
现行约定:**排序仍纯 RRF;闸门可用 denseDistance 绝对质量**,且 **不得** 再把主分/label 伪装成 L2。
|
||||
|
||||
### 5.4 后处理:统一流程,不要按 label 再分叉业务
|
||||
|
||||
```text
|
||||
candidates
|
||||
→ 每条 toQualityScore(...) ← 唯一认 label 的地方
|
||||
→ qualityScore + originalRank
|
||||
→ 统一:按 rank 排序 / 去重 / 每文档 chunk 上限 / return-n
|
||||
→ 统一:relevance_level、isLowQuality(只看 qualityScore)
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
CAND[candidates] --> QS[toQualityScore 每条]
|
||||
QS --> SORT[sort by originalRank ASC]
|
||||
SORT --> DEDUP[evidenceKey 去重]
|
||||
DEDUP --> CAP[max-chunks / return-n]
|
||||
CAP --> REL[relevance_level]
|
||||
CAP --> LOW[isLowQuality → filter retry]
|
||||
CAP --> OUT[EvidenceBlocks]
|
||||
```
|
||||
|
||||
可以记成:
|
||||
|
||||
> **Label 只活在进后处理之前的适配器里;后处理是 label-agnostic 的。**
|
||||
> **排序听 rank;闸门听 qualityScore(hybrid 可含 denseDistance)。**
|
||||
### 5.5 后处理还改不改?——要改,而且和归一化同一刀
|
||||
|
||||
后处理合理职责是 **裁剪与装配**,不是第二套检索:
|
||||
|
||||
| 保留 | 去掉或降级 |
|
||||
|------|------------|
|
||||
| evidenceKey 去重 | domain/entity/keyword **加分改序** |
|
||||
| max-chunks-per-document | contains 当相关度代理 |
|
||||
| return-n | PRECISE 强制 hint support |
|
||||
| excerpt 截断、EvidenceBlock | |
|
||||
| 统一 qualityScore 闸门 | |
|
||||
|
||||
L0 仍可:
|
||||
|
||||
- 导航:category filter(失败 unfiltered retry)
|
||||
- 解释:`hitReasons` 记 `l0_keyword_overlap` 等(**零分值**)
|
||||
|
||||
词面该不该高:交给 **BM25 子路 + RRF**。
|
||||
语义该不该近:交给 **dense 子路**(融合时已参与;dense-only mode 对照时单独看)。
|
||||
|
||||
### 5.6 mode=dense 还要不要
|
||||
|
||||
要,但定位清楚:
|
||||
|
||||
| 模式 | 定位 |
|
||||
|------|------|
|
||||
| **hybrid** | 线上主路径 / 默认 |
|
||||
| **dense** | 同库对照、评测、排障——看「去掉 BM25+RRF 后差在哪」 |
|
||||
|
||||
注意:
|
||||
|
||||
- hybrid **入库**数据完全适用于 dense 查询(每条都写了 `vector`)
|
||||
- hybrid **内部**仍有 dense 子路——那是融合的一部分,≠ `mode=dense`
|
||||
- 对照时固定 `retrieve-k` / `return-n` / filter / query 集,只切 mode
|
||||
- 优先比命中集合与排名;`relevance_level` 在 hybrid 下是序数 quality,慎作跨 mode 绝对值对比
|
||||
|
||||
---
|
||||
|
||||
## 6. 目标数据流(落地后)
|
||||
|
||||
```text
|
||||
VectorSearchService (mode=dense|hybrid)
|
||||
-> hits{ originalRank, score, scoreLabel=dense|hybrid, rawScore?, denseDistance? }
|
||||
-> KnowledgeDocumentRetriever / SearchPort
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
qualityScore = RetrievalScoreNormalizer.toQualityScore(...)
|
||||
sort by originalRank ASC
|
||||
evidenceKey dedup / max-chunks / return-n
|
||||
relevance_level & topSimilarity from qualityScore
|
||||
L0 overlap → hitReasons only
|
||||
-> ContextPack / Assembler / Projector
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[query] --> VSS[VectorSearchService]
|
||||
VSS -->|mode=dense| SD[searchDense]
|
||||
VSS -->|mode=hybrid| SH[searchHybrid + 可选 denseDistance]
|
||||
SD --> PORT[KnowledgeSearchPort / Adapter]
|
||||
SH --> PORT
|
||||
PORT --> POST[KnowledgeEvidencePostProcessor]
|
||||
POST --> PACK[ContextPacker]
|
||||
POST --> ASM[LookupResultAssembler]
|
||||
ASM --> PROJ[RagResultProjector]
|
||||
PROJ --> AGENT[Agent 可见契约]
|
||||
```
|
||||
|
||||
与上一代对比:
|
||||
|
||||
| 环节 | 上一代 | 现在 |
|
||||
|------|--------|------|
|
||||
| hybrid score | 常被 L2 覆盖 | 保持融合侧;label=`hybrid` |
|
||||
| BM25-only | maxL2 + `bm25_only_*` | 普通 hybrid hit,quality 看 rank |
|
||||
| 归一化 | 一律当 L2 | 按 label 唯一转换 |
|
||||
| 排序 | finalScore(含 boost) | originalRank |
|
||||
| L0 关键词 | +分改序 | 仅解释 |
|
||||
| 质量闸门 | baseScore(L2 兼容) | qualityScore |
|
||||
|
||||
---
|
||||
|
||||
## 7. 行为变化:必须说清楚的协议调整
|
||||
|
||||
这是**有意的行为变化**(对内检索质量语义;Agent ACI 字段名可不变):
|
||||
|
||||
1. hybrid 下证据顺序更贴近 **RRF**,不再被 contains 抬到前面
|
||||
2. 「词面很准、dense 略远」的命中,不再被默认打成低质占位
|
||||
3. `relevance_level` / category unfiltered retry 的触发分布可能变化
|
||||
4. hybrid 的 quality 是**本轮序数分**,跨 query 绝对值不可比;阈值可能需后续标定
|
||||
5. dense 对照模式:质量仍走 L2 归一化,行为更接近旧 dense 主路径
|
||||
|
||||
未改:
|
||||
|
||||
- Agent 可见字段结构(evidence 列表、relevance 枚举名等)
|
||||
- Milvus hybrid schema / 不必为本次 rebuild
|
||||
- `mode=dense` 开关本身
|
||||
|
||||
---
|
||||
|
||||
## 8. 和上一篇文章的衔接:阶段进度
|
||||
|
||||
对照 `RAG排序-多路召回与RRF.md` 的推进顺序:
|
||||
|
||||
| 阶段 | 内容 | 状态(截至 2026-07-28) |
|
||||
|------|------|-------------------------|
|
||||
| Phase 0 | chunk 去重、retrieve-k/return-n、身份 | **已落地** |
|
||||
| Phase 1~2 | 多路 + RRF;真 BM25 hybrid | **已落地**(库内 hybrid,非 app-layer 伪融合) |
|
||||
| 分数职责分离 | 排序 vs 可用性/质量闸门 | **本次收口**(qualityScore 统一) |
|
||||
| 去掉 L0 当裁判 | 关键词不改主序 | **本次收口** |
|
||||
| Phase 3 | 可插拔模型 Rerank | **未做**(候选池与评测闭环仍优先) |
|
||||
| 邻块 / query rewrite | 上下文与问句改写 | **未做** |
|
||||
|
||||
因此,本次文章不是推翻上一篇,而是补上上一篇写到「融合之后」却还没写完的半截:
|
||||
|
||||
> 融合解决「谁进来、谁先排」;
|
||||
> 归一化与后处理决定「算不算够好、会不会被规则再次打乱」。
|
||||
|
||||
---
|
||||
|
||||
## 9. 实现锚点(便于对照代码)
|
||||
|
||||
| 组件 | 职责 |
|
||||
|------|------|
|
||||
| `RetrievalScoreLabels` | `dense` / `hybrid` + 旧别名 canonicalize |
|
||||
| `RetrievalScoreNormalizer` | 唯一 `toQualityScore` |
|
||||
| `MilvusHybridKnowledgeStore` | 发射 label;hybrid 不 L2 覆盖 |
|
||||
| `VectorSearchService` | mode 路由;SearchResult 契约注释 |
|
||||
| `KnowledgeEvidencePostProcessor` | rank 保序、去 boost 改序、quality 闸门 |
|
||||
| 架构 §6.0 | mode 用途:hybrid 主路径 / dense 对照 |
|
||||
|
||||
验证(单测,非 live E2E):
|
||||
|
||||
- normalizer:L2 边界、rank 单调、别名
|
||||
- post-process:rank 不被 keyword 打乱;caps/return-n
|
||||
- lookup tool:保序;context pack 元数据仍在
|
||||
|
||||
已知未验证:真实 Milvus 联调对照、阈值标定。
|
||||
|
||||
---
|
||||
|
||||
## 10. 实践清单:以后别再踩的坑
|
||||
|
||||
1. **不要**为了复用旧 `normalizeL2`,把 hybrid 结果伪装成 L2。
|
||||
2. **不要**在已经 BM25 hybrid 之后,再用 L0 contains 大额加分改主序。
|
||||
3. **不要**把 `bm25_only` 当成第三种检索模式。
|
||||
4. **不要**把 hybrid 内部的 dense 子路,和 `mode=dense` 整次查询混为一谈。
|
||||
5. **要**让 label 差异停在适配器;后处理只认 qualityScore + rank。
|
||||
6. **要**用 dense mode 做召回对照,而不是第二套长期线上策略。
|
||||
7. **要**接受:hybrid 序数 quality 与绝对阈值之间,需要观测后再调,而不是再发明一层伪装。
|
||||
8. **下一步再考虑**精排模型——在契约掰直、有固定 query 回归集之后。
|
||||
|
||||
---
|
||||
|
||||
## 11. 结语
|
||||
|
||||
混合检索落地,解决的是「漏」和「跨路硬加分」里很大一块。
|
||||
但若后处理仍活在 L2 + 关键词 boost 的旧世界,hybrid 买到的 RRF 序和质量信号会被悄悄改写,甚至惩罚「只在 BM25 路很强」的好证据。
|
||||
|
||||
这次收口的核心就三句:
|
||||
|
||||
```text
|
||||
1. 一级 label 只有 dense / hybrid
|
||||
2. 唯一 toQualityScore;后处理统一、保 rank
|
||||
3. L0 关键词可以解释,不可以再当排序裁判
|
||||
```
|
||||
|
||||
它不是 RAG 的终点,而是 hybrid 从「能跑」变成「分数语义自洽」的必要一步。
|
||||
在此之后,评测闭环、阈值标定、邻块与精排,才有干净的基线可谈。
|
||||
|
||||
---
|
||||
|
||||
## 附录 A:术语
|
||||
|
||||
| 术语 | 含义 |
|
||||
|------|------|
|
||||
| L2 | 欧氏距离;dense ANN 常用;越小越相似 |
|
||||
| RRF | Reciprocal Rank Fusion;用名次融合多路,不融合原始分 |
|
||||
| scoreLabel | 一级分数语义:`dense` \| `hybrid` |
|
||||
| qualityScore | 归一化后的 0~1 质量分(越大越好),供等级与低质闸门 |
|
||||
| originalRank | 检索返回名次;后处理排序权威 |
|
||||
| mode=dense | 整次只跑 dense ANN(对照) |
|
||||
| mode=hybrid | dense+BM25+RRF(主路径) |
|
||||
| L0 | query hint / 可选 category filter;不作事实证据、不改主序 |
|
||||
|
||||
## 附录 B:相关材料
|
||||
|
||||
| 材料 | 路径 |
|
||||
|------|------|
|
||||
| 多路与 RRF 讨论 | `docs/RAG排序-多路召回与RRF.md` |
|
||||
| 当前架构 | `mvp/architecture/RAG知识检索架构.md` |
|
||||
| OpenSpec 归档 | `openspec/changes/archive/2026-07-28-rag-quality-score-unify/` |
|
||||
| 主规格 | `openspec/specs/rag-retrieval-quality-score/spec.md` |
|
||||
| devflow | `devflow/projects/2026-07-28-rag-quality-score-unify/` |
|
||||
@@ -0,0 +1,163 @@
|
||||
# RAG 审计补丁 E2E:`step_id` + query(含业务 FALLBACK 样本)
|
||||
|
||||
**日期**:2026-07-28
|
||||
**状态**:审计字段 live 验收记录
|
||||
**关联主文档**:[一次诊断到底发生了什么(SUCCESS 全流程)](一次诊断全流程-E2E导读.md)
|
||||
|
||||
> 主文档只保留 **SUCCESS 完整诊断** 与 **现行审计能力说明**。
|
||||
> 本页单独记录:改造后的一次 live 验收——**审计字段 PASS**,业务因 Milvus 空结果走了 **FALLBACK**。
|
||||
|
||||
---
|
||||
|
||||
## 1. 样本身份
|
||||
|
||||
| 项 | 值 |
|
||||
|----|----|
|
||||
| `session_id` | `e2e-audit-20260728171858` |
|
||||
| `run_id` | `0be605f6-e036-40d3-b364-11b73672241f` |
|
||||
| 入口 | `POST /api/chat`(与 SUCCESS 样例同构的 RAG-only 约束题) |
|
||||
| 冷启动 | 含 `AgentStepAuditTracker` 等补丁后的 `mvn spring-boot:run` |
|
||||
| 业务结局 | `release_outcome=FALLBACK`,`content_type=SAFE_FALLBACK` |
|
||||
| 工具 | 1× `lookup_knowledge` |
|
||||
|
||||
本地产物(若仍在):`target/e2e-audit-sse.txt`、`target/e2e-audit-session.txt`、`target/e2e-audit-run.txt`。
|
||||
|
||||
---
|
||||
|
||||
## 2. 验收目标 vs 非目标
|
||||
|
||||
| 目标 | 是否本页重点 |
|
||||
|------|----------------|
|
||||
| `tool_invocation.step_id` = 发出 tool_call 的 `agent_step.id` | **是** |
|
||||
| `input_params` 含安全 `query` 预览 | **是** |
|
||||
| Trace / timeline 带回 `step_id` | **是** |
|
||||
| 业务必须 SUCCESS | **否**(本 run 为 FALLBACK,归因环境) |
|
||||
|
||||
---
|
||||
|
||||
## 3. 审计结果:PASS
|
||||
|
||||
### 3.1 `agent_step`
|
||||
|
||||
| id | step_index | has_tool_call | 说明 |
|
||||
|----|------------|---------------|------|
|
||||
| **962** | 0 | 1 | 发出 `lookup_knowledge` |
|
||||
| 963 | 1 | 0 | 无证据后的收尾轮 |
|
||||
|
||||
### 3.2 `tool_invocation`(id=860)
|
||||
|
||||
| 字段 | 值 |
|
||||
|------|-----|
|
||||
| `step_id` | **962**(= step0) |
|
||||
| `tool_name` | `lookup_knowledge` |
|
||||
| `success` | 1(工具跑完;无证据也算执行成功) |
|
||||
| `search_mode` | hybrid |
|
||||
| `evidence_status` | `NO_EVIDENCE` |
|
||||
| `candidate_count` | 0 |
|
||||
| `relevance_level` | null |
|
||||
|
||||
**`input_params` 实值:**
|
||||
|
||||
```json
|
||||
{
|
||||
"query": "MySQL connection pool exhausted HikariCP diagnosis",
|
||||
"step_id": 962,
|
||||
"tool_call_id": "call_00_LireJzbiFEsbZ9ZVvgYp8330",
|
||||
"request_bytes": 62
|
||||
}
|
||||
```
|
||||
|
||||
### 3.3 Trace API
|
||||
|
||||
- `toolInvocations[0].stepId = 962`
|
||||
- `inputParams.query` 有值
|
||||
- timeline `TOOL_INVOCATION.details.step_id = 962`
|
||||
|
||||
### 3.4 对照表
|
||||
|
||||
| 检查项 | 预期 | 实际 | 判定 |
|
||||
|--------|------|------|------|
|
||||
| `tool_invocation.step_id` | = agent_step.id | 962 | **PASS** |
|
||||
| `input_params.query` | 有预览 | 有 | **PASS** |
|
||||
| `input_params.step_id` | 与列一致 | 962 | **PASS** |
|
||||
| Trace `stepId` | 非空 | 962 | **PASS** |
|
||||
| timeline `step_id` | 非空 | 962 | **PASS** |
|
||||
| `search_mode` | hybrid | hybrid | **PASS** |
|
||||
| 业务 outcome | (非本页 KPI) | FALLBACK | 见 §4 |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
S0[agent_step 962<br/>r1 tool_call] --> TI[tool_invocation 860<br/>step_id=962]
|
||||
TI --> IP[input_params.query]
|
||||
TI --> TR[Trace / timeline]
|
||||
TI --> RAG[hybrid candidate_count=0]
|
||||
RAG --> FB[FALLBACK<br/>NO_EVIDENCE]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 业务 FALLBACK 原因(与审计无关)
|
||||
|
||||
| 现象 | 说明 |
|
||||
|------|------|
|
||||
| 日志 | `found=false`,`evidenceBlocks=0` |
|
||||
| audit | `attempts[].usable=false`,`candidate_count=0` |
|
||||
| warm 检索 | 同期 `GET /api/search/similar` 亦失败 |
|
||||
| 日志噪音 | Milvus channel 未正确 shutdown 等提示(环境/客户端生命周期) |
|
||||
|
||||
**结论**:hybrid **路径进了**,但当次 **0 候选** → 无证据可写报告 → `SAFE_FALLBACK` / `INSUFFICIENT_EVIDENCE`。
|
||||
这不否定 `step_id` / `query` 落库;完整 SUCCESS 业务故事见主文档。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[同一 RAG-only 题] --> LK[lookup_knowledge]
|
||||
LK --> AUD[审计: step_id + query PASS]
|
||||
LK --> HIT{有候选?}
|
||||
HIT -->|SUCCESS 主文档 run| OK[PRECISE → DIAGNOSIS_REPORT]
|
||||
HIT -->|本页 run| NO[0 候选 → FALLBACK]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. 实现索引(补丁代码)
|
||||
|
||||
| 组件 | 职责 |
|
||||
|------|------|
|
||||
| `AgentStepAuditTracker` | run 级 bind/current/clear `agent_step.id` |
|
||||
| `HarnessAgentAuditHook` | `beforeModel` 落 step 后 `bind` |
|
||||
| `ToolBoundary.auditSafely` | 读 tracker,带上 `stepId` + `requestJson` |
|
||||
| `JpaToolInvocationAuditSink` | 写 `step_id` 列;安全展开 `input_params` |
|
||||
| `TraceAuditEvents.toolInvocation` | timeline details 的 `step_id` |
|
||||
| `JpaChatRunStore.finish` | `clear(runId)` |
|
||||
|
||||
**`input_params` 安全规则摘要**:
|
||||
|
||||
- 始终:`tool_call_id`、`request_bytes`;有则:`step_id`
|
||||
- 顶层 string/number/boolean/纯字符串数组;文本 ≤160 字符
|
||||
- 键名含 password/token/secret/apikey 等 → 不落库
|
||||
|
||||
更完整的字段词典与 SUCCESS 阶段拆解见主文档 §3.4 / §3.4.1。
|
||||
|
||||
---
|
||||
|
||||
## 6. 复盘 SQL
|
||||
|
||||
```bash
|
||||
python scripts/query_mysql.py "SELECT id, step_id, tool_name, relevance_level, CAST(input_params AS CHAR) AS inputp, LEFT(CAST(retrieval_details AS CHAR), 500) AS details FROM tool_invocation WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f'"
|
||||
|
||||
python scripts/query_mysql.py "SELECT id, step_index, agent_name, has_tool_call, token_count FROM agent_step WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f' ORDER BY step_index"
|
||||
|
||||
python scripts/query_mysql.py "SELECT run_id, status, release_outcome, tool_call_count, total_token_count FROM diagnosis_run WHERE run_id = '0be605f6-e036-40d3-b364-11b73672241f'"
|
||||
```
|
||||
|
||||
```http
|
||||
GET /api/diagnosis/e2e-audit-20260728171858/trace?runId=0be605f6-e036-40d3-b364-11b73672241f
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. 修订记录
|
||||
|
||||
| 日期 | 说明 |
|
||||
|------|------|
|
||||
| 2026-07-28 | 从主 walkthrough 拆出:专门记录 audit 验收 run(FALLBACK 业务 + step_id/query PASS) |
|
||||
@@ -39,6 +39,18 @@
|
||||
|
||||
> 现在的 K、现有的 L0/L1、现有的规则加分,下一步到底该扩召回、该融合,还是该上精排?
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
P[排序/质量问题] --> A{候选池 K 多大?}
|
||||
A -->|K 很小 3~5| B[先扩召回 / 修去重 / 轻规则]
|
||||
A -->|K 中等 15~30| C[多路 + RRF + 可选轻精排]
|
||||
A -->|K 很大 50+| D[强 rerank 才划算]
|
||||
|
||||
P --> E{跨路分数?}
|
||||
E -->|原始分硬加| F[尺度不同 · 易玄学]
|
||||
E -->|RRF 名次投票| G[推荐]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 现状解剖:有“重排”,不等于有“强 Rerank”
|
||||
@@ -57,6 +69,17 @@ Agent query
|
||||
-> RagResultProjector # 投影成 Agent 可见契约
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
AQ[Agent query] --> L0[KnowledgeQueryTransformer<br/>L0 hint / category filter]
|
||||
L0 --> L1[KnowledgeDocumentRetriever<br/>L1 向量 topK]
|
||||
L1 --> POST[KnowledgeEvidencePostProcessor<br/>归一化 + 规则 boost + 去重]
|
||||
POST --> PACK[KnowledgeContextPacker]
|
||||
PACK --> ASM[LookupResultAssembler]
|
||||
ASM --> PROJ[RagResultProjector]
|
||||
PROJ --> AGENT[Agent 可见结果]
|
||||
```
|
||||
|
||||
其中“重排”发生在后处理阶段,名字也常叫 rerank,但实现通常是:
|
||||
|
||||
```text
|
||||
@@ -69,6 +92,20 @@ finalScore = baseScore
|
||||
再按 finalScore 降序
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
BASE[baseScore<br/>向量相似度] --> SUM[finalScore]
|
||||
D[+ domain]
|
||||
E[+ entity]
|
||||
K[+ keyword]
|
||||
S[+ source_type]
|
||||
D --> SUM
|
||||
E --> SUM
|
||||
K --> SUM
|
||||
S --> SUM
|
||||
SUM --> SORT[按 finalScore 降序]
|
||||
```
|
||||
|
||||
同时会留下 `rerankTrace`(base/final score、boost reasons),便于内部审计。
|
||||
|
||||
### 2.2 这套做法解决了什么
|
||||
@@ -110,6 +147,18 @@ finalScore = baseScore
|
||||
|
||||
排序策略必须和 K 匹配。可以先用下面这张表做决策:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
K{retrieve-k / 候选规模}
|
||||
K -->|3~5| S1[修去重 · 轻规则<br/>双路径 dense 融合]
|
||||
K -->|15~30| S2[多路召回 + RRF<br/>可选轻精排]
|
||||
K -->|50+| S3[强 Cross-Encoder / 托管 Rerank]
|
||||
|
||||
S1 -.->|先别上| X1[商业 Rerank]
|
||||
S2 -.->|收益有限| X2[只继续调 keyword boost]
|
||||
S3 -.->|避免| X3[无评测堆模型]
|
||||
```
|
||||
|
||||
| 召回规模 K | 更适合做什么 | 不太值得先做什么 |
|
||||
|---|---|---|
|
||||
| 3 ~ 5 | 修去重、轻规则、双路径 dense 融合 | Cross-Encoder / 商业 Rerank |
|
||||
@@ -157,6 +206,15 @@ topK = 3,召回、排序、返回都是 3
|
||||
若结果差,再 unfiltered 重试
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[query + L0 category?] --> F[filtered dense]
|
||||
F --> BAD{结果差/空?}
|
||||
BAD -->|是| U[unfiltered retry]
|
||||
BAD -->|否| USE[采用 filtered]
|
||||
U --> REP[常见:整锅替换 filtered]
|
||||
```
|
||||
|
||||
这能工作,但常见实现是 **串行整锅替换**:
|
||||
|
||||
- retry 成功后,直接丢掉第一次 filtered 的全部结果
|
||||
@@ -170,6 +228,16 @@ topK = 3,召回、排序、返回都是 3
|
||||
去重合并后一起排序
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Q[query] --> A[路A filtered dense]
|
||||
Q --> B[路B unfiltered dense]
|
||||
A --> M[去重合并]
|
||||
B --> M
|
||||
M --> RRF[RRF / 统一排序]
|
||||
RRF --> OUT[topN]
|
||||
```
|
||||
|
||||
#### 为什么值得做?
|
||||
|
||||
因为两路解决的是不同失败模式:
|
||||
@@ -333,6 +401,23 @@ RRF(d) = Σ 1 / (k + rank_i(d))
|
||||
- 多路都靠前的候选,融合分自然更高
|
||||
- 只在一路偶然靠前的候选,不会单靠绝对分尺度“爆掉”
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph paths["各路有序结果"]
|
||||
P1[dense ranks]
|
||||
P2[BM25 ranks]
|
||||
P3[filtered ranks · 可选]
|
||||
end
|
||||
|
||||
P1 --> RRF["RRF(d) = Σ 1/(k + rank_i)"]
|
||||
P2 --> RRF
|
||||
P3 --> RRF
|
||||
RRF --> OUT[融合序 · 不依赖原始分尺度]
|
||||
|
||||
L2[L2 原分] -.->|不直接相加| X[避免]
|
||||
BM[BM25 原分] -.-> X
|
||||
```
|
||||
|
||||
### 6.2 为什么适合 RAG 多路融合
|
||||
|
||||
RRF 特别适合下面这种现实约束:
|
||||
@@ -553,8 +638,29 @@ Query
|
||||
Agent projection
|
||||
```
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[Query] --> U[dense unfiltered top20]
|
||||
Q --> F[dense filtered top10]
|
||||
Q --> B[bm25/keyword top10]
|
||||
U --> DEDUP[chunk 级去重<br/>docId#chunkIndex]
|
||||
F --> DEDUP
|
||||
B --> DEDUP
|
||||
DEDUP --> RRF[RRF / 加权 RRF]
|
||||
RRF --> RR[轻规则或模型精排 top5]
|
||||
RR --> CAP[每文档 chunk 上限 + pack]
|
||||
CAP --> AG[Agent projection]
|
||||
```
|
||||
|
||||
### 8.2 分阶段推进
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
P0[Phase0<br/>chunk 去重<br/>retrieve-k/return-n] --> P1[Phase1<br/>filtered+unfiltered RRF]
|
||||
P1 --> P2[Phase2<br/>BM25 跨维度]
|
||||
P2 --> P3[Phase3<br/>可插拔精排]
|
||||
```
|
||||
|
||||
#### Phase 0:先修前提
|
||||
|
||||
否则后面多路都会被吞:
|
||||
@@ -0,0 +1,583 @@
|
||||
# RAG 离线评测:讨论、设计与落地
|
||||
|
||||
**日期**:2026-07-28
|
||||
**范围**:`eval/rag-retrieval` 离线 baseline、fixture 生成、与 hybrid/quality 主路径对齐
|
||||
**读者**:要维护或扩展知识库回归评测的工程同学
|
||||
**关联实现 / 变更**:
|
||||
|
||||
- 目录:`eval/rag-retrieval/`
|
||||
- 脚本:`scripts/eval_rag_retrieval.py`、`generate_rag_lookup_snapshots.ps1`、`prepare_rag_eval_seed.ps1`
|
||||
- OpenSpec / devflow:`rag-eval-hybrid-baseline`(已归档)
|
||||
- 前置:`docs/RAG-Hybrid质量分与后处理.md`、`docs/RAG-Agent如何读relevance_level.md`
|
||||
|
||||
---
|
||||
|
||||
## 1. 为什么要单独谈评测
|
||||
|
||||
hybrid、chunk 去重、qualityScore 统一之后,工程上仍缺一块:
|
||||
|
||||
> **改检索之后,用什么可重复的信号判断「变好了还是变坏了」?**
|
||||
|
||||
完整诊断 E2E(多工具 + 最终回答)太重、太噪。需要一层**只盯 `lookup_knowledge` 召回与管道行为**的回归。
|
||||
|
||||
本文汇总讨论中形成的:
|
||||
|
||||
1. 离线评测是什么、不是什么
|
||||
2. Golden / Fixture 怎么设计、比什么
|
||||
3. 项目现状是否符合定义
|
||||
4. 改造复杂度与 sm-flow 落地(含 apply 中发现的闸门问题)
|
||||
5. 指标、报告、纪律
|
||||
|
||||
---
|
||||
|
||||
## 2. 离线评测:概念边界
|
||||
|
||||
### 2.1 离线 vs 在线
|
||||
|
||||
| 说法 | 含义 |
|
||||
|------|------|
|
||||
| **在线** | 真跑检索:embedding、Milvus、完整 `lookup_knowledge` |
|
||||
| **离线** | **不再访问检索栈**;用事先冻住的结果快照,和标准答案比对 |
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph online["在线(贵、真、偶发)"]
|
||||
S[Seed 语料] --> G[真 lookup / 检索管道]
|
||||
G --> F[写入 Fixtures]
|
||||
end
|
||||
|
||||
subgraph offline["离线(便宜、稳、可 CI)"]
|
||||
C[Golden cases] --> E[比对脚本]
|
||||
F2[已提交的 Fixtures] --> E
|
||||
E --> R[报告 / baseline / diff]
|
||||
end
|
||||
|
||||
F -.->|提交入库| F2
|
||||
```
|
||||
|
||||
可以记成:
|
||||
|
||||
```text
|
||||
在线:考试现场答题(环境会变)
|
||||
离线:用标准答卷复印件批改(环境冻结)
|
||||
```
|
||||
|
||||
### 2.2 离线适合 / 不适合回答的问题
|
||||
|
||||
**适合:**
|
||||
|
||||
- 契约有没有破(结构、关键字段、行为路径)
|
||||
- 在「同一份检索结果」假设下,期望文档/关键词/attempt 是否仍满足
|
||||
- 相对上一版 baseline 的回归 diff
|
||||
|
||||
**不适合单独承担:**
|
||||
|
||||
- hybrid 是否比 dense 更好 → 需要**同一时期**在线双跑
|
||||
- 阈值 0.75 是否合适 → 需要在线统计 level 分布
|
||||
- 换 embedding 后召回如何 → 必须重刷 fixture 或 live
|
||||
- Agent 最终诊断对不对 → 诊断 E2E
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Q1[契约 / 期望回归] --> OFF[离线 baseline]
|
||||
Q2[当前召回是否正确] --> ON[在线生成 fixture 或 live]
|
||||
Q3[dense vs hybrid 增益] --> CMP[同期双 mode 对照]
|
||||
Q4[诊断是否正确] --> E2E[诊断 harness]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 三块积木
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph golden_box["① Golden(薄)"]
|
||||
GQ[query]
|
||||
GE[期望:doc / keyword / attempt / …]
|
||||
end
|
||||
|
||||
subgraph fixture_box["② Fixture(冻)"]
|
||||
FM[meta: caseId, searchMode, kbScope, time]
|
||||
FL[lookupResult 结构化输出]
|
||||
end
|
||||
|
||||
subgraph judge_box["③ 裁判(规则)"]
|
||||
A[按 golden 字段断言]
|
||||
M[衍生 Hit / rank / 通过率]
|
||||
REP[json + md + 可选 diff]
|
||||
end
|
||||
|
||||
golden_box -->|caseId 对齐| judge_box
|
||||
fixture_box --> judge_box
|
||||
```
|
||||
|
||||
| 积木 | 是什么 | 不是什么 |
|
||||
|------|--------|----------|
|
||||
| **Golden** | 问什么 + **应该**怎样 | 不是整包线上成功 JSON 原样当期望 |
|
||||
| **Fixture** | 某次跑完**实际**怎样 | 不是每次离线评测都要重跑检索 |
|
||||
| **裁判** | 关键字段比对 + 指标 + 报告 | 不是两个大 JSON deep equal |
|
||||
|
||||
时间线:
|
||||
|
||||
```text
|
||||
① 设计 Golden
|
||||
② (改检索 / 换库 / 换 mode 时)在线跑 → 写/更新 Fixtures
|
||||
③ 日常:Fixtures × Golden → 报告(多数时候只做这一步)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 为什么不全量字段对比
|
||||
|
||||
讨论中的直觉「分数不固定」是对的,但原因不止这一条:
|
||||
|
||||
| 原因 | 说明 |
|
||||
|------|------|
|
||||
| 分数不稳定 | L2 / RRF / embedding 会漂 |
|
||||
| 实现细节会变 | trace 结构、reason 文案、时间戳、tool_call_id |
|
||||
| 截断与预算会变 | excerpt 长度、packedText |
|
||||
| 目标是「对不对」 | 不是字节级一致 |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
F[Fixture JSON] --> K[只抽取关键字段]
|
||||
G[Golden 期望] --> K
|
||||
K --> P{断言}
|
||||
P -->|通过| OK[Pass + 指标]
|
||||
P -->|失败| FAIL[Fail + 原因列表]
|
||||
|
||||
F -.->|不做| FULL[全量 deep equal]
|
||||
```
|
||||
|
||||
**全量对比**偶尔可用于「紧挨着两次生成器输出的工程 diff」,那不是 golden 质量标准。
|
||||
|
||||
---
|
||||
|
||||
## 5. Golden / Fixture 字段设计
|
||||
|
||||
### 5.1 Golden:一行长什么样(分层)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
ID[身份: caseId, scenario/tags, notes]
|
||||
IN[输入: query]
|
||||
P0[P0 命中: expectedDocIds/Sources, expectedKeywords]
|
||||
P1[P1 行为/展示: attempt, fallback, status, breadcrumb]
|
||||
P2[P2 细粒度: evidenceKey, minCount, firstRankMax]
|
||||
P3[P3 观察: relevanceLevel — 慎作硬门禁]
|
||||
NEG[负例: mustNotDocIds/Sources]
|
||||
|
||||
ID --> IN --> P0 --> P1
|
||||
P1 --> P2
|
||||
P1 --> NEG
|
||||
P1 --> P3
|
||||
```
|
||||
|
||||
| 优先级 | 比什么 | 作用 |
|
||||
|--------|--------|------|
|
||||
| P0 | docId / source、keywords | 召回对不对、段是否有用 |
|
||||
| P1 | selectedAttempt、fallback、evidenceStatus | 路径有没有坏 |
|
||||
| P1 | mustNot* | 硬负例 / decoy |
|
||||
| P2 | chunk / evidenceKey、条数、首条相关 rank | 身份与排序 |
|
||||
| P3 | relevance_level | 观察用;hybrid 下易松 |
|
||||
|
||||
**原则:** 期望对齐「用户/Agent 可感知的对错」,少锁实现细节。
|
||||
|
||||
### 5.2 Fixture:最少保留什么
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph fix["fixture"]
|
||||
META["meta<br/>caseId, query, retrievedAt<br/>searchMode, kbScope?"]
|
||||
RES["result / lookupResult<br/>有序 evidence[]<br/>attempt / fallback / status<br/>contextPack? / traces?"]
|
||||
DBG["debug 可选<br/>rawScore, denseDistance…"]
|
||||
end
|
||||
|
||||
META --> RES
|
||||
RES --> DBG
|
||||
```
|
||||
|
||||
每条 evidence 最少:
|
||||
|
||||
```text
|
||||
docId 或可对齐的 source
|
||||
excerpt / content(关键词断言需要)
|
||||
顺序 = rank(数组下标即可)
|
||||
evidenceKey / chunkIndex(多 chunk case 需要)
|
||||
title / breadcrumb(按需)
|
||||
```
|
||||
|
||||
**故意不锁:** score 全文、完整 trace、tool_call_id、packedText 全文(除非单独立项)。
|
||||
|
||||
### 5.3 Golden → Fixture 取值对照
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph g["Golden"]
|
||||
g1[expectedDocIds]
|
||||
g2[expectedKeywords]
|
||||
g3[expectedSelectedAttempt]
|
||||
g4[expectedFallbackReason]
|
||||
g5[mustNotDocIds]
|
||||
end
|
||||
|
||||
subgraph f["Fixture"]
|
||||
f1[evidence[].docId/source]
|
||||
f2[evidence[].excerpt 拼接]
|
||||
f3[retrievalTrace.selectedAttempt]
|
||||
f4[retrievalTrace.fallbackReason]
|
||||
f5[evidence 全表扫描]
|
||||
end
|
||||
|
||||
g1 --> f1
|
||||
g2 --> f2
|
||||
g3 --> f3
|
||||
g4 --> f4
|
||||
g5 --> f5
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. 指标与报告
|
||||
|
||||
### 6.1 两层指标
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph gate["门禁主信号"]
|
||||
PASS[逐 case Pass/Fail]
|
||||
RATE[通过率 / 按 tag 通过率]
|
||||
end
|
||||
|
||||
subgraph quality["质量刻度(报告展示)"]
|
||||
HIT[Hit@n / hitLevel strong·medium·weak·miss]
|
||||
RANK[first relevant rank / MRR]
|
||||
KW[keyword coverage]
|
||||
FB[fallback rate]
|
||||
LV[level 直方图 — 观察]
|
||||
end
|
||||
|
||||
PASS --> RATE
|
||||
HIT --> RANK
|
||||
```
|
||||
|
||||
| 指标 | 含义 |
|
||||
|------|------|
|
||||
| **Pass/Fail** | golden 声明的 expected* 是否全部满足 |
|
||||
| **hitLevel** | strong / medium / weak / miss(项目已有) |
|
||||
| **Recall@K** | strong+medium 算命中 |
|
||||
| **firstExpectedRank** | 第一条期望文档的排名 |
|
||||
| **fallback / attempt** | 行为路径 |
|
||||
| **level 分布** | 宜观察,hybrid 下慎作硬门禁 |
|
||||
|
||||
### 6.2 报告长什么样
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
EVAL[离线评测] --> J[baseline.json<br/>机器可读]
|
||||
EVAL --> M[baseline.md<br/>人读表格]
|
||||
EVAL --> D[baseline-diff.*<br/>相对上一版]
|
||||
|
||||
J --> CI[CI / 脚本解析]
|
||||
M --> HUM[人看失败原因]
|
||||
D --> REV[改代码还是改期望]
|
||||
```
|
||||
|
||||
工程上的「得出结果」=:
|
||||
|
||||
1. **门禁**:must-pass 是否全绿
|
||||
2. **诊断**:谁红、红在哪类断言
|
||||
3. **趋势**:相对旧 baseline 变好还是变差
|
||||
|
||||
---
|
||||
|
||||
## 7. 项目现状审计(改造前)
|
||||
|
||||
讨论结论:**模型符合定义,内容偏旧(约 70%)**。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph ok["已符合"]
|
||||
A1[golden × fixture × key-field]
|
||||
A2[seed + kb_scope=rag-eval]
|
||||
A3[Hit 分层 + recall + baseline diff]
|
||||
A4[离线不连库]
|
||||
end
|
||||
|
||||
subgraph gap["缺口"]
|
||||
B1[生成器仍传 vector-store.mode=spring]
|
||||
B2[fixture 无 searchMode/kbScope]
|
||||
B3[无 dense/hybrid 双目录对照]
|
||||
B4[无 mustNot / chunk 硬期望]
|
||||
B5[快照停在 boost 改序时代]
|
||||
end
|
||||
|
||||
ok --> gap
|
||||
```
|
||||
|
||||
| 维度 | 符合度 |
|
||||
|------|--------|
|
||||
| 三件套架构 | 高 |
|
||||
| 关键字段比对 | 高 |
|
||||
| 报告 / diff | 高 |
|
||||
| 与 hybrid 主路径同步 | 低(改造前) |
|
||||
| chunk / 负例 / mode 矩阵 | 弱或无 |
|
||||
|
||||
---
|
||||
|
||||
## 8. 改造策略与复杂度
|
||||
|
||||
### 8.1 两刀切分
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
K1["第一刀(已落地)<br/>search.mode 接线<br/>fixture meta<br/>README<br/>重刷 hybrid baseline"]
|
||||
K2["第二刀(未做)<br/>fixtures/hybrid vs dense<br/>对照表<br/>golden tags/mustNot"]
|
||||
|
||||
K1 --> DONE[可门禁当前主路径]
|
||||
K2 --> CMP[可回答 hybrid 增益]
|
||||
```
|
||||
|
||||
**复杂度判断:中低。** 不必重写框架;成本在联调环境与 baseline 纪律,不在算法。
|
||||
|
||||
| 工作 | 复杂度 |
|
||||
|------|--------|
|
||||
| 改生成参数 / meta / README | 低 |
|
||||
| seed + 重刷 fixture | 中低(看环境) |
|
||||
| 双目录对照 | 中低(第二刀) |
|
||||
| 换框架 / LLM judge | 高(不建议现在) |
|
||||
|
||||
### 8.2 sm-flow 落地范围(已确认)
|
||||
|
||||
- **仅第一刀**
|
||||
- **接线必交**;fixture 刷新尽力(本次环境可用,已刷绿)
|
||||
|
||||
Change:`rag-eval-hybrid-baseline`(已归档)。
|
||||
|
||||
---
|
||||
|
||||
## 9. 落地后的主链路(当前)
|
||||
|
||||
### 9.1 日常离线
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
GC[golden-cases.json] --> PY[eval_rag_retrieval.py]
|
||||
FX[fixtures/*.json] --> PY
|
||||
PY --> BR[reports/baseline.*]
|
||||
```
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
### 9.2 改检索后的完整环
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant Eng as 工程师
|
||||
participant Seed as prepare_rag_eval_seed
|
||||
participant Gen as generate_rag_lookup_snapshots
|
||||
participant Tool as LookupKnowledgeTool
|
||||
participant Off as eval_rag_retrieval.py
|
||||
participant Git as 仓库 baseline
|
||||
|
||||
Eng->>Seed: 导入 seed-docs (kb_scope=rag-eval)
|
||||
Seed-->>Eng: MySQL/L0/Milvus 就绪
|
||||
Eng->>Gen: -SearchMode hybrid
|
||||
Gen->>Tool: 每条 golden.query
|
||||
Tool-->>Gen: LookupResult
|
||||
Gen->>Gen: 写 fixture + searchMode/kbScope
|
||||
Eng->>Off: fixtures × golden
|
||||
Off-->>Eng: pass/fail + 指标
|
||||
Eng->>Git: 意图变更则更新 baseline(带 diff 原因)
|
||||
```
|
||||
|
||||
### 9.3 生成器配置(改造后)
|
||||
|
||||
| 参数 | 默认 | 含义 |
|
||||
|------|------|------|
|
||||
| `SearchMode` | `hybrid` | `retrieval.search.mode` |
|
||||
| `KbScope` | `rag-eval` | 评测语料隔离 |
|
||||
| (已删除) | — | `vector-store.mode=spring\|sdk` |
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 # hybrid
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures ... -SkipEval # 对照用
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Apply 中发现的关键问题:filter fallback 与 quality 闸门
|
||||
|
||||
### 10.1 现象
|
||||
|
||||
重刷 hybrid fixtures 后,`chat-l0-filter-fallback` 变红:
|
||||
|
||||
- 只命中 decoy(`overfilter-decoy`)
|
||||
- `selectedAttempt=FILTERED_VECTOR`,**没有** unfiltered retry
|
||||
- 根因:hybrid **纯 rank→quality** 时 rank1 恒为 ~1.0 → `isLowQuality` 永不成立
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph before["纯 rank quality(有问题)"]
|
||||
H1[hybrid RRF 序] --> R1[rank1 quality=1.0]
|
||||
R1 --> N1[isLowQuality=false]
|
||||
N1 --> X1[不 retry · decoy 留下]
|
||||
end
|
||||
|
||||
subgraph after["denseDistance 闸门(已修)"]
|
||||
H2[hybrid RRF 序 · 排序不变] --> D2[并行 dense 填 denseDistance]
|
||||
D2 --> Q2[quality = L2 归一化]
|
||||
Q2 --> L2{top quality < 0.5?}
|
||||
L2 -->|是| RET[UNFILTERED_VECTOR_RETRY]
|
||||
L2 -->|否| KEEP[保留 filtered 结果]
|
||||
end
|
||||
```
|
||||
|
||||
### 10.2 设计取舍(必须记清)
|
||||
|
||||
| 信号 | 用途 |
|
||||
|------|------|
|
||||
| **RRF / originalRank** | **排序权威**(谁在前) |
|
||||
| **denseDistance → quality** | **绝对质量闸门**(要不要 retry、level 档) |
|
||||
| **不**再:用 L2 覆盖 hybrid 主分 / label | 避免回到「伪装成 L2」 |
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph sort["排序"]
|
||||
RRF[RRF 返回序]
|
||||
end
|
||||
|
||||
subgraph gate["质量闸门"]
|
||||
L2[dense L2 若有]
|
||||
RK[rank 回退若无 dense]
|
||||
L2 --> QS[qualityScore]
|
||||
RK --> QS
|
||||
end
|
||||
|
||||
RRF --> LIST[evidence 列表顺序]
|
||||
QS --> LV[relevance_level]
|
||||
QS --> FB[isLowQuality → filter fallback]
|
||||
```
|
||||
|
||||
这与 quality 统一文的精神一致:**排序与闸门分信号**;只是 hybrid 闸门不能**只**靠序数分。
|
||||
|
||||
验证(归档时):
|
||||
|
||||
- offline **7/7 pass**
|
||||
- fallback case:`UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`
|
||||
|
||||
---
|
||||
|
||||
## 11. Baseline 纪律
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
RED[离线变红] --> Q{实现退步还是预期变了?}
|
||||
Q -->|退步| CODE[改代码]
|
||||
Q -->|预期变了| GOLD[改 golden / 重刷 fixture]
|
||||
GOLD --> NOTE[写清原因 · 更新 baseline]
|
||||
Q -->|禁止| BLIND[不看 diff 整锅覆盖]
|
||||
```
|
||||
|
||||
README 原话仍然成立:fixture 对不上,要么修链路,要么改期望——**二选一要显式**。
|
||||
|
||||
---
|
||||
|
||||
## 12. 与完整评测体系的位置
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph L1["L1 离线契约 — 已有且已对齐 hybrid"]
|
||||
OFF[fixtures × golden]
|
||||
end
|
||||
|
||||
subgraph L2["L2 在线召回 — 生成器已接线"]
|
||||
LIVE[seed → snapshot hybrid]
|
||||
end
|
||||
|
||||
subgraph L3["L3 对照与标定 — 部分未做"]
|
||||
DD[dense vs hybrid 双目录表]
|
||||
CAL[level/阈值直方图标定]
|
||||
end
|
||||
|
||||
subgraph L4["L4 诊断 E2E — 另一套"]
|
||||
DIAG[多工具 · 最终回答]
|
||||
end
|
||||
|
||||
L1 --> L2
|
||||
L2 --> L3
|
||||
L2 -.-> L4
|
||||
```
|
||||
|
||||
| 层 | 状态 |
|
||||
|----|------|
|
||||
| L1 离线 | **已落地**,hybrid fixtures + baseline 绿 |
|
||||
| L2 在线生成 | **已接线**,本机已成功重刷 |
|
||||
| L3 双 mode 对照目录 | **未做**(第二刀) |
|
||||
| L4 诊断 E2E | 独立 harness,非本文 |
|
||||
|
||||
---
|
||||
|
||||
## 13. 实践清单
|
||||
|
||||
1. **日常**:只跑离线 `eval_rag_retrieval.py`。
|
||||
2. **改检索 / 索引 / mode / 闸门**:seed → hybrid 生成 → 离线 → 看 diff 再更新 baseline。
|
||||
3. **对照 dense**:`-SearchMode dense` 指到另一 fixtures 目录(第二刀可产品化报表)。
|
||||
4. **Golden** 锁业务真值;**score / 完整 trace** 默认不锁。
|
||||
5. **relevance_level** 先观察,慎作硬门禁。
|
||||
6. **排序听 RRF**;**retry/level 听绝对 quality(dense L2)**。
|
||||
7. 评测语料固定 `kb_scope=rag-eval`,勿绑生产杂库。
|
||||
|
||||
---
|
||||
|
||||
## 14. 结语
|
||||
|
||||
离线评测不是「再造一个复杂平台」,而是:
|
||||
|
||||
```text
|
||||
固定问题(Golden)
|
||||
× 冻结答卷(Fixture)
|
||||
× 关键字段裁判
|
||||
→ 可 diff 的报告
|
||||
```
|
||||
|
||||
项目原本骨架正确;本轮补上了 **hybrid 时代的生成接线、fixture meta、baseline 重刷**,并在真实跑通时修正了 **「序数 quality 杀死 filter fallback」** 的闸门设计。
|
||||
|
||||
下一有价值的增量是 **dense/hybrid 同期对照表(第二刀)** 与 **level 分布标定**,而不是换评测框架。
|
||||
|
||||
---
|
||||
|
||||
## 附录 A:目录与命令速查
|
||||
|
||||
| 路径 | 作用 |
|
||||
|------|------|
|
||||
| `eval/rag-retrieval/cases/golden-cases.json` | Golden |
|
||||
| `eval/rag-retrieval/fixtures/*.json` | Fixtures(含 searchMode) |
|
||||
| `eval/rag-retrieval/seed-docs/` | 评测语料 |
|
||||
| `eval/rag-retrieval/reports/baseline.*` | 离线基线报告 |
|
||||
| `scripts/eval_rag_retrieval.py` | 离线裁判 |
|
||||
| `scripts/generate_rag_lookup_snapshots.ps1` | 在线生成 fixture |
|
||||
| `scripts/prepare_rag_eval_seed.ps1` | 导入 seed |
|
||||
|
||||
```powershell
|
||||
# 离线
|
||||
python scripts\eval_rag_retrieval.py
|
||||
|
||||
# 完整刷新(需 embedding + Milvus 等)
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid
|
||||
```
|
||||
|
||||
## 附录 B:相关文档
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| `eval/rag-retrieval/README.md` | 操作说明(以仓库为准) |
|
||||
| `docs/RAG-Hybrid质量分与后处理.md` | quality / 后处理 |
|
||||
| `docs/RAG-Agent如何读relevance_level.md` | Agent 如何读 level |
|
||||
| `docs/RAG排序-多路召回与RRF.md` | 多路与 RRF |
|
||||
| `devflow/projects/2026-07-28-rag-eval-hybrid-baseline/` | 本 change 档案 |
|
||||
| `openspec/specs/rag-eval-offline-baseline/spec.md` | 主规格 |
|
||||
File diff suppressed because it is too large
Load Diff
+89
-143
@@ -1,38 +1,65 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
This directory contains the offline retrieval baseline for the RAG refactor.
|
||||
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
|
||||
|
||||
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
||||
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
||||
evidence keywords before changing L0 behavior, query augmentation, evidence
|
||||
post-processing, or Spring AI VectorStore integration.
|
||||
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
|
||||
|
||||
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
|
||||
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
|
||||
|
||||
Related design notes:
|
||||
|
||||
- `docs/RAG-Hybrid质量分与后处理.md`
|
||||
- `docs/RAG-Agent如何读relevance_level.md`
|
||||
- `mvp/architecture/RAG知识检索架构.md` §6
|
||||
|
||||
## Offline vs live
|
||||
|
||||
| Layer | What | Needs live stack? |
|
||||
|-------|------|-------------------|
|
||||
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
|
||||
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
|
||||
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
|
||||
|
||||
Daily CI / local quick check: **offline only**.
|
||||
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
|
||||
|
||||
## Layout
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
||||
fixtures/*.json Saved retrieval fixtures for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/baseline-diff.* Optional diff reports
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
cases/golden-cases.json Fixed queries + expectations
|
||||
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
|
||||
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
|
||||
reports/baseline.json|md Last accepted offline report
|
||||
reports/baseline-diff.* Optional diff vs previous report
|
||||
```
|
||||
|
||||
## Seed Docs + Import/Reindex
|
||||
## Fixture shape (minimum)
|
||||
|
||||
The live-tool eval uses canonical seed documents so the real
|
||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
||||
```text
|
||||
caseId
|
||||
query
|
||||
retrievedAt
|
||||
searchMode # hybrid | dense (required on newly generated fixtures)
|
||||
kbScope # e.g. rag-eval when generation used a scope
|
||||
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
|
||||
```
|
||||
|
||||
Seed documents live in:
|
||||
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
|
||||
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
|
||||
|
||||
Older fixtures may omit `searchMode`; regenerate to attach meta.
|
||||
|
||||
## Seed docs + import
|
||||
|
||||
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
```
|
||||
|
||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
||||
Frontmatter example:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
Import or reindex the seed docs through the real upload pipeline:
|
||||
Import via real upload pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
||||
Milvus chunks.
|
||||
Isolation:
|
||||
|
||||
`kb_scope` isolates eval data:
|
||||
- App default may leave `retrieval.kb-scope` empty (all docs).
|
||||
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
|
||||
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
|
||||
|
||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
||||
without `kb_scope` remain searchable;
|
||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
||||
vector retrieval both use only the canonical eval seed docs;
|
||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
||||
boundary.
|
||||
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
|
||||
|
||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
||||
their frontmatter keywords matched the query.
|
||||
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
## Offline run (no live stack)
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Custom paths are also supported:
|
||||
Custom paths:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Generate Fixtures From LookupKnowledgeTool
|
||||
|
||||
Use the snapshot generator when fixtures should reflect the real
|
||||
`LookupKnowledgeTool` pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
For the intended live loop, run seed import first:
|
||||
## Generate fixtures (live stack)
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
python scripts\eval_rag_retrieval.py
|
||||
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
|
||||
```
|
||||
|
||||
The script runs a Spring test harness:
|
||||
|
||||
```text
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
||||
```
|
||||
|
||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
||||
|
||||
Custom paths are supported:
|
||||
Dense baseline snapshot (same seed, comparison only):
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
||||
-Fixtures eval\rag-retrieval\fixtures `
|
||||
-RetrievedAt 2026-07-06T00:00:00Z
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
|
||||
```
|
||||
|
||||
The generator is disabled in normal test runs. It only executes when
|
||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
||||
depends on the configured runtime retrieval stack.
|
||||
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
|
||||
|
||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
||||
signal: either the golden expectations need to be adjusted to the current
|
||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
||||
|
||||
## Modular RAG Contract
|
||||
|
||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
||||
`lookup_knowledge` output:
|
||||
Maven equivalent:
|
||||
|
||||
```text
|
||||
lookupResult.evidenceBlocks
|
||||
lookupResult.contextPack
|
||||
lookupResult.retrievalTrace
|
||||
lookupResult.rerankTrace
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
|
||||
-Drag.snapshot.enabled=true \
|
||||
-Dretrieval.kb-scope=rag-eval \
|
||||
-Dretrieval.search.mode=hybrid \
|
||||
test
|
||||
```
|
||||
|
||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
||||
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
|
||||
|
||||
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
|
||||
|
||||
## Golden assertions
|
||||
|
||||
Supported expectation fields include:
|
||||
|
||||
- `expectedSources` / `expectedDocIds`
|
||||
- `expectedBreadcrumbs`
|
||||
- `expectedKeywords`
|
||||
- `expectedBreadcrumbs` / `expectedKeywords`
|
||||
- `expectedSelectedAttempt`
|
||||
- `expectedFallbackReason`
|
||||
- `expectedFallbackReasons`
|
||||
- `expectedFallbackReason` / `expectedFallbackReasons`
|
||||
- `expectedEvidenceStatus`
|
||||
- `expectedContextSources`
|
||||
- `expectedRerankTopSource`
|
||||
|
||||
This lets the baseline catch regressions such as losing the expected evidence
|
||||
source, skipping context packing, changing the selected retrieval attempt, or
|
||||
breaking the filtered-vector to unfiltered-retry fallback.
|
||||
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
|
||||
|
||||
## Baseline Diff
|
||||
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
|
||||
|
||||
To compare a freshly generated report against an existing baseline:
|
||||
## Baseline diff
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
|
||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||
```
|
||||
|
||||
The diff reports aggregate regressions and case-level changes for:
|
||||
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
|
||||
Non-zero exit on case failure or regression in diff mode.
|
||||
|
||||
- pass rate, recall@K, strong hit rate, miss count
|
||||
- pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
## Hit levels
|
||||
|
||||
The command exits non-zero when a case fails or the diff contains a regression.
|
||||
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
|
||||
- `medium`: expected document found, coverage incomplete
|
||||
- `weak`: keyword hit without expected document
|
||||
- `miss`: neither
|
||||
|
||||
## Hit Levels
|
||||
`Recall@K` counts `strong` + `medium`.
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
||||
- `weak`: expected evidence keyword is found, but expected document is missing.
|
||||
- `miss`: expected document and expected evidence are not found.
|
||||
## Optional live smoke (post-reindex)
|
||||
|
||||
`Recall@K` counts `strong` and `medium` as retrieved.
|
||||
|
||||
## Scope
|
||||
|
||||
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
||||
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
||||
not a claim that live production retrieval accuracy is complete.
|
||||
|
||||
## Live Post-Reindex Acceptance
|
||||
|
||||
When embedding input changes, existing vectors do not update by themselves. For
|
||||
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
||||
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
||||
semantic signal.
|
||||
|
||||
Use this optional live acceptance flow after the application is running and the
|
||||
knowledge base has been reindexed:
|
||||
After reindex, with app up:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
|
||||
```
|
||||
|
||||
Custom service URL and output paths are supported:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py \
|
||||
--base-url http://127.0.0.1:9900 \
|
||||
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
||||
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The script calls:
|
||||
|
||||
```text
|
||||
GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
|
||||
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
"caseId" : "aiops-payment-latency-alert",
|
||||
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Payment Latency",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "Alert Scope Control",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Service Alerts",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Service Alerts",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.015384615398943424,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 2]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Service Alerts\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2108,
|
||||
"includedSources" : [ "payment-service-latency", "aiops-alert-scope-control", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"hitReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 236,
|
||||
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"categoryFilter" : "aiops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "aiops" ],
|
||||
"matched_keywords" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||
"entities" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||
"l0_titles" : [ "Payment Service Latency Alert" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"entities": ["payment-service", "HighLatency"],
|
||||
"l0_titles": ["Payment Service Latency Alert Playbook"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"categoryFilter" : "aiops",
|
||||
"candidateCount" : 5,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 969,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7736010700464249
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 11,
|
||||
"topScore": 0.84,
|
||||
"topSimilarity": 0.84
|
||||
}
|
||||
]
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.7736010700464249,
|
||||
"finalScore" : 0.7736010700464249,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.4683566689491272,
|
||||
"finalScore" : 0.4683566689491272,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.4981400966644287,
|
||||
"finalScore" : 0.4981400966644287,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.3814886808395386,
|
||||
"finalScore" : 0.3814886808395386,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "payment-service-latency",
|
||||
"baseScore": 0.84,
|
||||
"finalScore": 1.29,
|
||||
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.68,
|
||||
"finalScore": 0.78,
|
||||
"boostReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 5,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,65 +1,122 @@
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 188,
|
||||
"includedSources": ["aiops-alert-scope-control"],
|
||||
"omittedSources": []
|
||||
"caseId" : "aiops-prometheus-alert-scope",
|
||||
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "Alert Scope Control",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Service Alerts",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Service Alerts",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Payment Latency",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.03076923079788685,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 2]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Service Alerts\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2069,
|
||||
"includedSources" : [ "aiops-alert-scope-control", "payment-service-latency", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"entities": ["alert payload"],
|
||||
"l0_titles": ["AIOps Alert Scope Control"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"categoryFilter" : "aiops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "aiops" ],
|
||||
"matched_keywords" : [ "alert payload", "unrelated active alerts" ],
|
||||
"entities" : [ "alert payload", "unrelated active alerts" ],
|
||||
"l0_titles" : [ "AIOps Alert Scope Control" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 1,
|
||||
"usable": true,
|
||||
"durationMs": 8,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"categoryFilter" : "aiops",
|
||||
"candidateCount" : 5,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 1623,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7561411112546921
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "aiops-alert-scope-control",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.7561411112546921,
|
||||
"finalScore" : 0.7561411112546921,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.503810703754425,
|
||||
"finalScore" : 0.503810703754425,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.5772626996040344,
|
||||
"finalScore" : 0.5772626996040344,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.36770421266555786,
|
||||
"finalScore" : 0.36770421266555786,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 5,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,88 @@
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
"caseId" : "chat-diagnosis-flow",
|
||||
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "incident-diagnosis-flow",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "incident-diagnosis-flow#chunk-1",
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"title" : "Diagnosis Flow",
|
||||
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "incident-diagnosis-flow",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "incident-diagnosis-flow#chunk-0",
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.016129031777381897,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: incident-diagnosis-flow\ntitle: Diagnosis Flow\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.\n\n[Evidence 2]\nsource: incident-diagnosis-flow\ntitle: AIOps\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1032,
|
||||
"includedSources" : [ "incident-diagnosis-flow", "incident-diagnosis-flow" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 192,
|
||||
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"categoryFilter" : "ops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "ops" ],
|
||||
"matched_keywords" : [ "standard troubleshooting flow", "application incident" ],
|
||||
"entities" : [ "standard troubleshooting flow", "application incident" ],
|
||||
"l0_titles" : [ "Incident Diagnosis Flow" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["collect evidence", "verify", "remediation"],
|
||||
"entities": ["application incident"],
|
||||
"l0_titles": ["Incident Diagnosis Flow"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"categoryFilter" : "ops",
|
||||
"candidateCount" : 2,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 1540,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.6828859150409698
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 10,
|
||||
"topScore": 0.82,
|
||||
"topSimilarity": 0.82
|
||||
}
|
||||
]
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"baseScore" : 0.6828859150409698,
|
||||
"finalScore" : 0.6828859150409698,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"baseScore" : 0.3166210651397705,
|
||||
"finalScore" : 0.3166210651397705,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.82,
|
||||
"finalScore": 1.07,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.55,
|
||||
"finalScore": 0.55,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 2,
|
||||
"evidenceBlockCount" : 2,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
"caseId" : "chat-l0-domain-hint",
|
||||
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "L0",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## L0",
|
||||
"score" : 0.0317460335791111,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Chunking",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Chunking",
|
||||
"score" : 0.015625,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 2]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## L0\n\n[Evidence 4]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1965,
|
||||
"includedSources" : [ "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 219,
|
||||
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"categoryFilter" : "rag",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "rag" ],
|
||||
"matched_keywords" : [ "L0 keyword matching", "final retrieval result" ],
|
||||
"entities" : [ "L0 keyword matching", "final retrieval result" ],
|
||||
"l0_titles" : [ "RAG L0 Domain Entity Hint" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"categoryFilter" : "rag",
|
||||
"candidateCount" : 6,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 850,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.6438122987747192
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.6438122987747192,
|
||||
"finalScore" : 0.6438122987747192,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.41426247358322144,
|
||||
"finalScore" : 0.41426247358322144,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.3964804410934448,
|
||||
"finalScore" : 0.3964804410934448,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.28678786754608154,
|
||||
"finalScore" : 0.28678786754608154,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"baseScore": 0.75,
|
||||
"finalScore": 0.9,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 6,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,91 +1,149 @@
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"title": "RAG L0 Filter Fallback",
|
||||
"breadcrumb": "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"score": 0.83,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
"caseId" : "chat-l0-filter-fallback",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-l0-filter-fallback",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-filter-fallback#chunk-2",
|
||||
"source" : "rag-l0-filter-fallback",
|
||||
"title" : "Unfiltered Retry",
|
||||
"breadcrumb" : "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:UNFILTERED_VECTOR_RETRY", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-filter-decoy",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-filter-decoy#chunk-1",
|
||||
"source" : "rag-l0-filter-decoy",
|
||||
"title" : "Approval Window",
|
||||
"breadcrumb" : "RAG > Fallback > Decoy",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:UNFILTERED_VECTOR_RETRY", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "L0",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## L0",
|
||||
"score" : 0.03125,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.03053613007068634,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-l0-filter-fallback\ntitle: Unfiltered Retry\nbreadcrumb: RAG > Fallback > Unfiltered Retry\nlayer: L1\nreasons: semantic_rank:1, attempt:UNFILTERED_VECTOR_RETRY, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.\n\n[Evidence 2]\nsource: rag-l0-filter-decoy\ntitle: Approval Window\nbreadcrumb: RAG > Fallback > Decoy\nlayer: L1\nreasons: semantic_rank:2, attempt:UNFILTERED_VECTOR_RETRY, l0_domain_overlap\ncontent:\n## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n## L0\n\n[Evidence 5]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:5, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2904,
|
||||
"includedSources" : [ "rag-l0-filter-fallback", "rag-l0-filter-decoy", "rag-l0-domain-entity-hint", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.",
|
||||
"score": 0.66,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 214,
|
||||
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : "overfilter-decoy",
|
||||
"selectedAttempt" : "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason" : "filtered_vector_low_quality",
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "overfilter-decoy" ],
|
||||
"matched_keywords" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||
"entities" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||
"l0_titles" : [ "RAG L0 Filter Decoy" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : "overfilter-decoy",
|
||||
"candidateCount" : 2,
|
||||
"usable" : false,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 722,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.47391992807388306
|
||||
}, {
|
||||
"name" : "UNFILTERED_VECTOR_RETRY",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : null,
|
||||
"candidateCount" : 20,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 2606,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7571567445993423
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 1,
|
||||
"usable": false,
|
||||
"durationMs": 7,
|
||||
"topScore": 1.35,
|
||||
"topSimilarity": 0.325
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-l0-filter-fallback",
|
||||
"baseScore" : 0.7571567445993423,
|
||||
"finalScore" : 0.7571567445993423,
|
||||
"boostReasons" : [ "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-l0-filter-decoy",
|
||||
"baseScore" : 0.47391992807388306,
|
||||
"finalScore" : 0.47391992807388306,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.5912361443042755,
|
||||
"finalScore" : 0.5912361443042755,
|
||||
"boostReasons" : [ ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.473749577999115,
|
||||
"finalScore" : 0.473749577999115,
|
||||
"boostReasons" : [ ]
|
||||
}, {
|
||||
"finalRank" : 5,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.4492502808570862,
|
||||
"finalScore" : 0.4492502808570862,
|
||||
"boostReasons" : [ ]
|
||||
} ]
|
||||
},
|
||||
{
|
||||
"name": "UNFILTERED_VECTOR_RETRY",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter": null,
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 13,
|
||||
"topScore": 0.83,
|
||||
"topSimilarity": 0.83
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"baseScore": 0.83,
|
||||
"finalScore": 1.08,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.66,
|
||||
"finalScore": 0.81,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 20,
|
||||
"evidenceBlockCount" : 5,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,88 @@
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
"caseId" : "chat-mysql-connection-pool",
|
||||
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "mysql-connection-pool",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "mysql-connection-pool#chunk-2",
|
||||
"source" : "mysql-connection-pool",
|
||||
"title" : "Connection Pool",
|
||||
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "mysql-connection-pool",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "mysql-connection-pool#chunk-1",
|
||||
"source" : "mysql-connection-pool",
|
||||
"title" : "MySQL",
|
||||
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## MySQL",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: mysql-connection-pool\ntitle: Connection Pool\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.\n\n[Evidence 2]\nsource: mysql-connection-pool\ntitle: MySQL\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## MySQL",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1135,
|
||||
"includedSources" : [ "mysql-connection-pool", "mysql-connection-pool" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 216,
|
||||
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"categoryFilter" : "database",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "database" ],
|
||||
"matched_keywords" : [ "MySQL connection pool" ],
|
||||
"entities" : [ "MySQL connection pool" ],
|
||||
"l0_titles" : [ "MySQL Connection Pool Runbook" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["Database", "MySQL"],
|
||||
"matched_keywords": ["connection pool", "HikariCP", "max_connections"],
|
||||
"entities": ["MySQL", "HikariCP"],
|
||||
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"categoryFilter" : "database",
|
||||
"candidateCount" : 3,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 5267,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.8114794194698334
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 12,
|
||||
"topScore": 0.86,
|
||||
"topSimilarity": 0.86
|
||||
}
|
||||
]
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "mysql-connection-pool",
|
||||
"baseScore" : 0.8114794194698334,
|
||||
"finalScore" : 0.8114794194698334,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "mysql-connection-pool",
|
||||
"baseScore" : 0.49219560623168945,
|
||||
"finalScore" : 0.49219560623168945,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.86,
|
||||
"finalScore": 1.11,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.61,
|
||||
"finalScore": 0.61,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 3,
|
||||
"evidenceBlockCount" : 2,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
"caseId" : "chat-rag-chunk-context",
|
||||
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Chunking",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Chunking",
|
||||
"score" : 0.01587301678955555,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-0",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "RAG",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# RAG",
|
||||
"score" : 0.015625,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 2]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 3]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: RAG\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# RAG",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1966,
|
||||
"includedSources" : [ "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
{
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 203,
|
||||
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
|
||||
"omittedSources": []
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"categoryFilter" : "rag",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "rag" ],
|
||||
"matched_keywords" : [ "split into multiple chunks", "retrieval context" ],
|
||||
"entities" : [ "split into multiple chunks", "retrieval context" ],
|
||||
"l0_titles" : [ "RAG Chunk Context Reconstruction" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"entities": ["chunk", "breadcrumb"],
|
||||
"l0_titles": ["RAG Chunk Context Reconstruction"],
|
||||
"l0_match_count": 1
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"categoryFilter" : "rag",
|
||||
"candidateCount" : 6,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 743,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7487991750240326
|
||||
} ]
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.79,
|
||||
"topSimilarity": 0.79
|
||||
}
|
||||
]
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.7487991750240326,
|
||||
"finalScore" : 0.7487991750240326,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.4840593934059143,
|
||||
"finalScore" : 0.4840593934059143,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.42978107929229736,
|
||||
"finalScore" : 0.42978107929229736,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.3859822154045105,
|
||||
"finalScore" : 0.3859822154045105,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.79,
|
||||
"finalScore": 1.04,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"baseScore": 0.72,
|
||||
"finalScore": 0.87,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"evidenceCandidateCount" : 6,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
{
|
||||
"generatedAt": "2026-07-28T06:48:56.439554+00:00",
|
||||
"baselineReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"currentReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"baselineCaseCount": 7,
|
||||
"currentCaseCount": 7,
|
||||
"baselinePassRate": 1.0,
|
||||
"currentPassRate": 0.8571,
|
||||
"baselineRecallAtK": 1.0,
|
||||
"currentRecallAtK": 0.8571,
|
||||
"regressionCount": 6,
|
||||
"improvementCount": 0,
|
||||
"changedCount": 3,
|
||||
"hasRegression": true,
|
||||
"items": [
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "passRate",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate passRate changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "recallAtK",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate recallAtK changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "strongHitRate",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate strongHitRate changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "passed",
|
||||
"baselineValue": "True",
|
||||
"currentValue": "False",
|
||||
"delta": -1.0,
|
||||
"message": "chat-l0-filter-fallback passed changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "hitLevel",
|
||||
"baselineValue": "strong",
|
||||
"currentValue": "weak",
|
||||
"delta": -2.0,
|
||||
"message": "chat-l0-filter-fallback hitLevel changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "firstExpectedRank",
|
||||
"baselineValue": "1",
|
||||
"currentValue": "-",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback firstExpectedRank changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "selectedAttempt",
|
||||
"baselineValue": "UNFILTERED_VECTOR_RETRY",
|
||||
"currentValue": "FILTERED_VECTOR",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback selectedAttempt changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "fallbackReason",
|
||||
"baselineValue": "filtered_vector_low_quality",
|
||||
"currentValue": "-",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback fallbackReason changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "rerankTopSource",
|
||||
"baselineValue": "rag-l0-filter-fallback",
|
||||
"currentValue": "rag-l0-filter-decoy",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback rerankTopSource changed"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
# RAG Retrieval Baseline Diff
|
||||
|
||||
Generated at: `2026-07-28T06:48:56.439554+00:00`
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Baseline cases | 7 |
|
||||
| Current cases | 7 |
|
||||
| Baseline pass rate | 1.0 |
|
||||
| Current pass rate | 0.8571 |
|
||||
| Baseline recall@K | 1.0 |
|
||||
| Current recall@K | 0.8571 |
|
||||
| Regressions | 6 |
|
||||
| Improvements | 0 |
|
||||
| Changed | 3 |
|
||||
|
||||
## Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
|---|---|---|---|---|---|---:|---|
|
||||
| REGRESSION | aggregate | | passRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate passRate changed |
|
||||
| REGRESSION | aggregate | | recallAtK | 1.0 | 0.8571 | -0.14290000000000003 | aggregate recallAtK changed |
|
||||
| REGRESSION | aggregate | | strongHitRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate strongHitRate changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | passed | True | False | -1.0 | chat-l0-filter-fallback passed changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | hitLevel | strong | weak | -2.0 | chat-l0-filter-fallback hitLevel changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | firstExpectedRank | 1 | - | | chat-l0-filter-fallback firstExpectedRank changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | selectedAttempt | UNFILTERED_VECTOR_RETRY | FILTERED_VECTOR | | chat-l0-filter-fallback selectedAttempt changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | fallbackReason | filtered_vector_low_quality | - | | chat-l0-filter-fallback fallbackReason changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | rerankTopSource | rag-l0-filter-fallback | rag-l0-filter-decoy | | chat-l0-filter-fallback rerankTopSource changed |
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"generatedAt": "2026-07-06T13:37:59.726351+00:00",
|
||||
"generatedAt": "2026-07-28T06:55:09.379164+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
@@ -28,7 +28,7 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:mysql-connection-pool",
|
||||
"2:incident-diagnosis-flow"
|
||||
"2:mysql-connection-pool"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"connection pool",
|
||||
@@ -41,7 +41,7 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"incident-diagnosis-flow"
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
@@ -57,7 +57,7 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:incident-diagnosis-flow",
|
||||
"2:rag-chunk-context-reconstruction"
|
||||
"2:incident-diagnosis-flow"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"collect evidence",
|
||||
@@ -70,7 +70,7 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"rag-chunk-context-reconstruction"
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
@@ -86,7 +86,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:payment-service-latency",
|
||||
"2:mysql-connection-pool"
|
||||
"2:aiops-alert-scope-control",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"p95 latency",
|
||||
@@ -99,7 +101,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"mysql-connection-pool"
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
@@ -114,7 +118,10 @@
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:aiops-alert-scope-control"
|
||||
"1:aiops-alert-scope-control",
|
||||
"2:payment-service-latency",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"payload",
|
||||
@@ -126,6 +133,9 @@
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
@@ -142,7 +152,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-chunk-context-reconstruction",
|
||||
"2:rag-breadcrumb-embedding-gap"
|
||||
"2:rag-l0-domain-entity-hint",
|
||||
"3:rag-chunk-context-reconstruction",
|
||||
"4:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"neighbor chunk",
|
||||
@@ -155,7 +167,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-breadcrumb-embedding-gap"
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
@@ -171,7 +185,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-domain-entity-hint",
|
||||
"2:rag-l0-l1-fusion-ranking"
|
||||
"2:rag-chunk-context-reconstruction",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"domain detector",
|
||||
@@ -184,7 +200,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-l1-fusion-ranking"
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
@@ -200,7 +218,10 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-fallback",
|
||||
"2:rag-l0-domain-entity-hint"
|
||||
"2:rag-l0-filter-decoy",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-l0-domain-entity-hint",
|
||||
"5:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"skip the l0 filter",
|
||||
@@ -213,7 +234,10 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-fallback",
|
||||
"rag-l0-domain-entity-hint"
|
||||
"rag-l0-filter-decoy",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-fallback",
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
Generated at: `2026-07-28T06:55:09.379164+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
@@ -24,10 +24,10 @@ Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | |
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-filter-decoy<br>3:rag-l0-domain-entity-hint<br>4:rag-l0-domain-entity-hint<br>5:rag-chunk-context-reconstruction | |
|
||||
|
||||
@@ -0,0 +1,245 @@
|
||||
{
|
||||
"generatedAt": "2026-07-28T06:48:56.421877+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
"caseCount": 7,
|
||||
"topK": 5,
|
||||
"passedCount": 6,
|
||||
"failedCount": 1,
|
||||
"passRate": 0.8571,
|
||||
"lookupResultCaseCount": 7,
|
||||
"strongHitCount": 6,
|
||||
"mediumHitCount": 0,
|
||||
"weakHitCount": 1,
|
||||
"missCount": 0,
|
||||
"recallAtK": 0.8571,
|
||||
"strongHitRate": 0.8571,
|
||||
"averageFirstHitRank": 1.0
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:mysql-connection-pool",
|
||||
"2:mysql-connection-pool"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"connection pool",
|
||||
"max_connections",
|
||||
"hikaricp"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:incident-diagnosis-flow",
|
||||
"2:incident-diagnosis-flow"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"collect evidence",
|
||||
"verify",
|
||||
"remediation"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:payment-service-latency",
|
||||
"2:aiops-alert-scope-control",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"p95 latency",
|
||||
"payment-service",
|
||||
"downstream dependency"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:aiops-alert-scope-control",
|
||||
"2:payment-service-latency",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"payload",
|
||||
"unrelated active alerts",
|
||||
"scope"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "aiops-alert-scope-control",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-chunk-context-reconstruction",
|
||||
"2:rag-l0-domain-entity-hint",
|
||||
"3:rag-chunk-context-reconstruction",
|
||||
"4:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"neighbor chunk",
|
||||
"same section",
|
||||
"breadcrumb"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-domain-entity-hint",
|
||||
"2:rag-chunk-context-reconstruction",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"domain detector",
|
||||
"entity extractor",
|
||||
"metadata filter"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "weak",
|
||||
"passed": false,
|
||||
"firstExpectedRank": null,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-decoy",
|
||||
"2:rag-l0-filter-decoy"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"low quality"
|
||||
],
|
||||
"breadcrumbMatched": false,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-decoy",
|
||||
"rag-l0-filter-decoy"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-decoy",
|
||||
"failedChecks": [
|
||||
"expected document not found",
|
||||
"selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR",
|
||||
"fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none>",
|
||||
"rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy",
|
||||
"expected context sources missing: rag-l0-filter-fallback"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-28T06:48:56.421877+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Cases | 7 |
|
||||
| Top K | 5 |
|
||||
| Passed | 6 |
|
||||
| Failed | 1 |
|
||||
| Pass rate | 0.8571 |
|
||||
| LookupResult fixtures | 7 |
|
||||
| Recall@K | 0.8571 |
|
||||
| Strong hit rate | 0.8571 |
|
||||
| Strong hits | 6 |
|
||||
| Medium hits | 0 |
|
||||
| Weak hits | 1 |
|
||||
| Misses | 0 |
|
||||
| Average first hit rank | 1.0 |
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||
| chat-l0-filter-fallback | chat | false | weak | FILTERED_VECTOR | | supported | | 1:rag-l0-filter-decoy<br>2:rag-l0-filter-decoy | expected document not found<br>selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR<br>fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none><br>rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy<br>expected context sources missing: rag-l0-filter-fallback |
|
||||
@@ -0,0 +1,394 @@
|
||||
# RAG 检索可观测性、审计与 Trace(现行)
|
||||
|
||||
**更新日期**:2026-07-28
|
||||
**状态**:当前可运行
|
||||
**关联**:`lookup_knowledge`、Harness `ToolBoundary`、`tool_invocation`、`DiagnosisTraceService`、离线 eval
|
||||
|
||||
---
|
||||
|
||||
## 1. 三层边界
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph A["A. 请求内 Trace"]
|
||||
LR[LookupResult<br/>retrievalTrace / rerankTrace / relevanceLevel]
|
||||
end
|
||||
|
||||
subgraph B["B. 持久化审计 + Trace API"]
|
||||
TI[tool_invocation 表]
|
||||
DT[diagnosis_trace 事件摘要]
|
||||
API["GET /api/diagnosis/{sessionId}/trace"]
|
||||
end
|
||||
|
||||
subgraph C["C. 质量回归"]
|
||||
EV[eval/rag-retrieval offline baseline]
|
||||
end
|
||||
|
||||
subgraph agent["Agent 可见(非审计)"]
|
||||
RT[RagToolResult<br/>evidence + optional relevance_level]
|
||||
end
|
||||
|
||||
LK[LookupKnowledgeTool] --> LR
|
||||
LR --> PROJ[RagResultProjector]
|
||||
PROJ --> RT
|
||||
LR --> BOUND[ToolBoundary audit]
|
||||
BOUND --> TI
|
||||
BOUND --> DT
|
||||
TI --> API
|
||||
DT --> API
|
||||
EV -.->|不替代运行时 Trace| LK
|
||||
```
|
||||
|
||||
| 层 | 完善度 | 说明 |
|
||||
|----|--------|------|
|
||||
| A 请求内 | 高 | attempt / fallback / quality 齐全 |
|
||||
| B 持久化 + Trace API | 中高 | RAG 富字段入 `tool_invocation`,经 Trace API 回放 |
|
||||
| C 离线 eval | 高 | hybrid fixtures 回归 |
|
||||
|
||||
**Agent 看到的不是完整 Trace。** 完整检索轨迹在 A/B;Agent 只拿投影后的证据契约。
|
||||
|
||||
---
|
||||
|
||||
## 2. 端到端:从 lookup 到 Trace API
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant Agent
|
||||
participant Adapter as RagToolAdapter
|
||||
participant Bound as ToolBoundary
|
||||
participant Tool as LookupKnowledgeTool
|
||||
participant Sink as JpaToolInvocationAuditSink
|
||||
participant DB as tool_invocation
|
||||
participant Trace as DiagnosisTraceService
|
||||
participant API as GET .../trace
|
||||
|
||||
Agent->>Adapter: lookup_knowledge(query)
|
||||
Adapter->>Bound: execute(legacy, projector)
|
||||
Bound->>Tool: execute(query)
|
||||
Tool-->>Bound: raw LookupResult JSON
|
||||
Note over Tool: 内含 retrievalTrace / rerankTrace / evidenceBlocks
|
||||
Bound->>Bound: project → RagToolResult
|
||||
Bound->>Sink: AuditEvent + rawResultJson + agentResultJson
|
||||
Sink->>Sink: RagLookupAuditEnricher
|
||||
Sink->>DB: 富字段行
|
||||
Bound-->>Agent: 投影后 agent_result(无完整 trace)
|
||||
|
||||
API->>Trace: sessionId + optional runId
|
||||
Trace->>DB: find tool_invocation by run/session
|
||||
Trace-->>API: DiagnosisTraceResponse.toolInvocations[]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. A 层:请求内 Trace(`LookupResult`)
|
||||
|
||||
一次成功的 `lookup_knowledge` 内部出口是 **`LookupResult`**(比 Agent 契约更富)。
|
||||
|
||||
### 3.1 结构总览
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
LR[LookupResult]
|
||||
LR --> F[found]
|
||||
LR --> EB[evidenceBlocks[]]
|
||||
LR --> CP[contextPack]
|
||||
LR --> RT[retrievalTrace]
|
||||
LR --> RR[rerankTrace]
|
||||
LR --> RL[relevanceLevel]
|
||||
LR --> CH[completenessHint]
|
||||
LR --> CNT[evidenceCandidateCount / evidenceBlockCount]
|
||||
|
||||
RT --> ATT[attempts[]]
|
||||
RT --> SEL[selectedAttempt]
|
||||
RT --> FB[fallbackReason]
|
||||
RT --> HINT[queryHints L0]
|
||||
```
|
||||
|
||||
| 字段 | 含义 |
|
||||
|------|------|
|
||||
| `found` | 是否有可用证据块 |
|
||||
| `evidenceBlocks` | 后处理后的 chunk 级证据(含 evidenceKey、source、content…) |
|
||||
| `contextPack` | 字符预算打包文本(内部/审计用) |
|
||||
| `retrievalTrace` | **检索路径 Trace**(见下) |
|
||||
| `rerankTrace` | 后处理排序/quality 痕迹(现多为保序后的 quality) |
|
||||
| `relevanceLevel` | PRECISE / REFERENCE / null |
|
||||
| `completenessHint` | 给模型的天花板提示文案 |
|
||||
|
||||
### 3.2 `retrievalTrace`(检索路径)
|
||||
|
||||
| 字段 | 含义 |
|
||||
|------|------|
|
||||
| `originalQuery` | 原始查询 |
|
||||
| `rewrittenQuery` | L0/变换后用于检索的 query |
|
||||
| `categoryFilter` | 首次过滤的 category(可 null) |
|
||||
| `selectedAttempt` | 最终采用的 attempt 名 |
|
||||
| `fallbackReason` | 如 `filtered_vector_low_quality`;未降级为 null |
|
||||
| `evidenceStatus` | 内部:`supported` / `no_evidence` 等 |
|
||||
| `queryHints` | L0:domains、keywords、entities、l0_match_count… |
|
||||
| `attempts[]` | 每次检索尝试快照 |
|
||||
|
||||
**常见 `selectedAttempt`:**
|
||||
|
||||
| 值 | 含义 |
|
||||
|----|------|
|
||||
| `FILTERED_VECTOR` | 带 category 的首次检索即采用 |
|
||||
| `UNFILTERED_VECTOR` | 无 category,直接全库检索 |
|
||||
| `UNFILTERED_VECTOR_RETRY` | filtered 低质/无证据后去掉 category 重试 |
|
||||
|
||||
**单次 `attempts[]` 元素:**
|
||||
|
||||
| 字段 | 含义 |
|
||||
|------|------|
|
||||
| `name` | attempt 名 |
|
||||
| `query` | 该次实际检索句 |
|
||||
| `categoryFilter` | 该次 filter |
|
||||
| `candidateCount` | 召回候选数 |
|
||||
| `usable` | 后处理阈值后是否可用 |
|
||||
| `topScore` / `topSimilarity` | 引擎分 / 归一化 quality(0~1) |
|
||||
| `durationMs` | 耗时 |
|
||||
| `errorMessage` | 失败时 |
|
||||
|
||||
### 3.3 一次典型路径(含 filter fallback)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[query] --> L0[L0 hint → 可选 categoryFilter]
|
||||
L0 --> A1[attempt FILTERED_VECTOR]
|
||||
A1 --> PQ{isLowQuality?}
|
||||
PQ -->|否| USE1[selectedAttempt = FILTERED_VECTOR]
|
||||
PQ -->|是| A2[attempt UNFILTERED_VECTOR_RETRY]
|
||||
A2 --> USE2[selectedAttempt = RETRY<br/>fallbackReason = low_quality / no_evidence]
|
||||
USE1 --> POST[PostProcess · evidenceBlocks · relevanceLevel]
|
||||
USE2 --> POST
|
||||
POST --> LR[LookupResult 完整 Trace]
|
||||
```
|
||||
|
||||
### 3.4 与 Agent 投影的关系
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LR[LookupResult 全量 Trace] --> PROJ[RagResultProjector]
|
||||
PROJ --> AG[RagToolResult]
|
||||
AG --> F1[evidence_status]
|
||||
AG --> F2[evidence excerpt]
|
||||
AG --> F3[relevance_level 可选]
|
||||
AG --> F4[truncated / returned_count]
|
||||
|
||||
LR -.->|不投影| X1[retrievalTrace]
|
||||
LR -.->|不投影| X2[rerankTrace]
|
||||
LR -.->|不投影| X3[raw scores / contextPack 全文]
|
||||
```
|
||||
|
||||
人/系统要「为什么这样检索」→ 看 **A 全量** 或 **B 落库摘要**,不要只看 Agent 字段。
|
||||
|
||||
---
|
||||
|
||||
## 4. B 层:持久化 + Trace API
|
||||
|
||||
### 4.1 写入路径
|
||||
|
||||
| 组件 | 职责 |
|
||||
|------|------|
|
||||
| `ToolBoundary` | 执行后发 `ToolInvocationAuditEvent`(含 raw LookupResult JSON + agent JSON) |
|
||||
| `RagLookupAuditEnricher` | 从 LookupResult 抽有界 RAG 字段 |
|
||||
| `JpaToolInvocationAuditSink` | 写入 `tool_invocation` |
|
||||
| `TraceAuditEvents.toolInvocation` | 另写一条 diagnosis_trace 摘要事件(不含全文 LookupResult) |
|
||||
|
||||
### 4.2 `tool_invocation` 列(RAG)
|
||||
|
||||
| 列 | lookup_knowledge | 其它工具 |
|
||||
|----|------------------|----------|
|
||||
| `tool_name` | `lookup_knowledge` | 各自工具名 |
|
||||
| `retrieval_layer` | 通常 `L1` | `HARNESS` |
|
||||
| `relevance_level` | **PRECISE / REFERENCE / …** | **null**(不再写 evidence_status) |
|
||||
| `l0_match_count` | queryHints | null |
|
||||
| `l1_match_count` | evidence 块数等 | null |
|
||||
| `is_truncated` | 投影 truncated | false |
|
||||
| `retrieval_details` | JSON `rag_lookup_v1` | 通用 status 元数据 |
|
||||
| `output_preview` | level/attempt 摘要 | status=… |
|
||||
| `duration_ms` / `success` | 有 | 有 |
|
||||
|
||||
### 4.3 `retrieval_details`(rag_lookup_v1)示例
|
||||
|
||||
```json
|
||||
{
|
||||
"audit_schema": "rag_lookup_v1",
|
||||
"search_mode": "hybrid",
|
||||
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallback_reason": "filtered_vector_low_quality",
|
||||
"category_filter": "overfilter-decoy",
|
||||
"evidence_keys": ["doc#chunk-0"],
|
||||
"sources": ["doc"],
|
||||
"evidence_candidate_count": 8,
|
||||
"evidence_block_count": 2,
|
||||
"l0_hints": { "domains": ["mysql"], "matched_keywords": ["pool"] },
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"category_filter": "overfilter-decoy",
|
||||
"candidate_count": 2,
|
||||
"usable": false,
|
||||
"top_similarity": 0.3,
|
||||
"duration_ms": 12
|
||||
},
|
||||
{
|
||||
"name": "UNFILTERED_VECTOR_RETRY",
|
||||
"candidate_count": 5,
|
||||
"usable": true,
|
||||
"top_similarity": 0.9,
|
||||
"duration_ms": 20
|
||||
}
|
||||
],
|
||||
"truncated": false,
|
||||
"returned_count": 2,
|
||||
"evidence_status": "EVIDENCE_FOUND",
|
||||
"invocation_status": "READY",
|
||||
"tool_call_id": "call-…"
|
||||
}
|
||||
```
|
||||
|
||||
**默认不落库:** 原始 query 全文、chunk 正文 excerpt、完整 rerankTrace(体积与隐私)。
|
||||
|
||||
### 4.4 Trace API:人怎么读 RAG
|
||||
|
||||
**接口:**
|
||||
|
||||
```http
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId={runId}
|
||||
```
|
||||
|
||||
**实现:** `DiagnosisTraceController` → `DiagnosisTraceService.getTrace`
|
||||
按 `sessionId`(可选精确 `runId`)拉 run、steps、**toolInvocations**、摘要等。
|
||||
|
||||
**响应中与 RAG 相关的核心块:** `DiagnosisTraceResponse.toolInvocations[]`
|
||||
|
||||
| API 字段 | 来源列 | 读法 |
|
||||
|----------|--------|------|
|
||||
| `toolName` | `tool_name` | 是否为 `lookup_knowledge` |
|
||||
| `retrievalLayer` | `retrieval_layer` | L1 / HARNESS |
|
||||
| `relevanceLevel` | `relevance_level` | RAG 粗相关度(非 evidence_status) |
|
||||
| `l0MatchCount` / `l1MatchCount` | 同名列 | L0/L1 规模提示 |
|
||||
| `truncated` | `is_truncated` | 证据是否被投影截断 |
|
||||
| `outputPreview` | `output_preview` | 一行摘要(level/attempt…) |
|
||||
| `retrievalDetails` | 解析自 `retrieval_details` | **RAG Trace 主阵地** |
|
||||
| `retrievalDetailsRaw` | 原始 JSON 字符串 | 调试 |
|
||||
| `durationMs` / `success` / `errorMessage` | 同名列 | 耗时与成败 |
|
||||
| `inputParams` | 通常仅 tool_call_id、request_bytes | **不含完整 query**(有意) |
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
API["GET /api/diagnosis/{sessionId}/trace"] --> SVC[DiagnosisTraceService]
|
||||
SVC --> ROW[tool_invocation 行]
|
||||
ROW --> T1[列: relevanceLevel, L0/L1 count, layer…]
|
||||
ROW --> T2[retrievalDetails Map]
|
||||
T2 --> D1[search_mode]
|
||||
T2 --> D2[selected_attempt / fallback_reason]
|
||||
T2 --> D3[attempts[] / evidence_keys]
|
||||
T2 --> D4[evidence_status 契约状态]
|
||||
```
|
||||
|
||||
### 4.5 读 Trace 的推荐顺序(排查「这次知识库怎么检的」)
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
S1[找到 toolName=lookup_knowledge 的 invocation] --> S2{success?}
|
||||
S2 -->|否| E[看 errorMessage / evidence_status]
|
||||
S2 -->|是| S3[看 retrievalDetails.search_mode]
|
||||
S3 --> S4[看 selected_attempt + fallback_reason]
|
||||
S4 --> S5[看 attempts[] 每次 candidate_count / top_similarity / usable]
|
||||
S5 --> S6[看 evidence_keys / sources]
|
||||
S6 --> S7[看 relevanceLevel 列]
|
||||
S7 --> S8[需要原文?看 Agent 侧 evidence 或当时 canonical 存储 · 审计默认无 excerpt]
|
||||
```
|
||||
|
||||
| 现象 | 优先看 |
|
||||
|------|--------|
|
||||
| 为何走了 retry | `fallback_reason` + 两次 `attempts` |
|
||||
| 是否 hybrid | `search_mode` |
|
||||
| 滤错域 | `category_filter` + L0 domains |
|
||||
| 相关度档 | 列 `relevanceLevel`(PRECISE/REFERENCE) |
|
||||
| 返回了哪些块 | `evidence_keys` / `sources`(无正文) |
|
||||
| Agent 是否被截断 | `truncated` / `returned_count` |
|
||||
|
||||
### 4.6 diagnosis_trace 事件 vs tool_invocation 行
|
||||
|
||||
| 通道 | 内容 | 用途 |
|
||||
|------|------|------|
|
||||
| `tool_invocation` 行 | RAG 富字段完整摘要 | **主审计/回放** |
|
||||
| `diagnosis_trace` 中 `TOOL_INVOCATION` | tool_call_id、status、字节数、`has_raw_result` 等薄摘要 | 时间线事件,**不含**完整 retrieval_details |
|
||||
|
||||
查 RAG 细节以 **`toolInvocations[].retrievalDetails`** 为准。
|
||||
|
||||
---
|
||||
|
||||
## 5. 与 Agent / Eval 的边界
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph human["人 / 运维 / 评测"]
|
||||
TRACE[Trace API]
|
||||
EVAL[Offline eval]
|
||||
end
|
||||
|
||||
subgraph model["模型"]
|
||||
AGENT[RagToolResult only]
|
||||
end
|
||||
|
||||
TI[(tool_invocation)] --> TRACE
|
||||
FX[fixtures] --> EVAL
|
||||
PROJ[Projector] --> AGENT
|
||||
```
|
||||
|
||||
| 消费者 | 能看到 |
|
||||
|--------|--------|
|
||||
| Agent | evidence + 可选 relevance_level,无 attempt 细节 |
|
||||
| Trace API | 落库摘要:mode/attempt/fallback/keys/level… |
|
||||
| Offline eval | 冻结 fixture 全量 LookupResult(含 trace),与 golden 比对 |
|
||||
|
||||
---
|
||||
|
||||
## 6. 与旧文档差异
|
||||
|
||||
| 旧(archive `retrieval-observability`) | 现 |
|
||||
|----------------------------------------|-----|
|
||||
| `vector-store.mode` 多后端 | `search_mode` dense\|hybrid,单一 V2 store |
|
||||
| sink 理想化未落地 | `RagLookupAuditEnricher` + 列回填 |
|
||||
| `relevance_level` 混用 evidence_status | **列仅 RAG 等级**;契约状态在 details |
|
||||
| 未写清 Trace API 读法 | 本文 §4.4–4.5 |
|
||||
|
||||
---
|
||||
|
||||
## 7. 代码锚点
|
||||
|
||||
| 职责 | 类 / 路径 |
|
||||
|------|-----------|
|
||||
| 内建 Trace | `LookupKnowledgeTool`、`RetrievalTrace`、`LookupResult` |
|
||||
| 投影 | `RagResultProjector`、`RagToolResult` |
|
||||
| 审计事件 | `ToolInvocationAuditEvent`、`ToolBoundary` |
|
||||
| 富化 | `RagLookupAuditEnricher` |
|
||||
| 落库 | `JpaToolInvocationAuditSink`、`ToolInvocation` |
|
||||
| Trace API | `DiagnosisTraceController`、`DiagnosisTraceService`、`DiagnosisTraceResponse.ToolInvocationTrace` |
|
||||
| 离线回归 | `eval/rag-retrieval/`、`scripts/eval_rag_retrieval.py` |
|
||||
|
||||
---
|
||||
|
||||
## 8. 已知限制
|
||||
|
||||
- 持久化 **不存** 完整 query/excerpt(有意);要正文需 Agent 侧证据或其它存储
|
||||
- `dedup_reason` 列可能仍为空
|
||||
- 非 `lookup_knowledge` 工具仍为薄审计
|
||||
- **历史** `tool_invocation` 行可能仍把 evidence_status 写进 `relevance_level`(旧 sink)
|
||||
- `diagnosis_trace` 时间线事件不替代 `retrieval_details`
|
||||
|
||||
---
|
||||
|
||||
## 9. 相关文档
|
||||
|
||||
| 文档 | 内容 |
|
||||
|------|------|
|
||||
| `mvp/architecture/RAG知识检索架构.md` | 检索主架构 |
|
||||
| `docs/RAG-Agent如何读relevance_level.md` | Agent 如何读 level |
|
||||
| `docs/RAG-Hybrid质量分与后处理.md` | quality / 排序闸门 |
|
||||
| `docs/RAG离线评测-基线设计.md` | 离线评测(非运行时 Trace) |
|
||||
| `mvp/architecture/session-trace-lifecycle.md` | 会话/run Trace 总览(若存在) |
|
||||
+55
-7
@@ -168,7 +168,7 @@ sparse_vector -> SPARSE_INVERTED_INDEX + BM25
|
||||
```yaml
|
||||
retrieval:
|
||||
search:
|
||||
mode: hybrid # dense | hybrid
|
||||
mode: hybrid # dense | hybrid(见 6.0 用途约定)
|
||||
hybrid:
|
||||
rrf-k: 60
|
||||
kb-scope: ""
|
||||
@@ -178,7 +178,29 @@ rag:
|
||||
max-chunks-per-document: 2
|
||||
```
|
||||
|
||||
### 6.1 dense
|
||||
### 6.0 模式用途约定(保留双 mode 的原因)
|
||||
|
||||
知识库 **只维护一套** dense + BM25 schema 数据(默认 collection `biz`)。
|
||||
`retrieval.search.mode` 切换的是**同库上的查询算法**,不是两套互斥索引、也不是两套写入路径。
|
||||
|
||||
| 模式 | 定位 | 说明 |
|
||||
|---|---|---|
|
||||
| **hybrid** | **线上主路径 / 默认** | dense ANN + 服务端 BM25 + RRF;`lookup_knowledge` 正式召回只认此模式 |
|
||||
| **dense** | **对照 / 评测 / 排障** | 仅 dense ANN,用于和 hybrid 对比召回效果(命中文档/chunk、排名差异等) |
|
||||
|
||||
约定:
|
||||
|
||||
1. 生产配置保持 `mode: hybrid`;不要把 dense 当成第二套长期并行的线上策略。
|
||||
2. 需要看「去掉 BM25+RRF 后召回差在哪」时,临时切 `mode: dense`,其它参数(`retrieve-k`、`return-n`、category filter、query 集)尽量固定,再切回 hybrid。
|
||||
3. hybrid 入库的数据 **完全适用于** dense-only 查询:每条 chunk 都写了 `vector`;dense 模式只是不使用 `sparse_vector` / BM25 子路。
|
||||
4. 代码里 `@Value` 在配置缺失时的兜底仍可能是 `dense`(历史兼容);**以 `application.yml` 的 hybrid 为准**。若做回归,确认运行配置而不是只看注解默认值。
|
||||
|
||||
不建议的用法:
|
||||
|
||||
- 按请求/按租户在 dense 与 hybrid 之间当产品功能随意切换(当前也无稳定的 per-call mode 覆盖)。
|
||||
- 把 dense 模式的相关度表现直接当成 hybrid 的最终质量结论(hybrid 排序信 RRF,后处理分数仍多 L2 兼容,见下节)。
|
||||
|
||||
### 6.1 dense(对照基线)
|
||||
|
||||
```text
|
||||
query
|
||||
@@ -187,7 +209,9 @@ query
|
||||
-> topK
|
||||
```
|
||||
|
||||
### 6.2 hybrid(当前默认)
|
||||
仅走 `vector` 字段的 L2 ANN。用于基线对比,不作为正式主路径。
|
||||
|
||||
### 6.2 hybrid(当前默认 / 主路径)
|
||||
|
||||
```text
|
||||
query
|
||||
@@ -197,11 +221,18 @@ query
|
||||
-> topK fused hits
|
||||
```
|
||||
|
||||
阈值兼容:
|
||||
说明:hybrid **内部**的 dense 子路是融合的一部分,与配置项 mode=dense(整次检索只跑单路 ANN)不是同一概念。
|
||||
|
||||
- 后处理仍按 dense 兼容 L2 距离做 `normalizeL2`
|
||||
- hybrid 命中若能从并行 dense 结果拿到同 id 的 L2,则回填该分数
|
||||
- 仅 BM25 命中、无 dense 分时,按弱相关处理,避免虚高 `PRECISE`
|
||||
分数与后处理(quality 统一,2026-07-28):
|
||||
|
||||
- 一级 scoreLabel 仅 **dense | hybrid**(旧别名 canonicalize)。
|
||||
- **dense**:score = L2;qualityScore = 1 - clamp(L2)/maxL2Distance。
|
||||
- **hybrid**:返回序 = RRF 序;qualityScore 由 **本轮 rank 线性映射**(不把 RRF 原分当 L2;不做 dense L2 回填覆盖主分;无 m25_only_* 一级 label)。
|
||||
- 后处理:**统一**消费 qualityScore;排序主序 = originalRank;**不做** L0 关键词/domain contains 加分改序(重叠仅可写 hitReasons 解释)。
|
||||
-
|
||||
elevance_level / category 低质 unfiltered retry:只看 top qualityScore 与阈值。
|
||||
- 实现:RetrievalScoreNormalizer、KnowledgeEvidencePostProcessor;详见 OpenSpec
|
||||
ag-quality-score-unify。
|
||||
|
||||
### 6.3 category filter 与降级
|
||||
|
||||
@@ -266,6 +297,18 @@ Agent 只看到有界 `RagToolResult`:
|
||||
|
||||
完整内部结果仍在 `LookupResult` 中,供审计与调试使用。
|
||||
|
||||
### 9.1 Trace 与审计(现行入口)
|
||||
|
||||
请求内 `retrievalTrace` / 落库 `tool_invocation` / Trace API 读法见:
|
||||
|
||||
**[RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)**
|
||||
|
||||
要点:
|
||||
|
||||
- Agent **看不到**完整 retrievalTrace;人通过 `GET /api/diagnosis/{sessionId}/trace` 的 `toolInvocations[].retrievalDetails` 回放。
|
||||
- `relevance_level` 列存 RAG 等级(PRECISE/REFERENCE);`evidence_status` 在 details JSON。
|
||||
- 默认审计不落原始 query 全文与 excerpt 正文。
|
||||
|
||||
## 10. 写入与重建
|
||||
|
||||
### 10.1 日常写入
|
||||
@@ -326,3 +369,8 @@ POST /api/knowledge/rebuild-hybrid?confirm=REBUILD
|
||||
- L0 关键词匹配仍较粗,只作 hint,不作主召回
|
||||
- 尚未做邻块上下文自动扩展、cross-encoder rerank、真 query rewrite
|
||||
- `totalVectors` 统计接口仍可能返回 0,不代表 collection 为空;以 rebuild/init 结果与检索命中为准
|
||||
- `retrieval.search.mode=dense` 仅作召回对照,不是第二套主路径
|
||||
- 同一 hybrid schema 数据可被 dense / hybrid 两种查询复用;从纯旧 dense-only collection 升级必须 rebuild
|
||||
- hybrid 质量闸门优先用 `denseDistance` 绝对 L2;无 dense 时 rank 回退;排序仍跟 RRF
|
||||
- 后处理不再用 L0 关键词 boost 改序;词面信号以库内 BM25+RRF 为准
|
||||
- Trace/审计细节与限制见 [RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)
|
||||
@@ -12,8 +12,10 @@
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 |
|
||||
| [diagnosis-information-gain-stop-architecture.md](diagnosis-information-gain-stop-architecture.md) | 已实施的信息增益评价、Harness 饱和检测、Draft 合同失败降级与证据不足停止设计 |
|
||||
| [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md) | 当前 `lookup_knowledge` 检索:MilvusClientV2 dense+BM25 hybrid、chunk 证据身份、重建运维 |
|
||||
| [RAG知识检索架构.md](RAG知识检索架构.md) | 当前 `lookup_knowledge` 检索:MilvusClientV2 dense+BM25 hybrid、chunk 证据身份、重建运维 |
|
||||
| [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) | RAG Trace / 审计:请求内 retrievalTrace、tool_invocation 富字段、Trace API 读法 |
|
||||
|
||||
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。其中旧 RAG 描述(Spring AI VectorStore 主路径 + Milvus SDK fallback)已被当前 hybrid 实现取代,请以 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md) 为准。
|
||||
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。其中旧 RAG 描述(Spring AI VectorStore 主路径 + Milvus SDK fallback)已被当前 hybrid 实现取代,请以 [RAG知识检索架构.md](RAG知识检索架构.md) 为准。检索可观测与 Trace 以 [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) 为准(勿再依赖 archive 内旧 retrieval-observability)。
|
||||
|
||||
当前普通 Trace 与 Provider reasoning 审计使用独立存储和独立接口。Reasoning 访问控制、保留期限、加密要求以及真实 Provider/V015 验证仍由 ISS-015 跟踪,不能把“数据已分表”理解为“治理已经完成”。
|
||||
当前普通 Trace 与 LLM 步骤审计(`agent_reasoning_audit`:`reasoning_content` + `assistant_text`)使用独立存储和独立接口。
|
||||
DeepSeek thinking 捕获路径与 V015–V017 字段已 live 验证(2026-07-28)。Reasoning 访问控制、保留期限、加密要求仍由 ISS-015 跟踪,不能把“数据已分表 + 能抓到 thinking”理解为“治理已经完成”。
|
||||
|
||||
@@ -56,14 +56,36 @@ sequenceDiagram
|
||||
|
||||
PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。
|
||||
|
||||
## 5. Provider Reasoning 审计
|
||||
## 5. Provider Reasoning 与 Assistant 正文审计
|
||||
|
||||
`HarnessAgentAuditHook` 在每次模型步骤结束后检查 `AssistantMessage` metadata。当前识别 `reasoning_content`、`reasoningContent`、`reasoning` 和 `thinking`,但只接受 Provider 实际返回的非空文本:
|
||||
`HarnessAgentAuditHook` 在每次模型步骤结束后写入独立表 `agent_reasoning_audit`,**同时**尝试捕获:
|
||||
|
||||
- 有内容时写入 `agent_reasoning_audit`,单条最多保留 32000 个字符,并记录 UTF-8 `content_bytes`。
|
||||
- 无内容时写入 `reasoning_available=false`、`reasoning_content=NULL`、`content_bytes=0`,不得根据最终回答反推或生成 reasoning。
|
||||
- `agent_step.thought` 始终为空;步骤表只记录 message count、roles、是否有文本、Tool names、reasoning availability 和字节数等 metadata。
|
||||
- 普通 `diagnosis_trace_event` 的 `AGENT_MODEL_STEP` 只记录 reasoning availability/bytes,不保存 reasoning 原文。
|
||||
- Reasoning 只用于受限审计,不进入 Agent 后续上下文,不参与 EvidenceGuard、SemanticGuard 或 Release Policy 的事实判断。
|
||||
| 列 | 含义 |
|
||||
|---|---|
|
||||
| `reasoning_content` | Provider thinking / CoT |
|
||||
| `assistant_text` | 本步 assistant 可见正文,和/或 tool-call **计划**(不含 tool 结果) |
|
||||
| `content_source` | `PROVIDER_REASONING+ASSISTANT_TEXT` 等组合标记 |
|
||||
| `content_bytes` | 两列截断后 UTF-8 字节合计 |
|
||||
|
||||
当前查询隔离已经实现,完整访问治理和真实 Provider 行为验证仍属于 ISS-015。
|
||||
### 捕获路径(DeepSeek 生产)
|
||||
|
||||
当前 Chat 为 Spring AI 原生 `DeepSeekChatModel`(`deepseek-v4-flash`)。API 返回的 `message.reasoning_content` 被映射到 **`DeepSeekAssistantMessage.getReasoningContent()`**,而不是普通 `AssistantMessage.metadata`。Hook 优先读该专用字段,再反射 `getReasoningContent()`,最后才回退 metadata 键(`reasoning_content` / `thinking` 等)。
|
||||
|
||||
只接受 Provider 实际返回的非空文本;不得根据最终回答反推或生成伪 reasoning。
|
||||
|
||||
### 边界
|
||||
|
||||
- 单字段最多保留 32000 字符;tool **结果**只在 `tool_invocation`。
|
||||
- `agent_step.model_*` 仍为有界 metadata(含 `reasoning_available` / bytes / `content_source`)。
|
||||
- `agent_step.thought` 为兼容镜像:优先 reasoning,否则 assistant 正文;完整双字段以 `agent_reasoning_audit` 为准。
|
||||
- 普通 Timeline 的 `AGENT_MODEL_STEP` 不保存 reasoning/assistant 原文。
|
||||
- Reasoning / assistant 审计原文只用于受限审计 API,不进入 Agent 后续上下文,不参与 EvidenceGuard、SemanticGuard 或 Release Policy 的事实判断。
|
||||
|
||||
### 运行级结论读出
|
||||
|
||||
Run 结束时 `JpaChatRunStore` 从安全发布 JSON 提取 `diagnosis_run.conclusion`(与 `query` 并列),便于 Trace/DB 直接读结论;它不是 Provider thinking。
|
||||
|
||||
### 验证与治理
|
||||
|
||||
- **已 live 验证(2026-07-28)**:DeepSeek thinking 模式下 `reasoning_available=true`,且 reasoning 与 assistant 可同时非空。
|
||||
- 查询隔离已实现;访问控制、保留期限、加密等完整治理仍属 ISS-015。
|
||||
|
||||
@@ -7,7 +7,7 @@
|
||||
|
||||
SuperBizAgent 是面向故障诊断的可追踪 Agent 应用。当前系统只保留一个拥有 Tool loop 的 `Diagnosis Agent`;Harness 负责确定性的预算、取消、工具边界、证据验真、语义审查和安全发布。
|
||||
|
||||
知识检索当前为显式 `lookup_knowledge` 工具 + 单一 MilvusClientV2 后端(dense / dense+BM25 hybrid)。详细链路见 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md)。
|
||||
知识检索当前为显式 `lookup_knowledge` 工具 + 单一 MilvusClientV2 后端(dense / dense+BM25 hybrid)。详细链路见 [RAG知识检索架构.md](RAG知识检索架构.md)。
|
||||
|
||||
## 2. 分层
|
||||
|
||||
@@ -86,7 +86,7 @@ Agent
|
||||
- 证据按 chunk 级 `evidenceKey` 去重;Agent 侧 `document_id` 为 chunk 级身份。
|
||||
- 知识库全量重建:`python scripts/rebuild_hybrid_knowledge.py --confirm REBUILD`,默认操作 collection `biz`。
|
||||
|
||||
完整 schema、模式、重建与历史差异见 [rag-knowledge-retrieval-architecture.md](rag-knowledge-retrieval-architecture.md)。
|
||||
完整 schema、模式、重建与历史差异见 [RAG知识检索架构.md](RAG知识检索架构.md)。
|
||||
|
||||
## 5. Trace 与持久化
|
||||
|
||||
@@ -100,20 +100,20 @@ chat_session(sessionId)
|
||||
```
|
||||
|
||||
- `chat_session` 是 JPA Run 目录与多轮 metadata,不保存完整对话历史。
|
||||
- `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果和预算汇总真理源。
|
||||
- `agent_step` 只保存模型步骤 metadata,不保存 Prompt、消息正文、模型正文或 Thought。
|
||||
- `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果、预算汇总真理源;另含与 `query` 并列的提取字段 `conclusion`(业务结论读出,非 thinking)。
|
||||
- `agent_step` 保存模型步骤有界 metadata;`thought` 可为 reasoning/assistant 的兼容镜像,完整双字段不在此表。
|
||||
- `tool_invocation` 只保存 Tool durable audit metadata;完整调用由 Redis canonical store 短期保存。
|
||||
- `diagnosis_trace_event` 是追加式统一 Timeline,记录 Run、Routing、Agent、Tool、Evidence、Semantic 和 Release 生命周期事件;`details` 只能保存有界安全 metadata。
|
||||
- `agent_reasoning_audit` 与普通 Trace 分表,只保存 Provider 实际返回的 reasoning 或明确的 unavailable 记录;reasoning 不属于事实证据。
|
||||
- `agent_reasoning_audit` 与普通 Trace 分表,按模型步骤保存 Provider `reasoning_content` 与 `assistant_text`(及 `content_source`),或明确的 unavailable;二者均不属于事实证据。
|
||||
|
||||
## 6. 公开 API
|
||||
|
||||
当前诊断执行入口只有 `POST /api/chat`。诊断审计读取分为:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId={runId}`:普通 Trace,返回 Run、步骤、Tool metadata 和统一 Timeline,不返回 reasoning 原文。
|
||||
- `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`:独立 reasoning 审计读取,`runId` 必填并校验其属于 path `sessionId`。
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId={runId}`:普通 Trace,返回 Run(含 `query`/`conclusion`)、步骤、Tool metadata 和统一 Timeline,**不**返回 reasoning/assistant 原文。
|
||||
- `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`:独立 LLM 步骤审计读取(`reasoningContent` + `assistantText`),`runId` 必填并校验其属于 path `sessionId`。
|
||||
|
||||
Reasoning endpoint 是敏感审计面,不属于普通业务 API。当前已完成数据和查询隔离;认证授权、保留期限、加密要求及真实 Provider 验证仍由 ISS-015 收敛。Feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。
|
||||
Reasoning endpoint 是敏感审计面,不属于普通业务 API。数据和查询隔离、DeepSeek thinking 捕获路径已 live 验证;认证授权、保留期限、加密要求仍由 ISS-015 收敛。Feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。
|
||||
|
||||
与知识检索相关的独立 API:
|
||||
|
||||
@@ -124,9 +124,9 @@ Reasoning endpoint 是敏感审计面,不属于普通业务 API。当前已完
|
||||
|
||||
## 7. 安全边界
|
||||
|
||||
- 普通 SSE、Trace、Evidence Snapshot、业务结果和应用日志不输出或保存 reasoning 原文。
|
||||
- 仅当 Provider 在模型 metadata 中实际返回 reasoning 时,审计 Hook 才将其截断后写入独立表;Provider 未返回时不得伪造。
|
||||
- Reasoning 不能作为事实证据,也不能绕过 EvidenceGuard 或 SemanticGuard。
|
||||
- 普通 SSE、普通 Trace 的 steps/timeline、Evidence Snapshot、业务结果和应用日志不输出或保存 reasoning / assistant 审计原文。
|
||||
- 审计 Hook 仅当 Provider **实际返回** thinking(DeepSeek:`DeepSeekAssistantMessage.reasoningContent`;其它:metadata 键)时写入 `reasoning_content`;未返回时不得伪造。`assistant_text` 来自本步 assistant 可见输出与 tool-call 计划,不含 tool 结果。
|
||||
- Reasoning / assistant 审计原文不能作为事实证据,也不能绕过 EvidenceGuard 或 SemanticGuard。
|
||||
- 不向 Agent 暴露 Redis、canonical key、完整 Tool 请求/响应或数据库凭据。
|
||||
- EvidenceGuard 只接受当前 Run 的 READY canonical invocation。
|
||||
- SemanticGuard 无 Tool、无记忆、无回调主 Agent 能力。
|
||||
|
||||
@@ -37,7 +37,8 @@ disconnect、timeout 与 send failure 通过同一个 `ChatRunControl` 请求取
|
||||
|
||||
Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当前 Run EvidenceGuard 验真。
|
||||
|
||||
普通 Trace 不读取 `agent_reasoning_audit`,也不返回 reasoning 原文。Agent 模型步骤和 Timeline 只暴露 `reasoning_available`、`reasoning_bytes` 等有界 metadata。
|
||||
普通 Trace 不读取 `agent_reasoning_audit` 的 **原文**。Agent 模型步骤和 Timeline 只暴露 `reasoning_available`、`reasoning_bytes` / `assistant_bytes`、`content_source` 等有界 metadata。
|
||||
普通 Trace 的 `run` 可返回 `query` 与提取后的 `conclusion`(业务结论读出,非 thinking)。
|
||||
|
||||
## 4. PreviousTurn
|
||||
|
||||
@@ -45,11 +46,11 @@ Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当
|
||||
|
||||
## 5. Reasoning 审计查询
|
||||
|
||||
`GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}` 提供独立 reasoning 审计读取:
|
||||
`GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}` 提供独立 LLM 步骤审计读取:
|
||||
|
||||
- `runId` 必填,服务端先验证 Run 存在且属于 path `sessionId`,禁止跨 Session 串读。
|
||||
- 结果按 `step_index` 返回 Agent、reasoning availability、受限原文、字节数和创建时间。
|
||||
- Provider 未返回 reasoning 时仍保留 unavailable 记录,以区分“没有返回”与“审计遗漏”。
|
||||
- Reasoning 数据不回流到 PreviousTurn,不进入普通 Trace、SSE、Evidence Snapshot 或发布结果。
|
||||
- 结果按 `step_index` 返回:`reasoningAvailable`、`reasoningContent`、`assistantText`、`contentSource`、`contentBytes`、创建时间等。
|
||||
- Provider 未返回 reasoning 时仍保留记录(`reasoningAvailable=false`,可仍有 `assistantText`),以区分“没有 thinking”与“审计遗漏”。
|
||||
- Reasoning / assistant 审计原文不回流到 PreviousTurn,不进入普通 Trace 的 steps/timeline 原文、SSE、Evidence Snapshot 或发布结果。
|
||||
|
||||
该端点属于敏感审计面。当前完成了分表、独立查询和归属校验;身份认证、权限模型、保留期限、加密及真实 Provider/V015 验证尚未完成,由 ISS-015 阶段 3 收敛。
|
||||
该端点属于敏感审计面。分表、独立查询、归属校验,以及 DeepSeek 真实 thinking 捕获路径(`DeepSeekAssistantMessage.reasoningContent`)与 V015–V017 迁移已验证;身份认证、权限模型、保留期限、加密仍由 ISS-015 阶段 3 收敛。
|
||||
|
||||
@@ -2,11 +2,14 @@
|
||||
|
||||
- [ ] SSE metadata 的 `session_id`、`run_id` 非空且与数据库完全一致。
|
||||
- [ ] `diagnosis_run.intent=DIAGNOSIS`,status/release_outcome 与 done outcome 一致。
|
||||
- [ ] `diagnosis_run.conclusion` 与发布结论一致(SUCCESS 时非空短结论;FALLBACK 可为 type/message 摘要)。
|
||||
- [ ] `agent_step.agent_name` 只出现 `diagnosis_agent`。
|
||||
- [ ] AgentStep model_input/model_output 只含 metadata,thought 为空。
|
||||
- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内。
|
||||
- [ ] ToolInvocation input/output/retrieval details 不含 SQL、日志 query、raw response 或 evidence body。
|
||||
- [ ] AgentStep `model_input`/`model_output` 只含有界 metadata(可含 `reasoning_available`/`content_source`/bytes);**不含** reasoning/assistant 原文。
|
||||
- [ ] `agent_step.thought` 若非空,应为 reasoning 或 assistant 的兼容镜像,不得替代 `/trace/reasoning` 双字段验收。
|
||||
- [ ] `/trace/reasoning`:DeepSeek thinking 开启时应有 `reasoningAvailable=true` 且 `reasoningContent` 非空;`assistantText` 有正文和/或 tool-call 计划;`contentSource` 合理;**无** tool 结果体。
|
||||
- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内;`step_id` 能挂到对应 `agent_step.id`。
|
||||
- [ ] ToolInvocation input/output/retrieval details 不含 SQL 全文、raw response 或 evidence body 泄露(query 等允许的有界字段除外)。
|
||||
- [ ] Run 的模型/Tool/Token/字节预算均未超过集中配置。
|
||||
- [ ] content 只出现一次并来自 Release Policy;failure 与 content 互斥。
|
||||
- [ ] `logs/application.log` 不含 Prompt、Thought、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。
|
||||
- [ ] query_logs 标记为 Mock;query_mysql 只使用隔离只读 datasource contract。
|
||||
- [ ] `logs/application.log` 不含 Prompt、完整 reasoning 原文、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。
|
||||
- [ ] query_logs 标记为 Mock;query_mysql 仅在配置了隔离 datasource 时注册。
|
||||
|
||||
@@ -40,8 +40,8 @@ Fallback 必须在不暴露 Prompt、原始 Draft、原始 Tool 载荷和内部
|
||||
|---|---|---|---|
|
||||
| 阶段 1:Diagnosis Agent 硬停止策略 | 已完成 | ISS-016 已实现 Tool Scope 归一化与去重、`GAINED/NO_GAIN` 信息增益、连续 `NO_GAIN` 饱和停止、连续协议错误 `PROGRESS_PROTOCOL_VIOLATED` 受控停止、可修正协议 observation、受控 Fallback、回归测试和真实 E2E | 无;阶段 1 停止策略已收口 |
|
||||
| 阶段 2:Evidence Repair Schema | 未完成 | Repair 仍保持无 Tool、有限重试、失败后安全降级;Diagnosis Agent 最终输出已使用真实 `DiagnosisDraft` Schema | Evidence Repair 自身尚未注入真实 `DiagnosisDraft` JSON Schema,仍主要依赖 Prompt 文本约束和严格解析 |
|
||||
| 阶段 3:Reasoning 审计验证与治理 | 部分完成 | V015、独立 `agent_reasoning_audit`、审计 Hook、`reasoning_available=false`、受限查询接口和 `sessionId + runId` 归属校验已实现 | 真实 Provider reasoning metadata 行为、V015 真实迁移验收、访问控制、保留期限和加密要求尚未收敛 |
|
||||
| 阶段 4:Fallback 信息质量与最终 E2E | 大部分完成 | 已实现有界 `observed_facts`、`validation_issues`、证据范围/缺口和差异化 Fallback;已完成 Diagnosis SUCCESS、信息不足 FALLBACK、Trace/Tool/Token 对账 E2E | reasoning unavailable 场景及 Reasoning Audit 跨表精确核验尚未完成;全部失败路径仍需最终综合验收 |
|
||||
| 阶段 3:Reasoning 审计验证与治理 | 大部分完成 | V015–V017;`reasoning_content` + `assistant_text` + `content_source`;Hook 优先读 `DeepSeekAssistantMessage.getReasoningContent()`;受限查询返回双字段;`sessionId + runId` 归属校验;2026-07-28 DeepSeek live E2E 两步均 `reasoning_available=true` 且双字段非空;V016 `diagnosis_run.conclusion` | 访问控制、保留期限、加密;非 DeepSeek Provider 回归;reasoning unavailable 专项场景归档 |
|
||||
| 阶段 4:Fallback 信息质量与最终 E2E | 大部分完成 | 已实现有界 `observed_facts`、`validation_issues`、证据范围/缺口和差异化 Fallback;已完成 Diagnosis SUCCESS、信息不足 FALLBACK、Trace/Tool/Token 对账 E2E;SUCCESS 路径已跨表核验 reasoning 双字段与 run.conclusion | reasoning unavailable 场景归档;全部失败路径仍需最终综合验收 |
|
||||
|
||||
本表是当前进度事实,下面各阶段条目仍保留为完整目标。ISS-016 完成的是阶段 1 和阶段 4 的主体能力,并新增模型 Token 与 Tool 拒绝审计;它不替代阶段 2 的 Repair Schema,也不代表阶段 3 的 Reasoning 治理已经完成。
|
||||
|
||||
@@ -70,18 +70,20 @@ Fallback 必须在不暴露 Prompt、原始 Draft、原始 Tool 载荷和内部
|
||||
- 保持 Repair 无 Tool、最多一次、失败即安全 Fallback 的现有边界。
|
||||
- 覆盖合法修复、Schema 非法、解析失败和二次 EvidenceGuard 失败。
|
||||
|
||||
### 阶段 3:Reasoning 审计验证与治理(部分完成)
|
||||
### 阶段 3:Reasoning 审计验证与治理(大部分完成)
|
||||
|
||||
- 使用真实 Provider 验证 reasoning metadata 的键和返回行为。
|
||||
- Provider 不返回 reasoning 时写入 `reasoning_available=false`,不得伪造内容。
|
||||
- 验证 V015 数据库迁移和 `sessionId + runId` 精确 reasoning 查询。
|
||||
- 明确 reasoning 审计接口的访问控制、保留期限和加密要求。
|
||||
- [x] 使用真实 DeepSeek Provider 验证 thinking 返回路径:专用字段 `DeepSeekAssistantMessage.reasoningContent`(非 metadata)。
|
||||
- [x] 同表同时落库 `assistant_text`(正文 / tool-call 计划)与 `content_source`(V017)。
|
||||
- [x] Provider 不返回 reasoning 时写入 `reasoning_available=false`,不得伪造内容(单测覆盖;live unavailable 场景可再归档)。
|
||||
- [x] 验证 V015–V017 迁移与 `sessionId + runId` 精确 reasoning 查询(含双字段 API)。
|
||||
- [ ] 明确 reasoning 审计接口的访问控制、保留期限和加密要求。
|
||||
|
||||
### 阶段 4:Fallback 信息质量与最终 E2E(大部分完成)
|
||||
|
||||
- 工具成功但无法构造 verified snapshot 时,返回有界 `observed_facts` 和 `validation_issues`。
|
||||
- 确保普通 Trace 只记录 `reasoning_available/reasoning_bytes`,不返回 reasoning 原文。
|
||||
- 运行至少一个 Diagnosis SUCCESS、一个信息化 FALLBACK 和一个 reasoning unavailable 的真实 E2E。
|
||||
- 确保普通 Trace 只记录 `reasoning_available/reasoning_bytes` 等有界 metadata,不返回 reasoning/assistant 原文;`run.conclusion` 可为业务结论读出。
|
||||
- [x] Diagnosis SUCCESS 真实 E2E:SSE + Trace + Reasoning 双字段 + ToolInvocation + `diagnosis_run.conclusion` 对账(2026-07-28)。
|
||||
- [ ] 再归档一个 reasoning unavailable 与信息化 FALLBACK 的跨表精确核验样本。
|
||||
- 按 exact `sessionId + runId` 核对 SSE、Trace、Reasoning Audit、ToolInvocation 和 Run 终态。
|
||||
|
||||
## 6. 验收标准
|
||||
|
||||
@@ -1,13 +1,18 @@
|
||||
# Agent 推理审计表:agent_reasoning_audit
|
||||
|
||||
**状态**:当前独立敏感审计表;治理待 ISS-015 收敛
|
||||
**来源**:`V015__create_agent_reasoning_audit.sql`、`AgentReasoningAudit`
|
||||
**来源**:`V015__create_agent_reasoning_audit.sql`、`V017__add_content_source_to_agent_reasoning_audit.sql`、`AgentReasoningAudit`、`HarnessAgentAuditHook`
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_reasoning_audit` 保存模型 Provider 在 Agent 步骤 metadata 中实际返回的 reasoning 内容。它与普通 Trace、AgentStep、Evidence Snapshot 和业务发布结果物理分离,不能作为事实证据或诊断结论来源。
|
||||
`agent_reasoning_audit` 按 **Diagnosis Agent 每个模型步骤** 保存两类 LLM 面向文本,并与普通 Trace、AgentStep、Evidence Snapshot、业务发布结果物理分离:
|
||||
|
||||
Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没有返回”误判为“审计链路漏写”。系统不得从最终回答、Tool 调用或其他字段生成伪 reasoning。
|
||||
1. **Provider reasoning / thinking**(`reasoning_content`)
|
||||
2. **Assistant 可见正文**(`assistant_text`:最终 prose 和/或 tool-call **计划**)
|
||||
|
||||
它不能作为事实证据或诊断结论来源。Tool **结果**载荷不进本表(见 `tool_invocation`)。
|
||||
|
||||
Provider 未返回 reasoning 时仍写入记录(`reasoning_available=false`),防止把“没有返回”误判为“审计链路漏写”。系统不得从最终回答、Tool 调用或其他字段生成伪 reasoning。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -19,8 +24,10 @@ Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没
|
||||
| `step_index` | INT | 是 | Agent 模型步骤序号 |
|
||||
| `agent_name` | VARCHAR(64) | 是 | Agent 身份,当前为 `diagnosis_agent` |
|
||||
| `reasoning_available` | BOOLEAN | 是 | Provider 是否实际返回非空 reasoning |
|
||||
| `reasoning_content` | LONGTEXT | 否 | Provider reasoning 原文;Hook 当前最多保留 32000 个字符 |
|
||||
| `content_bytes` | INT | 是 | 截断后 reasoning 的 UTF-8 字节数;unavailable 时为 0 |
|
||||
| `reasoning_content` | LONGTEXT | 否 | Provider CoT / thinking 原文;Hook 单字段最多保留 32000 字符 |
|
||||
| `assistant_text` | LONGTEXT | 否 | 本步 assistant 可见正文;若有 tool_call,可附带 `tool_calls:` 计划预览(**不含** tool 结果) |
|
||||
| `content_source` | VARCHAR(64) | 否 | 本步写入摘要:`PROVIDER_REASONING+ASSISTANT_TEXT` / `PROVIDER_REASONING` / `ASSISTANT_TEXT` / `TOOL_CALL_PLAN` / `NONE` |
|
||||
| `content_bytes` | INT | 是 | 截断后 `reasoning_content` + `assistant_text` 的 UTF-8 字节合计 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间,默认当前时间 |
|
||||
|
||||
## 索引
|
||||
@@ -36,15 +43,36 @@ Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没
|
||||
- `session_id` 逻辑关联 `chat_session.session_id`。
|
||||
- 通过 `session_id + run_id + step_index` 与 `agent_step` 逻辑对应,不建立数据库外键。
|
||||
|
||||
## 写入规则
|
||||
## 写入规则(HarnessAgentAuditHook)
|
||||
|
||||
- Provider 返回非空 reasoning:`reasoning_available=true`,保存截断后的原文及实际 UTF-8 字节数。
|
||||
- Provider 未返回 reasoning:`reasoning_available=false`,`reasoning_content=NULL`,`content_bytes=0`。
|
||||
- `agent_step.thought` 继续保持为空;普通 Trace 仅保留 availability/bytes metadata。
|
||||
- Reasoning 不进入 SSE、PreviousTurn、Evidence Snapshot、发布结果或应用日志。
|
||||
### 捕获来源(按优先级)
|
||||
|
||||
1. **DeepSeek 生产主路径**:`DeepSeekAssistantMessage.getReasoningContent()`
|
||||
(Spring AI 将 API 的 `message.reasoning_content` 映射到该专用字段,**不是** `AssistantMessage.metadata`)
|
||||
2. 反射 `getReasoningContent()`(兼容序列化/子类)
|
||||
3. metadata 键兜底:`reasoning_content` / `reasoningContent` / `reasoning` / `thinking` / `reasoning_text` / `reasoningText`(单测与其它 Provider)
|
||||
|
||||
`assistant_text` 来自 `AssistantMessage.getText()`;若本步有 tool_calls,追加有界的 `tool_calls:` 名称与 args 预览。
|
||||
|
||||
### 落库语义
|
||||
|
||||
- 有 Provider reasoning:`reasoning_available=true`,写入截断后的 `reasoning_content`。
|
||||
- 无 Provider reasoning:`reasoning_available=false`,`reasoning_content=NULL`;仍可写入 `assistant_text`。
|
||||
- `content_source` 反映本步实际写入组合;两者皆空则为 `NONE`。
|
||||
- `content_bytes` 为两列截断后文本 UTF-8 字节之和。
|
||||
- **禁止**把 tool 执行结果、raw Tool response、用户 Prompt 全文写入本表。
|
||||
|
||||
### 与其它表的边界
|
||||
|
||||
- 普通 Trace / Timeline 只暴露 `reasoning_available`、`reasoning_bytes` / `assistant_bytes`、`content_source` 等有界 metadata,不返回 reasoning/assistant 原文。
|
||||
- `agent_step.model_input` / `model_output` 仍为有界 metadata。
|
||||
- `agent_step.thought` 可作为兼容镜像:优先 Provider reasoning,否则 assistant 正文;完整双字段以本表为准。
|
||||
- Reasoning / assistant 审计原文不进入 SSE、PreviousTurn、Evidence Snapshot、发布结果或应用日志。
|
||||
|
||||
## 查询与治理
|
||||
|
||||
当前独立接口为 `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`。`runId` 必填,服务端校验其属于 path `sessionId`,并按 `step_index` 返回记录。
|
||||
当前独立接口为 `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`。`runId` 必填,服务端校验其属于 path `sessionId`,并按 `step_index` 返回记录(含 `reasoningContent`、`assistantText`、`contentSource`)。
|
||||
|
||||
分表和归属校验已经实现,但不能等同于完整安全治理。身份认证、角色授权、保留/删除期限、静态与传输加密以及真实 Provider/V015 验证仍由 ISS-015 阶段 3 跟踪;治理完成前不应将该接口暴露给普通业务用户。
|
||||
**已验证(2026-07-28 live E2E,DeepSeek `deepseek-v4-flash`)**:thinking 模式下两步模型调用均可得到 `reasoning_available=true` 且 `content_source=PROVIDER_REASONING+ASSISTANT_TEXT`。
|
||||
|
||||
分表和归属校验已经实现,但不能等同于完整安全治理。身份认证、角色授权、保留/删除期限、静态与传输加密仍由 ISS-015 跟踪;治理完成前不应将该接口暴露给普通业务用户。
|
||||
|
||||
@@ -1,11 +1,13 @@
|
||||
# Agent 步骤表:agent_step
|
||||
|
||||
**状态**:当前表
|
||||
**来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep`
|
||||
**来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep`、`HarnessAgentAuditHook`
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。Provider reasoning 使用独立 `agent_reasoning_audit` 表,不复用历史 `thought` 字段。
|
||||
`agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存完整 Prompt、完整消息列表、Tool 结果体。
|
||||
|
||||
**Provider reasoning 与 assistant 正文的完整双字段** 在独立表 `agent_reasoning_audit`;本表只保留步骤级摘要与兼容镜像。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -17,8 +19,8 @@
|
||||
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||
| `agent_name` | VARCHAR(32) | 是 | 当前 Harness 写入固定为 `diagnosis_agent` |
|
||||
| `model_input` | TEXT | 否 | JSON metadata,仅包含 message count 与 roles |
|
||||
| `model_output` | TEXT | 否 | JSON metadata,仅包含 text presence 与 Tool names |
|
||||
| `thought` | TEXT | 否 | 当前 Harness 必须写空;字段仅保留历史兼容 |
|
||||
| `model_output` | TEXT | 否 | JSON metadata:`has_text`、`tool_names`、`reasoning_available`、`reasoning_bytes`、`assistant_bytes`、`content_source` |
|
||||
| `thought` | TEXT | 否 | 兼容镜像:优先 Provider reasoning,否则 assistant 正文;**完整双字段以 `agent_reasoning_audit` 为准** |
|
||||
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
|
||||
| `duration_ms` | INT | 否 | 本步骤耗时 |
|
||||
| `token_count` | INT | 否 | 本步骤 Token 消耗 |
|
||||
@@ -36,12 +38,14 @@
|
||||
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||
- `tool_invocation.step_id` 应关联当前模型步骤的 `agent_step.id`(由 `AgentStepAuditTracker` 在 beforeModel 绑定)。
|
||||
- `agent_reasoning_audit` 通过相同的 `session_id + run_id + step_index` 逻辑定位模型步骤,不建立数据库外键。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
|
||||
- 新 Trace 和验收读路径必须按 exact `run_id` 取数,避免同一 `sessionId` 多次运行混入。
|
||||
- 当前 Run 若出现 `diagnosis_agent` 之外的新写入,或 `thought` 非空,视为审计边界违规。
|
||||
- `model_output` 可保存 `has_text`、`tool_names`、`reasoning_available` 和 `reasoning_bytes` 等有界 metadata,不得保存 reasoning 原文。
|
||||
- 当前 Run 若出现 `diagnosis_agent` 之外的新写入,视为审计边界违规。
|
||||
- `model_output` **不得**保存 reasoning/assistant 原文;只允许有界 metadata。
|
||||
- `thought` 非空 **不再**视为违规:它是受限镜像,便于旧读路径一眼看到“本步在想什么/说什么”;敏感完整审计仍走 `/trace/reasoning`。
|
||||
- 需要同时查看 thinking 与 assistant 正文时,必须查 `agent_reasoning_audit` 或 reasoning Trace API。
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# 诊断运行表:diagnosis_run
|
||||
|
||||
**状态**:当前诊断运行主表
|
||||
**来源**:`V011__add_session_run_isolation.sql`、`DiagnosisRun`
|
||||
**来源**:`V011__add_session_run_isolation.sql`、`V012__add_chat_release_contract.sql`、`V016__add_conclusion_to_diagnosis_run.sql`、`DiagnosisRun`
|
||||
|
||||
## 定位
|
||||
|
||||
@@ -15,9 +15,10 @@
|
||||
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
|
||||
| `query` | TEXT | 是 | 本次 Chat 用户问题 |
|
||||
| `conclusion` | TEXT | 否 | 从安全发布内容提取的短结论,与 `query` 并列便于读出;非 reasoning 原文 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | 历史兼容字段;当前公开执行统一来自 Chat Harness |
|
||||
| `answer` | LONGTEXT | 否 | 安全发布的最终文本兼容字段 |
|
||||
| `answer` | LONGTEXT | 否 | 安全发布的最终内容 JSON/文本兼容字段 |
|
||||
| `intent` | VARCHAR(32) | 否 | `SYSTEM_CHAT`、`KNOWLEDGE_QUERY` 或 `DIAGNOSIS` |
|
||||
| `release_outcome` | VARCHAR(16) | 否 | `SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
|
||||
| `published_result` | JSON | 否 | Release Policy 允许发布的结构化安全结果 |
|
||||
@@ -55,4 +56,5 @@
|
||||
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
|
||||
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
|
||||
- 当前诊断发布结果以 `release_outcome + published_result` 为准,不得从旧 self-evaluation 推断 Release Policy 结果。
|
||||
- `diagnosis_run` 不保存 reasoning 原文;普通 Run/Trace 查询也不得通过聚合将其带出。
|
||||
- `conclusion` 由 `RunConclusionExtractor` 在 `finish` 时从 `answer`(安全发布 JSON)提取:优先 `report.conclusion.text`,否则 fallback 的 `type + message` 等;最多约 4000 字符。它是 **业务结论读出字段**,不是 Provider thinking。
|
||||
- `diagnosis_run` 不保存 reasoning 原文;普通 Run/Trace 查询也不得通过聚合 `agent_reasoning_audit` 将其带出。Trace 的 `run.conclusion` 可以返回上述短结论。
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
## Context
|
||||
|
||||
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
|
||||
- Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1.
|
||||
- Delivery baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1.
|
||||
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
|
||||
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
|
||||
As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.
|
||||
|
||||
This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checklist.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
|
||||
This change is **Delivery 1** from `docs/Milvus-Hybrid接入清单.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
|
||||
|
||||
## What Changes
|
||||
|
||||
@@ -47,7 +47,7 @@ This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checkl
|
||||
- Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
|
||||
- Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
|
||||
- Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids
|
||||
- Docs baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1
|
||||
- Docs baseline: `docs/Milvus-Hybrid接入清单.md` §1.1 Delivery 1
|
||||
|
||||
## Risks
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
Committed OpenSpec
|
||||
change: rag-eval-hybrid-baseline
|
||||
committed_at: 2026-07-28
|
||||
scale: standard-lean
|
||||
scope: knife-1-only
|
||||
acceptance: wiring-required; fixture-refresh-best-effort
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-28
|
||||
@@ -0,0 +1,13 @@
|
||||
# Brief: rag-eval-hybrid-baseline
|
||||
|
||||
## Background
|
||||
|
||||
Offline RAG eval structure is correct but generator/docs/fixtures predate hybrid search mode.
|
||||
|
||||
## Goals
|
||||
|
||||
Knife 1 only: `search.mode` wiring, fixture meta, README, best-effort fixture refresh.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Dual dense/hybrid fixture trees; golden mustNot/chunk/level hard gates; new frameworks.
|
||||
@@ -0,0 +1,111 @@
|
||||
# Decisions — rag-eval-hybrid-baseline
|
||||
|
||||
## sm-flow meta
|
||||
|
||||
- **Checkpoint**: Discover (in progress)
|
||||
- **Scale**: standard (lean) — eval harness alignment, multi-file, low prod risk
|
||||
- **Capability**: sm-flow built-in; openspec CLI `new change`; grill fallback (no external grill-with-docs runner)
|
||||
- **Slug**: `rag-eval-hybrid-baseline`
|
||||
- **Path**: `openspec/changes/rag-eval-hybrid-baseline/`
|
||||
|
||||
## Clarify summary
|
||||
|
||||
| Item | Content |
|
||||
|------|---------|
|
||||
| Problem | Offline eval model OK but wiring/fixtures/docs pre-hybrid; cannot gate current main path |
|
||||
| Goal | Knife-1: hybrid generator + meta + docs (+ refresh). Optional knife-2: dense/hybrid dual fixtures |
|
||||
| Touch | `scripts/generate_rag_lookup_snapshots.ps1`, snapshot test, eval README, fixtures/baseline, maybe `eval_rag_retrieval.py` |
|
||||
| Non-goals | New framework, LLM judge, prod retrieval redesign |
|
||||
|
||||
## Context summary
|
||||
|
||||
| Source | Conclusion | Into OpenSpec |
|
||||
|--------|------------|---------------|
|
||||
| Conversation design | Golden×fixture×key fields; not full JSON diff | Yes |
|
||||
| Current eval audit | ~70% aligned; dead spring mode; old fixtures | Yes |
|
||||
| `rag-quality-score-unify` | hybrid quality rank-based; don't hard-lock PRECISE | Yes |
|
||||
| `eval/rag-retrieval/README` | seed + kb_scope good; generator props stale | Yes |
|
||||
| Generator ps1 | `VectorStoreMode=spring` → must replace with search.mode | Yes |
|
||||
|
||||
**index**: hit rag-quality-score-unify / bm25-hybrid / chunk-identity archives.
|
||||
|
||||
## Question pool (grill)
|
||||
|
||||
| ID | Dim | Mode | Question | Status |
|
||||
|----|-----|------|----------|--------|
|
||||
| Q1 | 边界 | user-interview | 本 change 范围:仅第一刀,还是第一刀+第二刀(dense/hybrid 双目录对照)? | **已确认:仅第一刀** |
|
||||
| Q2 | 验收 | user-interview | Apply 时若本机无法连 embedding/Milvus 重刷 fixture,是否允许「只交接线+文档,fixture 刷新记未验证」? | **已确认:接线优先,刷新可未验证** |
|
||||
| Q3 | 术语 | evidence-driven | 生成器是否仍传 `vector-store.mode`? | **已查证:是** |
|
||||
| Q4 | 验收 | evidence-driven | 离线脚本是否已支持 Hit 分层与 baseline diff? | **已查证:是** |
|
||||
| Q5 | 边界 | evidence-driven | Golden 是否已有 mustNot/chunk key? | **已查证:无** |
|
||||
|
||||
### Q3–Q5 evidence
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-Dretrieval.vector-store.mode=$VectorStoreMode` default spring.
|
||||
- `eval_rag_retrieval.py`: strong/medium/weak/miss, recall, firstExpectedRank, compare-to diff.
|
||||
- `golden-cases.json`: doc/source/keyword/attempt/fallback; no mustNot, no evidenceKey expectations.
|
||||
|
||||
### Q1 用户确认
|
||||
|
||||
- **选择**: 仅第一刀(推荐)
|
||||
- **含义**: 生成器 `search.mode=hybrid`;fixture meta;README;能连环境则重刷。不做 dense/hybrid 双目录对照。
|
||||
|
||||
### Q2 用户确认
|
||||
|
||||
- **选择**: 接线优先,刷新可记未验证
|
||||
- **含义**: 脚本/测试/README/meta 必交付;fixtures/baseline 能刷则刷,不能刷则 acceptance 记未验证与补跑命令,不阻塞 apply 完成。
|
||||
|
||||
---
|
||||
|
||||
## Discover status
|
||||
|
||||
- [x] clarify
|
||||
- [x] context
|
||||
- [x] propose (`proposal.md`)
|
||||
- [x] grill (Q1–Q5 closed)
|
||||
|
||||
**Discover checkpoint: 完成。**
|
||||
|
||||
---
|
||||
|
||||
## Commit checkpoint
|
||||
|
||||
- **Capability**: sm-flow built-in specify/audit/commit; openspec status 4/4
|
||||
- **Cross-artifact**: brief/proposal → design → specs → tasks aligned (knife-1 only; Q1/Q2 reflected)
|
||||
- **Audit**: eval-only; L1 impact; no Agent ACI; live refresh best-effort per Q2
|
||||
- **Gate**: `.committed` written
|
||||
|
||||
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
|
||||
|
||||
**Next**: wait for explicit **Apply** authorization (e.g.「开始 apply / 实现」).
|
||||
|
||||
---
|
||||
|
||||
## Apply checkpoint
|
||||
|
||||
- **Capability**: openspec-apply-change + Committed tasks
|
||||
- **Authorization**: user「实现」
|
||||
|
||||
### Delivered (knife-1)
|
||||
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`: `-SearchMode hybrid|dense`, no `vector-store.mode`
|
||||
- `RagLookupSnapshotGeneratorTest`: `@DynamicPropertySource` for search.mode/kb-scope; fixture meta `searchMode`/`kbScope`
|
||||
- `eval/rag-retrieval/README.md` hybrid-era docs
|
||||
- Live refresh: seed OK → hybrid generate OK → offline **7/7 pass**, baseline updated
|
||||
|
||||
### Apply-discovered regression + fix
|
||||
|
||||
- **Issue**: pure rank→quality made hybrid rank1 always quality=1.0 → `isLowQuality` never true → L0 filter fallback case stuck on decoy (`FILTERED_VECTOR`).
|
||||
- **Fix**: hybrid still sorts by RRF order; optional parallel dense L2 stored as `denseDistance`; `toQualityScore(hybrid)` uses dense L2 for absolute gates when present (rank fallback if missing). Does **not** restore scoreLabel overwrite / boost re-rank.
|
||||
- **Verify**: fixture `chat-l0-filter-fallback` → `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`; offline passRate=1.0
|
||||
|
||||
### Commands run
|
||||
|
||||
```text
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||
```
|
||||
|
||||
**Apply checkpoint: 完成。** Ready for Archive when user requests.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Design: rag-eval-hybrid-baseline
|
||||
|
||||
## Context
|
||||
|
||||
Offline eval already implements golden × fixture × key-field checks. Production retrieval is hybrid (`retrieval.search.mode`) via `MilvusHybridKnowledgeStore`. Snapshot generation still injects removed `retrieval.vector-store.mode`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:** Wire snapshot generation to `search.mode`; emit fixture meta (`searchMode`, `kbScope`); document hybrid-era loop; refresh fixtures/baseline when env allows.
|
||||
|
||||
**Non-Goals:** Dual-mode fixture trees; golden mustNot/chunk/level hard gates; new eval framework; production retrieval changes.
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1 — Replace vector-store mode with search mode
|
||||
|
||||
| Before | After |
|
||||
|--------|--------|
|
||||
| `-Dretrieval.vector-store.mode=spring\|sdk` | `-Dretrieval.search.mode=hybrid\|dense` |
|
||||
| PS1 param `VectorStoreMode` | `SearchMode` default `hybrid` |
|
||||
|
||||
Java snapshot test does not need a Spring bean switch: `LookupKnowledgeTool` already honors global `retrieval.search.mode` via `VectorSearchService`. Only system property / process config must set the property before context loads (Maven `-D` + optional `properties` on `@SpringBootTest` if required).
|
||||
|
||||
### D2 — Fixture meta minimum
|
||||
|
||||
```text
|
||||
caseId, query, retrievedAt, searchMode, kbScope?, lookupResult
|
||||
```
|
||||
|
||||
- `searchMode`: actual mode used for generation.
|
||||
- `kbScope`: from `-Dretrieval.kb-scope` when non-empty.
|
||||
- Offline evaluator MAY ignore unknown meta fields (backward compatible).
|
||||
|
||||
### D3 — LookupResult payload
|
||||
|
||||
Continue serializing full `LookupResult` from tool. Prefer preserving any new block fields (`docId`, `evidenceKey`, `scoreLabel`) automatically via Jackson. No requirement to strip scores (offline does not hard-assert them).
|
||||
|
||||
### D4 — Acceptance if live refresh fails
|
||||
|
||||
Must deliver: ps1, test meta emission, README.
|
||||
Should attempt: seed + generate + eval.
|
||||
If blocked: do not fail the change; record commands and gap in acceptance/devflow.
|
||||
|
||||
### D5 — Baseline update policy
|
||||
|
||||
When fixtures refresh successfully: run offline eval; if intentional behavior change, update `reports/baseline.*` with diff review. Do not force green by weakening golden without note.
|
||||
|
||||
## Risks
|
||||
|
||||
| Risk | Mitigation |
|
||||
|------|------------|
|
||||
| Env cannot refresh fixtures | Q2: wiring-first acceptance |
|
||||
| Old fixtures fail offline after code drift | Document; refresh when possible; optional temporary note in README |
|
||||
| `@SpringBootTest` ignores late -D for some props | Set search.mode via test properties default hybrid + override from system property if needed |
|
||||
|
||||
## Interface impact
|
||||
|
||||
L1 — eval scripts, fixtures schema meta, docs. No Agent ACI.
|
||||
|
||||
## Audit
|
||||
|
||||
Eval-only pipeline; no new runtime module. Couples to existing `LookupKnowledgeTool` and config keys only.
|
||||
@@ -0,0 +1,66 @@
|
||||
# Change: Align RAG offline eval with hybrid + qualityScore era
|
||||
|
||||
## Why
|
||||
|
||||
`eval/rag-retrieval` already matches the offline model (golden × fixture × key-field checks × baseline/diff), but it is stuck on the pre-hybrid narrative:
|
||||
|
||||
- Snapshot generator still passes dead `retrieval.vector-store.mode=spring|sdk`.
|
||||
- Fixtures lack `searchMode` / scope meta; content still shows boost-style `hitReasons` and old score story.
|
||||
- No first-class dense vs hybrid fixture split for recall comparison.
|
||||
- Golden lacks optional hard-negatives / chunk keys / tags that the design discussion called out.
|
||||
|
||||
Without this, offline eval cannot gate the current main path (`retrieval.search.mode=hybrid`, V2 store, qualityScore post-process).
|
||||
|
||||
## What Changes
|
||||
|
||||
### Knife 1 (must) — make offline eval reflect current main path
|
||||
|
||||
1. **Generator wiring**
|
||||
- Replace `retrieval.vector-store.mode` with `retrieval.search.mode` (`hybrid` default; `dense` allowed).
|
||||
- Keep `-Dretrieval.kb-scope=rag-eval` (or configurable).
|
||||
- Update `scripts/generate_rag_lookup_snapshots.ps1` and any Java system-property docs/comments.
|
||||
|
||||
2. **Fixture meta**
|
||||
- Each fixture SHALL record at least: `caseId`, `query`, `retrievedAt`, `searchMode`, `kbScope` (when set), plus `lookupResult` payload.
|
||||
- Snapshot writer emits current `LookupResult` shape (evidence identity fields if already present on blocks).
|
||||
|
||||
3. **Refresh path**
|
||||
- Document and support: prepare seed → generate fixtures (hybrid) → `eval_rag_retrieval.py` → update baseline.
|
||||
- Refresh committed fixtures/baseline when live generation is available; if environment blocks live run, ship wiring + docs and record gap in acceptance.
|
||||
|
||||
4. **Docs**
|
||||
- Update `eval/rag-retrieval/README.md` to hybrid/quality narrative; remove spring vector-store as default.
|
||||
|
||||
### Out of scope this change (confirmed grill)
|
||||
|
||||
- Knife 2: `fixtures/hybrid` vs `fixtures/dense` dual layout and comparison report.
|
||||
- Golden extensions: `tags` / `mustNot*` / chunk keys / `expectedRelevanceLevel` hard gates.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Rewriting eval into a new framework or LLM-as-judge.
|
||||
- Full threshold calibration productization.
|
||||
- Neighbor chunks / query rewrite / cross-encoder.
|
||||
- Changing production retrieval code paths (except eval generator test harness props).
|
||||
- Forcing live Milvus E2E / fixture refresh when embedding/Milvus unavailable (**wiring+docs still complete**; refresh recorded as unverified).
|
||||
- Dense/hybrid dual fixture directories (later change).
|
||||
|
||||
## Context constraints
|
||||
|
||||
- Continues `rag-quality-score-unify`, `rag-bm25-hybrid-drop-sdk`, `rag-chunk-evidence-identity-dedup`.
|
||||
- Offline checker must remain dependency-free (no Milvus/LLM in `eval_rag_retrieval.py`).
|
||||
- Seed isolation via `kb_scope=rag-eval` stays.
|
||||
|
||||
## Impact
|
||||
|
||||
- **Interface**: L1/L2 docs + eval artifacts only; no Agent ACI change.
|
||||
- **Risk**: refreshed fixtures may change pass/fail vs old baseline — expect intentional baseline update with diff review.
|
||||
- **Scale**: **micro→standard lean** — multi-file scripts/docs/fixtures; no production architecture change. Use **standard** artifacts for clarity (`design` + `specs` + `tasks`).
|
||||
|
||||
## Success
|
||||
|
||||
- Generator defaults to hybrid search mode; dead vector-store mode flag gone.
|
||||
- Fixtures carry searchMode meta.
|
||||
- README describes correct offline/live loop.
|
||||
- Offline eval runs on refreshed or existing fixtures without requiring removed config keys.
|
||||
- (If knife 2) dual fixture roots documented and runnable.
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
# rag-eval-offline-baseline Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Snapshot generation SHALL use retrieval search mode
|
||||
|
||||
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
|
||||
|
||||
#### Scenario: Default hybrid generation
|
||||
|
||||
- **WHEN** the snapshot generator is invoked with default parameters
|
||||
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
|
||||
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
|
||||
|
||||
#### Scenario: Dense mode override for comparison runs
|
||||
|
||||
- **WHEN** the operator sets search mode to `dense`
|
||||
- **THEN** fixture generation SHALL use dense retrieval for that run
|
||||
|
||||
### Requirement: Generated fixtures SHALL record search meta
|
||||
|
||||
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
|
||||
|
||||
#### Scenario: Meta fields present
|
||||
|
||||
- **WHEN** a fixture is written for a golden case
|
||||
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
|
||||
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
|
||||
|
||||
### Requirement: Offline evaluation SHALL remain dependency-free
|
||||
|
||||
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
|
||||
|
||||
#### Scenario: Offline eval without live stack
|
||||
|
||||
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
|
||||
- **THEN** it SHALL produce pass/fail results using fixture contents only
|
||||
|
||||
### Requirement: Eval documentation SHALL describe the hybrid-era loop
|
||||
|
||||
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
|
||||
|
||||
#### Scenario: README main path
|
||||
|
||||
- **WHEN** an engineer follows the eval README happy path
|
||||
- **THEN** the documented default generation mode SHALL be hybrid search mode
|
||||
@@ -0,0 +1,34 @@
|
||||
# Tasks: rag-eval-hybrid-baseline
|
||||
|
||||
## 1. Generator wiring
|
||||
|
||||
- [x] 1.1 Update `scripts/generate_rag_lookup_snapshots.ps1`: replace `VectorStoreMode` / `vector-store.mode` with `SearchMode` default `hybrid` and `-Dretrieval.search.mode=...`
|
||||
- [x] 1.2 Keep `-Dretrieval.kb-scope` (default `rag-eval`); document `-SearchMode dense` override
|
||||
- [x] 1.3 Ensure snapshot test picks up `retrieval.search.mode` (system property and/or `@SpringBootTest` properties)
|
||||
|
||||
## 2. Fixture meta
|
||||
|
||||
- [x] 2.1 `RagLookupSnapshotGeneratorTest` writes `searchMode` and `kbScope` (when set) on each fixture
|
||||
- [x] 2.2 Confirm offline `eval_rag_retrieval.py` still loads fixtures (ignore extra meta)
|
||||
|
||||
## 3. Docs
|
||||
|
||||
- [x] 3.1 Rewrite `eval/rag-retrieval/README.md` hybrid-era loop; remove spring vector-store as default generation path
|
||||
- [x] 3.2 Note offline vs live responsibilities; point to qualityScore/hybrid main path briefly
|
||||
|
||||
## 4. Refresh attempt (best-effort)
|
||||
|
||||
- [x] 4.1 Attempt `prepare_rag_eval_seed` + hybrid snapshot generate + offline eval when environment allows
|
||||
- [x] 4.2 On success: update fixtures and `reports/baseline.*` if needed after diff review
|
||||
- [x] 4.3 On failure: record exact commands, error summary, and “unverified refresh” in change decisions/acceptance notes — do not block wiring delivery
|
||||
|
||||
## 5. Verify
|
||||
|
||||
- [x] 5.1 Static: grep shows no required `vector-store.mode` in snapshot generator path
|
||||
- [x] 5.2 Offline: `python scripts/eval_rag_retrieval.py` runs on committed fixtures (pass or documented baseline drift)
|
||||
|
||||
## 6. Apply-discovered fix (hybrid quality gate)
|
||||
|
||||
- [x] 6.1 Hybrid attaches optional `denseDistance` without overwriting RRF order/scoreLabel
|
||||
- [x] 6.2 `toQualityScore(hybrid)` prefers dense L2 for absolute gates; rank fallback if no dense
|
||||
- [x] 6.3 Regenerate fixtures; L0 filter fallback case green again; baseline updated
|
||||
@@ -0,0 +1 @@
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
Committed OpenSpec
|
||||
change: rag-quality-score-unify
|
||||
committed_at: 2026-07-28
|
||||
scale: standard
|
||||
interface_impact: L2
|
||||
gate: proposal+design+specs+tasks+cross-artifact+audit
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-28
|
||||
@@ -0,0 +1,20 @@
|
||||
# Brief: rag-quality-score-unify
|
||||
|
||||
## Background
|
||||
|
||||
True BM25 hybrid retrieval is live, but post-processing still normalizes as if every score were dense L2 and re-ranks with L0 keyword contains boosts. That splits ranking authority from quality gates and double-counts lexical signal.
|
||||
|
||||
## Goals
|
||||
|
||||
- Unify score labels to `dense` | `hybrid`.
|
||||
- Single `toQualityScore`; label-agnostic post-process.
|
||||
- Preserve retrieval rank; remove boost re-rank.
|
||||
- Hybrid quality = pure rank mapping (slice 1).
|
||||
|
||||
## Scope
|
||||
|
||||
Internal RAG pipeline: store emission, normalizer, evidence post-process, tests, architecture docs.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Fine re-rankers, query rewrite, neighbor chunks, schema rebuild, ACI field renames, removing dense comparison mode.
|
||||
@@ -0,0 +1,136 @@
|
||||
# Decisions — rag-quality-score-unify
|
||||
|
||||
## sm-flow meta
|
||||
|
||||
- **Checkpoint**: Discover(clarify + context + propose + grill)
|
||||
- **Scale**: standard
|
||||
- **Capability**: sm-flow 内置协议;openspec CLI `new change`;grill 使用内置协议(conversation-confirmed + evidence-driven),标注 fallback:未调用外部 `grill-with-docs` skill 文件执行器
|
||||
- **Slug**: `rag-quality-score-unify`
|
||||
- **OpenSpec path**: `openspec/changes/rag-quality-score-unify/`
|
||||
|
||||
## Clarify summary
|
||||
|
||||
| 项 | 内容 |
|
||||
|---|---|
|
||||
| 问题 | hybrid 已 RRF 融合,后处理仍 L2 伪装 + 关键词 boost 改序,质量信号不统一 |
|
||||
| 期望 | label 仅 dense/hybrid;toQualityScore 唯一归一化;后处理保 rank、去 boost 改序 |
|
||||
| 影响代码 | `MilvusHybridKnowledgeStore`, `VectorSearchService`, `KnowledgeEvidencePostProcessor`, retrieval 包新类, DTO 注释, 测试, 架构文档 |
|
||||
| 非目标 | 精排/rewrite/邻块、删 dense mode、改 ACI 字段结构、改 schema |
|
||||
|
||||
## Context summary (devflow)
|
||||
|
||||
| 来源 | 结论 | 需进 OpenSpec |
|
||||
|---|---|---|
|
||||
| `devflow/index.md` | rag-chunk-identity / bm25-hybrid / hybrid-rrf 均 archived | 是:承接不回退 |
|
||||
| `rag-bm25-hybrid-drop-sdk/decisions.md` | 曾要求 dense L2 enrichment 兼容阈值 | **是:本 change 废止该 decision** |
|
||||
| glossary | lookup_knowledge 为证据工具;不在此改 ACI 主结构 | 是:非目标 |
|
||||
| 架构文档 §6.0 | mode dense=对照,hybrid=主路径 | 是:保留 |
|
||||
|
||||
**index 使用状态**: 已命中相关 RAG 条目。
|
||||
|
||||
## Question pool (grill)
|
||||
|
||||
| ID | 维度 | 模式 | 问题 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | user-interview(对话已确认) | 一级 scoreLabel 是否只有 dense/hybrid,bm25_only 不作正式 label? | **已确认** |
|
||||
| Q2 | 边界 | user-interview(对话已确认) | 后处理是否去掉关键词 contains 加分改序,仅保 originalRank? | **已确认** |
|
||||
| Q3 | 边界 | user-interview(对话已确认) | 归一化是否唯一 toQualityScore;后处理 label-agnostic? | **已确认** |
|
||||
| Q4 | 验收 | user-interview(对话已确认) | dense mode 保留作召回对照;主路径 hybrid? | **已确认** |
|
||||
| Q5 | 技术 | evidence-driven | 当前代码是否仍 L2 回填 + boost 重排? | **已查证** |
|
||||
| Q6 | 技术 | evidence-driven | Agent ACI 是否暴露 scoreLabel? | **已查证** |
|
||||
| Q7 | 验收 | user-interview(对话已确认) | 接受 relevance_level / retry 分布变化? | **已确认** |
|
||||
| Q8 | 边界 | user-interview | hybrid quality 切片 1 是否采用**纯 rank 映射**(不做 max(rank,denseSim))? | **已确认** |
|
||||
|
||||
### Q1–Q4, Q7 用户确认摘录(本会话)
|
||||
|
||||
- Label:「应该只有 hybrid 和 dense」「bm25_only 不是第三种」→ 同意收成两种。
|
||||
- 归一化:「抽取抽象转换…后处理抽象统一」→ 同意。
|
||||
- 后处理:「关键词打分不合理」「可以,就按照这个」(去 boost 改序 + 归一化一起做)。
|
||||
- mode:保留 dense 作对照,写入架构 §6.0。
|
||||
- 行为变化:讨论中已说明 hybrid 顺序/等级/retry 会变,用户要求按该方案实施(经 sm-flow)。
|
||||
|
||||
### Q5 evidence-driven
|
||||
|
||||
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后仍 dense 回填 L2 / `bm25_only_no_dense`。
|
||||
- `KnowledgeEvidencePostProcessor.score`:`normalizeL2` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序。
|
||||
- `LookupKnowledgeTool`:`isLowQuality` 看 `topSimilarity`(来自 baseScore)。
|
||||
|
||||
### Q6 evidence-driven
|
||||
|
||||
- Agent 主契约 `RagToolResult` / projector 暴露 evidence 列表与 relevance_level,不依赖 scoreLabel 字符串;改 label 为内部/L2 影响。
|
||||
|
||||
### Q8 用户确认(2026-07-28)
|
||||
|
||||
- **问题原文**: hybrid 的 qualityScore(切片 1)采用哪种映射?
|
||||
- **用户选择**: 纯 rank 映射(推荐)
|
||||
- **确认状态**: 已确认
|
||||
- **实现约束**: `toQualityScore(hybrid)` = `rankToQuality(originalRank, batchSize)`;不看 RRF 原分量纲;不做 `max(rank, denseSim)`;不在 hybrid 路径为质量闸门再查/回填 dense L2。
|
||||
|
||||
---
|
||||
|
||||
## Discover status
|
||||
|
||||
- [x] clarify
|
||||
- [x] context
|
||||
- [x] propose (`proposal.md`)
|
||||
- [x] grill 完成(Q1–Q8 均已关闭)
|
||||
|
||||
**Discover checkpoint: 完成。**
|
||||
|
||||
---
|
||||
|
||||
## Commit checkpoint
|
||||
|
||||
### Capability
|
||||
|
||||
- specify: sm-flow 内置 + openspec status/instructions(fallback:按 template 手写 design/specs/tasks)
|
||||
- audit: sm-flow 内置协议(未调用外部 zoom-out)
|
||||
- commit gate: 文件完整性 + 一致性检查后写入 `.committed`
|
||||
|
||||
### Cross-artifact 对齐
|
||||
|
||||
| 链路 | 状态 |
|
||||
|---|---|
|
||||
| brief/proposal 目标范围 → design | 已对齐 |
|
||||
| design 决策(label/normalizer/保序/去 boost/纯 rank)→ specs | 已对齐 |
|
||||
| specs 可观察行为 → tasks 可执行切片 | 已对齐 |
|
||||
| decisions Q1–Q8 → proposal/design/specs | 已对齐 |
|
||||
|
||||
### Audit(≤5 句)
|
||||
|
||||
1. 链路仍是 Tool→Retriever→Store→Post→Pack→Project,无新外部系统。
|
||||
2. 分数所有权上收 store 发射 + normalizer;后处理只裁剪与质量闸门。
|
||||
3. 废止 bm25-hybrid 的「dense L2 enrichment」决策,属有意行为变化(L2 接口影响)。
|
||||
4. 风险主要是 hybrid 序数 quality 与阈值标定,已记入 design Risks。
|
||||
5. 不触及 Agent ACI 字段名与 Milvus schema。
|
||||
|
||||
### Commit gate checklist
|
||||
|
||||
- [x] proposal / design / specs / tasks / brief 存在
|
||||
- [x] 核心概念在 design 有对应
|
||||
- [x] design 关键决策在 tasks 有任务
|
||||
- [x] tasks 可验证(checkbox 纵向切片)
|
||||
- [x] 无未确认 user-interview
|
||||
- [x] `.committed` 已创建
|
||||
|
||||
**Commit checkpoint: 完成。Committed OpenSpec 就绪。**
|
||||
|
||||
**下一步**: 等待用户明确授权 **Apply**(例如「开始 apply / 实现」)。未授权前不改业务接线代码。
|
||||
|
||||
---
|
||||
|
||||
## Apply checkpoint
|
||||
|
||||
- **Capability**: openspec-apply-change + Committed OpenSpec tasks
|
||||
- **授权**: 用户「实现」
|
||||
- **完成**: tasks.md 全部勾选
|
||||
- **验证**:
|
||||
- `RetrievalScoreNormalizerTest` 4 passed
|
||||
- `KnowledgeEvidencePostProcessorTest` 6 passed
|
||||
- `LookupKnowledgeToolTest` 7 passed
|
||||
- `VectorSearchServiceTest` 2 passed
|
||||
- `VectorKnowledgeSearchAdapterHybridTest` 1 passed
|
||||
- **已知限制**: hybrid quality 为本轮 rank 序数映射,跨 query 绝对值不可比;阈值可能需后续标定
|
||||
- **行为变化**: 已落地(去 L2 回填、去 boost 改序、label dense/hybrid)
|
||||
|
||||
**Apply checkpoint: 完成。** 可进入 Archive(需用户确认是否 archive OpenSpec)。
|
||||
@@ -0,0 +1,134 @@
|
||||
# Design: rag-quality-score-unify
|
||||
|
||||
## Context
|
||||
|
||||
- Knowledge path already uses single `MilvusHybridKnowledgeStore` (MilvusClientV2) with dense + BM25 + RRF.
|
||||
- Chunk-level `evidenceKey` dedup and `retrieve-k` / `return-n` are landed.
|
||||
- Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
|
||||
- Prior design in `rag-bm25-hybrid-drop-sdk` required dense L2 enrichment for threshold compatibility — **this change supersedes that decision**.
|
||||
|
||||
Stakeholders: `lookup_knowledge` internal pipeline; Agent ACI field *names* unchanged; operators comparing `retrieval.search.mode=dense|hybrid`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
1. First-class `scoreLabel` values: only `dense` | `hybrid` (aliases canonicalize).
|
||||
2. Single `toQualityScore` adapter; post-process is label-agnostic.
|
||||
3. Preserve retrieval `originalRank` as sort authority; remove keyword/domain boost re-ranking.
|
||||
4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
|
||||
5. Keep `mode=dense` for offline recall comparison; production default remains hybrid.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Cross-encoder / query rewrite / neighbor chunks.
|
||||
- Schema rebuild or collection rename.
|
||||
- Changing Agent-facing ACI JSON field names.
|
||||
- Configurable `max(rank, denseSim)` quality (future).
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1 — Two labels only
|
||||
|
||||
| label | `score` meaning | quality mapping |
|
||||
|---|---|---|
|
||||
| `dense` | L2 distance (smaller better) | `1 - clamp(l2)/maxL2Distance` |
|
||||
| `hybrid` | engine fused score optional in `rawScore`; **not** used as L2 | `rankToQuality(originalRank, batchSize)` |
|
||||
|
||||
Canonicalize legacy strings: `l2_distance`→dense; `rrf_fused` / `bm25_only_*`→hybrid.
|
||||
|
||||
**Why not keep `bm25_only`:** it is not a search mode; it was a L2-fake patch. Hybrid path hits are all `hybrid`.
|
||||
|
||||
### D2 — Stop dense L2 overwrite on hybrid hits
|
||||
|
||||
`searchHybrid` SHALL:
|
||||
|
||||
1. Run `hybridSearch` + RRFRanker.
|
||||
2. Emit hits in RRF order with `scoreLabel=hybrid`, `originalRank=1..n`.
|
||||
3. Set `rawScore` from engine when present; `score` MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption.
|
||||
4. MUST NOT set `bm25_only_no_dense` or force `score=maxL2Distance` for threshold faking.
|
||||
5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).
|
||||
|
||||
`searchDense` SHALL emit `scoreLabel=dense` and L2 in `score`.
|
||||
|
||||
### D3 — `RetrievalScoreNormalizer` is the only label branch
|
||||
|
||||
```text
|
||||
qualityScore = RetrievalScoreNormalizer.toQualityScore(
|
||||
scoreLabel, score, originalRank, batchSize, maxL2Distance)
|
||||
```
|
||||
|
||||
- Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
|
||||
- Dense: existing L2 formula (behavior parity for dense mode).
|
||||
|
||||
Post-processor, `isLowQuality`, and `relevance_level` consume **only** `qualityScore` (exposed today as `baseScore` / `topSimilarity` fields for minimal DTO churn).
|
||||
|
||||
### D4 — Post-process sort and boosts
|
||||
|
||||
```text
|
||||
order = originalRank ASC, then stable evidenceKey
|
||||
// NO finalScore = base + 0.15 domain + ...
|
||||
```
|
||||
|
||||
- Remove additive boosts from sort key and from `finalScore` used for ordering.
|
||||
- Optional: if L0 hint string matches, append explanatory `hitReasons` only (e.g. `l0_keyword_overlap`) — zero score delta.
|
||||
- Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
|
||||
- `relevance_level`: compare top `qualityScore` to existing thresholds; **remove** `hasHintSupport` gate for PRECISE.
|
||||
- `isLowQuality` / category unfiltered retry: unchanged control flow, new score semantics.
|
||||
|
||||
### D5 — Trace fields
|
||||
|
||||
- `RerankTrace` may keep `baseScore`/`finalScore` names but both equal `qualityScore` when boosts are zero; `boostReasons` empty or explanation-only reasons without `:+0.xx` score deltas.
|
||||
- Prefer renaming comments to “quality trace”; no Agent contract change required.
|
||||
|
||||
### D6 — WIP files
|
||||
|
||||
Workspace may contain draft `RetrievalScoreLabels` / `RetrievalScoreNormalizer`. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.
|
||||
|
||||
## Interface impact
|
||||
|
||||
- **Level: L2** — internal DTO semantics (`score`, `scoreLabel`), post-process ordering and relevance distribution.
|
||||
- Agent ACI: field names stable; `relevance_level` *values distribution* may change (accepted behavioral change).
|
||||
|
||||
## Data flow (target)
|
||||
|
||||
```text
|
||||
VectorSearchService (mode dense|hybrid)
|
||||
-> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
|
||||
-> KnowledgeDocumentRetriever candidates
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
for each: qualityScore = toQualityScore(...)
|
||||
sort by originalRank
|
||||
dedup / caps / return-n
|
||||
relevance + topSimilarity from qualityScore
|
||||
-> pack / assemble / project
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute |
|
||||
| Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” |
|
||||
| PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains |
|
||||
| Tests assert boost re-order | Update `LookupKnowledgeToolTest.rerankUsesHintMatches...` |
|
||||
| Old label strings in sidecar/eval | canonicalize in normalizer |
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Deploy code; no Milvus schema migration.
|
||||
2. Default `retrieval.search.mode=hybrid` unchanged.
|
||||
3. Rollback: revert change; old L2-enrichment behavior returns.
|
||||
4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).
|
||||
|
||||
## Open Questions
|
||||
|
||||
- None for slice-1 (Q8 confirmed: pure rank).
|
||||
- Follow-up: threshold calibration after live traces; optional denseSim blend.
|
||||
|
||||
## Audit notes (inline)
|
||||
|
||||
Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
|
||||
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
|
||||
No new cross-module lifecycle. Couples only internal RAG pipeline.
|
||||
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.
|
||||
@@ -0,0 +1,80 @@
|
||||
# Change: Unify RAG quality score (dense/hybrid labels) and stop keyword boost re-rank
|
||||
|
||||
## Why
|
||||
|
||||
BM25 hybrid 已在库内完成 dense + BM25 + RRF 融合,但后处理仍:
|
||||
|
||||
1. 把 hybrid 结果**伪装成 L2** 再 `normalizeL2`(含 `bm25_only_no_dense` 弱分占位);
|
||||
2. 用 L0 domain/entity/keyword **contains 加分改主序**。
|
||||
|
||||
这导致:排序信号与质量闸门分裂;词面信号被 BM25 与后处理**双重计分**;「词面热、语义冷」的片段可能被抬到前面;hybrid 的 RRF 序被冲掉。
|
||||
|
||||
需要统一:**检索负责序,后处理只做 quality 归一化 + 裁剪装配**。
|
||||
|
||||
## What Changes
|
||||
|
||||
### 检索层(`MilvusHybridKnowledgeStore` / `VectorSearchService`)
|
||||
|
||||
- 一级 `scoreLabel` 仅两种:`dense` | `hybrid`(与 `retrieval.search.mode` 对齐)。
|
||||
- **废弃**正式一级 label:`l2_distance` / `rrf_fused` / `bm25_only_no_dense`(可读兼容映射到 dense/hybrid)。
|
||||
- `mode=dense`:`score` = L2 距离,`label=dense`,`originalRank` = ANN 序。
|
||||
- `mode=hybrid`:`label=hybrid`;**不再**用 dense L2 覆盖主 `score`;**不再**对 BM25-only 伪造 maxL2;`originalRank` = RRF 返回序;`rawScore` 可保留引擎融合分。
|
||||
- `mode=dense|hybrid` **保留**:hybrid 为线上主路径;dense 为同库对照/评测(已写入架构 §6.0)。
|
||||
|
||||
### 归一化(新)
|
||||
|
||||
- 新增唯一转换点 `RetrievalScoreNormalizer.toQualityScore(label, score, rank, batchSize, maxL2)` → `qualityScore ∈ [0,1]`(越大越好)。
|
||||
- `dense`:`1 - clamp(L2)/maxL2Distance`
|
||||
- `hybrid`:按 **rank** 映射(本轮 batch 线性),不把 RRF 原分当 L2 套公式。
|
||||
- Label 差异**只**在此消化。
|
||||
|
||||
### 后处理(`KnowledgeEvidencePostProcessor`)
|
||||
|
||||
- **统一流程**,只消费 `qualityScore` + `originalRank`(label-agnostic)。
|
||||
- **排序主序 = `originalRank` 升序**(保检索序);去掉 domain/entity/keyword/source_type **加分改序**。
|
||||
- L0 contains 匹配若保留,仅写入 `hitReasons` / trace 解释,**不参与 sort key、不加 finalScore**。
|
||||
- `relevance_level` / `isLowQuality` / category unfiltered retry:只看 top `qualityScore` 与既有阈值;**不再**要求 `hasHintSupport` 才能 PRECISE。
|
||||
- 保留:evidenceKey 去重、`max-chunks-per-document`、`return-n`、excerpt 截断、EvidenceBlock 装配。
|
||||
|
||||
### 文档 / 测试
|
||||
|
||||
- 更新 `mvp/architecture/RAG知识检索架构.md` §6 分数与后处理约定。
|
||||
- 单测:dense 路径 quality 与现 L2 归一化一致;hybrid 保序且不被 keyword 打乱;无 `bm25_only` 一级 label;旧 label 别名可 canonicalize。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Cross-encoder / listwise 精排、query rewrite、邻块扩展。
|
||||
- 删除 `mode=dense` 对照开关。
|
||||
- 改变 Agent 可见 ACI 字段结构(`evidence[]` / `relevance_level` 枚举名可不变;**分布会变**)。
|
||||
- 修改 Milvus schema / 强制全量 rebuild(本 change 不改 collection 结构)。
|
||||
- 上线可配 `max(rank, denseSim)` 混合 quality(可后续迭代;本 change 切片 1 用纯 rank 映射 hybrid)。
|
||||
|
||||
## Context constraints (from devflow)
|
||||
|
||||
- 承接 archived:`rag-chunk-evidence-identity-dedup`、`rag-bm25-hybrid-drop-sdk`、`rag-hybrid-search-rrf`、`modular-rag-pipeline`。
|
||||
- 单一知识后端仍为 `MilvusHybridKnowledgeStore`(V2);不得恢复 sdk/spring 主路径路由。
|
||||
- Agent 投影仍不暴露 raw fused score / 完整 contextPack 作为主契约(内部 LookupResult/trace 可保留调试字段)。
|
||||
- 历史 decision「Dense L2 enrichment for threshold compatibility」**本 change 有意废止**,改为 qualityScore 统一闸门。
|
||||
|
||||
## Impact
|
||||
|
||||
- **行为变化(对内检索质量语义)**:
|
||||
- hybrid 下证据顺序更贴近 RRF;
|
||||
- 词面命中不再被后处理 contains 二次抬序;
|
||||
- `relevance_level` 与 unfiltered retry 触发分布可能变化;
|
||||
- BM25-only 命中不再被标成 quality≈0。
|
||||
- **接口影响**:L2(内部 DTO/注释/scoreLabel 字符串约定);Agent ACI 字段名不变。
|
||||
- **风险**:hybrid rank→quality 为序数映射,绝对值不跨 query 可比;阈值 0.75/0.5 可能需后续观测再调(本 change 先沿用配置项)。
|
||||
|
||||
## Scale
|
||||
|
||||
- **standard**(多文件、有意行为变化、需 design + specs + tasks + 测试)。
|
||||
|
||||
## Depends on
|
||||
|
||||
- 已落地 hybrid schema + chunk evidenceKey(archived changes 如上)。
|
||||
- 对话已确认的设计口径(见 change `decisions.md`)。
|
||||
|
||||
## WIP note
|
||||
|
||||
- 工作区可能已有未接线的 `RetrievalScoreLabels` / `RetrievalScoreNormalizer` 草稿文件;apply 阶段以 **Committed OpenSpec** 为准接入或改写,不视为已完成实现。
|
||||
+94
@@ -0,0 +1,94 @@
|
||||
# rag-retrieval-quality-score Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Primary score labels SHALL be only dense or hybrid
|
||||
|
||||
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
|
||||
|
||||
#### Scenario: Dense mode labels hits as dense
|
||||
|
||||
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
|
||||
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
|
||||
- **AND** `score` SHALL be the dense L2 distance
|
||||
|
||||
#### Scenario: Hybrid mode labels hits as hybrid
|
||||
|
||||
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
|
||||
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
|
||||
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
|
||||
|
||||
#### Scenario: Legacy aliases canonicalize
|
||||
|
||||
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
|
||||
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
|
||||
|
||||
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
|
||||
|
||||
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
|
||||
|
||||
#### Scenario: No L2 enrichment overwrite
|
||||
|
||||
- **WHEN** hybrid search completes
|
||||
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
|
||||
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
|
||||
|
||||
### Requirement: Quality score SHALL be produced by a single normalizer
|
||||
|
||||
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
|
||||
|
||||
#### Scenario: Dense quality from L2
|
||||
|
||||
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
|
||||
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
|
||||
|
||||
#### Scenario: Hybrid quality from rank
|
||||
|
||||
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
|
||||
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
|
||||
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
|
||||
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
|
||||
|
||||
### Requirement: Post-process SHALL preserve retrieval rank order
|
||||
|
||||
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
|
||||
|
||||
#### Scenario: Keyword overlap does not promote lower rank
|
||||
|
||||
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
|
||||
- **AND** B matches more L0 keywords via string contains than A
|
||||
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
|
||||
|
||||
#### Scenario: Structural caps still apply after rank order
|
||||
|
||||
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
|
||||
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
|
||||
- **AND** `rag.return-n` SHALL still bound total blocks
|
||||
|
||||
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
|
||||
|
||||
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
|
||||
|
||||
#### Scenario: Low quality uses top qualityScore
|
||||
|
||||
- **WHEN** post-process finishes with at least one evidence block
|
||||
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
|
||||
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
|
||||
|
||||
#### Scenario: PRECISE does not require hint support
|
||||
|
||||
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
|
||||
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
|
||||
|
||||
### Requirement: Dense mode remains available for recall comparison
|
||||
|
||||
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
|
||||
|
||||
#### Scenario: Mode dense still callable
|
||||
|
||||
- **WHEN** mode is `dense`
|
||||
- **THEN** search SHALL call dense ANN only and label hits `dense`
|
||||
@@ -0,0 +1,39 @@
|
||||
# Tasks: rag-quality-score-unify
|
||||
|
||||
## 1. Score contract utilities
|
||||
|
||||
- [x] 1.1 Finalize `RetrievalScoreLabels` (`dense` / `hybrid` + canonicalize legacy aliases)
|
||||
- [x] 1.2 Finalize `RetrievalScoreNormalizer.toQualityScore` (dense L2 formula; hybrid pure rank map with batchSize)
|
||||
- [x] 1.3 Unit tests for normalizer: dense L2 edges; hybrid rank monotonicity; alias canonicalize
|
||||
|
||||
## 2. Store / search emission
|
||||
|
||||
- [x] 2.1 `searchDense`: emit `scoreLabel=dense`, L2 `score`, stable rank order
|
||||
- [x] 2.2 `searchHybrid`: emit `scoreLabel=hybrid`; keep RRF order as `originalRank`; stop dense L2 overwrite and `bm25_only_*` labels; no parallel dense probe for score rewrite
|
||||
- [x] 2.3 Update `VectorSearchService.SearchResult` / adapter comments so `score`+`scoreLabel` contract matches design
|
||||
- [x] 2.4 Ensure `KnowledgeDocumentRetriever` / `VectorKnowledgeSearchAdapter` propagate `scoreLabel`, `score`, `rawScore`, `originalRank` unchanged
|
||||
|
||||
## 3. Post-process
|
||||
|
||||
- [x] 3.1 `KnowledgeEvidencePostProcessor`: compute quality via normalizer; sort by `originalRank` ASC (stable tie-break)
|
||||
- [x] 3.2 Remove domain/entity/keyword/source_type additive boosts from ordering/`finalScore`
|
||||
- [x] 3.3 Optional: L0 overlap only as explanatory `hitReasons` (no score delta)
|
||||
- [x] 3.4 `relevance_level` / `isLowQuality` / `topSimilarity` use qualityScore only; drop `hasHintSupport` gate for PRECISE
|
||||
- [x] 3.5 Keep evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate
|
||||
|
||||
## 4. Tests
|
||||
|
||||
- [x] 4.1 Update `KnowledgeEvidencePostProcessorTest` for rank order + caps under new scoring
|
||||
- [x] 4.2 Update `LookupKnowledgeToolTest.rerankUsesHintMatchesAndContextPackPreservesMetadata` (no boost re-order; metadata/context pack still ok)
|
||||
- [x] 4.3 Adjust any tests asserting `l2_distance` / boost reasons `:+0.xx` as needed
|
||||
- [x] 4.4 Run targeted unit tests for touched classes
|
||||
|
||||
## 5. Docs
|
||||
|
||||
- [x] 5.1 Update `mvp/architecture/RAG知识检索架构.md` §6 score/post-process (replace L2-enrichment narrative)
|
||||
- [x] 5.2 Align `application.yml` comments if still describing L2-only post-process for hybrid
|
||||
|
||||
## 6. Verify
|
||||
|
||||
- [x] 6.1 Confirm no production path still sets `bm25_only_no_dense` or overwrites hybrid scores with L2 for thresholds
|
||||
- [x] 6.2 Note known limitation: hybrid quality is ordinal within batch; thresholds may need later calibration
|
||||
@@ -0,0 +1,50 @@
|
||||
# rag-eval-offline-baseline Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Snapshot generation SHALL use retrieval search mode
|
||||
|
||||
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
|
||||
|
||||
#### Scenario: Default hybrid generation
|
||||
|
||||
- **WHEN** the snapshot generator is invoked with default parameters
|
||||
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
|
||||
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
|
||||
|
||||
#### Scenario: Dense mode override for comparison runs
|
||||
|
||||
- **WHEN** the operator sets search mode to `dense`
|
||||
- **THEN** fixture generation SHALL use dense retrieval for that run
|
||||
|
||||
### Requirement: Generated fixtures SHALL record search meta
|
||||
|
||||
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
|
||||
|
||||
#### Scenario: Meta fields present
|
||||
|
||||
- **WHEN** a fixture is written for a golden case
|
||||
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
|
||||
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
|
||||
|
||||
### Requirement: Offline evaluation SHALL remain dependency-free
|
||||
|
||||
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
|
||||
|
||||
#### Scenario: Offline eval without live stack
|
||||
|
||||
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
|
||||
- **THEN** it SHALL produce pass/fail results using fixture contents only
|
||||
|
||||
### Requirement: Eval documentation SHALL describe the hybrid-era loop
|
||||
|
||||
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
|
||||
|
||||
#### Scenario: README main path
|
||||
|
||||
- **WHEN** an engineer follows the eval README happy path
|
||||
- **THEN** the documented default generation mode SHALL be hybrid search mode
|
||||
@@ -0,0 +1,104 @@
|
||||
# rag-retrieval-quality-score Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Primary score labels SHALL be only dense or hybrid
|
||||
|
||||
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
|
||||
|
||||
#### Scenario: Dense mode labels hits as dense
|
||||
|
||||
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
|
||||
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
|
||||
- **AND** `score` SHALL be the dense L2 distance
|
||||
|
||||
#### Scenario: Hybrid mode labels hits as hybrid
|
||||
|
||||
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
|
||||
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
|
||||
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
|
||||
|
||||
#### Scenario: Legacy aliases canonicalize
|
||||
|
||||
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
|
||||
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
|
||||
|
||||
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
|
||||
|
||||
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
|
||||
|
||||
#### Scenario: No L2 enrichment overwrite
|
||||
|
||||
- **WHEN** hybrid search completes
|
||||
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
|
||||
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
|
||||
|
||||
### Requirement: Quality score SHALL be produced by a single normalizer
|
||||
|
||||
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
|
||||
|
||||
#### Scenario: Dense quality from L2
|
||||
|
||||
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
|
||||
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
|
||||
|
||||
#### Scenario: Hybrid quality from rank
|
||||
|
||||
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
|
||||
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
|
||||
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
|
||||
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
|
||||
|
||||
### Requirement: Post-process SHALL preserve retrieval rank order
|
||||
|
||||
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
|
||||
|
||||
#### Scenario: Keyword overlap does not promote lower rank
|
||||
|
||||
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
|
||||
- **AND** B matches more L0 keywords via string contains than A
|
||||
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
|
||||
|
||||
#### Scenario: Structural caps still apply after rank order
|
||||
|
||||
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
|
||||
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
|
||||
- **AND** `rag.return-n` SHALL still bound total blocks
|
||||
|
||||
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
|
||||
|
||||
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
|
||||
|
||||
#### Scenario: Low quality uses top qualityScore
|
||||
|
||||
- **WHEN** post-process finishes with at least one evidence block
|
||||
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
|
||||
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
|
||||
|
||||
#### Scenario: PRECISE does not require hint support
|
||||
|
||||
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
|
||||
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
|
||||
|
||||
### Requirement: Dense mode remains available for recall comparison
|
||||
|
||||
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
|
||||
|
||||
#### Scenario: Mode dense still callable
|
||||
|
||||
- **WHEN** mode is `dense`
|
||||
- **THEN** search SHALL call dense ANN only and label hits `dense`
|
||||
|
||||
### Requirement: Hybrid quality gates MAY use optional dense distance without reordering
|
||||
|
||||
When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.
|
||||
|
||||
#### Scenario: Dense distance does not replace hybrid label
|
||||
|
||||
- **WHEN** a hybrid hit includes denseDistance
|
||||
- **THEN** scoreLabel SHALL still canonicalize to hybrid
|
||||
- **AND** post-process sort order SHALL still follow originalRank from hybrid results
|
||||
@@ -90,12 +90,12 @@
|
||||
<groupId>org.springframework.boot</groupId>
|
||||
<artifactId>spring-boot-starter-web</artifactId>
|
||||
</dependency>
|
||||
<dependency>
|
||||
<groupId>org.springframework.boot</groupId>
|
||||
<artifactId>spring-boot-devtools</artifactId>
|
||||
<scope>runtime</scope>
|
||||
<optional>true</optional>
|
||||
</dependency>
|
||||
<!--
|
||||
spring-boot-devtools removed on purpose.
|
||||
Classpath restart (restartedMain) recreates beans without reliably closing
|
||||
MilvusClientV2 gRPC channels, causing orphan channels and long hybrid RPC retries.
|
||||
Prefer full process restart: stop then `mvn spring-boot:run`.
|
||||
-->
|
||||
<dependency>
|
||||
<groupId>io.milvus</groupId>
|
||||
<artifactId>milvus-sdk-java</artifactId>
|
||||
|
||||
@@ -3,7 +3,9 @@ param(
|
||||
[string]$Fixtures = "eval\rag-retrieval\fixtures",
|
||||
[string]$RetrievedAt = "",
|
||||
[string]$KbScope = "rag-eval",
|
||||
[string]$VectorStoreMode = "spring",
|
||||
# hybrid = production main path; dense = same-collection recall baseline
|
||||
[ValidateSet("hybrid", "dense")]
|
||||
[string]$SearchMode = "hybrid",
|
||||
[switch]$SkipEval
|
||||
)
|
||||
|
||||
@@ -16,7 +18,7 @@ $mavenArgs = @(
|
||||
"-Drag.snapshot.cases=$Cases",
|
||||
"-Drag.snapshot.fixtures=$Fixtures",
|
||||
"-Dretrieval.kb-scope=$KbScope",
|
||||
"-Dretrieval.vector-store.mode=$VectorStoreMode"
|
||||
"-Dretrieval.search.mode=$SearchMode"
|
||||
)
|
||||
|
||||
if ($RetrievedAt -ne "") {
|
||||
@@ -25,10 +27,16 @@ if ($RetrievedAt -ne "") {
|
||||
|
||||
$mavenArgs += "test"
|
||||
|
||||
Write-Host "Generating RAG lookupResult fixtures from real LookupKnowledgeTool..."
|
||||
Write-Host "Generating RAG lookupResult fixtures (search.mode=$SearchMode, kb-scope=$KbScope)..."
|
||||
& mvn @mavenArgs
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
throw "Fixture generation failed with exit code $LASTEXITCODE"
|
||||
}
|
||||
|
||||
if (-not $SkipEval) {
|
||||
Write-Host "Running offline RAG retrieval baseline..."
|
||||
& python scripts\eval_rag_retrieval.py --cases $Cases --fixtures $Fixtures
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
throw "Offline eval failed with exit code $LASTEXITCODE"
|
||||
}
|
||||
}
|
||||
|
||||
@@ -4,6 +4,7 @@ import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.config.MysqlToolProperties.DataSourceProperties;
|
||||
import com.superbiz.agent.harness.agent.HarnessEvidenceTools;
|
||||
import com.superbiz.agent.harness.agent.DiagnosisAgentFactory;
|
||||
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
|
||||
import com.superbiz.agent.harness.agent.DiagnosisAgentLimits;
|
||||
import com.superbiz.agent.harness.agent.DiagnosisAgentUseCase;
|
||||
import com.superbiz.agent.harness.application.ChatApplicationUseCase;
|
||||
@@ -153,8 +154,9 @@ public class HarnessChatConfiguration {
|
||||
@Bean
|
||||
public ToolBoundary toolBoundary(DiagnosisHarnessCore core, ToolCallKeyFactory keyFactory,
|
||||
CanonicalInvocationStore store, ObjectMapper objectMapper, Clock clock,
|
||||
ToolInvocationAuditSink auditSink) {
|
||||
return new ToolBoundary(core, keyFactory, store, objectMapper, clock, auditSink);
|
||||
ToolInvocationAuditSink auditSink,
|
||||
AgentStepAuditTracker stepTracker) {
|
||||
return new ToolBoundary(core, keyFactory, store, objectMapper, clock, auditSink, stepTracker);
|
||||
}
|
||||
|
||||
@Bean
|
||||
@@ -240,8 +242,15 @@ public class HarnessChatConfiguration {
|
||||
|
||||
@Bean
|
||||
public HarnessEvidenceTools harnessEvidenceTools(RagToolAdapter rag, QueryLogsToolAdapter logs,
|
||||
MysqlToolAdapter mysql) {
|
||||
return HarnessEvidenceTools.fromAdapters(rag, logs, mysql);
|
||||
MysqlToolAdapter mysql,
|
||||
MysqlToolProperties mysqlToolProperties) {
|
||||
// query_mysql is only exposed when at least one logical datasource is configured.
|
||||
// Default application.yml has harness.mysql-tools.data-sources: {} — do not inject a dead tool.
|
||||
boolean mysqlEnabled = mysqlToolProperties != null
|
||||
&& mysqlToolProperties.getDataSources() != null
|
||||
&& mysqlToolProperties.getDataSources().values().stream()
|
||||
.anyMatch(ds -> ds != null && ds.getJdbcUrl() != null && !ds.getJdbcUrl().isBlank());
|
||||
return HarnessEvidenceTools.fromAdapters(rag, logs, mysqlEnabled ? mysql : null);
|
||||
}
|
||||
|
||||
@Bean
|
||||
@@ -250,10 +259,12 @@ public class HarnessChatConfiguration {
|
||||
AgentStepRepository steps,
|
||||
AgentReasoningAuditRepository reasoningAudits,
|
||||
DiagnosisTraceRecorder traceRecorder,
|
||||
ModelCallAuditor modelCallAuditor) {
|
||||
ModelCallAuditor modelCallAuditor,
|
||||
AgentStepAuditTracker stepTracker) {
|
||||
return new DiagnosisAgentFactory(chatModel, core, tools, mapper,
|
||||
List.of(new HarnessAgentAuditHook(
|
||||
steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits)),
|
||||
steps, mapper, DiagnosisAgentFactory.AGENT_NAME, traceRecorder, reasoningAudits,
|
||||
stepTracker)),
|
||||
traceRecorder, modelCallAuditor);
|
||||
}
|
||||
|
||||
|
||||
@@ -6,10 +6,15 @@ import org.slf4j.LoggerFactory;
|
||||
import org.springframework.context.annotation.Configuration;
|
||||
|
||||
/**
|
||||
* Milvus configuration notes.
|
||||
* Milvus 知识路径配置说明(无额外 Bean 装配)。
|
||||
*
|
||||
* <p>Knowledge RAG uses {@link MilvusHybridKnowledgeStore} (MilvusClientV2) exclusively.
|
||||
* Legacy {@code MilvusServiceClient} bean is no longer created for the knowledge path.</p>
|
||||
* <p>知识库 RAG 唯一实现:{@link MilvusHybridKnowledgeStore}({@code MilvusClientV2})。</p>
|
||||
* <ul>
|
||||
* <li>支持 dense 与 dense+BM25 {@code hybridSearch}+RRF。</li>
|
||||
* <li>不再为知识路径创建 legacy {@code MilvusServiceClient} Bean。</li>
|
||||
* <li>Spring AI {@code VectorStore} starter 仍可存在于 classpath,但只作 sidecar,
|
||||
* 不作 lookup_knowledge 主路径(starter 无 BM25 hybrid API)。</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Configuration
|
||||
public class MilvusConfig {
|
||||
|
||||
@@ -42,12 +42,32 @@ public class AgentReasoningAudit {
|
||||
@Column(name = "agent_name", nullable = false, length = 64)
|
||||
private String agentName;
|
||||
|
||||
/** True when provider returned a non-blank reasoning/thinking chain. */
|
||||
@Column(name = "reasoning_available", nullable = false)
|
||||
private Boolean reasoningAvailable;
|
||||
|
||||
/**
|
||||
* Provider chain-of-thought / reasoning_content when available.
|
||||
* Never stores tool result payloads.
|
||||
*/
|
||||
@Column(name = "reasoning_content", columnDefinition = "LONGTEXT")
|
||||
private String reasoningContent;
|
||||
|
||||
/**
|
||||
* Assistant-visible text for this model step (final prose and/or tool-call plan).
|
||||
* Tool result message bodies are not stored here.
|
||||
*/
|
||||
@Column(name = "assistant_text", columnDefinition = "LONGTEXT")
|
||||
private String assistantText;
|
||||
|
||||
/**
|
||||
* Summary of what was stored:
|
||||
* PROVIDER_REASONING+ASSISTANT_TEXT | PROVIDER_REASONING | ASSISTANT_TEXT | TOOL_CALL_PLAN | NONE
|
||||
*/
|
||||
@Column(name = "content_source", length = 64)
|
||||
private String contentSource;
|
||||
|
||||
/** UTF-8 byte length of reasoning_content + assistant_text combined. */
|
||||
@Column(name = "content_bytes", nullable = false)
|
||||
private Integer contentBytes;
|
||||
|
||||
|
||||
@@ -41,6 +41,10 @@ public class DiagnosisRun {
|
||||
@Column(name = "query", nullable = false, columnDefinition = "TEXT")
|
||||
private String query;
|
||||
|
||||
/** Extracted final conclusion text, parallel to query for simple readout. */
|
||||
@Column(name = "conclusion", columnDefinition = "TEXT")
|
||||
private String conclusion;
|
||||
|
||||
@Column(name = "status", length = 16)
|
||||
private String status = "PENDING";
|
||||
|
||||
|
||||
@@ -71,6 +71,8 @@ public class DiagnosisTraceResponse {
|
||||
private String runId;
|
||||
private String sessionId;
|
||||
private String query;
|
||||
/** Extracted conclusion text, parallel to query. */
|
||||
private String conclusion;
|
||||
private String status;
|
||||
private String agentFlow;
|
||||
private String intent;
|
||||
@@ -164,7 +166,11 @@ public class DiagnosisTraceResponse {
|
||||
private Integer stepIndex;
|
||||
private String agentName;
|
||||
private Boolean reasoningAvailable;
|
||||
/** Provider thinking / CoT when available. */
|
||||
private String reasoningContent;
|
||||
/** Assistant visible text and/or tool-call plan (no tool results). */
|
||||
private String assistantText;
|
||||
private String contentSource;
|
||||
private Integer contentBytes;
|
||||
private LocalDateTime createdAt;
|
||||
}
|
||||
|
||||
@@ -31,7 +31,7 @@ public class EvidencePostprocessResult {
|
||||
private RerankTrace rerankTrace;
|
||||
|
||||
/**
|
||||
* 排序第一名的 baseScore(0~1 相似度,不含规则 boost)。
|
||||
* 排序第一名(originalRank 最优)的 qualityScore(0~1,越大越好)。
|
||||
* 用于 isLowQuality 与 attempt.topSimilarity。
|
||||
*/
|
||||
private Double topSimilarity;
|
||||
|
||||
@@ -9,70 +9,59 @@ import java.util.Map;
|
||||
/**
|
||||
* L1 向量命中后、后处理前的统一候选结构。
|
||||
*
|
||||
* <p>由检索适配器从 {@code KnowledgeSearchHit} / 向量结果映射而来。
|
||||
* 后处理会基于它做归一化、规则 boost、chunk 级去重并生成 {@link EvidenceBlock}。</p>
|
||||
* <p>后处理:{@code RetrievalScoreNormalizer} → qualityScore;按 {@link #originalRank} 保序;
|
||||
* evidenceKey 去重 / 截断;不再用关键词 boost 改序。</p>
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RetrievedEvidenceCandidate {
|
||||
|
||||
/** 向量库记录 id。 */
|
||||
private String id;
|
||||
|
||||
/**
|
||||
* 文档级 id(metadata.docId 等)。
|
||||
* 用于每文档 chunk 上限;不等于 evidenceKey。
|
||||
*/
|
||||
private String docId;
|
||||
|
||||
/** 文档内切片序号;可能为空(老数据)。 */
|
||||
private Integer chunkIndex;
|
||||
|
||||
/**
|
||||
* 片段级去重/投影主键。
|
||||
* 通常为 docId#chunk-N,fallback 为 vector:{id}。
|
||||
*/
|
||||
private String evidenceKey;
|
||||
|
||||
/**
|
||||
* 来源标识(_source / source / filePath / docId 等)。
|
||||
* 可与同文档其他 chunk 重复;不再作为唯一去重键。
|
||||
*/
|
||||
private String source;
|
||||
|
||||
private String title;
|
||||
|
||||
private String breadcrumb;
|
||||
|
||||
/** chunk 正文原文(后处理前未截断或仅底层原样)。 */
|
||||
private String content;
|
||||
|
||||
/** 固定为 L1(向量层);预留多路召回标记。 */
|
||||
private String retrievalLayer;
|
||||
|
||||
/** 所属 attempt 名,如 FILTERED_VECTOR。 */
|
||||
private String retrievalAttempt;
|
||||
|
||||
/**
|
||||
* 兼容 L2 距离分(越小越相似),后处理会 normalize 成 baseScore。
|
||||
* 引擎主分:dense=L2;hybrid=融合分。量纲由 {@link #scoreLabel} 解释。
|
||||
*/
|
||||
private Double score;
|
||||
|
||||
/** 底层原始分。 */
|
||||
private Double rawScore;
|
||||
|
||||
/** rawScore 语义标签:l2_distance / similarity。 */
|
||||
/**
|
||||
* {@code dense} | {@code hybrid}(及可被 canonicalize 的历史别名)。
|
||||
*/
|
||||
private String scoreLabel;
|
||||
|
||||
/** 向量召回顺序(从 1 起),规则 rerank 前的名次。 */
|
||||
/**
|
||||
* 检索返回名次(从 1 起)——后处理排序权威。
|
||||
*/
|
||||
private Integer originalRank;
|
||||
|
||||
/**
|
||||
* 扁平化 metadata(string map)。
|
||||
* 可能含 docId、chunkIndex、category、kb_scope 等。
|
||||
* hybrid 命中可选的 dense L2,仅供质量闸门;不参与排序。
|
||||
*/
|
||||
private Double denseDistance;
|
||||
|
||||
private Map<String, String> metadata;
|
||||
|
||||
/** 初步命中原因,后处理会追加 boost reasons。 */
|
||||
/**
|
||||
* 初步命中原因;后处理可追加 L0 重叠解释(无分值)。
|
||||
*/
|
||||
private List<String> hitReasons;
|
||||
}
|
||||
|
||||
@@ -31,6 +31,10 @@ public final class HarnessEvidenceTools {
|
||||
private final List<ToolCallback> callbacks;
|
||||
private final Map<String, EvidenceToolInvoker> invokers;
|
||||
|
||||
/**
|
||||
* @param mysqlInvoker optional; when null, {@code query_mysql} is not registered
|
||||
* (no logical datasource configured / tool unavailable).
|
||||
*/
|
||||
public HarnessEvidenceTools(EvidenceToolInvoker ragInvoker,
|
||||
EvidenceToolInvoker logsInvoker,
|
||||
EvidenceToolInvoker mysqlInvoker) {
|
||||
@@ -39,28 +43,38 @@ public final class HarnessEvidenceTools {
|
||||
Objects.requireNonNull(ragInvoker, "ragInvoker must not be null"));
|
||||
registered.put(AgentToolContracts.QUERY_LOGS,
|
||||
Objects.requireNonNull(logsInvoker, "logsInvoker must not be null"));
|
||||
registered.put(AgentToolContracts.QUERY_MYSQL,
|
||||
Objects.requireNonNull(mysqlInvoker, "mysqlInvoker must not be null"));
|
||||
if (mysqlInvoker != null) {
|
||||
registered.put(AgentToolContracts.QUERY_MYSQL, mysqlInvoker);
|
||||
}
|
||||
this.invokers = Map.copyOf(registered);
|
||||
this.callbacks = List.of(
|
||||
definition(AgentToolContracts.LOOKUP_KNOWLEDGE,
|
||||
AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolCall.class),
|
||||
definition(AgentToolContracts.QUERY_LOGS,
|
||||
AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsToolCall.class),
|
||||
definition(AgentToolContracts.QUERY_MYSQL,
|
||||
|
||||
List<ToolCallback> built = new java.util.ArrayList<>();
|
||||
built.add(definition(AgentToolContracts.LOOKUP_KNOWLEDGE,
|
||||
AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolCall.class));
|
||||
built.add(definition(AgentToolContracts.QUERY_LOGS,
|
||||
AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsToolCall.class));
|
||||
if (mysqlInvoker != null) {
|
||||
built.add(definition(AgentToolContracts.QUERY_MYSQL,
|
||||
AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolCall.class));
|
||||
}
|
||||
this.callbacks = List.copyOf(built);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param mysqlAdapter optional; omit registration when null or when no datasources are wired
|
||||
*/
|
||||
public static HarnessEvidenceTools fromAdapters(RagToolAdapter ragAdapter,
|
||||
QueryLogsToolAdapter logsAdapter,
|
||||
MysqlToolAdapter mysqlAdapter) {
|
||||
Objects.requireNonNull(ragAdapter, "ragAdapter must not be null");
|
||||
Objects.requireNonNull(logsAdapter, "logsAdapter must not be null");
|
||||
Objects.requireNonNull(mysqlAdapter, "mysqlAdapter must not be null");
|
||||
EvidenceToolInvoker mysql = mysqlAdapter == null
|
||||
? null
|
||||
: bridge(AgentToolContracts.QUERY_MYSQL, mysqlAdapter::execute);
|
||||
return new HarnessEvidenceTools(
|
||||
bridge(AgentToolContracts.LOOKUP_KNOWLEDGE, ragAdapter::execute),
|
||||
bridge(AgentToolContracts.QUERY_LOGS, logsAdapter::execute),
|
||||
bridge(AgentToolContracts.QUERY_MYSQL, mysqlAdapter::execute));
|
||||
mysql);
|
||||
}
|
||||
|
||||
public List<ToolCallback> callbacks() {
|
||||
|
||||
+24
-4
@@ -6,18 +6,21 @@ import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.fasterxml.jackson.databind.ObjectReader;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
|
||||
import com.superbiz.agent.harness.audit.ModelCallComponent;
|
||||
import com.superbiz.agent.harness.audit.RunConclusionExtractor;
|
||||
import com.superbiz.agent.harness.contract.IntentType;
|
||||
import com.superbiz.agent.harness.contract.PreviousTurn;
|
||||
import com.superbiz.agent.harness.contract.PublishedResult;
|
||||
import com.superbiz.agent.harness.contract.ReleaseOutcome;
|
||||
import com.superbiz.agent.harness.core.RunBudgetUsage;
|
||||
import com.superbiz.agent.harness.core.RunContext;
|
||||
import com.superbiz.agent.harness.audit.ModelCallComponent;
|
||||
import com.superbiz.agent.repository.ChatSessionRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.lang.Nullable;
|
||||
import org.springframework.stereotype.Component;
|
||||
import org.springframework.transaction.annotation.Transactional;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.List;
|
||||
@@ -32,19 +35,28 @@ public class JpaChatRunStore implements ChatRunStore {
|
||||
private final ObjectMapper objectMapper;
|
||||
private final ObjectReader publishedReader;
|
||||
private final PublishedResultPolicy publishedPolicy;
|
||||
private final AgentStepAuditTracker stepTracker;
|
||||
|
||||
public JpaChatRunStore(ChatSessionRepository chatSessions,
|
||||
DiagnosisRunRepository runs,
|
||||
ObjectMapper objectMapper) {
|
||||
this(chatSessions, runs, objectMapper,
|
||||
new PublishedResultPolicy(PreviousTurnLimits.defaults()));
|
||||
new PublishedResultPolicy(PreviousTurnLimits.defaults()), null);
|
||||
}
|
||||
|
||||
public JpaChatRunStore(ChatSessionRepository chatSessions,
|
||||
DiagnosisRunRepository runs,
|
||||
ObjectMapper objectMapper,
|
||||
PublishedResultPolicy publishedPolicy) {
|
||||
this(chatSessions, runs, objectMapper, publishedPolicy, null);
|
||||
}
|
||||
|
||||
@Autowired
|
||||
public JpaChatRunStore(ChatSessionRepository chatSessions,
|
||||
DiagnosisRunRepository runs,
|
||||
ObjectMapper objectMapper,
|
||||
PublishedResultPolicy publishedPolicy) {
|
||||
PublishedResultPolicy publishedPolicy,
|
||||
@Nullable AgentStepAuditTracker stepTracker) {
|
||||
this.chatSessions = Objects.requireNonNull(chatSessions, "chatSessions must not be null");
|
||||
this.runs = Objects.requireNonNull(runs, "runs must not be null");
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||
@@ -52,6 +64,7 @@ public class JpaChatRunStore implements ChatRunStore {
|
||||
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
|
||||
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
|
||||
this.publishedPolicy = Objects.requireNonNull(publishedPolicy, "publishedPolicy must not be null");
|
||||
this.stepTracker = stepTracker;
|
||||
}
|
||||
|
||||
@Override
|
||||
@@ -110,11 +123,13 @@ public class JpaChatRunStore implements ChatRunStore {
|
||||
String safeContentJson, PublishedResult publishedResult, int durationMs) {
|
||||
Objects.requireNonNull(context, "context must not be null");
|
||||
Objects.requireNonNull(outcome, "outcome must not be null");
|
||||
try {
|
||||
DiagnosisRun run = requiredRun(context.runId());
|
||||
run.setIntent(intent);
|
||||
run.setReleaseOutcome(outcome);
|
||||
run.setStatus(status(outcome));
|
||||
run.setAnswer(safeContentJson);
|
||||
run.setConclusion(RunConclusionExtractor.extract(objectMapper, safeContentJson));
|
||||
run.setPublishedResult(outcome == ReleaseOutcome.SUCCESS && intent == IntentType.DIAGNOSIS
|
||||
&& publishedResult != null ? write(publishedResult) : null);
|
||||
run.setTotalDurationMs(Math.max(0, durationMs));
|
||||
@@ -132,6 +147,11 @@ public class JpaChatRunStore implements ChatRunStore {
|
||||
chatSessions.save(session);
|
||||
});
|
||||
}
|
||||
} finally {
|
||||
if (stepTracker != null) {
|
||||
stepTracker.clear(context.runId());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private Optional<PreviousTurn> readPreviousTurn(DiagnosisRun run) {
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
package com.superbiz.agent.harness.audit;
|
||||
|
||||
import org.springframework.stereotype.Component;
|
||||
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Binds the in-flight {@code agent_step.id} for a run so tool audits can set {@code step_id}.
|
||||
*
|
||||
* <p>Lifecycle: {@link #bind} on Agent {@code beforeModel} (step row created). Binding stays
|
||||
* until the next {@code beforeModel} for the same run, covering tool execution that happens
|
||||
* after {@code afterModel} emits tool_calls. {@link #clear} on run finish is optional cleanup.</p>
|
||||
*/
|
||||
@Component
|
||||
public final class AgentStepAuditTracker {
|
||||
|
||||
private final ConcurrentHashMap<String, Long> currentStepIdByRun = new ConcurrentHashMap<>();
|
||||
|
||||
public void bind(String runId, Long stepId) {
|
||||
if (runId == null || runId.isBlank() || stepId == null) {
|
||||
return;
|
||||
}
|
||||
currentStepIdByRun.put(runId, stepId);
|
||||
}
|
||||
|
||||
public Long currentStepId(String runId) {
|
||||
if (runId == null || runId.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return currentStepIdByRun.get(runId);
|
||||
}
|
||||
|
||||
public void clear(String runId) {
|
||||
if (runId == null || runId.isBlank()) {
|
||||
return;
|
||||
}
|
||||
currentStepIdByRun.remove(runId);
|
||||
}
|
||||
}
|
||||
@@ -7,32 +7,53 @@ import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.AgentReasoningAudit;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.repository.AgentReasoningAuditRepository;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.deepseek.DeepSeekAssistantMessage;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.lang.reflect.Method;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.Objects;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Persists per-model-step audit for the diagnosis agent.
|
||||
*
|
||||
* <p>LLM-facing content policy for {@link AgentReasoningAudit}:</p>
|
||||
* <ul>
|
||||
* <li>{@code reasoning_content}: provider thinking/CoT when present</li>
|
||||
* <li>{@code assistant_text}: assistant visible text and/or tool-call plan</li>
|
||||
* <li>Tool <em>result</em> payloads are never stored here (use {@code tool_invocation})</li>
|
||||
* </ul>
|
||||
*/
|
||||
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
|
||||
public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
|
||||
private static final Logger log = LoggerFactory.getLogger(HarnessAgentAuditHook.class);
|
||||
private static final int MAX_TEXT_CHARS = 32_000;
|
||||
|
||||
public static final String SOURCE_BOTH = "PROVIDER_REASONING+ASSISTANT_TEXT";
|
||||
public static final String SOURCE_PROVIDER = "PROVIDER_REASONING";
|
||||
public static final String SOURCE_ASSISTANT = "ASSISTANT_TEXT";
|
||||
public static final String SOURCE_TOOL_PLAN = "TOOL_CALL_PLAN";
|
||||
public static final String SOURCE_NONE = "NONE";
|
||||
|
||||
private final AgentStepRepository repository;
|
||||
private final ObjectMapper objectMapper;
|
||||
private final String agentName;
|
||||
private final DiagnosisTraceRecorder traceRecorder;
|
||||
private final AgentReasoningAuditRepository reasoningRepository;
|
||||
private final AgentStepAuditTracker stepTracker;
|
||||
private final ConcurrentHashMap<String, Integer> stepCounters = new ConcurrentHashMap<>();
|
||||
private final ConcurrentHashMap<String, PendingStep> pendingSteps = new ConcurrentHashMap<>();
|
||||
|
||||
@@ -42,12 +63,19 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
|
||||
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
|
||||
String agentName, DiagnosisTraceRecorder traceRecorder) {
|
||||
this(repository, objectMapper, agentName, traceRecorder, null);
|
||||
this(repository, objectMapper, agentName, traceRecorder, null, null);
|
||||
}
|
||||
|
||||
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
|
||||
String agentName, DiagnosisTraceRecorder traceRecorder,
|
||||
AgentReasoningAuditRepository reasoningRepository) {
|
||||
this(repository, objectMapper, agentName, traceRecorder, reasoningRepository, null);
|
||||
}
|
||||
|
||||
public HarnessAgentAuditHook(AgentStepRepository repository, ObjectMapper objectMapper,
|
||||
String agentName, DiagnosisTraceRecorder traceRecorder,
|
||||
AgentReasoningAuditRepository reasoningRepository,
|
||||
AgentStepAuditTracker stepTracker) {
|
||||
this.repository = Objects.requireNonNull(repository, "repository must not be null");
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||
if (agentName == null || agentName.isBlank()) {
|
||||
@@ -56,6 +84,7 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
this.agentName = agentName;
|
||||
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
|
||||
this.reasoningRepository = reasoningRepository;
|
||||
this.stepTracker = stepTracker;
|
||||
}
|
||||
|
||||
@Override
|
||||
@@ -82,6 +111,9 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
.hasToolCall(false)
|
||||
.build());
|
||||
pending = new PendingStep(saved.getId(), pending.startedNanos(), pending.input());
|
||||
if (stepTracker != null && saved.getId() != null) {
|
||||
stepTracker.bind(identity.runId(), saved.getId());
|
||||
}
|
||||
} catch (RuntimeException exception) {
|
||||
log.warn("Failed to persist AgentStep audit before model: agent={}", agentName);
|
||||
}
|
||||
@@ -106,14 +138,15 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
? List.of()
|
||||
: assistant.getToolCalls().stream().map(AssistantMessage.ToolCall::name)
|
||||
.distinct().sorted().toList();
|
||||
ReasoningContent reasoning = reasoningContent(assistant);
|
||||
Map<String, Object> output = outputMetadata(assistant, toolNames, reasoning);
|
||||
LlmTurnContent turn = llmTurnContent(assistant, toolNames);
|
||||
Map<String, Object> output = outputMetadata(assistant, toolNames, turn);
|
||||
int durationMs = durationMillis(pending.startedNanos());
|
||||
try {
|
||||
AgentStep step = pending.id() == null ? null : repository.findById(pending.id()).orElse(null);
|
||||
if (step != null) {
|
||||
step.setModelOutput(write(output));
|
||||
step.setThought(null);
|
||||
// Keep thought aligned with provider reasoning when present; else assistant text.
|
||||
step.setThought(firstNonBlank(turn.reasoningContent(), turn.assistantText()));
|
||||
step.setHasToolCall(!toolNames.isEmpty());
|
||||
step.setDurationMs(durationMs);
|
||||
repository.save(step);
|
||||
@@ -121,7 +154,7 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
} catch (RuntimeException exception) {
|
||||
log.warn("Failed to complete AgentStep audit: agent={}", agentName);
|
||||
}
|
||||
persistReasoning(identity, stepIndex, reasoning);
|
||||
persistReasoning(identity, stepIndex, turn);
|
||||
traceRecorder.record(TraceAuditEvents.agentModelStep(
|
||||
identity.sessionId(), identity.runId(), agentName, stepIndex,
|
||||
durationMs, pending.input(), output));
|
||||
@@ -139,32 +172,148 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
}
|
||||
|
||||
private Map<String, Object> outputMetadata(AssistantMessage assistant, List<String> toolNames,
|
||||
ReasoningContent reasoning) {
|
||||
LlmTurnContent turn) {
|
||||
Map<String, Object> metadata = new LinkedHashMap<>();
|
||||
metadata.put("has_text", assistant != null
|
||||
&& assistant.getText() != null && !assistant.getText().isBlank());
|
||||
metadata.put("tool_names", toolNames);
|
||||
metadata.put("reasoning_available", reasoning.available());
|
||||
metadata.put("reasoning_bytes", reasoning.bytes());
|
||||
metadata.put("reasoning_available", turn.reasoningAvailable());
|
||||
metadata.put("reasoning_bytes", turn.reasoningBytes());
|
||||
metadata.put("assistant_bytes", turn.assistantBytes());
|
||||
metadata.put("content_source", turn.contentSource());
|
||||
return metadata;
|
||||
}
|
||||
|
||||
private ReasoningContent reasoningContent(AssistantMessage assistant) {
|
||||
if (assistant == null || assistant.getMetadata() == null) {
|
||||
return ReasoningContent.empty();
|
||||
}
|
||||
for (String key : List.of("reasoning_content", "reasoningContent", "reasoning", "thinking")) {
|
||||
Object value = assistant.getMetadata().get(key);
|
||||
if (value instanceof CharSequence text && !text.toString().isBlank()) {
|
||||
String content = bounded(text.toString());
|
||||
return new ReasoningContent(content, true,
|
||||
content.getBytes(StandardCharsets.UTF_8).length);
|
||||
}
|
||||
}
|
||||
return ReasoningContent.empty();
|
||||
/**
|
||||
* Captures both provider reasoning and assistant-visible text.
|
||||
* Tool <em>results</em> are never included.
|
||||
*/
|
||||
static LlmTurnContent llmTurnContent(AssistantMessage assistant, List<String> toolNames) {
|
||||
String reasoning = providerReasoning(assistant);
|
||||
String assistantBody = assistantText(assistant);
|
||||
String toolPlan = toolCallPlan(assistant, toolNames);
|
||||
|
||||
String assistantCombined = joinNonBlank("\n", assistantBody, toolPlan);
|
||||
boolean hasReasoning = hasText(reasoning);
|
||||
boolean hasAssistant = hasText(assistantCombined);
|
||||
|
||||
String source;
|
||||
if (hasReasoning && hasAssistant) {
|
||||
source = SOURCE_BOTH;
|
||||
} else if (hasReasoning) {
|
||||
source = SOURCE_PROVIDER;
|
||||
} else if (hasText(assistantBody)) {
|
||||
source = SOURCE_ASSISTANT;
|
||||
} else if (hasText(toolPlan)) {
|
||||
source = SOURCE_TOOL_PLAN;
|
||||
} else {
|
||||
source = SOURCE_NONE;
|
||||
}
|
||||
|
||||
private void persistReasoning(AuditIdentity identity, int stepIndex, ReasoningContent reasoning) {
|
||||
String reasoningBound = bound(reasoning);
|
||||
String assistantBound = bound(assistantCombined);
|
||||
int bytes = utf8Bytes(reasoningBound) + utf8Bytes(assistantBound);
|
||||
return new LlmTurnContent(
|
||||
hasReasoning,
|
||||
reasoningBound,
|
||||
assistantBound,
|
||||
source,
|
||||
utf8Bytes(reasoningBound),
|
||||
utf8Bytes(assistantBound),
|
||||
bytes);
|
||||
}
|
||||
|
||||
/**
|
||||
* DeepSeek puts CoT on {@link DeepSeekAssistantMessage#getReasoningContent()},
|
||||
* <em>not</em> on {@link AssistantMessage#getMetadata()}. Older docs/tests used metadata keys;
|
||||
* keep those as fallback for mocks and non-DeepSeek providers.
|
||||
*/
|
||||
static String providerReasoning(AssistantMessage assistant) {
|
||||
if (assistant == null) {
|
||||
return null;
|
||||
}
|
||||
// 1) Native DeepSeek message field (primary path in production)
|
||||
if (assistant instanceof DeepSeekAssistantMessage deepSeek) {
|
||||
String nativeReasoning = blankToNull(deepSeek.getReasoningContent());
|
||||
if (nativeReasoning != null) {
|
||||
return nativeReasoning;
|
||||
}
|
||||
}
|
||||
// 2) Reflective getReasoningContent() for subclasses / reloaded types
|
||||
String reflective = invokeReasoningGetter(assistant);
|
||||
if (reflective != null) {
|
||||
return reflective;
|
||||
}
|
||||
// 3) Metadata keys (tests / other providers)
|
||||
Map<String, Object> metadata = assistant.getMetadata();
|
||||
if (metadata != null) {
|
||||
for (String key : List.of(
|
||||
"reasoning_content", "reasoningContent", "reasoning", "thinking",
|
||||
"reasoning_text", "reasoningText")) {
|
||||
Object value = metadata.get(key);
|
||||
if (value instanceof CharSequence text) {
|
||||
String trimmed = blankToNull(text.toString());
|
||||
if (trimmed != null) {
|
||||
return trimmed;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static String invokeReasoningGetter(AssistantMessage assistant) {
|
||||
try {
|
||||
Method method = assistant.getClass().getMethod("getReasoningContent");
|
||||
Object value = method.invoke(assistant);
|
||||
return value instanceof CharSequence text ? blankToNull(text.toString()) : null;
|
||||
} catch (ReflectiveOperationException ignored) {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private static String blankToNull(String value) {
|
||||
if (value == null) {
|
||||
return null;
|
||||
}
|
||||
String trimmed = value.trim();
|
||||
return trimmed.isEmpty() ? null : trimmed;
|
||||
}
|
||||
|
||||
private static String assistantText(AssistantMessage assistant) {
|
||||
if (assistant == null || assistant.getText() == null || assistant.getText().isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return assistant.getText().trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Records which tools the model decided to call (names + arg preview), not tool outputs.
|
||||
*/
|
||||
private static String toolCallPlan(AssistantMessage assistant, List<String> toolNames) {
|
||||
if (assistant == null || assistant.getToolCalls() == null || assistant.getToolCalls().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
List<String> lines = new ArrayList<>();
|
||||
lines.add("tool_calls:");
|
||||
for (AssistantMessage.ToolCall call : assistant.getToolCalls()) {
|
||||
if (call == null) {
|
||||
continue;
|
||||
}
|
||||
String name = call.name() == null ? "?" : call.name();
|
||||
String args = call.arguments() == null ? "" : call.arguments().trim();
|
||||
if (args.length() > 500) {
|
||||
args = args.substring(0, 500) + "...";
|
||||
}
|
||||
lines.add("- " + name + (args.isEmpty() ? "" : " args=" + args));
|
||||
}
|
||||
if (lines.size() == 1 && toolNames != null && !toolNames.isEmpty()) {
|
||||
lines.add("- " + String.join(", ", toolNames));
|
||||
}
|
||||
return lines.size() <= 1 ? null : String.join("\n", lines);
|
||||
}
|
||||
|
||||
private void persistReasoning(AuditIdentity identity, int stepIndex, LlmTurnContent turn) {
|
||||
if (reasoningRepository == null) {
|
||||
return;
|
||||
}
|
||||
@@ -174,18 +323,51 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
.runId(identity.runId())
|
||||
.stepIndex(stepIndex)
|
||||
.agentName(agentName)
|
||||
.reasoningAvailable(reasoning.available())
|
||||
.reasoningContent(reasoning.content())
|
||||
.contentBytes(reasoning.bytes())
|
||||
.reasoningAvailable(turn.reasoningAvailable())
|
||||
.reasoningContent(turn.reasoningContent())
|
||||
.assistantText(turn.assistantText())
|
||||
.contentSource(turn.contentSource())
|
||||
.contentBytes(turn.totalBytes())
|
||||
.build());
|
||||
} catch (RuntimeException exception) {
|
||||
log.warn("Failed to persist reasoning audit: agent={}, step={}", agentName, stepIndex);
|
||||
}
|
||||
}
|
||||
|
||||
private static String bounded(String value) {
|
||||
int maxChars = 32_000;
|
||||
return value.length() <= maxChars ? value : value.substring(0, maxChars);
|
||||
private static String bound(String value) {
|
||||
if (value == null) {
|
||||
return null;
|
||||
}
|
||||
return value.length() <= MAX_TEXT_CHARS ? value : value.substring(0, MAX_TEXT_CHARS);
|
||||
}
|
||||
|
||||
private static int utf8Bytes(String value) {
|
||||
return value == null ? 0 : value.getBytes(StandardCharsets.UTF_8).length;
|
||||
}
|
||||
|
||||
private static String joinNonBlank(String sep, String a, String b) {
|
||||
boolean ha = hasText(a);
|
||||
boolean hb = hasText(b);
|
||||
if (ha && hb) {
|
||||
return a + sep + b;
|
||||
}
|
||||
if (ha) {
|
||||
return a;
|
||||
}
|
||||
if (hb) {
|
||||
return b;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String a, String b) {
|
||||
if (hasText(a)) {
|
||||
return a;
|
||||
}
|
||||
if (hasText(b)) {
|
||||
return b;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private AuditIdentity identity(RunnableConfig config) {
|
||||
@@ -233,10 +415,14 @@ public final class HarnessAgentAuditHook extends MessagesModelHook {
|
||||
private record AuditIdentity(String sessionId, String runId) {
|
||||
}
|
||||
|
||||
private record ReasoningContent(String content, boolean available, int bytes) {
|
||||
private static ReasoningContent empty() {
|
||||
return new ReasoningContent(null, false, 0);
|
||||
}
|
||||
record LlmTurnContent(
|
||||
boolean reasoningAvailable,
|
||||
String reasoningContent,
|
||||
String assistantText,
|
||||
String contentSource,
|
||||
int reasoningBytes,
|
||||
int assistantBytes,
|
||||
int totalBytes) {
|
||||
}
|
||||
|
||||
private record PendingStep(Long id, long startedNanos, Map<String, Object> input) {
|
||||
|
||||
@@ -1,66 +1,247 @@
|
||||
package com.superbiz.agent.harness.audit;
|
||||
|
||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.harness.contract.InvocationStatus;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import org.springframework.beans.factory.ObjectProvider;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.stereotype.Component;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Iterator;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Persists tool invocation audit rows.
|
||||
*
|
||||
* <p>For {@code lookup_knowledge}, enriches RAG-specific columns from internal LookupResult
|
||||
* JSON when present. {@code relevance_level} stores RAG PRECISE/REFERENCE (not evidence_status).
|
||||
* {@code evidence_status} remains in {@code retrieval_details} for all tools.</p>
|
||||
*/
|
||||
@Component
|
||||
public final class JpaToolInvocationAuditSink implements ToolInvocationAuditSink {
|
||||
|
||||
/** Max chars stored for free-text query / sql / topic style fields. */
|
||||
private static final int MAX_TEXT_PREVIEW = 160;
|
||||
private static final int MAX_REQUEST_KEYS = 12;
|
||||
private static final Set<String> SENSITIVE_KEY_FRAGMENTS = Set.of(
|
||||
"password", "passwd", "secret", "token", "apikey", "api_key",
|
||||
"authorization", "credential", "private_key", "access_key");
|
||||
|
||||
private final ToolInvocationRepository repository;
|
||||
private final ObjectMapper objectMapper;
|
||||
private final DiagnosisTraceRecorder traceRecorder;
|
||||
|
||||
public JpaToolInvocationAuditSink(ToolInvocationRepository repository, ObjectMapper objectMapper) {
|
||||
this(repository, objectMapper, DiagnosisTraceRecorder.noop());
|
||||
}
|
||||
private final RagLookupAuditEnricher ragEnricher;
|
||||
|
||||
@Autowired
|
||||
public JpaToolInvocationAuditSink(ToolInvocationRepository repository, ObjectMapper objectMapper,
|
||||
DiagnosisTraceRecorder traceRecorder) {
|
||||
public JpaToolInvocationAuditSink(ToolInvocationRepository repository,
|
||||
ObjectMapper objectMapper,
|
||||
DiagnosisTraceRecorder traceRecorder,
|
||||
ObjectProvider<RagLookupAuditEnricher> ragEnricherProvider) {
|
||||
this.repository = Objects.requireNonNull(repository, "repository must not be null");
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
|
||||
this.ragEnricher = ragEnricherProvider == null ? null : ragEnricherProvider.getIfAvailable();
|
||||
}
|
||||
|
||||
/** Test helper with explicit enricher (may be null). */
|
||||
static JpaToolInvocationAuditSink forTest(ToolInvocationRepository repository,
|
||||
ObjectMapper objectMapper,
|
||||
DiagnosisTraceRecorder traceRecorder,
|
||||
RagLookupAuditEnricher ragEnricher) {
|
||||
return new JpaToolInvocationAuditSink(repository, objectMapper, traceRecorder, ragEnricher, true);
|
||||
}
|
||||
|
||||
private JpaToolInvocationAuditSink(ToolInvocationRepository repository,
|
||||
ObjectMapper objectMapper,
|
||||
DiagnosisTraceRecorder traceRecorder,
|
||||
RagLookupAuditEnricher ragEnricher,
|
||||
boolean testMarker) {
|
||||
this.repository = Objects.requireNonNull(repository, "repository must not be null");
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||
this.traceRecorder = Objects.requireNonNull(traceRecorder, "traceRecorder must not be null");
|
||||
this.ragEnricher = ragEnricher;
|
||||
}
|
||||
|
||||
@Override
|
||||
public void record(ToolInvocationAuditEvent event) {
|
||||
Objects.requireNonNull(event, "event must not be null");
|
||||
traceRecorder.record(TraceAuditEvents.toolInvocation(event));
|
||||
|
||||
Map<String, Object> details = baseResultMetadata(event);
|
||||
String retrievalLayer = "HARNESS";
|
||||
Integer l0 = null;
|
||||
Integer l1 = null;
|
||||
String relevanceLevel = null;
|
||||
Boolean truncated = false;
|
||||
String dedupReason = null;
|
||||
String outputPreview = "status=%s,evidence_status=%s".formatted(
|
||||
event.status(), event.evidenceStatus());
|
||||
|
||||
if (ragEnricher != null && ragEnricher.supports(event.toolName())) {
|
||||
RagLookupAuditEnricher.Enrichment enrichment =
|
||||
ragEnricher.enrich(event.rawResultJson(), event.agentResultJson());
|
||||
if (enrichment.retrievalDetails() != null) {
|
||||
details.putAll(enrichment.retrievalDetails());
|
||||
}
|
||||
if (enrichment.retrievalLayer() != null && !enrichment.retrievalLayer().isBlank()) {
|
||||
retrievalLayer = enrichment.retrievalLayer();
|
||||
}
|
||||
l0 = enrichment.l0MatchCount();
|
||||
l1 = enrichment.l1MatchCount();
|
||||
// RAG semantic level only — never overwrite with evidence_status enum names
|
||||
relevanceLevel = enrichment.relevanceLevel();
|
||||
if (enrichment.truncated() != null) {
|
||||
truncated = enrichment.truncated();
|
||||
}
|
||||
dedupReason = enrichment.dedupReason();
|
||||
if (enrichment.outputPreview() != null && !enrichment.outputPreview().isBlank()) {
|
||||
outputPreview = enrichment.outputPreview()
|
||||
+ ",status=" + event.status()
|
||||
+ ",evidence_status=" + event.evidenceStatus();
|
||||
}
|
||||
}
|
||||
|
||||
// Always keep harness evidence_status in details (column relevance_level is RAG-only when set)
|
||||
details.put("evidence_status", event.evidenceStatus().name());
|
||||
details.put("invocation_status", event.status().name());
|
||||
|
||||
repository.save(ToolInvocation.builder()
|
||||
.sessionId(event.sessionId())
|
||||
.runId(event.runId())
|
||||
.stepId(event.stepId())
|
||||
.toolName(event.toolName())
|
||||
.inputParams(write(inputMetadata(event)))
|
||||
.outputPreview("status=%s,evidence_status=%s".formatted(
|
||||
event.status(), event.evidenceStatus()))
|
||||
.outputPreview(outputPreview)
|
||||
.outputLength(event.agentResultBytes())
|
||||
.retrievalLayer("HARNESS")
|
||||
.isTruncated(false)
|
||||
.relevanceLevel(event.evidenceStatus().name())
|
||||
.retrievalDetails(write(resultMetadata(event)))
|
||||
.retrievalLayer(retrievalLayer)
|
||||
.l0MatchCount(l0)
|
||||
.l1MatchCount(l1)
|
||||
.isTruncated(Boolean.TRUE.equals(truncated))
|
||||
.relevanceLevel(relevanceLevel)
|
||||
.dedupReason(dedupReason)
|
||||
.retrievalDetails(write(details))
|
||||
.durationMs(event.durationMs())
|
||||
.success(event.status() == InvocationStatus.READY)
|
||||
.errorMessage(event.errorCode())
|
||||
.build());
|
||||
}
|
||||
|
||||
private Map<String, Object> inputMetadata(ToolInvocationAuditEvent event) {
|
||||
/**
|
||||
* Bounded request audit: always tool_call_id + request_bytes; optionally step_id and
|
||||
* safe scalar fields from request JSON (query/topic/region/limit/...). Never stores
|
||||
* password/token-like keys or nested blobs wholesale.
|
||||
*/
|
||||
Map<String, Object> inputMetadata(ToolInvocationAuditEvent event) {
|
||||
Map<String, Object> metadata = new LinkedHashMap<>();
|
||||
metadata.put("tool_call_id", event.toolCallId());
|
||||
metadata.put("request_bytes", event.requestBytes());
|
||||
if (event.stepId() != null) {
|
||||
metadata.put("step_id", event.stepId());
|
||||
}
|
||||
appendSafeRequestFields(metadata, event.requestJson());
|
||||
return metadata;
|
||||
}
|
||||
|
||||
private Map<String, Object> resultMetadata(ToolInvocationAuditEvent event) {
|
||||
private void appendSafeRequestFields(Map<String, Object> metadata, String requestJson) {
|
||||
if (requestJson == null || requestJson.isBlank()) {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(requestJson);
|
||||
if (root == null || !root.isObject()) {
|
||||
return;
|
||||
}
|
||||
int added = 0;
|
||||
Iterator<Map.Entry<String, JsonNode>> fields = root.fields();
|
||||
while (fields.hasNext() && added < MAX_REQUEST_KEYS) {
|
||||
Map.Entry<String, JsonNode> entry = fields.next();
|
||||
String key = entry.getKey();
|
||||
if (key == null || key.isBlank() || isSensitiveKey(key)) {
|
||||
continue;
|
||||
}
|
||||
JsonNode value = entry.getValue();
|
||||
if (value == null || value.isNull()) {
|
||||
continue;
|
||||
}
|
||||
if (value.isTextual()) {
|
||||
String text = value.asText();
|
||||
if (text == null || text.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
metadata.put(key, truncate(text, MAX_TEXT_PREVIEW));
|
||||
if (text.length() > MAX_TEXT_PREVIEW) {
|
||||
metadata.put(key + "_truncated", true);
|
||||
metadata.put(key + "_chars", text.length());
|
||||
}
|
||||
added++;
|
||||
} else if (value.isNumber()) {
|
||||
metadata.put(key, value.numberValue());
|
||||
added++;
|
||||
} else if (value.isBoolean()) {
|
||||
metadata.put(key, value.booleanValue());
|
||||
added++;
|
||||
} else if (value.isArray() && isStringArray(value)) {
|
||||
List<String> items = new ArrayList<>();
|
||||
for (int i = 0; i < value.size() && items.size() < 8; i++) {
|
||||
JsonNode item = value.get(i);
|
||||
if (item != null && item.isTextual() && !item.asText().isBlank()) {
|
||||
items.add(truncate(item.asText(), 64));
|
||||
}
|
||||
}
|
||||
if (!items.isEmpty()) {
|
||||
metadata.put(key, items);
|
||||
added++;
|
||||
}
|
||||
}
|
||||
// objects / mixed arrays intentionally omitted
|
||||
}
|
||||
} catch (Exception ignored) {
|
||||
metadata.put("request_parse", "failed");
|
||||
}
|
||||
}
|
||||
|
||||
private static boolean isStringArray(JsonNode value) {
|
||||
if (value == null || !value.isArray() || value.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
for (JsonNode n : value) {
|
||||
if (n == null || !n.isTextual()) {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
return true;
|
||||
}
|
||||
|
||||
private static boolean isSensitiveKey(String key) {
|
||||
String normalized = key.toLowerCase(Locale.ROOT);
|
||||
for (String fragment : SENSITIVE_KEY_FRAGMENTS) {
|
||||
if (normalized.contains(fragment)) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
private static String truncate(String value, int maxChars) {
|
||||
if (value == null) {
|
||||
return null;
|
||||
}
|
||||
if (value.length() <= maxChars) {
|
||||
return value;
|
||||
}
|
||||
return value.substring(0, maxChars);
|
||||
}
|
||||
|
||||
private Map<String, Object> baseResultMetadata(ToolInvocationAuditEvent event) {
|
||||
Map<String, Object> metadata = new LinkedHashMap<>();
|
||||
metadata.put("tool_call_id", event.toolCallId());
|
||||
metadata.put("status", event.status().name());
|
||||
|
||||
@@ -0,0 +1,335 @@
|
||||
package com.superbiz.agent.harness.audit;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Component;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
|
||||
/**
|
||||
* Builds durable RAG audit fields from internal {@code LookupResult} JSON
|
||||
* (and optional projected agent result) without dumping full excerpts/traces.
|
||||
*/
|
||||
@Component
|
||||
public class RagLookupAuditEnricher {
|
||||
|
||||
public static final String TOOL_LOOKUP_KNOWLEDGE = "lookup_knowledge";
|
||||
|
||||
private static final int MAX_EVIDENCE_KEYS = 12;
|
||||
private static final int MAX_HINT_VALUES = 8;
|
||||
private static final int PREVIEW_CHARS = 160;
|
||||
|
||||
private final ObjectMapper objectMapper;
|
||||
private final String searchMode;
|
||||
|
||||
public RagLookupAuditEnricher(
|
||||
ObjectMapper objectMapper,
|
||||
@Value("${retrieval.search.mode:hybrid}") String searchMode) {
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper");
|
||||
this.searchMode = searchMode == null || searchMode.isBlank()
|
||||
? "hybrid"
|
||||
: searchMode.trim().toLowerCase(Locale.ROOT);
|
||||
}
|
||||
|
||||
public boolean supports(String toolName) {
|
||||
return TOOL_LOOKUP_KNOWLEDGE.equals(toolName);
|
||||
}
|
||||
|
||||
/**
|
||||
* @param rawLookupJson internal LookupResult JSON (may be null on hard failures)
|
||||
* @param agentResultJson projected RagToolResult JSON (may be null)
|
||||
*/
|
||||
public Enrichment enrich(String rawLookupJson, String agentResultJson) {
|
||||
Map<String, Object> details = new LinkedHashMap<>();
|
||||
details.put("audit_schema", "rag_lookup_v1");
|
||||
details.put("search_mode", searchMode);
|
||||
|
||||
String retrievalLayer = "L1";
|
||||
Integer l0MatchCount = null;
|
||||
Integer l1MatchCount = null;
|
||||
String relevanceLevel = null;
|
||||
Boolean truncated = null;
|
||||
String dedupReason = null;
|
||||
String outputPreview = null;
|
||||
|
||||
try {
|
||||
if (rawLookupJson != null && !rawLookupJson.isBlank()) {
|
||||
JsonNode root = objectMapper.readTree(rawLookupJson);
|
||||
if (root != null && root.isObject()) {
|
||||
relevanceLevel = text(root, "relevanceLevel");
|
||||
if (relevanceLevel == null) {
|
||||
relevanceLevel = text(root, "relevance_level");
|
||||
}
|
||||
|
||||
JsonNode blocks = root.get("evidenceBlocks");
|
||||
if (blocks != null && blocks.isArray()) {
|
||||
l1MatchCount = blocks.size();
|
||||
List<String> keys = new ArrayList<>();
|
||||
List<String> sources = new ArrayList<>();
|
||||
String layer = null;
|
||||
for (int i = 0; i < blocks.size() && keys.size() < MAX_EVIDENCE_KEYS; i++) {
|
||||
JsonNode b = blocks.get(i);
|
||||
if (b == null || !b.isObject()) {
|
||||
continue;
|
||||
}
|
||||
String key = firstText(b, "evidenceKey", "evidence_key");
|
||||
if (key != null) {
|
||||
keys.add(key);
|
||||
}
|
||||
String source = text(b, "source");
|
||||
if (source != null && sources.size() < MAX_EVIDENCE_KEYS) {
|
||||
sources.add(source);
|
||||
}
|
||||
if (layer == null) {
|
||||
layer = text(b, "retrievalLayer");
|
||||
}
|
||||
}
|
||||
if (!keys.isEmpty()) {
|
||||
details.put("evidence_keys", keys);
|
||||
}
|
||||
if (!sources.isEmpty()) {
|
||||
details.put("sources", sources);
|
||||
}
|
||||
if (layer != null && !layer.isBlank()) {
|
||||
retrievalLayer = layer;
|
||||
}
|
||||
}
|
||||
|
||||
Integer candidateCount = intVal(root, "evidenceCandidateCount");
|
||||
if (candidateCount != null) {
|
||||
details.put("evidence_candidate_count", candidateCount);
|
||||
if (l1MatchCount == null) {
|
||||
l1MatchCount = candidateCount;
|
||||
}
|
||||
}
|
||||
Integer blockCount = intVal(root, "evidenceBlockCount");
|
||||
if (blockCount != null) {
|
||||
details.put("evidence_block_count", blockCount);
|
||||
}
|
||||
|
||||
String completenessHint = text(root, "completenessHint");
|
||||
if (completenessHint != null) {
|
||||
details.put("completeness_hint", truncate(completenessHint, PREVIEW_CHARS));
|
||||
}
|
||||
|
||||
JsonNode trace = root.get("retrievalTrace");
|
||||
if (trace != null && trace.isObject()) {
|
||||
putText(details, "selected_attempt", text(trace, "selectedAttempt"));
|
||||
putText(details, "fallback_reason", text(trace, "fallbackReason"));
|
||||
putText(details, "retrieval_evidence_status", text(trace, "evidenceStatus"));
|
||||
putText(details, "category_filter", text(trace, "categoryFilter"));
|
||||
// query texts intentionally omitted from durable audit by default (PII/size)
|
||||
|
||||
JsonNode hints = trace.get("queryHints");
|
||||
if (hints != null && hints.isObject()) {
|
||||
Integer l0 = intVal(hints, "l0_match_count");
|
||||
if (l0 == null) {
|
||||
l0 = arraySize(hints.get("domains"));
|
||||
}
|
||||
l0MatchCount = l0;
|
||||
Map<String, Object> hintSnap = new LinkedHashMap<>();
|
||||
putLimitedList(hintSnap, "domains", hints.get("domains"));
|
||||
putLimitedList(hintSnap, "matched_keywords", hints.get("matched_keywords"));
|
||||
if (hintSnap.isEmpty()) {
|
||||
putLimitedList(hintSnap, "matched_keywords", hints.get("matchedKeywords"));
|
||||
}
|
||||
if (!hintSnap.isEmpty()) {
|
||||
details.put("l0_hints", hintSnap);
|
||||
}
|
||||
}
|
||||
|
||||
JsonNode attempts = trace.get("attempts");
|
||||
if (attempts != null && attempts.isArray()) {
|
||||
List<Map<String, Object>> attemptSnap = new ArrayList<>();
|
||||
for (JsonNode a : attempts) {
|
||||
if (a == null || !a.isObject()) {
|
||||
continue;
|
||||
}
|
||||
Map<String, Object> row = new LinkedHashMap<>();
|
||||
putText(row, "name", text(a, "name"));
|
||||
putText(row, "category_filter", text(a, "categoryFilter"));
|
||||
if (a.has("candidateCount") && a.get("candidateCount").canConvertToInt()) {
|
||||
row.put("candidate_count", a.get("candidateCount").asInt());
|
||||
}
|
||||
if (a.has("usable") && a.get("usable").isBoolean()) {
|
||||
row.put("usable", a.get("usable").asBoolean());
|
||||
}
|
||||
if (a.has("topSimilarity") && a.get("topSimilarity").isNumber()) {
|
||||
row.put("top_similarity", a.get("topSimilarity").asDouble());
|
||||
}
|
||||
if (a.has("durationMs") && a.get("durationMs").canConvertToInt()) {
|
||||
row.put("duration_ms", a.get("durationMs").asInt());
|
||||
}
|
||||
if (!row.isEmpty()) {
|
||||
attemptSnap.add(row);
|
||||
}
|
||||
}
|
||||
if (!attemptSnap.isEmpty()) {
|
||||
details.put("attempts", attemptSnap);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
JsonNode pack = root.get("contextPack");
|
||||
if (pack != null && pack.isObject()) {
|
||||
if (pack.has("usedChars") && pack.get("usedChars").canConvertToInt()) {
|
||||
details.put("context_used_chars", pack.get("usedChars").asInt());
|
||||
}
|
||||
if (pack.has("charBudget") && pack.get("charBudget").canConvertToInt()) {
|
||||
details.put("context_char_budget", pack.get("charBudget").asInt());
|
||||
}
|
||||
}
|
||||
|
||||
// compact preview for list UIs
|
||||
outputPreview = buildPreview(relevanceLevel, details);
|
||||
}
|
||||
}
|
||||
} catch (Exception ignored) {
|
||||
details.put("enrich_error", "lookup_result_parse_failed");
|
||||
}
|
||||
|
||||
try {
|
||||
if (agentResultJson != null && !agentResultJson.isBlank()) {
|
||||
JsonNode agent = objectMapper.readTree(agentResultJson);
|
||||
if (agent != null && agent.isObject()) {
|
||||
if (agent.has("truncated") && agent.get("truncated").isBoolean()) {
|
||||
truncated = agent.get("truncated").asBoolean();
|
||||
details.put("truncated", truncated);
|
||||
}
|
||||
if (relevanceLevel == null) {
|
||||
relevanceLevel = text(agent, "relevance_level");
|
||||
if (relevanceLevel == null) {
|
||||
relevanceLevel = text(agent, "relevanceLevel");
|
||||
}
|
||||
}
|
||||
if (agent.has("returned_count") && agent.get("returned_count").canConvertToInt()) {
|
||||
details.put("returned_count", agent.get("returned_count").asInt());
|
||||
}
|
||||
}
|
||||
}
|
||||
} catch (Exception ignored) {
|
||||
details.put("agent_result_parse", "failed");
|
||||
}
|
||||
|
||||
if (outputPreview == null) {
|
||||
outputPreview = buildPreview(relevanceLevel, details);
|
||||
}
|
||||
|
||||
return new Enrichment(
|
||||
retrievalLayer,
|
||||
l0MatchCount,
|
||||
l1MatchCount,
|
||||
relevanceLevel,
|
||||
truncated,
|
||||
dedupReason,
|
||||
details,
|
||||
outputPreview
|
||||
);
|
||||
}
|
||||
|
||||
private static String buildPreview(String relevanceLevel, Map<String, Object> details) {
|
||||
String attempt = details.get("selected_attempt") == null ? null : String.valueOf(details.get("selected_attempt"));
|
||||
String fallback = details.get("fallback_reason") == null ? null : String.valueOf(details.get("fallback_reason"));
|
||||
StringBuilder sb = new StringBuilder("lookup_knowledge");
|
||||
if (relevanceLevel != null) {
|
||||
sb.append(" level=").append(relevanceLevel);
|
||||
}
|
||||
if (attempt != null) {
|
||||
sb.append(" attempt=").append(attempt);
|
||||
}
|
||||
if (fallback != null) {
|
||||
sb.append(" fallback=").append(fallback);
|
||||
}
|
||||
Object keys = details.get("evidence_keys");
|
||||
if (keys instanceof List<?> list) {
|
||||
sb.append(" keys=").append(list.size());
|
||||
}
|
||||
return truncate(sb.toString(), PREVIEW_CHARS);
|
||||
}
|
||||
|
||||
private static void putLimitedList(Map<String, Object> target, String key, JsonNode node) {
|
||||
if (node == null || !node.isArray() || node.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
List<String> values = new ArrayList<>();
|
||||
for (int i = 0; i < node.size() && values.size() < MAX_HINT_VALUES; i++) {
|
||||
JsonNode n = node.get(i);
|
||||
if (n != null && n.isTextual() && !n.asText().isBlank()) {
|
||||
values.add(n.asText());
|
||||
}
|
||||
}
|
||||
if (!values.isEmpty()) {
|
||||
target.put(key, values);
|
||||
}
|
||||
}
|
||||
|
||||
private static void putText(Map<String, Object> map, String key, String value) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
map.put(key, value);
|
||||
}
|
||||
}
|
||||
|
||||
private static String firstText(JsonNode node, String... fields) {
|
||||
for (String f : fields) {
|
||||
String v = text(node, f);
|
||||
if (v != null) {
|
||||
return v;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static String text(JsonNode node, String field) {
|
||||
if (node == null || field == null || !node.has(field) || node.get(field).isNull()) {
|
||||
return null;
|
||||
}
|
||||
String v = node.get(field).asText(null);
|
||||
return v == null || v.isBlank() ? null : v;
|
||||
}
|
||||
|
||||
private static Integer intVal(JsonNode node, String field) {
|
||||
if (node == null || !node.has(field) || node.get(field).isNull()) {
|
||||
return null;
|
||||
}
|
||||
JsonNode n = node.get(field);
|
||||
if (n.isIntegralNumber() || n.canConvertToInt()) {
|
||||
return n.asInt();
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static Integer arraySize(JsonNode node) {
|
||||
if (node != null && node.isArray()) {
|
||||
return node.size();
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static String truncate(String value, int max) {
|
||||
if (value == null) {
|
||||
return null;
|
||||
}
|
||||
if (value.length() <= max) {
|
||||
return value;
|
||||
}
|
||||
return value.substring(0, max) + "...";
|
||||
}
|
||||
|
||||
public record Enrichment(
|
||||
String retrievalLayer,
|
||||
Integer l0MatchCount,
|
||||
Integer l1MatchCount,
|
||||
String relevanceLevel,
|
||||
Boolean truncated,
|
||||
String dedupReason,
|
||||
Map<String, Object> retrievalDetails,
|
||||
String outputPreview
|
||||
) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,114 @@
|
||||
package com.superbiz.agent.harness.audit;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
|
||||
/**
|
||||
* Extracts a short human-readable conclusion string from released safe content JSON
|
||||
* for {@code diagnosis_run.conclusion}.
|
||||
*/
|
||||
public final class RunConclusionExtractor {
|
||||
|
||||
private static final int MAX_CHARS = 4_000;
|
||||
|
||||
private RunConclusionExtractor() {
|
||||
}
|
||||
|
||||
public static String extract(ObjectMapper objectMapper, String safeContentJson) {
|
||||
if (safeContentJson == null || safeContentJson.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(safeContentJson);
|
||||
if (root == null || !root.isObject()) {
|
||||
return bound(safeContentJson.trim());
|
||||
}
|
||||
// DIAGNOSIS_REPORT payload shape stored in answer
|
||||
String fromReport = text(root.path("report").path("conclusion").path("text"));
|
||||
if (fromReport != null) {
|
||||
return bound(fromReport);
|
||||
}
|
||||
// nested content envelope (defensive)
|
||||
String nested = text(root.path("payload").path("report").path("conclusion").path("text"));
|
||||
if (nested != null) {
|
||||
return bound(nested);
|
||||
}
|
||||
// SAFE_FALLBACK envelope: {"fallback":{type,message,...}}
|
||||
String fromFallback = fallbackConclusion(root.path("fallback"));
|
||||
if (fromFallback != null) {
|
||||
return bound(fromFallback);
|
||||
}
|
||||
String nestedFb = fallbackConclusion(root.path("payload").path("fallback"));
|
||||
if (nestedFb != null) {
|
||||
return bound(nestedFb);
|
||||
}
|
||||
// bare SafeFallback object (defensive)
|
||||
if (root.hasNonNull("type") || root.has("message") || root.has("conclusion")) {
|
||||
String bare = fallbackConclusion(root);
|
||||
if (bare != null) {
|
||||
return bound(bare);
|
||||
}
|
||||
}
|
||||
// knowledge / plain answer shapes
|
||||
String plain = text(root.path("answer"));
|
||||
if (plain != null) {
|
||||
return bound(plain);
|
||||
}
|
||||
return null;
|
||||
} catch (Exception ignored) {
|
||||
return bound(safeContentJson.trim());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a short conclusion from a SafeFallback-shaped node.
|
||||
* Prefer explicit conclusion text, else "type: message", else either alone.
|
||||
*/
|
||||
private static String fallbackConclusion(JsonNode node) {
|
||||
if (node == null || node.isMissingNode() || node.isNull() || !node.isObject()) {
|
||||
return null;
|
||||
}
|
||||
String conclusion = text(node.path("conclusion"));
|
||||
if (conclusion != null) {
|
||||
return conclusion;
|
||||
}
|
||||
String type = enumOrText(node.path("type"));
|
||||
String message = text(node.path("message"));
|
||||
if (type != null && message != null) {
|
||||
return type + ": " + message;
|
||||
}
|
||||
if (message != null) {
|
||||
return message;
|
||||
}
|
||||
return type;
|
||||
}
|
||||
|
||||
private static String enumOrText(JsonNode node) {
|
||||
if (node == null || node.isMissingNode() || node.isNull()) {
|
||||
return null;
|
||||
}
|
||||
if (node.isTextual() || node.isNumber() || node.isBoolean()) {
|
||||
String value = node.asText();
|
||||
return value == null || value.isBlank() ? null : value.trim();
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static String text(JsonNode node) {
|
||||
if (node == null || node.isMissingNode() || node.isNull()) {
|
||||
return null;
|
||||
}
|
||||
if (!node.isTextual()) {
|
||||
return null;
|
||||
}
|
||||
String value = node.asText();
|
||||
return value == null || value.isBlank() ? null : value.trim();
|
||||
}
|
||||
|
||||
private static String bound(String value) {
|
||||
if (value == null) {
|
||||
return null;
|
||||
}
|
||||
return value.length() <= MAX_CHARS ? value : value.substring(0, MAX_CHARS);
|
||||
}
|
||||
}
|
||||
@@ -5,6 +5,17 @@ import com.superbiz.agent.harness.contract.InvocationStatus;
|
||||
|
||||
import java.util.Objects;
|
||||
|
||||
/**
|
||||
* Durable tool-invocation audit event.
|
||||
*
|
||||
* <p>{@code rawResultJson} is optional internal executor output (e.g. full {@code LookupResult}
|
||||
* before Agent projection). It is used only to enrich durable RAG fields and must not be
|
||||
* echoed wholesale into agent-facing views.</p>
|
||||
*
|
||||
* <p>{@code stepId} links to {@code agent_step.id} when the in-flight Agent step is known.
|
||||
* {@code requestJson} is the tool request envelope body used only to derive bounded
|
||||
* {@code input_params} (e.g. query preview) — never dump secrets wholesale.</p>
|
||||
*/
|
||||
public record ToolInvocationAuditEvent(
|
||||
String sessionId,
|
||||
String runId,
|
||||
@@ -15,7 +26,11 @@ public record ToolInvocationAuditEvent(
|
||||
String errorCode,
|
||||
int durationMs,
|
||||
int requestBytes,
|
||||
int agentResultBytes) {
|
||||
int agentResultBytes,
|
||||
String rawResultJson,
|
||||
String agentResultJson,
|
||||
Long stepId,
|
||||
String requestJson) {
|
||||
|
||||
public ToolInvocationAuditEvent {
|
||||
requireText(sessionId, "sessionId");
|
||||
@@ -35,6 +50,38 @@ public record ToolInvocationAuditEvent(
|
||||
}
|
||||
}
|
||||
|
||||
/** Backward-compatible constructor without raw/agent JSON / step / request body. */
|
||||
public ToolInvocationAuditEvent(String sessionId,
|
||||
String runId,
|
||||
String toolCallId,
|
||||
String toolName,
|
||||
InvocationStatus status,
|
||||
EvidenceStatus evidenceStatus,
|
||||
String errorCode,
|
||||
int durationMs,
|
||||
int requestBytes,
|
||||
int agentResultBytes) {
|
||||
this(sessionId, runId, toolCallId, toolName, status, evidenceStatus, errorCode,
|
||||
durationMs, requestBytes, agentResultBytes, null, null, null, null);
|
||||
}
|
||||
|
||||
/** Backward-compatible constructor with raw/agent JSON only. */
|
||||
public ToolInvocationAuditEvent(String sessionId,
|
||||
String runId,
|
||||
String toolCallId,
|
||||
String toolName,
|
||||
InvocationStatus status,
|
||||
EvidenceStatus evidenceStatus,
|
||||
String errorCode,
|
||||
int durationMs,
|
||||
int requestBytes,
|
||||
int agentResultBytes,
|
||||
String rawResultJson,
|
||||
String agentResultJson) {
|
||||
this(sessionId, runId, toolCallId, toolName, status, evidenceStatus, errorCode,
|
||||
durationMs, requestBytes, agentResultBytes, rawResultJson, agentResultJson, null, null);
|
||||
}
|
||||
|
||||
private static void requireText(String value, String name) {
|
||||
if (value == null || value.isBlank()) {
|
||||
throw new IllegalArgumentException(name + " must not be blank");
|
||||
|
||||
@@ -129,9 +129,14 @@ public final class TraceAuditEvents {
|
||||
details.put("evidence_status", tool.evidenceStatus().name());
|
||||
details.put("request_bytes", tool.requestBytes());
|
||||
details.put("agent_result_bytes", tool.agentResultBytes());
|
||||
details.put("has_raw_result", tool.rawResultJson() != null && !tool.rawResultJson().isBlank());
|
||||
if (tool.stepId() != null) {
|
||||
details.put("step_id", tool.stepId());
|
||||
}
|
||||
if (tool.errorCode() != null) {
|
||||
details.put("error_code", tool.errorCode());
|
||||
}
|
||||
// Do not embed raw LookupResult / agent JSON here — durable RAG fields go to tool_invocation.
|
||||
return new DiagnosisTraceAuditEvent(
|
||||
tool.sessionId(), tool.runId(), TracePhase.TOOL, TraceEventType.TOOL_INVOCATION,
|
||||
tool.status() == com.superbiz.agent.harness.contract.InvocationStatus.READY
|
||||
|
||||
@@ -3,9 +3,9 @@ package com.superbiz.agent.harness.tool.boundary;
|
||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.harness.audit.AgentStepAuditTracker;
|
||||
import com.superbiz.agent.harness.audit.ToolInvocationAuditEvent;
|
||||
import com.superbiz.agent.harness.audit.ToolInvocationAuditSink;
|
||||
import com.superbiz.agent.harness.contract.EvidenceStatus;
|
||||
import com.superbiz.agent.harness.core.BudgetExceededException;
|
||||
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||
import com.superbiz.agent.harness.core.RunAbortedException;
|
||||
@@ -19,10 +19,10 @@ import com.superbiz.agent.harness.tool.store.ToolCallKeyFactory;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.time.Clock;
|
||||
import java.time.Duration;
|
||||
import java.time.Instant;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.Objects;
|
||||
|
||||
public final class ToolBoundary {
|
||||
@@ -35,13 +35,14 @@ public final class ToolBoundary {
|
||||
private final ObjectMapper objectMapper;
|
||||
private final Clock clock;
|
||||
private final ToolInvocationAuditSink auditSink;
|
||||
private final AgentStepAuditTracker stepTracker;
|
||||
|
||||
public ToolBoundary(DiagnosisHarnessCore core,
|
||||
ToolCallKeyFactory keyFactory,
|
||||
CanonicalInvocationStore store,
|
||||
ObjectMapper objectMapper,
|
||||
Clock clock) {
|
||||
this(core, keyFactory, store, objectMapper, clock, ToolInvocationAuditSink.noop());
|
||||
this(core, keyFactory, store, objectMapper, clock, ToolInvocationAuditSink.noop(), null);
|
||||
}
|
||||
|
||||
public ToolBoundary(DiagnosisHarnessCore core,
|
||||
@@ -50,12 +51,23 @@ public final class ToolBoundary {
|
||||
ObjectMapper objectMapper,
|
||||
Clock clock,
|
||||
ToolInvocationAuditSink auditSink) {
|
||||
this(core, keyFactory, store, objectMapper, clock, auditSink, null);
|
||||
}
|
||||
|
||||
public ToolBoundary(DiagnosisHarnessCore core,
|
||||
ToolCallKeyFactory keyFactory,
|
||||
CanonicalInvocationStore store,
|
||||
ObjectMapper objectMapper,
|
||||
Clock clock,
|
||||
ToolInvocationAuditSink auditSink,
|
||||
AgentStepAuditTracker stepTracker) {
|
||||
this.core = Objects.requireNonNull(core, "core must not be null");
|
||||
this.keyFactory = Objects.requireNonNull(keyFactory, "keyFactory must not be null");
|
||||
this.store = Objects.requireNonNull(store, "store must not be null");
|
||||
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||
this.clock = Objects.requireNonNull(clock, "clock must not be null");
|
||||
this.auditSink = Objects.requireNonNull(auditSink, "auditSink must not be null");
|
||||
this.stepTracker = stepTracker;
|
||||
}
|
||||
|
||||
public ToolBoundaryResult execute(RunContext context,
|
||||
@@ -63,12 +75,12 @@ public final class ToolBoundary {
|
||||
ToolExecutor executor,
|
||||
ToolResultProjector projector) {
|
||||
Instant startedAt = clock.instant();
|
||||
ToolBoundaryResult result = executeCanonical(context, request, executor, projector);
|
||||
auditSafely(context, request, result, startedAt);
|
||||
return result;
|
||||
ExecutionOutcome outcome = executeCanonical(context, request, executor, projector);
|
||||
auditSafely(context, request, outcome.result(), outcome.rawResponse(), startedAt);
|
||||
return outcome.result();
|
||||
}
|
||||
|
||||
private ToolBoundaryResult executeCanonical(RunContext context,
|
||||
private ExecutionOutcome executeCanonical(RunContext context,
|
||||
ToolCallRequestEnvelope request,
|
||||
ToolExecutor executor,
|
||||
ToolResultProjector projector) {
|
||||
@@ -79,23 +91,23 @@ public final class ToolBoundary {
|
||||
core.beforeToolCall(context, request.toolName());
|
||||
long requestBytes = store.limits().utf8Bytes(request.requestJson());
|
||||
if (requestBytes > store.limits().maxRecordBytes()) {
|
||||
return errorAndNoRecord(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
|
||||
return ExecutionOutcome.of(errorAndNoRecord(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE));
|
||||
}
|
||||
core.reserveRunBytes(context, requestBytes);
|
||||
store.begin(key, CanonicalToolInvocation.projecting(
|
||||
request.toolCallId(), request.runId(), request.toolName(),
|
||||
request.requestJson(), clock.instant()));
|
||||
} catch (DuplicateInvocationException e) {
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.DUPLICATE_TOOL_CALL);
|
||||
return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.DUPLICATE_TOOL_CALL));
|
||||
} catch (RunAbortedException | BudgetExceededException e) {
|
||||
return ToolBoundaryResult.error(toolCallId,
|
||||
return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId,
|
||||
e instanceof BudgetExceededException
|
||||
? ToolBoundaryErrorCode.BUDGET_EXHAUSTED
|
||||
: ToolBoundaryErrorCode.RUN_INACTIVE);
|
||||
: ToolBoundaryErrorCode.RUN_INACTIVE));
|
||||
} catch (IllegalArgumentException e) {
|
||||
return ToolBoundaryResult.error(toolCallId, classifyPreflightError(e));
|
||||
return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, classifyPreflightError(e)));
|
||||
} catch (CanonicalStoreException e) {
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.STORE_ERROR);
|
||||
return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.STORE_ERROR));
|
||||
}
|
||||
|
||||
String rawResponse;
|
||||
@@ -107,7 +119,7 @@ public final class ToolBoundary {
|
||||
}
|
||||
} catch (Exception e) {
|
||||
markErrorSafely(key, null, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR);
|
||||
return ExecutionOutcome.of(ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR));
|
||||
}
|
||||
|
||||
try {
|
||||
@@ -115,13 +127,16 @@ public final class ToolBoundary {
|
||||
core.reserveRunBytes(context, store.limits().utf8Bytes(rawResponse));
|
||||
} catch (ResultTooLargeException e) {
|
||||
markErrorSafely(key, null, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE), rawResponse);
|
||||
} catch (BudgetExceededException e) {
|
||||
markErrorSafely(key, null, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED), rawResponse);
|
||||
} catch (RunAbortedException e) {
|
||||
markErrorSafely(key, null, ToolBoundaryErrorCode.RUN_INACTIVE);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE), rawResponse);
|
||||
}
|
||||
|
||||
ProjectedToolResult projected;
|
||||
@@ -133,29 +148,35 @@ public final class ToolBoundary {
|
||||
}
|
||||
} catch (Exception e) {
|
||||
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.PROJECTION_ERROR);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.PROJECTION_ERROR);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.PROJECTION_ERROR), rawResponse);
|
||||
}
|
||||
|
||||
try {
|
||||
store.limits().validateAgentResult(projected.agentResult());
|
||||
core.reserveRunBytes(context, store.limits().utf8Bytes(projected.agentResult()));
|
||||
store.markReady(key, rawResponse, projected.agentResult(), projected.evidenceStatus(), clock.instant());
|
||||
return ToolBoundaryResult.ready(toolCallId, projected.agentResult(), projected.evidenceStatus());
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.ready(toolCallId, projected.agentResult(), projected.evidenceStatus()),
|
||||
rawResponse);
|
||||
} catch (ResultTooLargeException e) {
|
||||
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RESULT_TOO_LARGE), rawResponse);
|
||||
} catch (BudgetExceededException e) {
|
||||
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.BUDGET_EXHAUSTED), rawResponse);
|
||||
} catch (RunAbortedException e) {
|
||||
markErrorSafely(key, rawResponse, ToolBoundaryErrorCode.RUN_INACTIVE);
|
||||
return ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE);
|
||||
return new ExecutionOutcome(
|
||||
ToolBoundaryResult.error(toolCallId, ToolBoundaryErrorCode.RUN_INACTIVE), rawResponse);
|
||||
} catch (CanonicalStoreException e) {
|
||||
ToolBoundaryErrorCode code = e instanceof ResultTooLargeException
|
||||
? ToolBoundaryErrorCode.RESULT_TOO_LARGE
|
||||
: ToolBoundaryErrorCode.PROJECTION_ERROR;
|
||||
markErrorSafely(key, rawResponse, code);
|
||||
return ToolBoundaryResult.error(toolCallId, code);
|
||||
return new ExecutionOutcome(ToolBoundaryResult.error(toolCallId, code), rawResponse);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -222,6 +243,7 @@ public final class ToolBoundary {
|
||||
private void auditSafely(RunContext context,
|
||||
ToolCallRequestEnvelope request,
|
||||
ToolBoundaryResult result,
|
||||
String rawResponse,
|
||||
Instant startedAt) {
|
||||
if (context == null || request == null || result == null
|
||||
|| !context.runId().equals(request.runId())
|
||||
@@ -230,11 +252,16 @@ public final class ToolBoundary {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
Long stepId = stepTracker == null ? null : stepTracker.currentStepId(context.runId());
|
||||
auditSink.record(new ToolInvocationAuditEvent(
|
||||
context.sessionId(), context.runId(), request.toolCallId(), request.toolName(),
|
||||
result.status(), result.evidenceStatus(), result.errorCode(),
|
||||
saturatingInt(Math.max(0L, Duration.between(startedAt, clock.instant()).toMillis())),
|
||||
utf8Bytes(request.requestJson()), utf8Bytes(result.agentResult())));
|
||||
utf8Bytes(request.requestJson()), utf8Bytes(result.agentResult()),
|
||||
rawResponse,
|
||||
result.agentResult(),
|
||||
stepId,
|
||||
request.requestJson()));
|
||||
} catch (RuntimeException exception) {
|
||||
log.warn("Failed to persist Tool durable audit: tool={}, status={}",
|
||||
request.toolName(), result.status());
|
||||
@@ -249,6 +276,12 @@ public final class ToolBoundary {
|
||||
return value >= Integer.MAX_VALUE ? Integer.MAX_VALUE : (int) value;
|
||||
}
|
||||
|
||||
private record ExecutionOutcome(ToolBoundaryResult result, String rawResponse) {
|
||||
static ExecutionOutcome of(ToolBoundaryResult result) {
|
||||
return new ExecutionOutcome(result, null);
|
||||
}
|
||||
}
|
||||
|
||||
private static final class InvalidToolCallIdException extends IllegalArgumentException {
|
||||
private InvalidToolCallIdException(Throwable cause) {
|
||||
super("Invalid tool call ID", cause);
|
||||
|
||||
@@ -184,6 +184,7 @@ public class DiagnosisTraceService {
|
||||
.runId(run.getRunId())
|
||||
.sessionId(run.getSessionId())
|
||||
.query(run.getQuery())
|
||||
.conclusion(run.getConclusion())
|
||||
.status(run.getStatus())
|
||||
.agentFlow(run.getAgentFlow())
|
||||
.intent(run.getIntent() == null ? null : run.getIntent().name())
|
||||
@@ -239,6 +240,8 @@ public class DiagnosisTraceService {
|
||||
.agentName(audit.getAgentName())
|
||||
.reasoningAvailable(audit.getReasoningAvailable())
|
||||
.reasoningContent(audit.getReasoningContent())
|
||||
.assistantText(audit.getAssistantText())
|
||||
.contentSource(audit.getContentSource())
|
||||
.contentBytes(audit.getContentBytes())
|
||||
.createdAt(audit.getCreatedAt())
|
||||
.build();
|
||||
|
||||
@@ -86,6 +86,7 @@ public class KnowledgeDocumentRetriever {
|
||||
.rawScore(hit.rawScore())
|
||||
.scoreLabel(hit.scoreLabel())
|
||||
.originalRank(hit.originalRank())
|
||||
.denseDistance(hit.denseDistance())
|
||||
.metadata(hit.metadata() == null ? java.util.Map.of() : hit.metadata())
|
||||
.hitReasons(List.of("semantic_rank:" + hit.originalRank(), "attempt:" + attemptName))
|
||||
.build());
|
||||
|
||||
@@ -6,6 +6,7 @@ import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import com.superbiz.agent.service.retrieval.EvidenceIdentity;
|
||||
import com.superbiz.agent.service.retrieval.RetrievalScoreNormalizer;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
@@ -20,28 +21,25 @@ import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* 检索后处理:分数归一化、规则 rerank、chunk 级证据组装、相关等级判定。
|
||||
* 检索后处理:统一 qualityScore、保检索序、chunk 级装配、相关等级判定。
|
||||
*
|
||||
* <h3>处理步骤</h3>
|
||||
* <ol>
|
||||
* <li>把候选 L2 距离归一成 0~1 的 baseScore</li>
|
||||
* <li>用 L0 hint 做规则加分,得到 finalScore</li>
|
||||
* <li>按 finalScore 降序排序</li>
|
||||
* <li>按 evidenceKey 去重</li>
|
||||
* <li>按 maxChunksPerDocument 截断同文档 chunk</li>
|
||||
* <li>按 return-n 截断最终 evidence 条数</li>
|
||||
* <li>根据 top baseScore + hint 支撑计算 relevanceLevel</li>
|
||||
* <li>{@link RetrievalScoreNormalizer#toQualityScore} → qualityScore(唯一 label 分支)</li>
|
||||
* <li>按 {@code originalRank} 升序(检索权威序;不做关键词 boost 改序)</li>
|
||||
* <li>evidenceKey 去重 / maxChunksPerDocument / return-n</li>
|
||||
* <li>用 top qualityScore 定 relevanceLevel / 低质闸门字段</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>L0 domain/keyword 命中仅写入解释性 {@code hitReasons},不改变分数与主序。</p>
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeEvidencePostProcessor {
|
||||
|
||||
private static final String LEVEL_PRECISE = "PRECISE";
|
||||
private static final String LEVEL_HIGHLY_RELEVANT = "HIGHLY_RELEVANT";
|
||||
private static final String LEVEL_REFERENCE = "REFERENCE";
|
||||
|
||||
private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档";
|
||||
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
|
||||
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
|
||||
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
@@ -53,22 +51,21 @@ public class KnowledgeEvidencePostProcessor {
|
||||
@Value("${retrieval.normalization.reference-threshold:0.5}")
|
||||
private double referenceThreshold = 0.5;
|
||||
|
||||
/** 同一 docId 最多保留的 chunk 数。 */
|
||||
@Value("${rag.max-chunks-per-document:2}")
|
||||
private int maxChunksPerDocument = 2;
|
||||
|
||||
/**
|
||||
* 后处理后最多返回的 evidence 条数。
|
||||
* 0 或负数表示不在此层截断(仍可能被 projector 预算截断)。
|
||||
*/
|
||||
@Value("${rag.return-n:5}")
|
||||
private int returnN = 5;
|
||||
|
||||
public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) {
|
||||
List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates;
|
||||
int batchSize = safeCandidates.size();
|
||||
|
||||
List<ScoredCandidate> ranked = safeCandidates.stream()
|
||||
.map(candidate -> score(query, candidate))
|
||||
.sorted(Comparator.comparingDouble(ScoredCandidate::finalScore).reversed())
|
||||
.map(candidate -> score(query, candidate, batchSize))
|
||||
.sorted(Comparator
|
||||
.comparingInt((ScoredCandidate s) -> rankOrMax(s.candidate().getOriginalRank()))
|
||||
.thenComparing(s -> resolveEvidenceKey(s.candidate()), Comparator.nullsLast(String::compareTo)))
|
||||
.toList();
|
||||
|
||||
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
|
||||
@@ -104,15 +101,15 @@ public class KnowledgeEvidencePostProcessor {
|
||||
traceItems.add(RerankTrace.Item.builder()
|
||||
.finalRank(finalRank++)
|
||||
.source(candidate.getSource())
|
||||
.baseScore(scored.baseScore())
|
||||
.finalScore(scored.finalScore())
|
||||
.boostReasons(scored.boostReasons())
|
||||
.baseScore(scored.qualityScore())
|
||||
.finalScore(scored.qualityScore())
|
||||
.boostReasons(scored.explainReasons())
|
||||
.build());
|
||||
}
|
||||
|
||||
List<EvidenceBlock> blocks = new ArrayList<>(deduped.values());
|
||||
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).baseScore();
|
||||
RelevanceAssessment assessment = computeRelevance(query, ranked);
|
||||
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).qualityScore();
|
||||
RelevanceAssessment assessment = computeRelevance(ranked);
|
||||
return EvidencePostprocessResult.builder()
|
||||
.candidateCount(safeCandidates.size())
|
||||
.evidenceBlockCount(blocks.size())
|
||||
@@ -132,12 +129,11 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return topSimilarity == null || topSimilarity < referenceThreshold;
|
||||
}
|
||||
|
||||
/**
|
||||
* @deprecated 保留给旧测试/调用;新路径请用 {@link RetrievalScoreNormalizer#l2ToQuality}。
|
||||
*/
|
||||
public double normalizeL2(Double l2Score) {
|
||||
if (l2Score == null) {
|
||||
return 0.0;
|
||||
}
|
||||
double clamped = Math.min(l2Score, maxL2Distance);
|
||||
return Math.max(0.0, 1.0 - clamped / maxL2Distance);
|
||||
return RetrievalScoreNormalizer.l2ToQuality(l2Score, maxL2Distance);
|
||||
}
|
||||
|
||||
public double getReferenceThreshold() {
|
||||
@@ -157,7 +153,7 @@ public class KnowledgeEvidencePostProcessor {
|
||||
.retrievalLayer(candidate.getRetrievalLayer())
|
||||
.content(truncate(candidate.getContent(), 800))
|
||||
.score(candidate.getScore())
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons()))
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.explainReasons()))
|
||||
.build();
|
||||
}
|
||||
|
||||
@@ -177,33 +173,29 @@ public class KnowledgeEvidencePostProcessor {
|
||||
if (docId != null) {
|
||||
return docId;
|
||||
}
|
||||
// No docId: do not collapse unrelated fallback keys under one bucket.
|
||||
return evidenceKey;
|
||||
}
|
||||
|
||||
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate) {
|
||||
double baseScore = normalizeL2(candidate.getScore());
|
||||
double finalScore = baseScore;
|
||||
List<String> boosts = new ArrayList<>();
|
||||
|
||||
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate, int batchSize) {
|
||||
double quality = RetrievalScoreNormalizer.toQualityScore(
|
||||
candidate.getScoreLabel(),
|
||||
candidate.getScore(),
|
||||
candidate.getOriginalRank(),
|
||||
batchSize,
|
||||
maxL2Distance,
|
||||
candidate.getDenseDistance());
|
||||
List<String> explain = new ArrayList<>();
|
||||
// L0 重叠仅解释,不改变 quality / 排序
|
||||
if (matchesAny(candidate, query.getDomainHints())) {
|
||||
finalScore += 0.15;
|
||||
boosts.add("domain_match:+0.15");
|
||||
explain.add("l0_domain_overlap");
|
||||
}
|
||||
if (matchesAny(candidate, query.getEntities())) {
|
||||
finalScore += 0.20;
|
||||
boosts.add("entity_match:+0.20");
|
||||
explain.add("l0_entity_overlap");
|
||||
}
|
||||
if (matchesAny(candidate, query.getMatchedKeywords())) {
|
||||
finalScore += 0.10;
|
||||
boosts.add("keyword_match:+0.10");
|
||||
explain.add("l0_keyword_overlap");
|
||||
}
|
||||
if (isPreferredSourceType(candidate)) {
|
||||
finalScore += 0.05;
|
||||
boosts.add("source_type:+0.05");
|
||||
}
|
||||
|
||||
return new ScoredCandidate(candidate, baseScore, finalScore, boosts);
|
||||
return new ScoredCandidate(candidate, quality, explain);
|
||||
}
|
||||
|
||||
private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) {
|
||||
@@ -225,42 +217,21 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return false;
|
||||
}
|
||||
|
||||
private boolean isPreferredSourceType(RetrievedEvidenceCandidate candidate) {
|
||||
Map<String, String> metadata = candidate.getMetadata();
|
||||
if (metadata == null || metadata.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
String type = firstNonBlank(metadata.get("source_type"), metadata.get("documentType"), metadata.get("type"));
|
||||
if (type == null) {
|
||||
return false;
|
||||
}
|
||||
String normalized = type.toLowerCase(Locale.ROOT);
|
||||
return normalized.contains("runbook") || normalized.contains("guide") || normalized.contains("case");
|
||||
}
|
||||
|
||||
private RelevanceAssessment computeRelevance(KnowledgeQuery query, List<ScoredCandidate> ranked) {
|
||||
private RelevanceAssessment computeRelevance(List<ScoredCandidate> ranked) {
|
||||
if (ranked.isEmpty()) {
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
ScoredCandidate top = ranked.get(0);
|
||||
if (top.baseScore() >= highlyRelevantThreshold && hasHintSupport(query, top)) {
|
||||
double top = ranked.get(0).qualityScore();
|
||||
if (top >= highlyRelevantThreshold) {
|
||||
// 不再要求 hasHintSupport;rank/L2 quality 足够即高相关/精准
|
||||
return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE);
|
||||
}
|
||||
if (top.baseScore() >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
|
||||
}
|
||||
if (top.baseScore() >= referenceThreshold) {
|
||||
if (top >= referenceThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
|
||||
private boolean hasHintSupport(KnowledgeQuery query, ScoredCandidate top) {
|
||||
return matchesAny(top.candidate(), query.getDomainHints())
|
||||
|| matchesAny(top.candidate(), query.getEntities())
|
||||
|| matchesAny(top.candidate(), query.getMatchedKeywords());
|
||||
}
|
||||
|
||||
private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) {
|
||||
Set<String> reasons = new LinkedHashSet<>();
|
||||
if (existing.getHitReasons() != null) {
|
||||
@@ -276,13 +247,13 @@ public class KnowledgeEvidencePostProcessor {
|
||||
}
|
||||
}
|
||||
|
||||
private List<String> mergeReasons(List<String> base, List<String> boosts) {
|
||||
private List<String> mergeReasons(List<String> base, List<String> extra) {
|
||||
Set<String> merged = new LinkedHashSet<>();
|
||||
if (base != null) {
|
||||
merged.addAll(base);
|
||||
}
|
||||
if (boosts != null) {
|
||||
merged.addAll(boosts);
|
||||
if (extra != null) {
|
||||
merged.addAll(extra);
|
||||
}
|
||||
return new ArrayList<>(merged);
|
||||
}
|
||||
@@ -294,8 +265,8 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
return EvidenceIdentity.firstNonBlank(values);
|
||||
private static int rankOrMax(Integer rank) {
|
||||
return rank == null || rank < 1 ? Integer.MAX_VALUE : rank;
|
||||
}
|
||||
|
||||
private String nullToEmpty(String value) {
|
||||
@@ -303,9 +274,8 @@ public class KnowledgeEvidencePostProcessor {
|
||||
}
|
||||
|
||||
private record ScoredCandidate(RetrievedEvidenceCandidate candidate,
|
||||
double baseScore,
|
||||
double finalScore,
|
||||
List<String> boostReasons) {
|
||||
double qualityScore,
|
||||
List<String> explainReasons) {
|
||||
}
|
||||
|
||||
private record RelevanceAssessment(String level, String hint) {
|
||||
|
||||
@@ -23,8 +23,14 @@ import java.util.Map;
|
||||
/**
|
||||
* 向量索引写入服务(RAG 入库侧)。
|
||||
*
|
||||
* <p>写入单一后端 {@link MilvusHybridKnowledgeStore}(dense + BM25 search_text)。
|
||||
* 不再使用 legacy {@code MilvusServiceClient} insert/delete。</p>
|
||||
* <p>唯一后端 {@link MilvusHybridKnowledgeStore}(Milvus SDK v2):</p>
|
||||
* <ul>
|
||||
* <li>dense:应用侧 embedding → 字段 {@code vector}</li>
|
||||
* <li>BM25:{@link #buildSearchText} → 字段 {@code search_text};
|
||||
* sparse 由 collection 上 BM25 Function 自动生成,本类不写 sparse</li>
|
||||
* </ul>
|
||||
* <p>不再使用 legacy {@code MilvusServiceClient} insert/delete,
|
||||
* 也不走 Spring AI {@code VectorStore#add}(starter 无 hybrid schema/BM25 Function)。</p>
|
||||
*/
|
||||
@Service
|
||||
public class VectorIndexService {
|
||||
@@ -118,12 +124,13 @@ public class VectorIndexService {
|
||||
for (int i = 0; i < chunks.size(); i++) {
|
||||
DocumentChunk chunk = chunks.get(i);
|
||||
try {
|
||||
// dense embedding 与 BM25 search_text 同源(title/path 增强)
|
||||
List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk));
|
||||
Map<String, Object> metadata = buildMetadata(path.toString(), chunk, chunks.size());
|
||||
knowledgeStore.upsertChunk(
|
||||
chunk.getContent(),
|
||||
buildSearchText(chunk),
|
||||
vector,
|
||||
chunk.getContent(), // 返回原文
|
||||
buildSearchText(chunk), // BM25 语料;sparse 由 Milvus Function 生成
|
||||
vector, // dense 向量
|
||||
metadata,
|
||||
chunk.getChunkIndex());
|
||||
logger.info("分片 {}/{} 索引成功", i + 1, chunks.size());
|
||||
@@ -153,12 +160,13 @@ public class VectorIndexService {
|
||||
for (int i = 0; i < chunks.size(); i++) {
|
||||
DocumentChunk chunk = chunks.get(i);
|
||||
try {
|
||||
// dense embedding 与 BM25 search_text 同源(title/path 增强)
|
||||
List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk));
|
||||
Map<String, Object> metadata = buildDocumentMetadata(docId, chunk, chunks.size(), category, frontmatter);
|
||||
knowledgeStore.upsertChunk(
|
||||
chunk.getContent(),
|
||||
buildSearchText(chunk),
|
||||
vector,
|
||||
chunk.getContent(), // 返回原文
|
||||
buildSearchText(chunk), // BM25 语料;sparse 由 Milvus Function 生成
|
||||
vector, // dense 向量
|
||||
metadata,
|
||||
chunk.getChunkIndex());
|
||||
logger.info("文档分块 {}/{} 索引成功,docId: {}", i + 1, chunks.size(), docId);
|
||||
@@ -212,12 +220,18 @@ public class VectorIndexService {
|
||||
return metadata;
|
||||
}
|
||||
|
||||
/**
|
||||
* Dense embedding 输入。与 {@link #buildSearchText} 同源,保证 dense/BM25 看到同一增强文本。
|
||||
*/
|
||||
static String buildEmbeddingText(DocumentChunk chunk) {
|
||||
return buildSearchText(chunk);
|
||||
}
|
||||
|
||||
/**
|
||||
* Text used for BM25 {@code search_text} and dense embedding.
|
||||
* 构造写入 Milvus 的检索文本(BM25 {@code search_text},并复用为 dense embedding 输入)。
|
||||
*
|
||||
* <p>在正文前拼接 title / breadcrumb,提高「按标题或路径关键词」的 BM25 命中率,
|
||||
* 同时让 dense 向量也编码结构信息。无标题路径时退回纯 content。</p>
|
||||
*/
|
||||
static String buildSearchText(DocumentChunk chunk) {
|
||||
String content = trimToEmpty(chunk.getContent());
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.service.milvus.MilvusHybridKnowledgeStore;
|
||||
import com.superbiz.agent.service.retrieval.RetrievalScoreLabels;
|
||||
import lombok.Getter;
|
||||
import lombok.Setter;
|
||||
import org.slf4j.Logger;
|
||||
@@ -13,11 +14,18 @@ import java.util.List;
|
||||
import java.util.Locale;
|
||||
|
||||
/**
|
||||
* Knowledge vector retrieval facade.
|
||||
* 知识库向量检索门面(lookup_knowledge / RAG 召回入口)。
|
||||
*
|
||||
* <p><b>Single backend:</b> {@link MilvusHybridKnowledgeStore} (Milvus Java SDK v2).
|
||||
* Legacy {@code MilvusServiceClient} search and Spring AI VectorStore routing for
|
||||
* {@code lookup_knowledge} have been removed.</p>
|
||||
* <p><b>唯一后端:</b>{@link MilvusHybridKnowledgeStore}(Milvus Java SDK v2)。</p>
|
||||
*
|
||||
* <h3>模式切换</h3>
|
||||
* <p>{@code retrieval.search.mode}(同库查询算法,非两套写入):</p>
|
||||
* <ul>
|
||||
* <li>{@code hybrid} —— 线上主路径:dense + 服务端 BM25 + RRF</li>
|
||||
* <li>{@code dense} —— 对照/评测:仅 dense ANN</li>
|
||||
* </ul>
|
||||
* <p>命中 {@link SearchResult#scoreLabel} 仅为 {@link RetrievalScoreLabels#DENSE} /
|
||||
* {@link RetrievalScoreLabels#HYBRID}。质量分由后处理 {@code RetrievalScoreNormalizer} 统一计算。</p>
|
||||
*/
|
||||
@Service
|
||||
public class VectorSearchService {
|
||||
@@ -31,7 +39,7 @@ public class VectorSearchService {
|
||||
private VectorEmbeddingService embeddingService;
|
||||
|
||||
/**
|
||||
* dense | hybrid
|
||||
* 检索模式:{@code hybrid}(主路径)| {@code dense}(召回对照)。
|
||||
*/
|
||||
@Value("${retrieval.search.mode:dense}")
|
||||
private String searchMode = "dense";
|
||||
@@ -53,18 +61,34 @@ public class VectorSearchService {
|
||||
return knowledgeStore.searchDense(query, queryVector, topK, category);
|
||||
}
|
||||
|
||||
/**
|
||||
* 单条召回结果。列表顺序即检索权威序(adapter 赋 originalRank=1..n)。
|
||||
*
|
||||
* <ul>
|
||||
* <li>{@code scoreLabel=dense}:{@link #score} = L2 距离(越小越好)</li>
|
||||
* <li>{@code scoreLabel=hybrid}:{@link #score}/{@link #rawScore} = 引擎融合分;
|
||||
* 后处理 quality 主要按 rank 映射,不把 score 当 L2</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Setter
|
||||
@Getter
|
||||
public static class SearchResult {
|
||||
private String id;
|
||||
private String content;
|
||||
/**
|
||||
* Compatibility score for post-process normalizeL2.
|
||||
* Dense path: L2 distance. Hybrid path: dense L2 when available.
|
||||
* 引擎主分:dense=L2;hybrid=融合分(量纲由 scoreLabel 解释)。
|
||||
*/
|
||||
private float score;
|
||||
/** 引擎原始分(与 score 同源或更细,便于调试)。 */
|
||||
private Double rawScore;
|
||||
/** {@link RetrievalScoreLabels#DENSE} 或 {@link RetrievalScoreLabels#HYBRID}。 */
|
||||
private String scoreLabel;
|
||||
/**
|
||||
* Optional dense L2 for the same id (hybrid path only).
|
||||
* Used for absolute quality / low-quality gates; does <b>not</b> replace sort order.
|
||||
*/
|
||||
private Double denseDistance;
|
||||
/** metadata JSON 字符串(docId、source、title…)。 */
|
||||
private String metadata;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -5,6 +5,7 @@ import com.google.gson.JsonObject;
|
||||
import com.superbiz.agent.config.MilvusProperties;
|
||||
import com.superbiz.agent.constant.MilvusConstants;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import com.superbiz.agent.service.retrieval.RetrievalScoreLabels;
|
||||
import io.milvus.common.clientenum.FunctionType;
|
||||
import io.milvus.v2.client.ConnectConfig;
|
||||
import io.milvus.v2.client.MilvusClientV2;
|
||||
@@ -41,10 +42,35 @@ import java.util.Map;
|
||||
import java.util.UUID;
|
||||
|
||||
/**
|
||||
* Single knowledge vector backend (Milvus Java SDK v2).
|
||||
* 知识库向量后端(Milvus Java SDK v2)—— dense + BM25 混合检索的唯一实现。
|
||||
*
|
||||
* <p>Supports dense ANN and dense+BM25 hybrid search via {@code hybridSearch} + {@link RRFRanker}.
|
||||
* Legacy {@code MilvusServiceClient} search is not used.</p>
|
||||
* <h3>为什么不用 Spring AI {@code spring-ai-starter-vector-store-milvus}</h3>
|
||||
* <ul>
|
||||
* <li>Spring AI Milvus starter(截至 2.0.0 / 1.1.8)只封装 dense {@code similaritySearch}。</li>
|
||||
* <li>底层仍是 V1 {@code MilvusServiceClient} + 单路 {@code SearchParam},无 {@code hybridSearch} /
|
||||
* BM25 Function / {@link RRFRanker}。</li>
|
||||
* <li>真混合检索(dense ANN + 服务端 BM25 sparse,再 RRF 融合)必须走 Milvus SDK v2,
|
||||
* 见 {@link #searchHybrid}。</li>
|
||||
* </ul>
|
||||
*
|
||||
* <h3>Collection schema(默认名 {@code biz})</h3>
|
||||
* <pre>
|
||||
* id VarChar PK
|
||||
* content VarChar —— 原文,返回给上层
|
||||
* search_text VarChar+analyzer —— BM25 输入文本(可含 title/path 增强)
|
||||
* sparse_vector SparseFloatVector —— 由 BM25 Function 从 search_text 自动生成,写入时不必填
|
||||
* vector FloatVector —— dense 向量(应用侧 embedding)
|
||||
* metadata JSON —— docId / source / category / kb_scope 等
|
||||
* </pre>
|
||||
*
|
||||
* <h3>检索模式</h3>
|
||||
* <ul>
|
||||
* <li>{@link #searchDense}:单路 L2 ANN;{@code scoreLabel=dense}。</li>
|
||||
* <li>{@link #searchHybrid}:dense + BM25 + 服务端 {@link RRFRanker};{@code scoreLabel=hybrid};
|
||||
* 返回序即 RRF 序,不再用 dense L2 覆盖主分。</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>配置入口:{@code milvus.collection}、{@code retrieval.search.mode}、{@code retrieval.hybrid.rrf-k}。</p>
|
||||
*/
|
||||
@Service
|
||||
public class MilvusHybridKnowledgeStore {
|
||||
@@ -52,11 +78,20 @@ public class MilvusHybridKnowledgeStore {
|
||||
private static final Logger log = LoggerFactory.getLogger(MilvusHybridKnowledgeStore.class);
|
||||
private static final Gson GSON = new Gson();
|
||||
|
||||
/** 主键(稳定 UUID,由 source + chunkIndex 派生,便于幂等重写)。 */
|
||||
public static final String FIELD_ID = "id";
|
||||
/** 返回给 LLM / 上层的原文 chunk。 */
|
||||
public static final String FIELD_CONTENT = "content";
|
||||
/**
|
||||
* BM25 输入字段。写入明文;Milvus 侧 analyzer + BM25 Function 生成 {@link #FIELD_SPARSE}。
|
||||
* 通常比 content 多带 title/path 等检索增强词。
|
||||
*/
|
||||
public static final String FIELD_SEARCH_TEXT = "search_text";
|
||||
/** 稀疏向量字段;由 BM25 Function 自动产出,insert 时不要手动填。 */
|
||||
public static final String FIELD_SPARSE = "sparse_vector";
|
||||
/** Dense 向量字段(应用侧 EmbeddingModel 生成)。 */
|
||||
public static final String FIELD_DENSE = "vector";
|
||||
/** 业务元数据 JSON(过滤、证据身份、展示用)。 */
|
||||
public static final String FIELD_METADATA = "metadata";
|
||||
|
||||
private final MilvusProperties milvusProperties;
|
||||
@@ -64,12 +99,14 @@ public class MilvusHybridKnowledgeStore {
|
||||
@Value("${milvus.collection:biz}")
|
||||
private String collectionName = "biz";
|
||||
|
||||
/**
|
||||
* RRF 平滑参数 k:score(d) = Σ 1/(k + rank_i(d))。
|
||||
* k 越大,各路排名差异被压得越平;默认 60 与常见 RRF 设定一致。
|
||||
*/
|
||||
@Value("${retrieval.hybrid.rrf-k:60}")
|
||||
private int rrfK = 60;
|
||||
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
/** 非空时追加 {@code metadata.kb_scope} 过滤,实现多知识域隔离。 */
|
||||
@Value("${retrieval.kb-scope:}")
|
||||
private String kbScope = "";
|
||||
|
||||
@@ -79,6 +116,10 @@ public class MilvusHybridKnowledgeStore {
|
||||
this.milvusProperties = milvusProperties;
|
||||
}
|
||||
|
||||
/**
|
||||
* 懒连接:首次调用时建连、确保 collection schema 存在并 load。
|
||||
* 线程安全;后续检索/写入复用同一 {@link MilvusClientV2}。
|
||||
*/
|
||||
public synchronized MilvusClientV2 client() {
|
||||
if (client == null) {
|
||||
client = connect();
|
||||
@@ -92,6 +133,21 @@ public class MilvusHybridKnowledgeStore {
|
||||
return collectionName;
|
||||
}
|
||||
|
||||
/**
|
||||
* 写入单个 chunk(dense + BM25 所需明文)。
|
||||
*
|
||||
* <p>只插入 {@code content / search_text / vector / metadata};
|
||||
* {@code sparse_vector} 由 collection 上的 BM25 Function 在服务端从 {@code search_text} 生成。</p>
|
||||
*
|
||||
* <p>id 由 {@code source|docId + chunkIndex} 的 nameUUID 派生,同一 chunk 重复写入会得到相同 id
|
||||
*(配合先 delete 再 insert 的上层逻辑实现覆盖)。</p>
|
||||
*
|
||||
* @param content 原文(返回字段)
|
||||
* @param searchText BM25 / 可与 dense embedding 同源的检索文本
|
||||
* @param denseVector 应用侧 embedding
|
||||
* @param metadata 须尽量带 {@code _source} 或 {@code docId},供 id 与过滤使用
|
||||
* @param chunkIndex 分片序号
|
||||
*/
|
||||
public void upsertChunk(String content,
|
||||
String searchText,
|
||||
List<Float> denseVector,
|
||||
@@ -110,6 +166,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
JsonObject row = new JsonObject();
|
||||
row.addProperty(FIELD_ID, id);
|
||||
row.addProperty(FIELD_CONTENT, content == null ? "" : content);
|
||||
// 仅写明文;sparse 由 BM25 Function(search_text -> sparse_vector) 自动生成
|
||||
row.addProperty(FIELD_SEARCH_TEXT, searchText == null ? "" : searchText);
|
||||
row.add(FIELD_DENSE, GSON.toJsonTree(denseVector));
|
||||
row.add(FIELD_METADATA, GSON.toJsonTree(metadata == null ? Map.of() : metadata));
|
||||
@@ -120,6 +177,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
.build());
|
||||
}
|
||||
|
||||
/** 按 metadata.docId 删除该文档全部 chunk(重建/覆盖前调用)。 */
|
||||
public void deleteByDocId(String docId) {
|
||||
if (docId == null || docId.isBlank()) {
|
||||
return;
|
||||
@@ -131,6 +189,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
.build());
|
||||
}
|
||||
|
||||
/** 按 metadata._source(规范化路径)删除,用于按文件路径重索引。 */
|
||||
public void deleteBySource(String sourcePath) {
|
||||
if (sourcePath == null || sourcePath.isBlank()) {
|
||||
return;
|
||||
@@ -144,8 +203,8 @@ public class MilvusHybridKnowledgeStore {
|
||||
}
|
||||
|
||||
/**
|
||||
* Drop the configured knowledge collection (if present) and recreate empty dense+BM25 schema.
|
||||
* Used by knowledge rebuild scripts. Existing vectors in this collection are destroyed.
|
||||
* 删除并重建当前知识 collection(空的 dense+BM25 schema)。
|
||||
* 供 {@code /api/knowledge/rebuild-hybrid} 与重建脚本使用;会销毁该 collection 全部向量。
|
||||
*/
|
||||
public synchronized Map<String, Object> dropAndRecreateCollection() {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
@@ -178,6 +237,10 @@ public class MilvusHybridKnowledgeStore {
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* 单路 dense ANN(L2)。
|
||||
* {@code score} = L2 距离(越小越好);{@code scoreLabel} = {@link RetrievalScoreLabels#DENSE}。
|
||||
*/
|
||||
public List<VectorSearchService.SearchResult> searchDense(String queryEmbeddingText,
|
||||
List<Float> queryVector,
|
||||
int topK,
|
||||
@@ -194,12 +257,22 @@ public class MilvusHybridKnowledgeStore {
|
||||
builder.filter(filter);
|
||||
}
|
||||
SearchResp resp = client().search(builder.build());
|
||||
return toSearchResults(resp, "l2_distance", false);
|
||||
return toSearchResults(resp, RetrievalScoreLabels.DENSE);
|
||||
}
|
||||
|
||||
/**
|
||||
* Dense + BM25 hybrid fused by RRF. Dense L2 scores are attached when the same id
|
||||
* appears in a parallel dense search so quality thresholds stay meaningful.
|
||||
* Dense + BM25 真混合检索(Milvus 服务端融合)。
|
||||
*
|
||||
* <ol>
|
||||
* <li>dense 子路:{@code vector},L2</li>
|
||||
* <li>BM25 子路:{@code sparse_vector} + {@link EmbeddedText}</li>
|
||||
* <li>{@link HybridSearchReq} + {@link RRFRanker} → 返回序即权威序</li>
|
||||
* </ol>
|
||||
*
|
||||
* <p>{@code scoreLabel=hybrid};{@code score}/{@code rawScore} 保留引擎融合分,
|
||||
* <b>不</b>用 dense L2 覆盖主分或改 label。可选并行 dense 探测仅填充
|
||||
* {@link VectorSearchService.SearchResult#setDenseDistance},供后处理绝对质量闸门
|
||||
* (如 L0 filter low-quality → unfiltered retry),排序仍以 RRF 返回序为准。</p>
|
||||
*/
|
||||
public List<VectorSearchService.SearchResult> searchHybrid(String queryText,
|
||||
List<Float> queryVector,
|
||||
@@ -236,37 +309,45 @@ public class MilvusHybridKnowledgeStore {
|
||||
.build();
|
||||
|
||||
SearchResp hybridResp = client().hybridSearch(hybridReq);
|
||||
List<VectorSearchService.SearchResult> fused = toSearchResults(hybridResp, "rrf_fused", true);
|
||||
|
||||
// Attach dense-compatible L2 when available.
|
||||
Map<String, Float> denseScores = new HashMap<>();
|
||||
try {
|
||||
for (VectorSearchService.SearchResult denseHit :
|
||||
searchDense(queryText, queryVector, pathTopK, category)) {
|
||||
if (denseHit.getId() != null) {
|
||||
denseScores.put(denseHit.getId(), denseHit.getScore());
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.warn("Dense score enrichment failed: {}", e.getMessage());
|
||||
}
|
||||
for (VectorSearchService.SearchResult hit : fused) {
|
||||
Float dense = denseScores.get(hit.getId());
|
||||
if (dense != null) {
|
||||
hit.setScore(dense);
|
||||
hit.setScoreLabel("l2_distance");
|
||||
} else {
|
||||
// BM25-only hit: treat as weak for legacy thresholds
|
||||
hit.setScore((float) maxL2Distance);
|
||||
hit.setScoreLabel("bm25_only_no_dense");
|
||||
}
|
||||
}
|
||||
List<VectorSearchService.SearchResult> fused = toSearchResults(hybridResp, RetrievalScoreLabels.HYBRID);
|
||||
attachDenseDistances(fused, queryText, queryVector, pathTopK, category);
|
||||
return fused;
|
||||
}
|
||||
|
||||
private List<VectorSearchService.SearchResult> toSearchResults(SearchResp resp,
|
||||
String scoreLabel,
|
||||
boolean fused) {
|
||||
/**
|
||||
* Attach dense L2 by id for quality gates only — never overwrites hybrid score/label/order.
|
||||
*/
|
||||
private void attachDenseDistances(List<VectorSearchService.SearchResult> fused,
|
||||
String queryText,
|
||||
List<Float> queryVector,
|
||||
int pathTopK,
|
||||
String category) {
|
||||
if (fused == null || fused.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
try {
|
||||
Map<String, Float> denseById = new HashMap<>();
|
||||
for (VectorSearchService.SearchResult denseHit :
|
||||
searchDense(queryText, queryVector, pathTopK, category)) {
|
||||
if (denseHit.getId() != null) {
|
||||
denseById.put(denseHit.getId(), denseHit.getScore());
|
||||
}
|
||||
}
|
||||
for (VectorSearchService.SearchResult hit : fused) {
|
||||
Float l2 = denseById.get(hit.getId());
|
||||
if (l2 != null) {
|
||||
hit.setDenseDistance(l2.doubleValue());
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.warn("Dense distance attach for hybrid quality gate failed: {}", e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 将 Milvus {@link SearchResp} 映射为上层结果;列表顺序即检索权威序(adapter 赋 originalRank)。
|
||||
*/
|
||||
private List<VectorSearchService.SearchResult> toSearchResults(SearchResp resp, String scoreLabel) {
|
||||
List<VectorSearchService.SearchResult> out = new ArrayList<>();
|
||||
if (resp == null || resp.getSearchResults() == null || resp.getSearchResults().isEmpty()) {
|
||||
return out;
|
||||
@@ -293,27 +374,16 @@ public class MilvusHybridKnowledgeStore {
|
||||
Float score = row.getScore();
|
||||
mapped.setRawScore(score == null ? null : score.doubleValue());
|
||||
mapped.setScoreLabel(scoreLabel);
|
||||
if (fused) {
|
||||
// temporary; may be overwritten with dense L2
|
||||
mapped.setScore(score == null ? (float) maxL2Distance : invertUnknownScore(score));
|
||||
} else {
|
||||
mapped.setScore(score == null ? (float) maxL2Distance : score);
|
||||
}
|
||||
// dense: L2;hybrid: 引擎融合分(后处理 quality 主要看 rank,不依赖此量纲)
|
||||
mapped.setScore(score == null ? 0f : score);
|
||||
out.add(mapped);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
private float invertUnknownScore(float score) {
|
||||
// RRF-like small scores: map higher better -> small L2-like distance
|
||||
double bounded = Math.max(0.0, Math.min(1.0, score));
|
||||
if (score > 1.0f) {
|
||||
// already distance-like
|
||||
return score;
|
||||
}
|
||||
return (float) ((1.0 - bounded) * maxL2Distance);
|
||||
}
|
||||
|
||||
/**
|
||||
* 组装标量过滤表达式:category、kb_scope(配置级)可叠加,用 {@code &&} 连接。
|
||||
*/
|
||||
private String buildFilter(String category) {
|
||||
List<String> parts = new ArrayList<>();
|
||||
String categoryFilter = trimToNull(category);
|
||||
@@ -352,6 +422,17 @@ public class MilvusHybridKnowledgeStore {
|
||||
return new MilvusClientV2(builder.build());
|
||||
}
|
||||
|
||||
/**
|
||||
* 若不存在则创建 dense+BM25 hybrid collection。
|
||||
*
|
||||
* <p>关键点:</p>
|
||||
* <ul>
|
||||
* <li>{@code search_text} 开启 analyzer,作为 BM25 语料。</li>
|
||||
* <li>{@link FunctionType#BM25}:input={@code search_text} → output={@code sparse_vector}。</li>
|
||||
* <li>dense:IVF_FLAT + L2;sparse:SPARSE_INVERTED_INDEX + BM25。</li>
|
||||
* </ul>
|
||||
* <p>已存在的 collection 不会改 schema;schema 变更需走 {@link #dropAndRecreateCollection()}。</p>
|
||||
*/
|
||||
private void ensureCollection(MilvusClientV2 milvusClient) {
|
||||
Boolean exists = milvusClient.hasCollection(HasCollectionReq.builder()
|
||||
.collectionName(collectionName)
|
||||
@@ -376,6 +457,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
.dataType(DataType.VarChar)
|
||||
.maxLength(MilvusConstants.CONTENT_MAX_LENGTH)
|
||||
.build());
|
||||
// BM25 语料字段:必须 enableAnalyzer,Function 才能从文本生成 sparse
|
||||
schema.addField(AddFieldReq.builder()
|
||||
.fieldName(FIELD_SEARCH_TEXT)
|
||||
.dataType(DataType.VarChar)
|
||||
@@ -395,6 +477,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
.fieldName(FIELD_METADATA)
|
||||
.dataType(DataType.JSON)
|
||||
.build());
|
||||
// 写入 search_text 时,Milvus 自动维护 sparse_vector(应用层 insert 不填 sparse)
|
||||
schema.addFunction(CreateCollectionReq.Function.builder()
|
||||
.functionType(FunctionType.BM25)
|
||||
.name("bm25_fn")
|
||||
@@ -446,6 +529,7 @@ public class MilvusHybridKnowledgeStore {
|
||||
}
|
||||
}
|
||||
|
||||
/** 过滤表达式字符串转义,防止引号打断 expr。 */
|
||||
private static String escapeFilter(String value) {
|
||||
return value.replace("\\", "\\\\").replace("\"", "\\\"");
|
||||
}
|
||||
|
||||
@@ -19,6 +19,8 @@ public record KnowledgeSearchHit(
|
||||
String source,
|
||||
String title,
|
||||
String breadcrumb,
|
||||
int originalRank
|
||||
int originalRank,
|
||||
/** Optional dense L2 for hybrid quality gates; null on dense-only hits. */
|
||||
Double denseDistance
|
||||
) {
|
||||
}
|
||||
|
||||
@@ -1,10 +1,20 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
/**
|
||||
* Retrieval mode for {@link KnowledgeSearchPort}.
|
||||
* Delivery 1 only requires {@link #DENSE}; hybrid arrives in a later change.
|
||||
* {@link KnowledgeSearchPort} 检索模式。
|
||||
*
|
||||
* <ul>
|
||||
* <li>{@link #DENSE} —— 单路向量 ANN(L2)</li>
|
||||
* <li>{@link #HYBRID} —— dense + Milvus 服务端 BM25 + RRF 融合</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>当前实际生效模式由全局配置 {@code retrieval.search.mode} 决定
|
||||
*(见 {@link com.superbiz.agent.service.VectorSearchService});
|
||||
* 请求里的 mode 预留作将来 per-call 覆盖,adapter 暂未按请求切换。</p>
|
||||
*/
|
||||
public enum KnowledgeSearchMode {
|
||||
/** 仅 dense 向量检索。 */
|
||||
DENSE,
|
||||
/** dense + BM25 hybrid(Milvus {@code hybridSearch} + RRF)。 */
|
||||
HYBRID
|
||||
}
|
||||
|
||||
@@ -3,8 +3,11 @@ package com.superbiz.agent.service.retrieval;
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Application boundary for knowledge semantic search.
|
||||
* Implementations may wrap VectorStore, hybrid engines, etc. without leaking SDK details upward.
|
||||
* 知识语义检索的应用边界端口。
|
||||
*
|
||||
* <p>实现可对接 dense / hybrid 等引擎,但不得向上层泄漏 SDK 类型。
|
||||
* 当前实现:{@link VectorKnowledgeSearchAdapter} → {@code VectorSearchService}
|
||||
* → {@code MilvusHybridKnowledgeStore}(Milvus SDK v2 dense 或 dense+BM25 RRF)。</p>
|
||||
*/
|
||||
public interface KnowledgeSearchPort {
|
||||
|
||||
|
||||
@@ -0,0 +1,45 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
/**
|
||||
* 检索结果一级 {@code scoreLabel} 约定。
|
||||
*
|
||||
* <p>只区分两种检索形态(与 {@code retrieval.search.mode} 对齐),
|
||||
* 不再使用 {@code bm25_only_*} 等作为正式一级 label。</p>
|
||||
*/
|
||||
public final class RetrievalScoreLabels {
|
||||
|
||||
/** dense-only ANN:{@code score} 为 L2 距离(越小越好)。 */
|
||||
public static final String DENSE = "dense";
|
||||
|
||||
/** hybrid(dense+BM25+RRF):{@code score}/raw 为融合侧信号;质量分主要看 rank。 */
|
||||
public static final String HYBRID = "hybrid";
|
||||
|
||||
private RetrievalScoreLabels() {
|
||||
}
|
||||
|
||||
/**
|
||||
* 将历史/别名 label 归一到 {@link #DENSE} 或 {@link #HYBRID}。
|
||||
* 未知或空 → dense(保守,按 L2 解释失败时 quality 偏低)。
|
||||
*/
|
||||
public static String canonicalize(String scoreLabel) {
|
||||
if (scoreLabel == null || scoreLabel.isBlank()) {
|
||||
return DENSE;
|
||||
}
|
||||
String label = scoreLabel.trim().toLowerCase();
|
||||
return switch (label) {
|
||||
case DENSE, "l2_distance", "l2" -> DENSE;
|
||||
case HYBRID, "rrf_fused", "rrf", "bm25_only_no_dense", "bm25_only" -> HYBRID;
|
||||
default -> label.contains("hybrid") || label.contains("rrf") || label.contains("bm25")
|
||||
? HYBRID
|
||||
: DENSE;
|
||||
};
|
||||
}
|
||||
|
||||
public static boolean isHybrid(String scoreLabel) {
|
||||
return HYBRID.equals(canonicalize(scoreLabel));
|
||||
}
|
||||
|
||||
public static boolean isDense(String scoreLabel) {
|
||||
return DENSE.equals(canonicalize(scoreLabel));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,71 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
/**
|
||||
* 检索分 → 统一 {@code qualityScore ∈ [0,1]}(越大越好)的唯一转换点。
|
||||
*
|
||||
* <p>后处理排序仍按 {@code originalRank};本类只负责质量闸门 / relevance 用分。</p>
|
||||
*
|
||||
* <ul>
|
||||
* <li>{@link RetrievalScoreLabels#DENSE}:{@code score} = L2 → {@code 1 - clamp(l2)/maxL2}</li>
|
||||
* <li>{@link RetrievalScoreLabels#HYBRID}:优先用可选 {@code denseDistance} 做绝对质量
|
||||
* (恢复 L0 filter low-quality 等闸门);无 dense 时回退 rank 映射</li>
|
||||
* </ul>
|
||||
*/
|
||||
public final class RetrievalScoreNormalizer {
|
||||
|
||||
private RetrievalScoreNormalizer() {
|
||||
}
|
||||
|
||||
/**
|
||||
* @param scoreLabel {@link RetrievalScoreLabels#DENSE} / {@link RetrievalScoreLabels#HYBRID}
|
||||
* @param score 引擎主分:dense=L2;hybrid=融合分(hybrid 质量不依赖其量纲)
|
||||
* @param originalRank 检索名次(1-based)
|
||||
* @param batchSize 本轮候选数(rank 回退映射用)
|
||||
* @param maxL2Distance L2 上界
|
||||
* @param denseDistance hybrid 命中上可选的 dense L2;dense 模式可传 null
|
||||
*/
|
||||
public static double toQualityScore(String scoreLabel,
|
||||
Double score,
|
||||
Integer originalRank,
|
||||
int batchSize,
|
||||
double maxL2Distance,
|
||||
Double denseDistance) {
|
||||
String label = RetrievalScoreLabels.canonicalize(scoreLabel);
|
||||
if (RetrievalScoreLabels.HYBRID.equals(label)) {
|
||||
if (denseDistance != null) {
|
||||
return l2ToQuality(denseDistance, maxL2Distance);
|
||||
}
|
||||
// BM25-only hybrid hit (no dense neighbor): conservative mid quality via rank
|
||||
return rankToQuality(originalRank, batchSize);
|
||||
}
|
||||
return l2ToQuality(score, maxL2Distance);
|
||||
}
|
||||
|
||||
/** Backward-compatible overload without denseDistance. */
|
||||
public static double toQualityScore(String scoreLabel,
|
||||
Double score,
|
||||
Integer originalRank,
|
||||
int batchSize,
|
||||
double maxL2Distance) {
|
||||
return toQualityScore(scoreLabel, score, originalRank, batchSize, maxL2Distance, null);
|
||||
}
|
||||
|
||||
public static double l2ToQuality(Double l2Score, double maxL2Distance) {
|
||||
if (l2Score == null) {
|
||||
return 0.0;
|
||||
}
|
||||
double max = maxL2Distance > 0 ? maxL2Distance : 2.0;
|
||||
double clamped = Math.min(Math.max(l2Score, 0.0), max);
|
||||
return Math.max(0.0, 1.0 - clamped / max);
|
||||
}
|
||||
|
||||
public static double rankToQuality(Integer originalRank, int batchSize) {
|
||||
int rank = originalRank == null || originalRank < 1 ? 1 : originalRank;
|
||||
int n = batchSize > 0 ? Math.max(batchSize, rank) : Math.max(rank, 1);
|
||||
if (n <= 1) {
|
||||
return 1.0;
|
||||
}
|
||||
double quality = 1.0 - (rank - 1) / (double) n;
|
||||
return Math.max(1.0 / n, Math.min(1.0, quality));
|
||||
}
|
||||
}
|
||||
+6
-5
@@ -10,11 +10,11 @@ import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* {@link KnowledgeSearchPort} adapter.
|
||||
* {@link KnowledgeSearchPort} 适配器:把向量检索结果映射为带 evidenceKey 的命中结构。
|
||||
*
|
||||
* <p>Delegates to {@link VectorSearchService}, which is backed solely by
|
||||
* Milvus V2 dense / dense+BM25 hybrid store. Mode selection lives in
|
||||
* {@code retrieval.search.mode}.</p>
|
||||
* <p>委托 {@link VectorSearchService}(背后仅 {@code MilvusHybridKnowledgeStore}):
|
||||
* dense 或 dense+BM25 hybrid 由配置 {@code retrieval.search.mode} 选择。
|
||||
* 本类负责 metadata 解析、docId/chunk 身份与 evidenceKey,不碰 SDK。</p>
|
||||
*/
|
||||
@Component
|
||||
public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort {
|
||||
@@ -76,7 +76,8 @@ public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort {
|
||||
source,
|
||||
EvidenceIdentity.metadataValue(metadata, "title"),
|
||||
EvidenceIdentity.metadataValue(metadata, "breadcrumb"),
|
||||
originalRank
|
||||
originalRank,
|
||||
result.getDenseDistance()
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@@ -167,17 +167,21 @@ rag:
|
||||
enabled: false
|
||||
content-preview-limit: 300
|
||||
|
||||
# 检索配置(单一 Milvus V2 后端;已移除 sdk/spring/auto 路由)
|
||||
# 检索配置
|
||||
# 知识主路径:Milvus Java SDK v2(MilvusHybridKnowledgeStore),非 Spring AI VectorStore starter。
|
||||
# 原因:starter(含 2.0.0)仅 dense similarity,无 hybridSearch / BM25 Function / RRFRanker。
|
||||
# 已移除 legacy sdk/spring/auto 多后端路由。
|
||||
retrieval:
|
||||
kb-scope: "" # empty means search all documents in hybrid collection
|
||||
kb-scope: "" # 非空则过滤 metadata.kb_scope;空=不过滤
|
||||
search:
|
||||
mode: hybrid # dense | hybrid (dense + BM25 RRF)
|
||||
# hybrid=线上主路径;dense=同库对照/评测/排障(非第二套线上策略)。见 mvp/architecture/rag-knowledge-retrieval-architecture.md §6.0
|
||||
mode: hybrid # dense=单路L2对照 | hybrid=dense+服务端BM25+RRF
|
||||
hybrid:
|
||||
rrf-k: 60
|
||||
rrf-k: 60 # RRF 平滑参数 k,score=Σ 1/(k+rank)
|
||||
normalization:
|
||||
max-l2-distance: 2.0 # L2 距离上界(BGE-M3 单位向量 = 2.0)
|
||||
highly-relevant-threshold: 0.75 # similarity >= 0.75 → HIGHLY_RELEVANT
|
||||
reference-threshold: 0.5 # similarity >= 0.5 → REFERENCE
|
||||
max-l2-distance: 2.0 # dense quality:L2 上界(单位向量 ≈ 2.0)
|
||||
highly-relevant-threshold: 0.75 # qualityScore >= 0.75 → PRECISE(hybrid 为序数分,见架构 §6)
|
||||
reference-threshold: 0.5 # qualityScore >= 0.5 → REFERENCE;低于则低质/可 unfiltered retry
|
||||
|
||||
# Prometheus 配置
|
||||
prometheus:
|
||||
@@ -213,6 +217,8 @@ logging:
|
||||
|
||||
# Agent-facing MySQL Tool uses independent logical datasources only.
|
||||
# Production entries are supplied by a dedicated profile and Secret injection.
|
||||
# When data-sources is empty (default), query_mysql is NOT registered on the Diagnosis Agent
|
||||
# (avoids a permanently broken tool that the model can still call).
|
||||
harness:
|
||||
chat:
|
||||
worker-core-pool-size: 2
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user