Compare commits
2
Commits
99d4f6f216
...
376ad0c241
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
376ad0c241 | ||
|
|
ac1f831903 |
@@ -11,6 +11,8 @@
|
||||
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
|
||||
| 2026-07-27 | rag-hybrid-search-rrf | Delivery 2:可配置 hybrid 检索与 RRF 多路融合(不绑旧 SDK)。 | RAG/hybrid/RRF | hybrid mode, RRF, KnowledgeSearchPort, filtered+unfiltered fusion, sparse-lite lexical | openspec/changes/archive/2026-07-27-rag-hybrid-search-rrf | archived |
|
||||
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
|
||||
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
|
||||
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# Acceptance: rag-chunk-evidence-identity-dedup
|
||||
|
||||
## 实现结果
|
||||
|
||||
Delivery 1 completed:
|
||||
|
||||
- chunk identity fields on candidates/evidence blocks
|
||||
- `KnowledgeSearchPort` + dense adapter
|
||||
- evidenceKey dedup + maxChunksPerDocument + return-n
|
||||
- retrieve-k on lookup tool
|
||||
- projector keeps same-source distinct chunks; `document_id` is chunk-scoped
|
||||
|
||||
## 验证
|
||||
|
||||
### 脚本验证
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,KnowledgeEvidencePostProcessorTest,RagResultProjectorTest" test
|
||||
```
|
||||
|
||||
结果:通过(exit 0)
|
||||
|
||||
### 静态验证
|
||||
|
||||
- Search/port wiring reviewed against OpenSpec tasks
|
||||
- No new legacy SDK dependency added
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
未运行(纯检索契约变更,无 UI)
|
||||
|
||||
### 未验证
|
||||
|
||||
- 全量 harness E2E / 真实 Milvus 联调(Delivery 2 前可补)
|
||||
- 生产配置默认 retrieve-k/return-n 调优
|
||||
|
||||
## 归档状态
|
||||
|
||||
- OpenSpec change ready to archive
|
||||
- User pre-authorized archive for sm-flow staged delivery
|
||||
|
||||
## 后续
|
||||
|
||||
- Delivery 2: milvus hybrid search (separate change)
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: rag-chunk-evidence-identity-dedup
|
||||
|
||||
## 背景
|
||||
|
||||
lookup_knowledge 同文档多 chunk 在后处理与投影阶段被 source 级去重吞掉。Hybrid 多路召回前必须先修证据身份与裁剪契约。
|
||||
|
||||
## 目标
|
||||
|
||||
- chunk 级 evidenceKey 身份
|
||||
- 按 evidenceKey 去重 + 每文档 chunk 上限
|
||||
- retrieve-k / return-n 分离
|
||||
- Projector 保留同 source 不同 chunk
|
||||
- 薄 KnowledgeSearchPort,为 Delivery 2 hybrid 铺路
|
||||
|
||||
## 范围
|
||||
|
||||
Delivery 1 only(见 `docs/milvus-hybrid-search-integration-checklist.md` §1.1)。
|
||||
|
||||
## 非目标
|
||||
|
||||
hybrid schema、BM25、删 SDK、session dedup、邻块重建、模型 rerank。
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
`openspec/changes/rag-chunk-evidence-identity-dedup`
|
||||
@@ -0,0 +1,102 @@
|
||||
# Decisions: rag-chunk-evidence-identity-dedup
|
||||
|
||||
## Capability sources
|
||||
|
||||
- sm-flow orchestration
|
||||
- OpenSpec fallback protocol (file-based propose/apply/archive) — external openspec-propose/apply skills used as reference; execution via sm-flow fallback
|
||||
- grill: fallback built-in protocol
|
||||
- audit: fallback built-in protocol
|
||||
|
||||
## Scale
|
||||
|
||||
standard
|
||||
|
||||
## Clarify
|
||||
|
||||
- Problem: same-document multi-chunk evidence collapsed by source-level dedup.
|
||||
- Outcome: Delivery 1 foundation before hybrid Delivery 2.
|
||||
- Slug: `rag-chunk-evidence-identity-dedup`
|
||||
- User authorized apply + archive in advance for sm-flow staged changes.
|
||||
|
||||
## Context
|
||||
|
||||
- Read: `devflow/glossary/CONTEXT.md`, modular-rag-pipeline brief, `openspec/specs/rag-knowledge-retrieval`, `rag-log-projections`, checklist doc §1.1
|
||||
- Constraints into OpenSpec:
|
||||
- L0 hint-only remains
|
||||
- Do not thicken legacy SDK path
|
||||
- Agent tool name/input stable
|
||||
- Hybrid out of scope this change
|
||||
|
||||
## Question pool (grill)
|
||||
|
||||
| # | Dimension | Mode | Question | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | evidence-driven | evidenceKey / document_id 语义? | Resolved: evidenceKey=chunk id; projected document_id=evidenceKey |
|
||||
| Q2 | 边界 | evidence-driven | Delivery 1 vs 2 边界? | Resolved: per checklist; no schema/hybrid/SDK delete |
|
||||
| Q3 | 验收 | evidence-driven | 如何验收多 chunk? | Resolved: unit tests multi-chunk keep + projector |
|
||||
| Q4 | 接口 | user-interview | document_id 改为 chunk 级是否可接受? | **Pre-authorized by user** via “apply/archive 直接授权” + prior design agreement on scheme A (document_id=evidenceKey). Recorded as accepted behavior change. |
|
||||
| Q5 | 技术 | evidence-driven | SearchPort 是否本 change 必须? | Resolved: thin port required as foundation |
|
||||
|
||||
### Evidence-driven conclusions (reported)
|
||||
|
||||
1. Current collapse points: `KnowledgeEvidencePostProcessor.sourceKey` and `RagResultProjector` source fallback.
|
||||
2. Metadata already has docId/chunkIndex on write path; not first-class on read path.
|
||||
3. Existing main-spec still says source-level dedup — this change intentionally deltas that requirement.
|
||||
|
||||
### User-interview
|
||||
|
||||
- Q4 accepted under prior design alignment (scheme A) and explicit apply authorization for this sm-flow run. No remaining open product preference questions for Delivery 1.
|
||||
|
||||
## Audit
|
||||
|
||||
Module chain:
|
||||
|
||||
```text
|
||||
LookupKnowledgeTool -> SearchPort -> Retriever -> PostProcessor -> Packer -> Assembler -> RagResultProjector
|
||||
```
|
||||
|
||||
Risks:
|
||||
|
||||
1. Agent payload growth — mitigated by return-n + maxChunksPerDocument + projector budgets.
|
||||
2. document_id semantic shift — documented L3 behavior change; tests updated.
|
||||
3. Old data without chunkIndex — vector id fallback.
|
||||
|
||||
No ADR conflict with modular RAG L0/L1 boundary.
|
||||
|
||||
## Cross-artifact alignment
|
||||
|
||||
| From | To | Status |
|
||||
|---|---|---|
|
||||
| brief goals | proposal | 已对齐 |
|
||||
| proposal scope | design decisions | 已对齐 |
|
||||
| design identity/dedup/port | specs | 已对齐 |
|
||||
| specs scenarios | tasks | 已对齐 |
|
||||
|
||||
## Interface impact
|
||||
|
||||
- L2 internal DTO
|
||||
- L3 Agent `document_id` chunk-scoped
|
||||
|
||||
## Commit gate
|
||||
|
||||
- proposal/design/specs/tasks present
|
||||
- no open user-interview blockers for Delivery 1
|
||||
- apply authorized by user at sm-flow start
|
||||
|
||||
## Pre-apply research
|
||||
|
||||
Reference files:
|
||||
|
||||
- `LookupKnowledgeTool.java`
|
||||
- `KnowledgeDocumentRetriever.java`
|
||||
- `KnowledgeEvidencePostProcessor.java`
|
||||
- `RagResultProjector.java`
|
||||
- `LookupKnowledgeToolTest.java`
|
||||
- `RagResultProjectorTest.java`
|
||||
- `docs/milvus-hybrid-search-integration-checklist.md`
|
||||
|
||||
Stack notes:
|
||||
|
||||
- No MQ/request envelope changes
|
||||
- Spring `@Value` config pattern for rag.* keys
|
||||
- Tests use ReflectionTestUtils + Mockito
|
||||
@@ -0,0 +1,16 @@
|
||||
# Evidence: rag-chunk-evidence-identity-dedup
|
||||
|
||||
## 代码证据(变更前)
|
||||
|
||||
- `KnowledgeEvidencePostProcessor.sourceKey` 使用 source/title 去重
|
||||
- `RagResultProjector` 用 source 回退 document_id 并 HashSet 去重
|
||||
- 写入路径 metadata 已有 docId/chunkIndex,读路径未一等化
|
||||
|
||||
## 规格证据
|
||||
|
||||
- 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正)
|
||||
- `docs/milvus-hybrid-search-integration-checklist.md` §1.1 定义 Delivery 1 地基
|
||||
|
||||
## 验证证据
|
||||
|
||||
- 单测覆盖 multi-chunk keep / true-dup merge / maxChunksPerDocument / projector same-source multi-chunk
|
||||
@@ -0,0 +1,23 @@
|
||||
# Acceptance: rag-hybrid-search-rrf
|
||||
|
||||
## Result
|
||||
|
||||
Hybrid mode implemented on KnowledgeSearchPort:
|
||||
|
||||
- dense unfiltered + dense filtered + lexical rank over union
|
||||
- RRF fusion by evidenceKey
|
||||
- dense-compatible score preserved for thresholds
|
||||
- default mode remains dense
|
||||
|
||||
## Verification
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=RrfFusionTest,VectorKnowledgeSearchAdapterHybridTest,LookupKnowledgeToolTest" test
|
||||
```
|
||||
|
||||
Pass.
|
||||
|
||||
## Residual
|
||||
|
||||
- True Milvus BM25/sparse schema + reindex still follow-up
|
||||
- Lexical path only ranks dense-recalled candidates (does not expand pure-term misses outside dense topK)
|
||||
@@ -0,0 +1,3 @@
|
||||
# Brief: rag-hybrid-search-rrf
|
||||
|
||||
Delivery 2 after chunk identity. Enable hybrid multi-path + RRF on KnowledgeSearchPort without legacy SDK hybrid API. True BM25 schema rebuild is staged follow-up; this change ships sparse-lite lexical ranking over dense candidate union + filtered/unfiltered dense fusion.
|
||||
@@ -0,0 +1,19 @@
|
||||
# Decisions: rag-hybrid-search-rrf
|
||||
|
||||
## Capability
|
||||
|
||||
sm-flow + OpenSpec fallback; apply pre-authorized.
|
||||
|
||||
## Depends
|
||||
|
||||
Delivery 1 archived.
|
||||
|
||||
## Grill (compressed, pre-authorized)
|
||||
|
||||
- Q: Full BM25 schema now? A: No — sparse-lite + RRF first; schema rebuild follow-up.
|
||||
- Q: Default mode? A: dense default; hybrid opt-in.
|
||||
- Q: Threshold score? A: keep dense-compatible L2 mapping.
|
||||
|
||||
## Design
|
||||
|
||||
Hybrid paths: dense unfiltered + dense filtered + lexical rank over union; RRF fuse by evidenceKey.
|
||||
@@ -0,0 +1,5 @@
|
||||
# Evidence
|
||||
|
||||
- Delivery 1 identity/port foundation required
|
||||
- RRF utility and hybrid adapter unit tests green
|
||||
- Lexical sparse-lite intentionally intermediate until BM25 schema
|
||||
@@ -0,0 +1,809 @@
|
||||
# Milvus Hybrid Search 接入对照清单
|
||||
|
||||
**日期**:2026-07-27
|
||||
**前提**:旧 Milvus SDK 直连检索路径后续废弃,不作为长期实现基础
|
||||
**目标**:在现有 `lookup_knowledge` pipeline 上接入 dense + sparse/BM25 混合检索,融合优先走服务端 RRF
|
||||
**关联文档**:
|
||||
- `docs/rag-ranking-multipath-retrieval-and-rrf.md`(排序与多路召回判断框架)
|
||||
- 本文后续实现讨论以本节 **「交付拆分:分块去重 + Hybrid 同规划」** 为基线
|
||||
|
||||
---
|
||||
|
||||
## 1. 结论先说
|
||||
|
||||
可以接,而且和前面讨论的多路召回 / RRF 高度一致。
|
||||
但当前项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是:
|
||||
|
||||
```text
|
||||
1. schema 只有 dense,没有 sparse/BM25 字段
|
||||
2. 写入只产 dense embedding
|
||||
3. 检索抽象仍以单路 similaritySearch 为中心
|
||||
4. 后处理仍承担了过多“伪融合”职责
|
||||
5. 证据去重粒度偏文档/source 级,同文档多 chunk 会被吞掉
|
||||
```
|
||||
|
||||
**接入原则:**
|
||||
|
||||
```text
|
||||
- 不继续加厚旧 SDK search 分支
|
||||
- 以“检索端口 + 写入端口”抽象为准
|
||||
- hybrid 融合尽量下沉到向量库(RRFRanker)
|
||||
- 应用层保留:filter 策略、chunk 去重、return-n、阈值、投影
|
||||
- 分块去重与 hybrid 同规划、分里程碑交付(先共用地基,再开 hybrid)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1.1 交付拆分:分块去重 + Hybrid 同规划
|
||||
|
||||
> 实现讨论基线。后续排期、拆 PR、评审范围,默认按本节两个交付理解。
|
||||
|
||||
### 判断
|
||||
|
||||
**可以一起做,而且应该绑在同一条改造主线上**;
|
||||
但不要理解成「一个 PR 把 hybrid 全做完」。
|
||||
|
||||
更准确的表述:
|
||||
|
||||
```text
|
||||
同一条演进线,两层交付:
|
||||
交付 1:共用地基(证据身份 + chunk 去重 + 检索裁剪 + 端口雏形)
|
||||
交付 2:hybrid(schema/写入/查询/RRF + 阈值校准)
|
||||
```
|
||||
|
||||
### 为什么必须同规划
|
||||
|
||||
两边改的是同一条链上的相邻环节:
|
||||
|
||||
```text
|
||||
检索命中
|
||||
-> 候选身份(docId / chunkIndex / evidenceKey) ← hybrid 要,去重也要
|
||||
-> 去重 / 单文档 chunk 上限 ← 分块去重
|
||||
-> 排序融合(现在规则 / 以后库内 RRF) ← hybrid
|
||||
-> return-n / Agent 投影
|
||||
```
|
||||
|
||||
若拆开且顺序错误:
|
||||
|
||||
| 只做一项 | 后果 |
|
||||
|---|---|
|
||||
| 只做 hybrid,不做 chunk 去重 | 多路召回更多同文档片段,仍被 source 级去重吞掉,**hybrid 收益被吃掉** |
|
||||
| 只做去重,完全不管候选/端口结构 | 能立刻改善,但接 hybrid 时往往还要再改一遍 DTO 与映射 |
|
||||
|
||||
因此:
|
||||
|
||||
> **分块去重不是 hybrid 的可选项,而是 hybrid 生效的前提。**
|
||||
> 设计上当一件事;代码上分两个可独立验证的里程碑。
|
||||
|
||||
### 必须放进同一批(交付 1 公共地基)
|
||||
|
||||
这些强烈建议同一波完成,作为后续实现讨论的最小必选范围:
|
||||
|
||||
| 项 | 原因 |
|
||||
|---|---|
|
||||
| 候选补 `docId` / `chunkIndex` / `evidenceKey` | 去重 key 与 hybrid hit 身份统一 |
|
||||
| 后处理按 `docId#chunkIndex`(或 vector id fallback)去重 | 修复「同文档多 chunk 被吞」 |
|
||||
| `maxChunksPerDocument` | 放开多 chunk 后防止单文档刷屏 |
|
||||
| `retrieve-k` / `return-n` 分离 | hybrid 扩召回时必需;现在 K=3 也不该三者混用 |
|
||||
| Projector 去重语义对齐 | 后处理放出的多 chunk,不能在投影阶段再按 `source` 砍成 1 条 |
|
||||
| `SearchHit` / `RetrievedEvidenceCandidate` 字段对齐 | 避免 hybrid 再引入第三套结果结构 |
|
||||
| (建议)`KnowledgeSearchPort` 雏形 | 检索调用面先稳定,后续只换实现 |
|
||||
|
||||
可称为:
|
||||
|
||||
```text
|
||||
「检索结果身份与裁剪契约」
|
||||
```
|
||||
|
||||
**不上 hybrid 也有独立价值**,并且为交付 2 铺路。
|
||||
|
||||
### 不要硬塞进交付 1 的同一 PR
|
||||
|
||||
可同规划、建议第二波(交付 2):
|
||||
|
||||
| 项 | 原因 |
|
||||
|---|---|
|
||||
| 新 collection + sparse/BM25 schema | 数据迁移/重灌,风险独立 |
|
||||
| 全量重索引 | 耗时长,需单独验证 |
|
||||
| 打开 `search.mode=hybrid` | 依赖 sparse 数据已就绪 |
|
||||
| fused score 阈值重标定 | 要 hybrid 跑起来后有样本 |
|
||||
| 删除旧 SDK 读路径 | 最后做,降低回滚成本 |
|
||||
|
||||
否则单个交付会同时碰:业务排序逻辑 + 数据迁移 + 基础设施,评审、回滚、评测都困难。
|
||||
|
||||
### 交付 1:chunk 级证据身份 + 去重 + 检索裁剪
|
||||
|
||||
**主题:** 让同一次检索内,同文档多个相关 chunk 能作为独立证据存活,并为 hybrid 统一 hit 模型。
|
||||
|
||||
**范围(实现讨论默认包含):**
|
||||
|
||||
```text
|
||||
1. RetrievedEvidenceCandidate / EvidenceBlock
|
||||
- 补 docId、chunkIndex、evidenceKey
|
||||
2. KnowledgeDocumentRetriever
|
||||
- 从 metadata 抽取 docId/chunkIndex
|
||||
- evidenceKey 规则:
|
||||
docId + "#chunk-" + chunkIndex
|
||||
fallback: "vector:" + id
|
||||
fallback: "rank:" + originalRank
|
||||
3. KnowledgeEvidencePostProcessor
|
||||
- 去重 key = evidenceKey(不再 source/title 优先)
|
||||
- maxChunksPerDocument(建议默认 2)
|
||||
- 同 key 才 merge;merge 不覆盖更高分 content
|
||||
4. RagResultProjector
|
||||
- 按 evidence 身份去重(chunk 级 document_id 或显式 chunk 身份)
|
||||
- 禁止再仅用 source 当“每文档一条”的唯一键
|
||||
5. 配置
|
||||
- rag.retrieve-k
|
||||
- rag.return-n
|
||||
- rag.max-chunks-per-document
|
||||
- 逐步弱化/废弃单一 rag.top-k 身兼多职
|
||||
6. (建议同批)KnowledgeSearchPort / SearchRequest / SearchHit 雏形
|
||||
- 即使底层暂时仍是 dense-only,调用面先稳定
|
||||
7. 单测
|
||||
- 同 doc 两 chunk 都保留
|
||||
- 同 doc+chunk 真重复只留一条
|
||||
- 超 maxChunksPerDocument 裁掉低分
|
||||
- projector 不再误杀同 source 不同 chunk
|
||||
```
|
||||
|
||||
**明确不包含:**
|
||||
|
||||
```text
|
||||
- sparse/BM25 schema
|
||||
- 全量重灌
|
||||
- hybridSearch 开关
|
||||
- 旧 SDK 删除
|
||||
```
|
||||
|
||||
**完成定义(交付 1 Done):**
|
||||
|
||||
```text
|
||||
[ ] 同文档多相关 chunk 可同时出现在内部 evidenceBlocks
|
||||
[ ] Agent 投影后仍能看到多于 1 条同文档片段(未超预算时)
|
||||
[ ] retrieve-k / return-n 可配置且行为可测
|
||||
[ ] 候选身份字段稳定,足够支撑后续 hybrid hit 映射
|
||||
[ ] 不依赖旧 SDK 新增逻辑
|
||||
```
|
||||
|
||||
### 交付 2:Hybrid 写入 + 查询
|
||||
|
||||
**主题:** 在交付 1 的身份/裁剪契约稳定后,打开 dense + sparse/BM25 与库内 RRF。
|
||||
|
||||
**范围:**
|
||||
|
||||
```text
|
||||
1. 新 collection schema(dense + sparse/BM25 + 必要标量字段)
|
||||
2. 写入 dense + sparse,doc_id 级删除与重灌
|
||||
3. KnowledgeSearchPort 实现 hybrid 模式
|
||||
4. ranker = RRF(默认)/ Weighted(可配路权)
|
||||
5. category filter 策略:
|
||||
- 唯一 domain 时 filtered hybrid
|
||||
- 低质时 unfiltered 兜底(或双路径轻量合并)
|
||||
6. 分数语义区分 fused/dense/sparse,重标定 found/relevance
|
||||
7. 回归评测与延迟对比
|
||||
8. 冻结并最终删除旧 SDK 读路径
|
||||
```
|
||||
|
||||
**完成定义(交付 2 Done):**
|
||||
|
||||
```text
|
||||
[ ] 可配置 dense | hybrid 切换
|
||||
[ ] hybrid 默认 RRF,术语类与语义类回归不回退
|
||||
[ ] filter 误杀有兜底
|
||||
[ ] 应用层仍按 chunk 身份去重,hybrid 多命中不会被 source 级逻辑误伤
|
||||
[ ] 新逻辑不再依赖旧 SDK search
|
||||
```
|
||||
|
||||
### 节奏与评审方式
|
||||
|
||||
```text
|
||||
规划:一件事(检索质量主线)
|
||||
设计评审:按交付 1 + 交付 2 两章看
|
||||
开发:
|
||||
先合并交付 1(可独立上线/验证)
|
||||
再做交付 2(数据迁移 + hybrid 开关)
|
||||
验收:
|
||||
交付 1 用“同文档多 chunk”用例
|
||||
交付 2 用“术语/语义/filter 兜底/延迟”用例
|
||||
```
|
||||
|
||||
### 反模式(实现讨论时直接否决)
|
||||
|
||||
```text
|
||||
❌ 一个大 PR:去重 + schema 重灌 + hybrid 开关 + 删 SDK
|
||||
❌ 先上 hybrid、后补 chunk 去重
|
||||
❌ 交付 1 仍按 source 去重,只把 hybrid 分数接进来
|
||||
❌ 为 hybrid 新建第三套与 candidate/EvidenceBlock 并行的结果模型长期共存
|
||||
❌ 在旧 SDK search 实现里继续堆 hybrid 细节作为长期方案
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 现状对照
|
||||
|
||||
| 层级 | 当前实现 | Hybrid 需要 |
|
||||
|---|---|---|
|
||||
| Collection | `id / vector / content / metadata` | 至少再有 sparse/BM25 文本检索能力 |
|
||||
| 写入 | `VectorIndexService` 只写 dense | 同步维护 dense + sparse/BM25 |
|
||||
| 检索门面 | `VectorSearchService`:`sdk \| spring \| auto` | 单端口:`search(query, options)`,内部可 hybrid |
|
||||
| 旧 SDK 路径 | `MilvusServiceClient.search` | **废弃,不再作为主实现** |
|
||||
| Spring AI 路径 | `VectorStore.similaritySearch` | 可作过渡 dense 读路径,但 hybrid 能力要单独确认/扩展 |
|
||||
| 后处理 | 规则 boost + source 去重 | 融合交给库;后处理做裁剪/等级/打包 |
|
||||
| Agent 投影 | `RagResultProjector` | 基本不动 |
|
||||
|
||||
当前关键文件:
|
||||
|
||||
```text
|
||||
写入:
|
||||
VectorIndexService
|
||||
DocumentChunkService
|
||||
VectorEmbeddingService
|
||||
MilvusClientFactory # schema/index 创建(旧)
|
||||
|
||||
读取:
|
||||
VectorSearchService # 门面,含 sdk/spring 路由
|
||||
KnowledgeDocumentRetriever
|
||||
KnowledgeEvidencePostProcessor
|
||||
LookupKnowledgeTool
|
||||
|
||||
配置:
|
||||
retrieval.vector-store.mode
|
||||
retrieval.kb-scope
|
||||
rag.top-k
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 目标架构(不绑旧 SDK)
|
||||
|
||||
```text
|
||||
┌─────────────────────────┐
|
||||
upload/init │ KnowledgeWritePort │
|
||||
chunk + embed -> │ - upsertChunks() │
|
||||
│ - deleteByDocId() │
|
||||
└───────────┬─────────────┘
|
||||
│
|
||||
▼
|
||||
Vector DB
|
||||
dense + sparse/BM25
|
||||
metadata filters
|
||||
▲
|
||||
┌───────────┴─────────────┐
|
||||
lookup_knowledge │ KnowledgeSearchPort │
|
||||
query + options-> │ - search() │
|
||||
│ - mode: DENSE/HYBRID │
|
||||
└───────────┬─────────────┘
|
||||
│
|
||||
▼
|
||||
KnowledgeDocumentRetriever
|
||||
│
|
||||
▼
|
||||
PostProcess(dedup/chunk cap/threshold/pack)
|
||||
│
|
||||
▼
|
||||
LookupResult / Projector
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- **Port** 是应用边界,实现可换成 Spring AI、Milvus 新客户端、或其他封装。
|
||||
- 旧 `MilvusServiceClient` 检索实现可以暂时留着,但 **新功能不要往里堆**。
|
||||
- Hybrid 是 `KnowledgeSearchPort` 的一种 mode,不是再开一套平行 tool。
|
||||
|
||||
---
|
||||
|
||||
## 4. Schema 改造清单
|
||||
|
||||
### 4.1 建议逻辑模型
|
||||
|
||||
```text
|
||||
id string PK # chunk 级唯一 id
|
||||
doc_id string # 文档 id(从 metadata 提升为一等字段更稳)
|
||||
chunk_index int
|
||||
content text/varchar # 原始 chunk 正文(给 BM25 / 返回)
|
||||
title string nullable
|
||||
breadcrumb string nullable
|
||||
category string nullable
|
||||
kb_scope string nullable
|
||||
dense_vector float vector # embedding(title/path/content)
|
||||
sparse_vector sparse vector # BM25 或 sparse embedding
|
||||
metadata json # 兼容扩展字段
|
||||
```
|
||||
|
||||
### 4.2 和现状差异
|
||||
|
||||
| 字段 | 现状 | 建议 |
|
||||
|---|---|---|
|
||||
| `vector` | 有 | 可改名 `dense_vector`,或保留别名兼容 |
|
||||
| `content` | 有,仅存储/返回 | 同时作为 BM25 输入文本 |
|
||||
| `sparse_vector` | 无 | **新增,hybrid 必需** |
|
||||
| `docId/chunkIndex` | 塞在 JSON metadata | 建议提升为可过滤/可排序字段 |
|
||||
| `category/kb_scope` | metadata JSON | 建议提升,filter 更稳 |
|
||||
|
||||
### 4.3 索引
|
||||
|
||||
```text
|
||||
dense_vector -> 向量索引(COSINE/IP/L2,与 embedding 一致)
|
||||
sparse_vector -> 稀疏倒排 / BM25 索引
|
||||
category/kb_scope/doc_id -> 标量过滤索引(如需要)
|
||||
```
|
||||
|
||||
### 4.4 迁移策略
|
||||
|
||||
不要幻想“只改 search 方法”:
|
||||
|
||||
1. **新建 collection 或新版本 collection**(推荐)
|
||||
2. 全量重灌知识库(dense + sparse)
|
||||
3. 双写一段时间(可选)
|
||||
4. 切换读路径到 hybrid
|
||||
5. 下线旧 collection / 旧 SDK 读路径
|
||||
|
||||
就地改老 collection 风险高:已有数据无 sparse,历史 metadata 形态也不统一。
|
||||
|
||||
---
|
||||
|
||||
## 5. 写入路径改造清单
|
||||
|
||||
### 5.1 需要动的职责
|
||||
|
||||
| 类/模块 | 现在 | 改造 |
|
||||
|---|---|---|
|
||||
| `DocumentChunkService` | 产出 chunk 正文/title/breadcrumb | 基本可复用 |
|
||||
| `VectorEmbeddingService` | 只做 dense embed | 保留;sparse/BM25 另算或交给库 |
|
||||
| `VectorIndexService` | 组装 metadata + insert dense | 升级为 write port 实现:dense+sparse 一并 upsert |
|
||||
| 删除逻辑 | 按 `metadata.docId` / `_source` 删 | 统一按 `doc_id` 删,避免路径不一致 |
|
||||
|
||||
### 5.2 写入时每条 chunk 必须具备
|
||||
|
||||
```text
|
||||
- dense_vector: embed(buildEmbeddingText(chunk))
|
||||
- sparse 输入: 建议用“可检索文本”
|
||||
title + breadcrumb + content
|
||||
而不是只丢 raw content
|
||||
- doc_id / chunk_index / category / kb_scope
|
||||
- 稳定 chunk id(doc_id + chunk_index 派生)
|
||||
```
|
||||
|
||||
### 5.3 注意
|
||||
|
||||
- embedding 文本可以继续拼 `Title/Path/Content`
|
||||
- **返回给 Agent 的 content 仍应是原文 chunk**,不要返回 embedding 拼接串
|
||||
- BM25 文本建议包含 title/breadcrumb,否则专有名词在标题里时字面路会弱
|
||||
|
||||
---
|
||||
|
||||
## 6. 检索路径改造清单
|
||||
|
||||
### 6.1 新检索端口(建议)
|
||||
|
||||
不要继续扩:
|
||||
|
||||
```text
|
||||
searchSimilarDocuments(query, topK, category)
|
||||
```
|
||||
|
||||
建议收敛成:
|
||||
|
||||
```text
|
||||
SearchRequest {
|
||||
query: string
|
||||
retrieveK: int # 例如 20
|
||||
returnN: int # 例如 5,可在后处理裁
|
||||
mode: DENSE | HYBRID
|
||||
categoryFilter?: string
|
||||
kbScope?: string
|
||||
ranker: RRF | WEIGHTED
|
||||
rrfK: int # 默认 60
|
||||
weights?: {dense, sparse}
|
||||
}
|
||||
|
||||
SearchHit {
|
||||
id, docId, chunkIndex
|
||||
content, title, breadcrumb
|
||||
source, category
|
||||
scores: {
|
||||
fused?, denseRank?, sparseRank?, raw?...
|
||||
}
|
||||
metadata
|
||||
}
|
||||
```
|
||||
|
||||
### 6.2 `VectorSearchService` 怎么演进
|
||||
|
||||
短期:
|
||||
|
||||
```text
|
||||
保留门面类名也可
|
||||
但内部:
|
||||
- 不再把 sdk 当长期分支
|
||||
- 增加 hybridSearch(...) 能力
|
||||
- mode 配置改为:
|
||||
dense | hybrid
|
||||
(spring 仅作 dense 兼容实现)
|
||||
```
|
||||
|
||||
中期:
|
||||
|
||||
```text
|
||||
VectorSearchService 实现 KnowledgeSearchPort
|
||||
旧 sdk 分支删除或仅 test/fallback 开关默认关
|
||||
```
|
||||
|
||||
### 6.3 Hybrid 查询语义
|
||||
|
||||
```text
|
||||
路 A: dense(query_embedding) limit=retrieveK
|
||||
路 B: bm25/sparse(query_text) limit=retrieveK
|
||||
可选过滤: category / kb_scope
|
||||
融合: RRFRanker(k=60) 或 WeightedRanker
|
||||
输出: top retrieveK/returnN
|
||||
```
|
||||
|
||||
对应我们之前的公式:
|
||||
|
||||
```text
|
||||
RRF_w(d) = Σ w_i / (k + rank_i(d))
|
||||
```
|
||||
|
||||
- 用 RRF:先不调权重
|
||||
- 用 Weighted:调的是 **dense/sparse 路权**,不是 keyword contains 加分
|
||||
|
||||
### 6.4 filtered + unfiltered 还要不要?
|
||||
|
||||
还要,但定位变了:
|
||||
|
||||
| 能力 | 放哪 |
|
||||
|---|---|
|
||||
| dense + bm25 融合 | **库内 hybrid** |
|
||||
| category filter 开/关 | 应用策略层,可变成两次 hybrid 或 filter 参数 |
|
||||
| chunk 去重 / 每文档上限 | 应用后处理 |
|
||||
| found / relevanceLevel | 应用后处理 |
|
||||
|
||||
推荐策略:
|
||||
|
||||
```text
|
||||
if 唯一 domain:
|
||||
hybrid(query, filter=category) # 主路
|
||||
若低质量:
|
||||
hybrid(query, filter=null) # 兜底
|
||||
或并行两条 hybrid 再做一次轻量合并
|
||||
else:
|
||||
hybrid(query, filter=null)
|
||||
```
|
||||
|
||||
注意:这里的“两条”是 **filter 策略双路径**,不是再手写一套 dense/bm25 融合。
|
||||
|
||||
---
|
||||
|
||||
## 7. 和现有 pipeline 的衔接(按类)
|
||||
|
||||
### 7.1 基本不动
|
||||
|
||||
| 类 | 原因 |
|
||||
|---|---|
|
||||
| `LookupKnowledgeTool` | 继续编排 transform → retrieve → post → pack |
|
||||
| `KnowledgeQueryTransformer` | 仍产 categoryFilter / hints |
|
||||
| `KnowledgeContextPacker` | 仍做字符预算 |
|
||||
| `LookupResultAssembler` | 仍组装内部结果 |
|
||||
| `RagToolAdapter` / `RagResultProjector` | Agent 契约保持稳定 |
|
||||
|
||||
### 7.2 要改
|
||||
|
||||
| 类 | 改什么 |
|
||||
|---|---|
|
||||
| `KnowledgeDocumentRetriever` | 调新 search port;透传 retrieveK/mode;把 docId/chunkIndex 提成候选一等字段 |
|
||||
| `KnowledgeEvidencePostProcessor` | 弱化“跨路融合”职责;保留 dedup、chunk cap、阈值、轻精排 |
|
||||
| `VectorSearchService` | 成为 hybrid 入口,去掉对旧 sdk 的依赖增长 |
|
||||
| `VectorIndexService` | 写入 dense+sparse,统一 doc_id 删除 |
|
||||
| schema 工厂/初始化 | 新 collection 定义与索引 |
|
||||
|
||||
### 7.3 后处理职责重新划分
|
||||
|
||||
**交给 Milvus hybrid:**
|
||||
|
||||
- dense/sparse 多路召回
|
||||
- RRF / weighted 融合
|
||||
- 基础 topK
|
||||
|
||||
**留给应用层:**
|
||||
|
||||
```text
|
||||
1. chunk 级去重(docId#chunkIndex)
|
||||
2. maxChunksPerDocument
|
||||
3. return-n 裁剪
|
||||
4. relevanceLevel / isLowQuality
|
||||
5. 可选轻规则精排(title/breadcrumb 命中)
|
||||
6. context pack 与 Agent 投影
|
||||
```
|
||||
|
||||
**应降级或删除的:**
|
||||
|
||||
```text
|
||||
把 keyword contains 大额加分当主排序器
|
||||
在应用层重复实现一套 dense+bm25 分数硬加
|
||||
```
|
||||
|
||||
L0 仍只做:
|
||||
|
||||
```text
|
||||
- 是否启用 category filter
|
||||
- 轻量精排特征
|
||||
- trace 解释
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. 配置建议(示意)
|
||||
|
||||
```properties
|
||||
# 检索模式:dense | hybrid
|
||||
retrieval.search.mode=hybrid
|
||||
|
||||
# 旧 sdk 读路径默认关闭(后续删除)
|
||||
retrieval.legacy-sdk.enabled=false
|
||||
|
||||
# 召回/返回分离
|
||||
rag.retrieve-k=20
|
||||
rag.return-n=5
|
||||
rag.max-chunks-per-document=2
|
||||
|
||||
# 融合
|
||||
retrieval.hybrid.ranker=rrf
|
||||
retrieval.hybrid.rrf-k=60
|
||||
# 若用 weighted:
|
||||
# retrieval.hybrid.ranker=weighted
|
||||
# retrieval.hybrid.weight.dense=1.0
|
||||
# retrieval.hybrid.weight.sparse=0.8
|
||||
|
||||
# filter 策略
|
||||
retrieval.filter.retry-unfiltered-on-low-quality=true
|
||||
retrieval.normalization.reference-threshold=0.5
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- `rag.top-k` 应逐步废弃,避免“召回/返回/展示”一个参数打天下
|
||||
- 阈值字段若 hybrid 后分数语义变化,需要重新校准,不能照搬旧 L2 经验值
|
||||
|
||||
---
|
||||
|
||||
## 9. 分数与阈值:hybrid 后要重标定
|
||||
|
||||
当前后处理默认假设:
|
||||
|
||||
```text
|
||||
score ≈ 兼容 L2 距离
|
||||
normalizeL2 后得到 0~1
|
||||
```
|
||||
|
||||
hybrid 后常见变化:
|
||||
|
||||
| 来源 | 语义 |
|
||||
|---|---|
|
||||
| dense raw | L2 / cosine |
|
||||
| sparse/BM25 raw | 另一套 |
|
||||
| fused RRF | 名次融合分,不是相似度概率 |
|
||||
|
||||
因此:
|
||||
|
||||
1. `SearchHit` 要区分 `fusedScore` / `denseScore` / `sparseScore`
|
||||
2. `isLowQuality` 不要直接拿 RRF 分当旧 L2 用
|
||||
3. 过渡期可:
|
||||
- 用“是否有命中 + 规则完整性”判断 found
|
||||
- 或只对 dense 分做阈值,RRF 只负责排序
|
||||
4. 重新用 15~30 条回归 query 标定
|
||||
|
||||
---
|
||||
|
||||
## 10. 分阶段落地(推荐)
|
||||
|
||||
> 与 **§1.1 交付 1 / 交付 2** 对齐。Phase 编号用于执行拆解;对外沟通优先用两个交付里程碑。
|
||||
|
||||
### 交付 1 对应 Phase
|
||||
|
||||
#### Phase 0:应用层前提(交付 1 核心)
|
||||
|
||||
```text
|
||||
[ ] chunk 级去重(不要 source 级吞 chunk)
|
||||
[ ] 候选暴露 docId / chunkIndex / evidenceKey
|
||||
[ ] retrieve-k / return-n 分离
|
||||
[ ] maxChunksPerDocument
|
||||
[ ] Projector 按证据身份去重(不再 source 唯一)
|
||||
[ ] 明确 legacy-sdk 读路径仅兼容、默认关或冻结
|
||||
[ ] 单测:同文档多 chunk / 真重复 / 单文档上限
|
||||
```
|
||||
|
||||
#### Phase 1:检索端口收敛(交付 1 建议同批或紧随)
|
||||
|
||||
```text
|
||||
[ ] 定义 KnowledgeSearchPort / SearchRequest / SearchHit
|
||||
[ ] VectorSearchService 适配该端口(先 dense-only 也可)
|
||||
[ ] KnowledgeDocumentRetriever 只依赖端口
|
||||
[ ] 单测用 fake search port,不再绑 SDK 细节
|
||||
```
|
||||
|
||||
**交付 1 出口:** 不上 hybrid 也可合并;hybrid 所需 hit 身份与裁剪契约已稳定。
|
||||
|
||||
### 交付 2 对应 Phase
|
||||
|
||||
#### Phase 2:写入与 schema 支持 sparse/BM25
|
||||
|
||||
```text
|
||||
[ ] 新 collection schema
|
||||
[ ] 写入 dense + sparse/BM25 文本
|
||||
[ ] doc_id 级删除与重灌
|
||||
[ ] 知识库全量重建脚本/任务
|
||||
```
|
||||
|
||||
#### Phase 3:打开 hybrid 读路径
|
||||
|
||||
```text
|
||||
[ ] search.mode=hybrid
|
||||
[ ] ranker=rrf
|
||||
[ ] category filter 策略接入
|
||||
[ ] low-quality 时 unfiltered 兜底
|
||||
[ ] trace 记录 dense/sparse/fused 信息(内部)
|
||||
[ ] 确认 hybrid 多命中仍走 chunk 级去重,不被 source 误伤
|
||||
```
|
||||
|
||||
#### Phase 4:瘦身后处理 + 下线旧路径
|
||||
|
||||
```text
|
||||
[ ] 规则 boost 降为轻精排或可关
|
||||
[ ] 删除/隔离旧 SDK search 实现
|
||||
[ ] 校准 found/relevance 阈值
|
||||
[ ] 回归评测与延迟对比
|
||||
```
|
||||
|
||||
**交付 2 出口:** dense|hybrid 可切换;RRF 默认可用;旧 SDK 检索不再被新逻辑依赖。
|
||||
|
||||
---
|
||||
|
||||
## 11. 类级改造对照表
|
||||
|
||||
| 类 | 优先级 | 动作 | 是否依赖旧 SDK |
|
||||
|---|---|---|---|
|
||||
| `KnowledgeDocumentRetriever` | P0 | 接新端口,透传 hybrid 选项,补 chunk 身份 | 否 |
|
||||
| `KnowledgeEvidencePostProcessor` | P0 | 去重改 chunk 级;融合职责外移 | 否 |
|
||||
| `LookupKnowledgeTool` | P1 | 使用 retrieve-k/return-n;保留 filter 降级策略 | 否 |
|
||||
| `VectorSearchService` | P0 | 增加 hybrid;冻结/移除 sdk 增长 | 实现可无 SDK |
|
||||
| `VectorIndexService` | P0 | dense+sparse 写入,doc_id 删除 | 实现可无 SDK |
|
||||
| `MilvusClientFactory` | P2 | 仅迁移期维护;新 schema 建议新模块 | 旧 |
|
||||
| `VectorEmbeddingService` | P1 | 继续 dense;不塞 hybrid 逻辑 | 否 |
|
||||
| `RagResultProjector` | P2 | 若 document_id 变 chunk 级,同步语义 | 否 |
|
||||
| Spring AI `VectorStore` | P2 | 可继续承载 dense;hybrid 需单独能力层 | 否 |
|
||||
|
||||
---
|
||||
|
||||
## 12. 测试清单
|
||||
|
||||
### 单元
|
||||
|
||||
```text
|
||||
[ ] RRF 融合结果顺序(可用 fixture,不连库)
|
||||
[ ] chunk 去重:同 doc 不同 chunk 都保留
|
||||
[ ] 同 doc 超过 maxChunksPerDocument 被裁
|
||||
[ ] category filter 低质时走 unfiltered
|
||||
[ ] SearchHit 字段映射:docId/chunkIndex/content
|
||||
```
|
||||
|
||||
### 集成 / 回归
|
||||
|
||||
```text
|
||||
[ ] 专有名词/错误码 query:hybrid 应优于 pure dense
|
||||
[ ] 换说法语义 query:hybrid 不低于 pure dense
|
||||
[ ] 错误 domain filter:unfiltered 兜底仍能找回
|
||||
[ ] 长文档多 chunk:返回不少于 2 个相关片段(若存在)
|
||||
[ ] 延迟:hybrid P95 可接受
|
||||
[ ] 重灌后旧 doc 删除干净,无幽灵 chunk
|
||||
```
|
||||
|
||||
### 兼容
|
||||
|
||||
```text
|
||||
[ ] mode=dense 仍可用(回滚开关)
|
||||
[ ] Agent 契约字段不破(evidence/document_id/excerpt)
|
||||
[ ] tool_invocation / trace 仍有 selectedAttempt 与基本检索信息
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 13. 明确不做的事
|
||||
|
||||
1. **继续在旧 SDK `search` 上叠 hybrid 细节当长期方案**
|
||||
2. **应用层把 dense raw 分和 BM25 raw 分直接相加**
|
||||
3. **用 L0 contains 大额加分替代库内 RRF**
|
||||
4. **只改查询、不重灌 sparse 数据**
|
||||
5. **hybrid 后仍拿旧 L2 阈值硬套 fused score**
|
||||
6. **让 Agent 直接依赖内部 fused/raw score 字段**(除非契约明确升级)
|
||||
|
||||
---
|
||||
|
||||
## 14. 和前序讨论的对齐
|
||||
|
||||
| 讨论结论 | 在本清单中的落点 |
|
||||
|---|---|
|
||||
| K=3 不必先上复杂 rerank | `retrieve-k=20, return-n=5` |
|
||||
| filtered + unfiltered 有价值 | hybrid 之上的 filter 策略双路径 |
|
||||
| BM25 是跨维度召回 | schema sparse/BM25 + hybrid 路 |
|
||||
| 跨路优先 RRF | 库内 `RRFRanker` |
|
||||
| `w_i` 是路权 | `WeightedRanker` / 配置 weight.dense/sparse |
|
||||
| L0 只做导航 | 仅影响 filter 与轻精排,不负责主融合 |
|
||||
| 旧 SDK 后续废弃 | 新开发只走 search/write port,不绑 SDK |
|
||||
|
||||
---
|
||||
|
||||
## 15. 最小可交付定义(MVP)
|
||||
|
||||
MVP 拆成两个可独立验收的里程碑(与 §1.1 一致)。
|
||||
|
||||
### MVP-1:分块去重与身份契约(交付 1)
|
||||
|
||||
```text
|
||||
1. 候选/证据具备 docId、chunkIndex、evidenceKey
|
||||
2. 后处理与投影均按 chunk 身份去重
|
||||
3. maxChunksPerDocument 生效
|
||||
4. retrieve-k / return-n 分离且可测
|
||||
5. 同文档多相关 chunk 在未超预算时可同时到达 Agent
|
||||
6. 不新增对旧 SDK 的依赖
|
||||
```
|
||||
|
||||
### MVP-2:Hybrid 接入(交付 2)
|
||||
|
||||
```text
|
||||
1. 新 collection 可写入 dense + BM25/sparse
|
||||
2. 知识库可全量重建
|
||||
3. lookup_knowledge 可通过配置切换 dense/hybrid
|
||||
4. hybrid 默认 RRF 融合
|
||||
5. 应用层 chunk 去重与 return-n 在 hybrid 下仍正确
|
||||
6. 旧 SDK 检索不再被新逻辑依赖
|
||||
7. 至少一套回归 query 证明:
|
||||
- 术语类 query 不回退
|
||||
- 语义类 query 不回退
|
||||
- filter 误杀有兜底
|
||||
```
|
||||
|
||||
只有 MVP-1 完成,才建议开始 MVP-2 的数据迁移与开关切换。
|
||||
|
||||
---
|
||||
|
||||
## 16. 建议的下一步实现顺序(动手时)
|
||||
|
||||
实现讨论与排期默认按此顺序:
|
||||
|
||||
```text
|
||||
1. 交付 1 / MVP-1
|
||||
- chunk 身份
|
||||
- chunk 去重 + maxChunksPerDocument
|
||||
- retrieve-k / return-n
|
||||
- Projector 对齐
|
||||
- SearchPort 雏形(建议)
|
||||
|
||||
2. 交付 2 / MVP-2
|
||||
- schema + 重灌
|
||||
- hybrid RRF 读路径
|
||||
- filter 兜底
|
||||
- 阈值校准
|
||||
- 下线旧 SDK 读路径
|
||||
```
|
||||
|
||||
**不要跳过交付 1 直接做 hybrid。**
|
||||
交付 1 不依赖旧 SDK,不阻塞后续 hybrid,且单独合并就有质量收益。
|
||||
|
||||
---
|
||||
|
||||
## 17. 后续实现讨论检查清单
|
||||
|
||||
开会或开 PR 前,用下面问题对齐范围:
|
||||
|
||||
```text
|
||||
[ ] 本次是交付 1、交付 2,还是仅其中子项?
|
||||
[ ] 是否改动了证据身份字段(docId/chunkIndex/evidenceKey)?
|
||||
[ ] 去重 key 是否仍存在 source 级路径?
|
||||
[ ] retrieve-k 与 return-n 是否仍混用 top-k?
|
||||
[ ] 是否把 schema 重灌/hybrid 开关误塞进交付 1?
|
||||
[ ] 是否有新增旧 SDK 依赖?
|
||||
[ ] 单测是否覆盖“同文档多 chunk”?
|
||||
[ ] 若已 hybrid:fused score 是否被误当成旧 L2 阈值?
|
||||
```
|
||||
@@ -0,0 +1,782 @@
|
||||
# 诊断 Agent 场景下的 RAG 排序:从 K=3 规则加分,到多路召回与 RRF
|
||||
|
||||
**日期**:2026-07-27
|
||||
**范围**:知识检索排序、多路召回、分数融合、Rerank 选型
|
||||
**读者**:需要在 Agent 系统里落地 RAG,而不是只做 Demo 问答的工程同学
|
||||
**关联实现**:`Lookup_knowledge` 模块化链路(L0 hint → L1 向量 → 规则后处理 → 投影)
|
||||
|
||||
---
|
||||
|
||||
## 1. 引言:RAG 排序常被高估,也常被低估
|
||||
|
||||
在诊断 Agent 里,知识库检索很少是“搜一下、答一句”那么简单。一次 `lookup_knowledge` 往往要在有限工具预算里,尽快给出可引用的证据片段。这时排序质量会直接影响:
|
||||
|
||||
- Agent 是否看得到正确 runbook / 排查步骤
|
||||
- 是否被错误 domain 的文档带偏
|
||||
- 是否在本就很少的 topK 里,把唯一有用的 chunk 挤掉
|
||||
|
||||
讨论排序时,团队里很容易出现两种极端:
|
||||
|
||||
1. **高估 Rerank**
|
||||
一提质量问题就上 Cross-Encoder、商业 Rerank、LLM listwise。
|
||||
但候选池只有 3 条时,精排只能在这 3 条里换座位,召不回的内容永远排不上来。
|
||||
|
||||
2. **低估融合**
|
||||
觉得“都是相似度,加一加就行”。
|
||||
但向量 L2、BM25、关键词加分根本不在同一尺度上,硬加权会把系统调成玄学。
|
||||
|
||||
本文基于一次针对真实诊断 Agent RAG 链路的讨论,整理一套可落地的判断框架:
|
||||
|
||||
```text
|
||||
1. 候选太少时,复杂 rerank 价值有限
|
||||
2. 多路召回解决“漏”,精排解决“噪”
|
||||
3. 跨路不要硬加原始分,优先 RRF 用名次投票
|
||||
4. L0 / 关键词适合做提示,不适合当最终裁判
|
||||
5. 权重 w_i 是“路权”,不是文档原始分
|
||||
```
|
||||
|
||||
目标不是证明某一种模型永远最优,而是回答工程上更常见的问题:
|
||||
|
||||
> 现在的 K、现有的 L0/L1、现有的规则加分,下一步到底该扩召回、该融合,还是该上精排?
|
||||
|
||||
---
|
||||
|
||||
## 2. 现状解剖:有“重排”,不等于有“强 Rerank”
|
||||
|
||||
### 2.1 一条典型主链路
|
||||
|
||||
以模块化 `lookup_knowledge` 为例,主路径大致是:
|
||||
|
||||
```text
|
||||
Agent query
|
||||
-> KnowledgeQueryTransformer # L0:domain / keyword hint,可选 category filter
|
||||
-> KnowledgeDocumentRetriever # L1:向量 topK
|
||||
-> KnowledgeEvidencePostProcessor # 归一化 + 规则 boost + 排序 + 去重
|
||||
-> KnowledgeContextPacker # 字符预算打包
|
||||
-> LookupResultAssembler
|
||||
-> RagResultProjector # 投影成 Agent 可见契约
|
||||
```
|
||||
|
||||
其中“重排”发生在后处理阶段,名字也常叫 rerank,但实现通常是:
|
||||
|
||||
```text
|
||||
baseScore = 向量距离归一化后的相似度
|
||||
finalScore = baseScore
|
||||
+ domain_match
|
||||
+ entity_match
|
||||
+ keyword_match
|
||||
+ source_type_prior
|
||||
再按 finalScore 降序
|
||||
```
|
||||
|
||||
同时会留下 `rerankTrace`(base/final score、boost reasons),便于内部审计。
|
||||
|
||||
### 2.2 这套做法解决了什么
|
||||
|
||||
它不是毫无意义:
|
||||
|
||||
- 在小候选池里,能把更像目标域、更像 runbook 的结果往前推
|
||||
- 有可解释的 boost 原因,方便 trace
|
||||
- 实现成本低,不引入额外模型服务
|
||||
|
||||
如果只是 Demo,或语料很小、query 很规范,这种轻规则重排往往够用。
|
||||
|
||||
### 2.3 它没解决什么
|
||||
|
||||
真正的问题通常不在“3 条里谁排第一”,而在:
|
||||
|
||||
1. **K 太小**
|
||||
`topK=3` 时,重排空间极小。复杂精排模型也只能在这 3 条里微调。
|
||||
|
||||
2. **规则加分不稳定**
|
||||
依赖 L0 词表和字符串 `contains`。短词、泛词、别名缺失都会让 boost 误触发或漏触发。
|
||||
|
||||
3. **L0 filter 可能误杀**
|
||||
唯一 domain 时加 category filter,能降噪;一旦 L0 判错域,正确文档可能根本进不了候选。
|
||||
|
||||
4. **去重粒度若偏文档级**
|
||||
同文档多个相关 chunk 可能被压成 1 条,召回了也会在后处理/投影阶段丢掉。
|
||||
|
||||
5. **Agent 侧看不到排序细节**
|
||||
投影层常只保留 excerpt 与少量元数据,score / hitReasons / retrievalTrace 被裁掉。这对安全边界合理,但对模型“判断有多相关”不友好。
|
||||
|
||||
一句话概括现状:
|
||||
|
||||
> 不是没有排序,而是在太小的候选池里,用不够稳的启发式做微调。
|
||||
|
||||
---
|
||||
|
||||
## 3. 关键判断:候选池大小决定策略上限
|
||||
|
||||
排序策略必须和 K 匹配。可以先用下面这张表做决策:
|
||||
|
||||
| 召回规模 K | 更适合做什么 | 不太值得先做什么 |
|
||||
|---|---|---|
|
||||
| 3 ~ 5 | 修去重、轻规则、双路径 dense 融合 | Cross-Encoder / 商业 Rerank |
|
||||
| 15 ~ 30 | 多路召回 + RRF + 轻精排 | 只继续调 keyword boost |
|
||||
| 50+ | 强 rerank 才划算 | 无评测地堆模型 |
|
||||
|
||||
背后的原因很简单:
|
||||
|
||||
```text
|
||||
Rerank 的价值来自:
|
||||
候选很多、噪声大 → 精排把好的顶到前面
|
||||
|
||||
如果好文档根本不在候选里:
|
||||
再贵的 rerank 也救不回来
|
||||
```
|
||||
|
||||
因此工程上应把两个参数拆开:
|
||||
|
||||
```text
|
||||
retrieve-k # 召回阶段多捞一些,例如 20
|
||||
return-n # 最终给 Agent 的条数,例如 3~5
|
||||
```
|
||||
|
||||
而不是始终:
|
||||
|
||||
```text
|
||||
topK = 3,召回、排序、返回都是 3
|
||||
```
|
||||
|
||||
**先让该进来的进来,再谈谁排前面。**
|
||||
|
||||
---
|
||||
|
||||
## 4. 多路召回:先分清“同模态双路径”和“跨维度多路”
|
||||
|
||||
“多路召回”不是只有 BM25 + 向量这一种。可以分层理解。
|
||||
|
||||
### 4.1 同模态双路径:filtered + unfiltered
|
||||
|
||||
很多系统已经有类似逻辑:
|
||||
|
||||
```text
|
||||
若 L0 给出唯一 category:
|
||||
先 filtered 向量检索
|
||||
若结果差,再 unfiltered 重试
|
||||
```
|
||||
|
||||
这能工作,但常见实现是 **串行整锅替换**:
|
||||
|
||||
- retry 成功后,直接丢掉第一次 filtered 的全部结果
|
||||
- 没有“两边好结果都保留,再统一排序”
|
||||
|
||||
更稳的做法是把它升级为并行/双路融合:
|
||||
|
||||
```text
|
||||
路 A:dense + category filter # 求准
|
||||
路 B:dense + 无 filter # 求全
|
||||
去重合并后一起排序
|
||||
```
|
||||
|
||||
#### 为什么值得做?
|
||||
|
||||
因为两路解决的是不同失败模式:
|
||||
|
||||
| 路径 | 优点 | 风险 |
|
||||
|---|---|---|
|
||||
| filtered | 更贴当前域,噪声少 | filter 错了会漏召回 |
|
||||
| unfiltered | 召回面宽,容错强 | 容易掺进其他域文档 |
|
||||
|
||||
典型场景:
|
||||
|
||||
- L0 误判成 `mysql`,真正文档在 `redis`
|
||||
→ 只走 filtered 会空或很差
|
||||
→ unfiltered 能救回
|
||||
- L0 判对了 `mysql`
|
||||
→ filtered 更干净
|
||||
→ unfiltered 可能带噪声
|
||||
|
||||
所以:
|
||||
|
||||
> filtered 求准,unfiltered 求全;融合比二选一更稳。
|
||||
|
||||
#### 它算多路吗?
|
||||
|
||||
算,但要说清楚边界:
|
||||
|
||||
```text
|
||||
这是同一 dense 检索器、不同过滤条件的双路径
|
||||
属于多路召回的子集
|
||||
还不是完整的跨维度多路
|
||||
```
|
||||
|
||||
它的主要价值是:
|
||||
|
||||
1. 降低 L0 category 误杀
|
||||
2. 保留“有 filter 时更准”的收益
|
||||
3. 避免 retry 整锅替换导致好结果被误删
|
||||
|
||||
即便暂时不上 BM25,只做这一步,也常常比继续调 keyword boost 更有效。
|
||||
|
||||
### 4.2 跨维度多路:Dense + BM25
|
||||
|
||||
更完整的多路,通常来自不同相关性维度:
|
||||
|
||||
| 路 | 擅长 | 不擅长 |
|
||||
|---|---|---|
|
||||
| Dense(向量) | 语义相近、换说法、同义表达 | 罕见专有名词可能漂 |
|
||||
| BM25 / 关键词 | 错误码、类名、配置键、告警名、精确术语 | 换一种说法就容易漏 |
|
||||
|
||||
例子:
|
||||
|
||||
```text
|
||||
用户说:连接池打满了
|
||||
文档写:HikariCP pending threads high
|
||||
→ dense 更容易搭上
|
||||
|
||||
用户说:SQLSTATE 08001
|
||||
文档标题就含 08001
|
||||
→ BM25 / 字面匹配往往更稳
|
||||
```
|
||||
|
||||
因此:
|
||||
|
||||
```text
|
||||
candidates = dense ∪ bm25
|
||||
```
|
||||
|
||||
不是“BM25 替代向量”,而是“补向量漏掉的字面命中”。
|
||||
|
||||
### 4.3 L0 能不能当一路召回?
|
||||
|
||||
可以,但要降权、控边界。
|
||||
|
||||
L0(文档 frontmatter 关键词 / domain hint)适合:
|
||||
|
||||
- 决定要不要启用 filtered 路
|
||||
- 提供精排时的弱特征
|
||||
- 写出可解释 trace
|
||||
|
||||
不适合:
|
||||
|
||||
- 关键词命中就直接当高置信事实证据
|
||||
- 用 L0 粗分主导最终排序
|
||||
|
||||
一句话:
|
||||
|
||||
> L0 是导航,不是裁判长。
|
||||
|
||||
---
|
||||
|
||||
## 5. 融合为什么难:不是不会加权,是分数不可比
|
||||
|
||||
多路召回之后,第一反应常常是:
|
||||
|
||||
```text
|
||||
final = 0.7 * dense_score + 0.3 * bm25_score
|
||||
```
|
||||
|
||||
这在课堂上好讲,在工程上很脆。
|
||||
|
||||
### 5.1 原始分为什么不能直接加
|
||||
|
||||
不同路的分数:
|
||||
|
||||
- Dense:L2 距离或 cosine similarity,分布随 embedding 模型和语料变化
|
||||
- BM25:另一套量纲,数值大小和 dense 完全不可比
|
||||
- 规则 boost:+0.15 / +0.20 这种启发式加分,更像人工偏好,不是校准概率
|
||||
|
||||
把它们直接线性相加,等于默认“0.1 的向量分提升”和“2.0 的 BM25 提升”可以交换。这个默认通常不成立。
|
||||
|
||||
### 5.2 规则 keyword boost 为什么不稳定
|
||||
|
||||
若最终分主要靠:
|
||||
|
||||
```text
|
||||
query.contains(keyword) || keyword.contains(query)
|
||||
```
|
||||
|
||||
再叠加固定加分,会出现:
|
||||
|
||||
- 短词误命中
|
||||
- 泛词普遍加分,区分度下降
|
||||
- 词表一改,线上排序整体漂移
|
||||
- 同义词没写进 frontmatter 就完全没帮助
|
||||
|
||||
所以:
|
||||
|
||||
> 现在的“权重”如果本质是关键词启发式加分,稳定性天然有限。
|
||||
|
||||
这不代表规则无用,而是应把它放对层:路内微调或精排特征,而不是跨路主融合器。
|
||||
|
||||
---
|
||||
|
||||
## 6. RRF:跨路融合时,优先用名次投票
|
||||
|
||||
### 6.1 核心思想
|
||||
|
||||
RRF(Reciprocal Rank Fusion)的关键思想是:
|
||||
|
||||
> 不管各路原始分是什么尺度,只看每条候选在各路里的名次,再投票。
|
||||
|
||||
基础公式:
|
||||
|
||||
```text
|
||||
RRF(d) = Σ 1 / (k + rank_i(d))
|
||||
```
|
||||
|
||||
其中:
|
||||
|
||||
| 符号 | 含义 |
|
||||
|---|---|
|
||||
| `d` | 某个候选文档或 chunk |
|
||||
| `i` | 第 i 路召回 |
|
||||
| `rank_i(d)` | d 在第 i 路中的名次(从 1 开始);未出现则该路贡献为 0 |
|
||||
| `k` | 常数,常用 60,用来缓和头部名次过强 |
|
||||
|
||||
直觉:
|
||||
|
||||
- 某路第 1 名:贡献约 `1/61`
|
||||
- 某路第 2 名:贡献约 `1/62`
|
||||
- 多路都靠前的候选,融合分自然更高
|
||||
- 只在一路偶然靠前的候选,不会单靠绝对分尺度“爆掉”
|
||||
|
||||
### 6.2 为什么适合 RAG 多路融合
|
||||
|
||||
RRF 特别适合下面这种现实约束:
|
||||
|
||||
```text
|
||||
dense 用 L2
|
||||
bm25 用 BM25
|
||||
filtered / unfiltered 虽同度量,但候选集合不同
|
||||
暂时没有可靠的分数校准器
|
||||
```
|
||||
|
||||
它把问题从:
|
||||
|
||||
```text
|
||||
如何把不可比的分数对齐?
|
||||
```
|
||||
|
||||
简化成:
|
||||
|
||||
```text
|
||||
各路是否都认为它靠前?
|
||||
```
|
||||
|
||||
### 6.3 加权 RRF:w_i 是路权,不是文档分
|
||||
|
||||
加权形式:
|
||||
|
||||
```text
|
||||
RRF_w(d) = Σ w_i / (k + rank_i(d))
|
||||
```
|
||||
|
||||
这里的 `w_i` 非常容易被误解。
|
||||
|
||||
#### 正确理解
|
||||
|
||||
```text
|
||||
w_i = 第 i 整路的话语权
|
||||
```
|
||||
|
||||
例如:
|
||||
|
||||
```text
|
||||
w_dense = 1.0
|
||||
w_bm25 = 1.5 # 想让 BM25 更重要,就提高这一路的 w
|
||||
w_filtered_dense = 1.0
|
||||
```
|
||||
|
||||
含义是:
|
||||
|
||||
- BM25 这一路投出的“名次票”更值钱
|
||||
- 并不是把某个文档的 BM25 原始分 12.7 直接拿来和向量分相加
|
||||
|
||||
#### 错误理解
|
||||
|
||||
```text
|
||||
先算 keyword boost 得到一个大杂烩分
|
||||
再把这个分塞进 RRF
|
||||
```
|
||||
|
||||
或:
|
||||
|
||||
```text
|
||||
RRF = f(L2原始分, BM25原始分, keyword加分)
|
||||
```
|
||||
|
||||
这都不是加权 RRF。
|
||||
|
||||
### 6.4 分层心智模型
|
||||
|
||||
建议始终按三层理解分数角色:
|
||||
|
||||
```text
|
||||
第 1 层:路内排序
|
||||
dense 路用向量分排序
|
||||
bm25 路用 BM25 分排序
|
||||
各用各的分,互不直接相加
|
||||
|
||||
第 2 层:跨路融合
|
||||
只吃各路 rank
|
||||
普通 RRF 或加权 RRF(w_i 为路权)
|
||||
|
||||
第 3 层:精排(可选)
|
||||
对融合后的 topM 再打分
|
||||
这里才适合规则特征或 Cross-Encoder
|
||||
```
|
||||
|
||||
一句话记住:
|
||||
|
||||
> 原始分只负责路内排名;RRF 只负责跨路投票;w_i 只调节哪一路更值钱;精排才做最终挑剔。
|
||||
|
||||
### 6.5 手算例子
|
||||
|
||||
设:
|
||||
|
||||
```text
|
||||
k = 60
|
||||
w_dense = 1.0
|
||||
w_bm25 = 1.5
|
||||
```
|
||||
|
||||
候选 A:
|
||||
|
||||
- dense 第 1 名
|
||||
- bm25 第 5 名
|
||||
|
||||
```text
|
||||
RRF(A)
|
||||
= 1.0/(60+1) + 1.5/(60+5)
|
||||
= 1/61 + 1.5/65
|
||||
≈ 0.01639 + 0.02308
|
||||
≈ 0.03947
|
||||
```
|
||||
|
||||
候选 B:
|
||||
|
||||
- dense 第 4 名
|
||||
- bm25 第 1 名
|
||||
|
||||
```text
|
||||
RRF(B)
|
||||
= 1.0/(60+4) + 1.5/(60+1)
|
||||
= 1/64 + 1.5/61
|
||||
≈ 0.01563 + 0.02459
|
||||
≈ 0.04022
|
||||
```
|
||||
|
||||
在这个设定下 B 更高,因为你主动提高了 BM25 路权,BM25 头名更吃香。
|
||||
|
||||
如果把 `w_bm25` 降到 `0.5`,同样名次下 dense 会重新占主导。
|
||||
这就是路权的意义:调的是整路话语权,不是某个文档的原始分公式。
|
||||
|
||||
### 6.6 w_i 怎么设
|
||||
|
||||
实操建议:
|
||||
|
||||
1. **起步全设 1.0**
|
||||
先看纯 RRF,不要一上来调花活。
|
||||
|
||||
2. **再按评测微调**
|
||||
- 专有名词/报错码总靠 BM25 才找得到,却总被 dense 压下去 → 提高 `w_bm25`(如 1.2~1.5)
|
||||
- BM25 噪声大、泛词乱入 → 降低 `w_bm25`(如 0.6~0.8)
|
||||
- filtered 很准但有时过窄 → 可与 unfiltered 同权,或略高一点
|
||||
|
||||
3. **经验范围**
|
||||
`w_i` 常见落在 `0.5 ~ 2.0`。
|
||||
若一路是 10、另一路是 0.1,基本等于放弃多路。
|
||||
|
||||
4. **没有回归集就不要谈“最优权重”**
|
||||
路权应来自离线评测,而不是长期拍脑袋。
|
||||
|
||||
---
|
||||
|
||||
## 7. 精排放在哪里:有了候选池,才配谈 Rerank
|
||||
|
||||
### 7.1 轻规则精排
|
||||
|
||||
在 RRF 融合出 top 10~15 后,可以用更稳的字段级规则做二次排序:
|
||||
|
||||
```text
|
||||
title 命中 > breadcrumb 命中 > body 覆盖
|
||||
精确术语 > 泛词
|
||||
runbook / case 来源轻微加分
|
||||
与已选证据过相似则降权(MMR 思路)
|
||||
```
|
||||
|
||||
注意:这里的规则特征,最好作用于 **融合后的候选精排**,而不是重新发明一套跨路原始分加法。
|
||||
|
||||
### 7.2 模型精排
|
||||
|
||||
当 `retrieve-k` 到 15~30,且评测证明融合后噪声仍高时,再考虑:
|
||||
|
||||
```text
|
||||
RRF topM
|
||||
-> Cross-Encoder / 托管 Rerank API
|
||||
-> 取 topN 给 Agent
|
||||
```
|
||||
|
||||
可选路线:
|
||||
|
||||
- 开源 reranker(如 bge-reranker 一类)
|
||||
- 云厂商 / Cohere 等托管 rerank
|
||||
- LLM listwise(贵且不稳,一般不当首选)
|
||||
|
||||
### 7.3 什么时候不要上模型 rerank
|
||||
|
||||
- 仍在 `K=3` 主路径上
|
||||
- 还没修 chunk 级去重
|
||||
- 还没有固定 query 回归集
|
||||
- 延迟和工具预算已经很紧
|
||||
|
||||
否则花的是精排成本,换不到召回质量。
|
||||
|
||||
---
|
||||
|
||||
## 8. 工程落地:比“换模型”更重要的顺序
|
||||
|
||||
### 8.1 建议的目标形态
|
||||
|
||||
```text
|
||||
Query
|
||||
├─ dense unfiltered top 20
|
||||
├─ dense filtered top 10 # 有唯一 domain 时
|
||||
└─ bm25 / keyword top 10 # 第二阶段
|
||||
│
|
||||
▼
|
||||
chunk 级去重(docId#chunkIndex)
|
||||
│
|
||||
▼
|
||||
RRF / 加权 RRF 融合
|
||||
│
|
||||
▼
|
||||
轻规则或模型精排,取 top 5
|
||||
│
|
||||
▼
|
||||
每文档 chunk 上限 + context pack
|
||||
│
|
||||
▼
|
||||
Agent projection
|
||||
```
|
||||
|
||||
### 8.2 分阶段推进
|
||||
|
||||
#### Phase 0:先修前提
|
||||
|
||||
否则后面多路都会被吞:
|
||||
|
||||
1. 去重 key 从“文档/source”改为 `docId + chunkIndex`(fallback 可用向量主键)
|
||||
2. 配置拆分:
|
||||
|
||||
```properties
|
||||
rag.retrieve-k=20
|
||||
rag.return-n=5
|
||||
rag.max-chunks-per-document=2
|
||||
```
|
||||
|
||||
3. 明确 baseScore(可用性阈值)与融合分/精排分(排序)职责分离
|
||||
|
||||
#### Phase 1:同模态双路径 + RRF
|
||||
|
||||
1. filtered dense 与 unfiltered dense 都产出候选
|
||||
2. 合并去重,不再整锅替换
|
||||
3. RRF 融合后取 topN
|
||||
4. trace 记录每路 rank 与 fused rank
|
||||
|
||||
这是最贴很多现有系统的一步,收益通常大于继续调 boost。
|
||||
|
||||
#### Phase 2:跨维度多路
|
||||
|
||||
1. 增加 BM25 / 关键词路
|
||||
2. 继续 RRF,必要时给 `w_bm25` 微调
|
||||
3. 增加字段级轻精排
|
||||
|
||||
#### Phase 3:可插拔模型 Rerank
|
||||
|
||||
```text
|
||||
interface Reranker {
|
||||
rerank(query, candidates, topN) -> rankedCandidates
|
||||
}
|
||||
```
|
||||
|
||||
实现可切换:
|
||||
|
||||
- `NoopReranker`
|
||||
- `RuleReranker`
|
||||
- `HttpCrossEncoderReranker`
|
||||
|
||||
让精排成为插件,而不是写死在业务里。
|
||||
|
||||
### 8.3 和 Agent 系统相关的额外约束
|
||||
|
||||
诊断 Agent 场景还有几个现实约束:
|
||||
|
||||
1. **工具预算有限**
|
||||
检索本身只是工具循环的一环,延迟不能无限涨。
|
||||
|
||||
2. **证据要可引用**
|
||||
最终给 Agent 的应是有界 excerpt,而不是内部全量 trace。
|
||||
|
||||
3. **投影会再裁一层**
|
||||
即便内部排序很细,Agent 可见字段仍可能只有 document/source/title/breadcrumb/excerpt。
|
||||
因此内部要保留完整 trace,外部保持契约稳定。
|
||||
|
||||
4. **安全发布与验真**
|
||||
排序再好,也不能绕过证据引用与 guard;RAG 优化的是“更可能拿到对的证据”,不是“让模型自由发挥”。
|
||||
|
||||
---
|
||||
|
||||
## 9. 评测:没有回归集,权重都是感觉
|
||||
|
||||
多路和 RRF 最怕“上线凭体感”。最少准备 15~30 条固定 query,覆盖:
|
||||
|
||||
- 标准故障词
|
||||
- 口语化换说法
|
||||
- 专有名词 / 错误码
|
||||
- 容易误判 domain 的问题
|
||||
- 同文档多 chunk 才完整的流程题
|
||||
- 负例:知识库本就没有答案
|
||||
|
||||
关注指标:
|
||||
|
||||
| 指标 | 看什么 |
|
||||
|---|---|
|
||||
| Recall@5 | 该出现的文档/chunk 是否进前 5 |
|
||||
| nDCG@5 或人工 0/1/2 | 排序是否把更相关的放前面 |
|
||||
| filter 误杀率 | 唯一 domain 是否经常害人 |
|
||||
| 同文档多 chunk 保留率 | 去重是否过粗 |
|
||||
| P95 延迟 | 多路是否打爆预算 |
|
||||
| 无证据正确率 | 不该有答案时是否老实说没有 |
|
||||
|
||||
路权 `w_i` 的调整,应建立在这些数上,而不是单次手工 query。
|
||||
|
||||
---
|
||||
|
||||
## 10. 反模式清单
|
||||
|
||||
下面这些做法看起来勤快,实际常把系统带偏:
|
||||
|
||||
1. **只在 K=3 上接昂贵 rerank**
|
||||
候选池不够,精排没有舞台。
|
||||
|
||||
2. **L2 和 BM25 直接加权相加**
|
||||
分数不可比,调参不可迁移。
|
||||
|
||||
3. **把 L0 contains 当最终裁判**
|
||||
词表质量绑死线上排序。
|
||||
|
||||
4. **文档级去重吞掉同文档多 chunk**
|
||||
多路召回也会在终点被自己吃掉。
|
||||
|
||||
5. **filtered 失败就整锅替换**
|
||||
丢掉本可保留的好结果。
|
||||
|
||||
6. **无评测调 w_i**
|
||||
今天的“最优权重”往往是过拟合某几条样例。
|
||||
|
||||
7. **L0 命中全文直接当高置信 evidence**
|
||||
导航信号被抬成事实,诊断场景尤其危险。
|
||||
|
||||
---
|
||||
|
||||
## 11. 可直接拿走的决策框架
|
||||
|
||||
遇到 RAG 排序问题时,按这个顺序问:
|
||||
|
||||
```text
|
||||
Q1. 正确答案是否经常连候选池都进不来?
|
||||
是 → 先扩召回 / 多路,不要先上复杂 rerank
|
||||
|
||||
Q2. 是否存在 filter 误杀?
|
||||
是 → filtered + unfiltered 双路径融合
|
||||
|
||||
Q3. 是否大量依赖专有名词、错误码、配置键?
|
||||
是 → 加 BM25 / 关键词路
|
||||
|
||||
Q4. 多路分数是否不可比?
|
||||
是 → RRF,而不是原始分硬加
|
||||
|
||||
Q5. 融合后 topM 仍噪声大,且 K 已经够大?
|
||||
是 → 再上规则精排或模型 rerank
|
||||
|
||||
Q6. 有没有固定回归集?
|
||||
没有 → 先建评测,再谈“最优权重”
|
||||
```
|
||||
|
||||
对应到一句话策略:
|
||||
|
||||
> **先扩召回,再用名次融合,最后才模型精排。**
|
||||
|
||||
---
|
||||
|
||||
## 12. 结语
|
||||
|
||||
RAG 排序讨论很容易变成模型名词竞赛。但在诊断 Agent 这种真实系统里,更常见的瓶颈是:
|
||||
|
||||
- 候选太少
|
||||
- 过滤过猛
|
||||
- 分数不可比
|
||||
- 去重过粗
|
||||
- 启发式加分承担了不该承担的最终裁决
|
||||
|
||||
RRF 的价值,不只是一个公式,而是一种工程态度:
|
||||
|
||||
```text
|
||||
承认各路分数不可比
|
||||
让每路先做好自己的排序
|
||||
再用名次投票决定谁更值得进入下游
|
||||
```
|
||||
|
||||
加权 RRF 也并不神秘:`w_i` 只是给整路调音量。
|
||||
想让 BM25 更有话语权,就提高 `w_bm25`;它不会、也不应该要求你先把 BM25 分和向量分校准到同一宇宙。
|
||||
|
||||
如果只记住三句:
|
||||
|
||||
1. **K=3 时,复杂 rerank 不是第一优先级。**
|
||||
2. **filtered + unfiltered 是值得做的同模态双路径;Dense + BM25 才是跨维度多路。**
|
||||
3. **跨路融合优先 RRF;L0 做导航,精排做挑剔,原始分不要跨路硬加。**
|
||||
|
||||
按这个脉络演进,通常比“继续把 keyword boost 调大一点”更接近稳定、可解释、可评测的 RAG 排序系统。
|
||||
|
||||
---
|
||||
|
||||
## 附录 A:术语对照
|
||||
|
||||
| 术语 | 含义 |
|
||||
|---|---|
|
||||
| L0 | 基于文档元数据/关键词的 query understanding 或弱召回 |
|
||||
| L1 | 向量语义召回 |
|
||||
| retrieve-k | 召回阶段候选数 |
|
||||
| return-n | 最终返回给 Agent 的证据数 |
|
||||
| baseScore | 向量相似度等主相关性分,常用于可用性阈值 |
|
||||
| finalScore / 精排分 | 用于排序的综合分 |
|
||||
| RRF | 基于名次的多路融合 |
|
||||
| w_i | 第 i 路在 RRF 中的路权 |
|
||||
| Rerank | 对已有候选做精排,不负责凭空召回新文档 |
|
||||
|
||||
## 附录 B:最小配置示例(示意)
|
||||
|
||||
```properties
|
||||
# 召回与返回分离
|
||||
rag.retrieve-k=20
|
||||
rag.return-n=5
|
||||
rag.max-chunks-per-document=2
|
||||
|
||||
# RRF
|
||||
rag.fusion.method=rrf
|
||||
rag.fusion.rrf-k=60
|
||||
rag.fusion.w-dense=1.0
|
||||
rag.fusion.w-dense-filtered=1.0
|
||||
rag.fusion.w-bm25=0.8
|
||||
|
||||
# 精排(先规则,后可插模型)
|
||||
rag.rerank.mode=rule # rule | model | off
|
||||
```
|
||||
|
||||
以上配置名是示意,重点在职责拆分,不在具体键名。
|
||||
|
||||
## 附录 C:和本文讨论直接对应的实现关注点
|
||||
|
||||
阅读或改造现有代码时,可重点核对:
|
||||
|
||||
1. 后处理是否把“规则 boost”命名成了 rerank,却未做候选扩展
|
||||
2. filtered 低质时是整锅替换,还是双路融合
|
||||
3. 去重 key 是 source/文档级,还是 chunk 级
|
||||
4. `topK` 是否同时承担召回、排序、返回三种职责
|
||||
5. 内部 `rerankTrace` 是否可观测,Agent 投影是否有意裁剪
|
||||
|
||||
这些点决定了:你写在黑板上的 RRF,能不能在系统里真正跑起来。
|
||||
@@ -0,0 +1,2 @@
|
||||
ready: 2026-07-27
|
||||
change: rag-chunk-evidence-identity-dedup
|
||||
@@ -0,0 +1,4 @@
|
||||
committed: 2026-07-27
|
||||
change: rag-chunk-evidence-identity-dedup
|
||||
scale: standard
|
||||
authorized-apply: user-preauthorized-sm-flow
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-27
|
||||
@@ -0,0 +1,122 @@
|
||||
# Design: RAG chunk evidence identity and dedup
|
||||
|
||||
## Context
|
||||
|
||||
- Modular RAG pipeline already exists (`KnowledgeQueryTransformer` → retriever → post-processor → packer → assembler → projector).
|
||||
- Delivery baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1.
|
||||
- Historical decision (modular RAG): L0 is hint-only; L1 is fact evidence. This change keeps that boundary.
|
||||
- Legacy Milvus SDK search path will be abandoned later; this change must not thicken SDK-specific logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals**
|
||||
|
||||
- Preserve multiple relevant chunks from the same document in one lookup.
|
||||
- Make evidence identity stable enough for future hybrid hits.
|
||||
- Separate retrieve width from return width.
|
||||
- Keep Agent tool name/input unchanged (`lookup_knowledge(query)`).
|
||||
|
||||
**Non-Goals**
|
||||
|
||||
- Hybrid / BM25 / sparse schema.
|
||||
- Cross-encoder rerank.
|
||||
- Deleting SDK mode.
|
||||
- Session dedup tracker.
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1. Evidence identity key
|
||||
|
||||
```text
|
||||
evidenceKey =
|
||||
if docId != null && chunkIndex != null:
|
||||
docId + "#chunk-" + chunkIndex
|
||||
else if vectorId != null:
|
||||
"vector:" + vectorId
|
||||
else:
|
||||
"rank:" + originalRank
|
||||
```
|
||||
|
||||
Rationale: works with current metadata (`docId`, `chunkIndex`) and degrades safely for older rows.
|
||||
|
||||
### D2. Dedup granularity
|
||||
|
||||
- Dedup key = `evidenceKey` only.
|
||||
- True duplicates (same key) merge hitReasons; keep higher-ranked content (first after score sort).
|
||||
- Do **not** merge different chunks of the same source into one content block.
|
||||
|
||||
### D3. Per-document cap
|
||||
|
||||
- After score sort, accept at most `rag.max-chunks-per-document` (default `2`) evidence blocks per `docId`.
|
||||
- If `docId` missing, treat each evidenceKey as its own document bucket.
|
||||
|
||||
### D4. retrieve-k / return-n
|
||||
|
||||
```properties
|
||||
rag.retrieve-k=20 # vector recall width
|
||||
rag.return-n=5 # max evidence blocks after post-process (before projector budget)
|
||||
rag.max-chunks-per-document=2
|
||||
```
|
||||
|
||||
Compatibility:
|
||||
|
||||
- If only legacy `rag.top-k` is set, use it as fallback for both until removed.
|
||||
- Prefer explicit retrieve-k/return-n when present.
|
||||
|
||||
### D5. Agent projection identity (behavior change)
|
||||
|
||||
- Projected `document_id` SHOULD be `evidenceKey` (chunk-scoped), not raw source.
|
||||
- This is intentional so EvidenceGuard references remain 1:1 with returned excerpts.
|
||||
- `source` remains human-readable document path/id and MAY repeat across chunks.
|
||||
- Marked as **observable behavior change** for Agent consumers and report references.
|
||||
|
||||
### D6. KnowledgeSearchPort (thin)
|
||||
|
||||
```text
|
||||
KnowledgeSearchPort.search(KnowledgeSearchRequest) -> List<KnowledgeSearchHit>
|
||||
```
|
||||
|
||||
- Default adapter delegates to existing `VectorSearchService.searchSimilarDocuments`.
|
||||
- Request carries query, topK, categoryFilter, mode placeholder (`DENSE` only in this change).
|
||||
- Hit carries id, content, score fields, metadata map, and extracted identity fields when available.
|
||||
- No hybrid implementation in this change.
|
||||
|
||||
## Module flow
|
||||
|
||||
```text
|
||||
LookupKnowledge
|
||||
-> transform (L0 hints)
|
||||
-> KnowledgeSearchPort.search(retrieveK, filter)
|
||||
-> map to RetrievedEvidenceCandidate (+ identity)
|
||||
-> post-process:
|
||||
score/sort
|
||||
evidenceKey dedup
|
||||
maxChunksPerDocument
|
||||
return-n truncate
|
||||
-> pack + assemble
|
||||
-> RagResultProjector (dedupe by evidence document_id=evidenceKey)
|
||||
```
|
||||
|
||||
## Interface impact
|
||||
|
||||
| Level | What |
|
||||
|---|---|
|
||||
| L2 | Internal DTO fields on candidate/EvidenceBlock |
|
||||
| L3 | Agent-facing `document_id` becomes chunk-scoped evidence id |
|
||||
|
||||
Migration/compat:
|
||||
|
||||
- In-repo guards/tests updated to accept chunk-scoped ids.
|
||||
- External human readers still see `source`/`title`.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Larger Agent payload | return-n + maxChunksPerDocument + existing projector budgets |
|
||||
| Missing chunkIndex in old data | vector id fallback keeps chunks distinct |
|
||||
| document_id semantic shift | documented; projector tests updated |
|
||||
|
||||
## Open questions
|
||||
|
||||
None remaining for Delivery 1. Hybrid belongs to next change.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Change: RAG chunk evidence identity and dedup
|
||||
|
||||
## Why
|
||||
|
||||
`lookup_knowledge` currently collapses same-document chunks at two layers:
|
||||
|
||||
1. `KnowledgeEvidencePostProcessor` dedupes by `source/title/breadcrumb`
|
||||
2. `RagResultProjector` dedupes by `document_id`, which usually falls back to `source`
|
||||
|
||||
As a result, L1 can recall multiple useful chunks from one document, but Agent often sees only one. This blocks multi-path / hybrid retrieval benefits and hurts long runbook completeness.
|
||||
|
||||
This change is **Delivery 1** from `docs/milvus-hybrid-search-integration-checklist.md`. Hybrid schema/search is **out of scope** and will be a separate change after this is archived.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add chunk-level evidence identity: `docId`, `chunkIndex`, `evidenceKey`
|
||||
- Extract identity in retrieval adapter from vector metadata
|
||||
- Deduplicate by `evidenceKey` (chunk identity), not document source
|
||||
- Cap chunks per document (`maxChunksPerDocument`, default 2)
|
||||
- Split `retrieve-k` and `return-n` (stop overloading single `top-k`)
|
||||
- Align Agent projection so same-source different chunks can both appear
|
||||
- Introduce a thin `KnowledgeSearchPort` so later hybrid can swap implementation without rewriting the pipeline
|
||||
- Keep L0 as hint-only; no BM25/sparse schema; no legacy SDK deletion in this change
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Milvus sparse/BM25 schema or reindex
|
||||
- Enabling hybrid search mode
|
||||
- Removing Milvus SDK path
|
||||
- Session-level RetrievedDocTracker restore
|
||||
- Neighbor chunk context reconstruction
|
||||
- Model reranker
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `rag-chunk-evidence-identity`: chunk-level identity, dedup, retrieve/return split, search port boundary
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `rag-knowledge-retrieval`: replace document-level evidence dedup requirement with chunk-level identity
|
||||
- `rag-log-projections` (RAG portion only): projection identity may be chunk-scoped `document_id`
|
||||
|
||||
## Impact
|
||||
|
||||
- Code: retrieval DTO/services, post-processor, lookup tool config, projector, tests
|
||||
- Agent-visible: more evidence items possible for same logical document when multiple chunks are relevant
|
||||
- Interface level: **L2/L3** — Agent `document_id` semantics become chunk-scoped evidence id (often `docId#chunk-N`); EvidenceGuard still validates against tool projection ids
|
||||
- Docs baseline: `docs/milvus-hybrid-search-integration-checklist.md` §1.1 Delivery 1
|
||||
|
||||
## Risks
|
||||
|
||||
- Agent context grows if many chunks pass; mitigated by `return-n` and `maxChunksPerDocument`
|
||||
- Existing tests assume source-level dedup; must update intentionally
|
||||
- Old indexes without `chunkIndex` need stable fallback keys (`vector:{id}`)
|
||||
+82
@@ -0,0 +1,82 @@
|
||||
# rag-chunk-evidence-identity Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Define chunk-level evidence identity, deduplication, retrieve/return separation, and the thin search port boundary used by `lookup_knowledge` before hybrid retrieval is introduced.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Retrieval candidates SHALL carry chunk-level identity
|
||||
|
||||
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
|
||||
|
||||
#### Scenario: Identity extracted from metadata
|
||||
|
||||
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
|
||||
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
|
||||
|
||||
#### Scenario: Fallback identity without chunk index
|
||||
|
||||
- **WHEN** chunk index is missing but vector id is present
|
||||
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
|
||||
|
||||
### Requirement: Evidence deduplication SHALL be chunk-scoped
|
||||
|
||||
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
|
||||
|
||||
#### Scenario: Same document different chunks are kept
|
||||
|
||||
- **WHEN** two candidates share the same source/docId but different chunk indexes
|
||||
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
|
||||
|
||||
#### Scenario: True duplicate keys merge without replacing higher-ranked content
|
||||
|
||||
- **WHEN** two candidates share the same `evidenceKey`
|
||||
- **THEN** post-processing SHALL keep a single block
|
||||
- **AND** SHALL preserve the higher-ranked content
|
||||
- **AND** MAY merge hit reasons
|
||||
|
||||
### Requirement: Post-processing SHALL enforce max chunks per document
|
||||
|
||||
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
|
||||
|
||||
#### Scenario: Excess chunks from one document are dropped
|
||||
|
||||
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
|
||||
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
|
||||
|
||||
### Requirement: Retrieval width and return width SHALL be separate
|
||||
|
||||
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
|
||||
|
||||
#### Scenario: retrieve-k widens recall without unbounded return
|
||||
|
||||
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
|
||||
- **THEN** the search port is asked for up to 20 candidates
|
||||
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
|
||||
|
||||
### Requirement: Agent projection SHALL preserve distinct chunk evidence
|
||||
|
||||
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
|
||||
|
||||
#### Scenario: Same source different evidence keys both project
|
||||
|
||||
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
|
||||
- **AND** projection budgets still allow both
|
||||
- **THEN** both SHALL appear in the projected evidence list
|
||||
|
||||
#### Scenario: Projected document_id is chunk-scoped
|
||||
|
||||
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
|
||||
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
|
||||
- **AND** `source` MAY still be the document path or doc id
|
||||
|
||||
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
|
||||
|
||||
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
|
||||
|
||||
#### Scenario: Default dense adapter
|
||||
|
||||
- **WHEN** search mode is dense-only (this change)
|
||||
- **THEN** the port implementation MAY delegate to the existing vector search facade
|
||||
- **AND** the retriever consumes port hits normalized into candidates with identity fields
|
||||
+24
@@ -0,0 +1,24 @@
|
||||
# rag-knowledge-retrieval Delta
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
|
||||
|
||||
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
|
||||
|
||||
#### Scenario: Chunk-level deduplication keeps distinct chunks
|
||||
|
||||
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
|
||||
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
|
||||
- **AND** SHALL NOT collapse them solely because source is equal
|
||||
|
||||
#### Scenario: Same evidenceKey collapses
|
||||
|
||||
- **WHEN** two candidates share the same evidenceKey
|
||||
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
|
||||
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
|
||||
|
||||
#### Scenario: Postprocess count tracking
|
||||
|
||||
- **WHEN** evidence post-processing completes
|
||||
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
|
||||
+19
@@ -0,0 +1,19 @@
|
||||
# rag-log-projections Delta (RAG portion)
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: RAG projection SHALL expose bounded document evidence only
|
||||
|
||||
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
|
||||
|
||||
#### Scenario: Project only bounded RAG evidence
|
||||
|
||||
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
|
||||
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
|
||||
|
||||
#### Scenario: Multiple chunks from one source may project
|
||||
|
||||
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
|
||||
- **AND** projection budgets allow them
|
||||
- **THEN** the projected evidence list SHALL include more than one item for that source
|
||||
- **AND** each item SHALL have a distinct `document_id`
|
||||
@@ -0,0 +1,43 @@
|
||||
# Tasks: rag-chunk-evidence-identity-dedup
|
||||
|
||||
## 1. Identity model
|
||||
|
||||
- [x] 1.1 Add `docId`, `chunkIndex`, `evidenceKey` to `RetrievedEvidenceCandidate` and `EvidenceBlock`
|
||||
- [x] 1.2 Implement shared identity helper (`docId#chunk-N` / `vector:{id}` / `rank:{n}`)
|
||||
- [x] 1.3 Extract identity in `KnowledgeDocumentRetriever` (or search-port mapper) from metadata
|
||||
|
||||
## 2. Search port boundary
|
||||
|
||||
- [x] 2.1 Add `KnowledgeSearchPort`, `KnowledgeSearchRequest`, `KnowledgeSearchHit`
|
||||
- [x] 2.2 Implement dense adapter delegating to `VectorSearchService`
|
||||
- [x] 2.3 Wire retriever to port; keep tool orchestration free of SDK details
|
||||
|
||||
## 3. Post-process dedup and caps
|
||||
|
||||
- [x] 3.1 Change dedup key to `evidenceKey`
|
||||
- [x] 3.2 Add `rag.max-chunks-per-document` (default 2)
|
||||
- [x] 3.3 Apply `rag.return-n` truncation after ranking/dedup/cap
|
||||
- [x] 3.4 Keep score-threshold / relevance behavior unchanged except ordering inputs
|
||||
|
||||
## 4. Lookup config
|
||||
|
||||
- [x] 4.1 Add `rag.retrieve-k` and `rag.return-n` with legacy `rag.top-k` fallback
|
||||
- [x] 4.2 Use retrieve-k for search port calls in `LookupKnowledgeTool`
|
||||
|
||||
## 5. Projector alignment
|
||||
|
||||
- [x] 5.1 Prefer evidenceKey / chunk-scoped id as projected `document_id`
|
||||
- [x] 5.2 Stop dropping second evidence solely because `source` matches
|
||||
- [x] 5.3 Keep existing budget/truncation behavior
|
||||
|
||||
## 6. Tests
|
||||
|
||||
- [x] 6.1 Update `LookupKnowledgeToolTest` source-dedup case to chunk-preserving behavior
|
||||
- [x] 6.2 Add post-processor tests: multi-chunk keep, true-dup merge, maxChunksPerDocument
|
||||
- [x] 6.3 Update/add `RagResultProjectorTest` for same-source multi-chunk projection
|
||||
- [x] 6.4 Add search-port adapter smoke test if practical
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Run targeted unit tests for lookup / post-processor / projector
|
||||
- [x] 7.2 Mark tasks complete and note any residual risks for Delivery 2
|
||||
@@ -0,0 +1 @@
|
||||
ready: 2026-07-27
|
||||
@@ -0,0 +1,3 @@
|
||||
committed: 2026-07-27
|
||||
change: rag-hybrid-search-rrf
|
||||
authorized-apply: user-preauthorized-sm-flow
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-27
|
||||
@@ -0,0 +1,58 @@
|
||||
# Design: rag-hybrid-search-rrf
|
||||
|
||||
## Context
|
||||
|
||||
- Delivery 1 archived: chunk evidenceKey, SearchPort, return caps.
|
||||
- Legacy SDK path will be abandoned later; hybrid must live on SearchPort.
|
||||
- Full sparse schema rebuild is operationally heavy; ship fusion first.
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1. Hybrid = multi-path + RRF on SearchPort
|
||||
|
||||
When `retrieval.search.mode=hybrid`:
|
||||
|
||||
```text
|
||||
paths:
|
||||
1) dense(query, filter=null) # always
|
||||
2) dense(query, filter=category) # if category present
|
||||
3) lexical rank over union of dense hits # sparse-lite
|
||||
fuse by evidenceKey using RRF(k)
|
||||
return topK fused hits
|
||||
```
|
||||
|
||||
### D2. RRF formula
|
||||
|
||||
```text
|
||||
score(d) = Σ w_i / (k + rank_i(d))
|
||||
```
|
||||
|
||||
Defaults: k=60, all w_i=1.0. Optional weights via config.
|
||||
|
||||
### D3. Score semantics
|
||||
|
||||
- `KnowledgeSearchHit.score` remains **compatible L2 distance from best dense hit** for post-process thresholds.
|
||||
- Fused RRF score is carried in metadata (`fusedScore`, `fusionRanks`) for trace, not as L2.
|
||||
|
||||
### D4. Lexical path (sparse-lite)
|
||||
|
||||
Until BM25 schema:
|
||||
|
||||
- Tokenize query (simple whitespace / non-alnum split, lower-case)
|
||||
- Score each candidate by term coverage over title+breadcrumb+content
|
||||
- Rank candidates for RRF path only
|
||||
- Not a replacement for true BM25 inverted index
|
||||
|
||||
### D5. Serial filter retry
|
||||
|
||||
Lookup tool keeps low-quality unfiltered retry as safety net even in hybrid, because hybrid already includes unfiltered dense; retry remains cheap no-op when already fused well.
|
||||
|
||||
## Non-goals now
|
||||
|
||||
- New Milvus collection fields
|
||||
- Reindex jobs
|
||||
- SDK hybridSearch API calls
|
||||
|
||||
## Risks
|
||||
|
||||
- Lexical path only ranks already-recalled dense candidates → does not expand pure-term misses outside dense topK. Acceptable intermediate; true BM25 later expands recall.
|
||||
@@ -0,0 +1,48 @@
|
||||
# Change: RAG hybrid search with RRF
|
||||
|
||||
## Why
|
||||
|
||||
Delivery 1 fixed chunk identity/dedup. Delivery 2 enables multi-path retrieval fused by RRF so category filters no longer whole-replace good results and lexical signals can complement dense ranks.
|
||||
|
||||
True Milvus sparse/BM25 schema migration remains staged: this change ships a **working hybrid mode on KnowledgeSearchPort** using multi-path dense (+ lexical rank path) and RRF, without thickening the legacy SDK API surface.
|
||||
|
||||
## What Changes
|
||||
|
||||
- `retrieval.search.mode=dense|hybrid` (default dense)
|
||||
- Hybrid search on `KnowledgeSearchPort`:
|
||||
- dense unfiltered path
|
||||
- dense filtered path when category present
|
||||
- lexical rank path over the candidate union (sparse-lite until BM25 schema lands)
|
||||
- RRF / optional weighted RRF fusion by evidenceKey
|
||||
- Preserve dense L2 scores for quality thresholds; RRF only orders
|
||||
- Lookup tool uses search mode; filter low-quality retry remains as safety net
|
||||
- Config for rrf-k and path weights
|
||||
- Tests for RRF fusion ordering and hybrid adapter behavior
|
||||
|
||||
## Non-goals
|
||||
|
||||
- Full Milvus BM25/sparse collection rebuild (documented follow-up)
|
||||
- Deleting legacy SDK mode entirely
|
||||
- Cross-encoder model rerank
|
||||
- Changing Agent tool schema name/input
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `rag-hybrid-search`: hybrid mode, multi-path recall, RRF fusion via search port
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `rag-chunk-evidence-identity`: hybrid hits must keep chunk identity
|
||||
- `rag-knowledge-retrieval`: filtered/unfiltered cooperation via fusion rather than only serial replace
|
||||
|
||||
## Impact
|
||||
|
||||
- Code: retrieval package, lookup wiring/config, tests
|
||||
- Runtime default remains dense unless hybrid enabled
|
||||
- Interface: L2 internal (search port behavior)
|
||||
|
||||
## Depends on
|
||||
|
||||
- Archived Delivery 1: `rag-chunk-evidence-identity-dedup`
|
||||
+44
@@ -0,0 +1,44 @@
|
||||
# rag-hybrid-search Specification
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Knowledge search SHALL support configurable dense and hybrid modes
|
||||
|
||||
The knowledge search port SHALL support `retrieval.search.mode` values `dense` and `hybrid`. Default SHALL be `dense` for backward-compatible behavior.
|
||||
|
||||
#### Scenario: Dense mode single path
|
||||
|
||||
- **WHEN** mode is dense
|
||||
- **THEN** the port SHALL perform a single dense vector search with the provided category filter
|
||||
|
||||
#### Scenario: Hybrid mode multi-path fusion
|
||||
|
||||
- **WHEN** mode is hybrid
|
||||
- **THEN** the port SHALL gather multiple retrieval paths and fuse them by RRF before returning hits
|
||||
|
||||
### Requirement: Hybrid fusion SHALL use RRF over evidence identities
|
||||
|
||||
Hybrid fusion SHALL score candidates by reciprocal rank fusion on evidenceKey identities and SHALL NOT require comparable raw scores across paths.
|
||||
|
||||
#### Scenario: Multi-path candidate rises with RRF
|
||||
|
||||
- **WHEN** a chunk ranks highly on more than one hybrid path
|
||||
- **THEN** its fused rank SHALL improve relative to a chunk that ranks highly on only one path
|
||||
|
||||
### Requirement: Hybrid SHALL preserve dense quality scores for thresholds
|
||||
|
||||
Fused ordering SHALL NOT replace the dense compatibility score used by post-process quality thresholds.
|
||||
|
||||
#### Scenario: Threshold score remains dense-compatible
|
||||
|
||||
- **WHEN** hybrid returns a hit that originated from dense search
|
||||
- **THEN** hit.score SHALL remain a dense-compatible distance/similarity mapping usable by existing normalizeL2 thresholds
|
||||
|
||||
### Requirement: Hybrid SHALL keep chunk identity fields
|
||||
|
||||
Hybrid hits SHALL populate docId, chunkIndex, and evidenceKey consistently with Delivery 1 identity rules.
|
||||
|
||||
#### Scenario: Fused hit keeps evidenceKey
|
||||
|
||||
- **WHEN** hybrid merges the same chunk from two paths
|
||||
- **THEN** the returned hit SHALL use one evidenceKey and retain chunk identity metadata
|
||||
@@ -0,0 +1,9 @@
|
||||
# Tasks: rag-hybrid-search-rrf
|
||||
|
||||
- [x] 1. Add RRF fusion utility
|
||||
- [x] 2. Extend KnowledgeSearchRequest/Hit metadata for fusion ranks
|
||||
- [x] 3. Implement hybrid multi-path search in VectorKnowledgeSearchAdapter
|
||||
- [x] 4. Wire retrieval.search.mode and rrf config
|
||||
- [x] 5. Pass mode from LookupKnowledgeTool / retriever
|
||||
- [x] 6. Unit tests for RRF and hybrid adapter ordering
|
||||
- [x] 7. Verify dense mode regression tests still pass
|
||||
@@ -0,0 +1,80 @@
|
||||
# rag-chunk-evidence-identity Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change rag-chunk-evidence-identity-dedup. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Retrieval candidates SHALL carry chunk-level identity
|
||||
|
||||
The knowledge retrieval pipeline SHALL attach stable identity fields to each retrieved candidate and resulting evidence block: document id when available, chunk index when available, and an `evidenceKey` derived by the identity rules in design.
|
||||
|
||||
#### Scenario: Identity extracted from metadata
|
||||
|
||||
- **WHEN** a vector hit includes metadata `docId` (or `doc_id`) and `chunkIndex` (or `chunk_index`)
|
||||
- **THEN** the candidate SHALL set `docId`, `chunkIndex`, and `evidenceKey` to `docId#chunk-{chunkIndex}`
|
||||
|
||||
#### Scenario: Fallback identity without chunk index
|
||||
|
||||
- **WHEN** chunk index is missing but vector id is present
|
||||
- **THEN** the candidate SHALL use `evidenceKey` of the form `vector:{id}`
|
||||
|
||||
### Requirement: Evidence deduplication SHALL be chunk-scoped
|
||||
|
||||
Post-processing SHALL treat two candidates as duplicates only when they share the same `evidenceKey`. Different chunks of the same source SHALL remain separate evidence blocks subject to per-document caps.
|
||||
|
||||
#### Scenario: Same document different chunks are kept
|
||||
|
||||
- **WHEN** two candidates share the same source/docId but different chunk indexes
|
||||
- **THEN** post-processing SHALL keep both as separate evidence blocks unless a per-document cap removes the lower-ranked one
|
||||
|
||||
#### Scenario: True duplicate keys merge without replacing higher-ranked content
|
||||
|
||||
- **WHEN** two candidates share the same `evidenceKey`
|
||||
- **THEN** post-processing SHALL keep a single block
|
||||
- **AND** SHALL preserve the higher-ranked content
|
||||
- **AND** MAY merge hit reasons
|
||||
|
||||
### Requirement: Post-processing SHALL enforce max chunks per document
|
||||
|
||||
The system SHALL limit accepted evidence blocks per document id using configuration `rag.max-chunks-per-document` with default 2.
|
||||
|
||||
#### Scenario: Excess chunks from one document are dropped
|
||||
|
||||
- **WHEN** more than N ranked chunks belong to the same docId and N equals the configured max
|
||||
- **THEN** only the top N by ranking score SHALL remain in evidence blocks
|
||||
|
||||
### Requirement: Retrieval width and return width SHALL be separate
|
||||
|
||||
The lookup pipeline SHALL use `rag.retrieve-k` for vector recall width and `rag.return-n` for maximum evidence blocks after post-processing. A legacy `rag.top-k` MAY act as fallback when the new keys are absent.
|
||||
|
||||
#### Scenario: retrieve-k widens recall without unbounded return
|
||||
|
||||
- **WHEN** `retrieve-k` is 20 and `return-n` is 5
|
||||
- **THEN** the search port is asked for up to 20 candidates
|
||||
- **AND** the assembled result contains at most 5 evidence blocks before Agent projection budgets
|
||||
|
||||
### Requirement: Agent projection SHALL preserve distinct chunk evidence
|
||||
|
||||
The RAG projector SHALL deduplicate Agent-facing evidence by chunk-scoped evidence identity. It SHALL NOT drop a second chunk solely because `source` matches a previous block.
|
||||
|
||||
#### Scenario: Same source different evidence keys both project
|
||||
|
||||
- **WHEN** two evidence blocks have different evidence keys (or chunk-scoped document ids) and the same source
|
||||
- **AND** projection budgets still allow both
|
||||
- **THEN** both SHALL appear in the projected evidence list
|
||||
|
||||
#### Scenario: Projected document_id is chunk-scoped
|
||||
|
||||
- **WHEN** an evidence block has evidenceKey `docA#chunk-2`
|
||||
- **THEN** the projected `document_id` SHALL be that evidenceKey (or an equivalent chunk-scoped id)
|
||||
- **AND** `source` MAY still be the document path or doc id
|
||||
|
||||
### Requirement: Lookup pipeline SHALL call a KnowledgeSearchPort boundary
|
||||
|
||||
Semantic candidate fetch SHALL go through a `KnowledgeSearchPort` abstraction rather than embedding new long-term SDK-specific hybrid logic into the tool orchestrator.
|
||||
|
||||
#### Scenario: Default dense adapter
|
||||
|
||||
- **WHEN** search mode is dense-only (this change)
|
||||
- **THEN** the port implementation MAY delegate to the existing vector search facade
|
||||
- **AND** the retriever consumes port hits normalized into candidates with identity fields
|
||||
|
||||
@@ -0,0 +1,46 @@
|
||||
# rag-hybrid-search Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change rag-hybrid-search-rrf. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Knowledge search SHALL support configurable dense and hybrid modes
|
||||
|
||||
The knowledge search port SHALL support `retrieval.search.mode` values `dense` and `hybrid`. Default SHALL be `dense` for backward-compatible behavior.
|
||||
|
||||
#### Scenario: Dense mode single path
|
||||
|
||||
- **WHEN** mode is dense
|
||||
- **THEN** the port SHALL perform a single dense vector search with the provided category filter
|
||||
|
||||
#### Scenario: Hybrid mode multi-path fusion
|
||||
|
||||
- **WHEN** mode is hybrid
|
||||
- **THEN** the port SHALL gather multiple retrieval paths and fuse them by RRF before returning hits
|
||||
|
||||
### Requirement: Hybrid fusion SHALL use RRF over evidence identities
|
||||
|
||||
Hybrid fusion SHALL score candidates by reciprocal rank fusion on evidenceKey identities and SHALL NOT require comparable raw scores across paths.
|
||||
|
||||
#### Scenario: Multi-path candidate rises with RRF
|
||||
|
||||
- **WHEN** a chunk ranks highly on more than one hybrid path
|
||||
- **THEN** its fused rank SHALL improve relative to a chunk that ranks highly on only one path
|
||||
|
||||
### Requirement: Hybrid SHALL preserve dense quality scores for thresholds
|
||||
|
||||
Fused ordering SHALL NOT replace the dense compatibility score used by post-process quality thresholds.
|
||||
|
||||
#### Scenario: Threshold score remains dense-compatible
|
||||
|
||||
- **WHEN** hybrid returns a hit that originated from dense search
|
||||
- **THEN** hit.score SHALL remain a dense-compatible distance/similarity mapping usable by existing normalizeL2 thresholds
|
||||
|
||||
### Requirement: Hybrid SHALL keep chunk identity fields
|
||||
|
||||
Hybrid hits SHALL populate docId, chunkIndex, and evidenceKey consistently with Delivery 1 identity rules.
|
||||
|
||||
#### Scenario: Fused hit keeps evidenceKey
|
||||
|
||||
- **WHEN** hybrid merges the same chunk from two paths
|
||||
- **THEN** the returned hit SHALL use one evidenceKey and retain chunk identity metadata
|
||||
|
||||
@@ -54,14 +54,23 @@ The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structu
|
||||
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
|
||||
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
|
||||
|
||||
#### Scenario: Duplicate source deduplication
|
||||
- **WHEN** L0 and L1 produce evidence with the same source identity
|
||||
- **THEN** the retrieval flow SHALL keep a single evidence block for that source
|
||||
- **AND** the evidence block SHALL preserve hit reasons from both retrieval paths when available
|
||||
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent. Duplicates are defined by chunk-level `evidenceKey` identity, not by document source alone.
|
||||
|
||||
#### Scenario: Chunk-level deduplication keeps distinct chunks
|
||||
|
||||
- **WHEN** L1 produces multiple candidates with the same source but different chunk identities
|
||||
- **THEN** the retrieval flow SHALL keep separate evidence blocks for those chunk identities subject to per-document caps
|
||||
- **AND** SHALL NOT collapse them solely because source is equal
|
||||
|
||||
#### Scenario: Same evidenceKey collapses
|
||||
|
||||
- **WHEN** two candidates share the same evidenceKey
|
||||
- **THEN** the retrieval flow SHALL keep a single evidence block for that identity
|
||||
- **AND** the evidence block SHALL preserve hit reasons from both paths when available
|
||||
|
||||
#### Scenario: Postprocess count tracking
|
||||
|
||||
- **WHEN** evidence post-processing completes
|
||||
- **THEN** the tool invocation details SHALL record candidate count and final evidence block count
|
||||
|
||||
@@ -248,3 +257,4 @@ The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_
|
||||
- **WHEN** `lookup_knowledge` returns no usable evidence
|
||||
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
|
||||
- **AND** it SHALL not include full evidence content as duplicated trace data
|
||||
|
||||
|
||||
@@ -3,18 +3,23 @@
|
||||
## Purpose
|
||||
|
||||
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: RAG projection SHALL expose bounded document evidence only
|
||||
|
||||
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
|
||||
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result. Projected evidence identity SHALL be chunk-scoped when chunk identity is available, allowing multiple excerpts from one logical source document.
|
||||
|
||||
#### Scenario: Project only bounded RAG evidence
|
||||
|
||||
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
|
||||
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
|
||||
|
||||
#### Scenario: Multiple chunks from one source may project
|
||||
|
||||
- **WHEN** the backend returns multiple evidence blocks with the same source and different chunk-scoped identities
|
||||
- **AND** projection budgets allow them
|
||||
- **THEN** the projected evidence list SHALL include more than one item for that source
|
||||
- **AND** each item SHALL have a distinct `document_id`
|
||||
|
||||
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
|
||||
|
||||
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
|
||||
@@ -41,3 +46,4 @@ Both adapters SHALL pass the framework `tool_call_id` and RunContext to the exis
|
||||
|
||||
- **WHEN** either the RAG or query-log adapter executes
|
||||
- **THEN** it passes the exact framework `tool_call_id` and RunContext through ToolBoundary without creating a second identifier, parallel store, raw response path, or legacy audit side effect
|
||||
|
||||
|
||||
@@ -8,17 +8,28 @@ import java.util.List;
|
||||
/**
|
||||
* 内部证据块(后处理输出,尚未投影到 Agent 契约)。
|
||||
*
|
||||
* <p>一条 block 通常对应一次向量命中的一个 chunk 内容(截断后)。
|
||||
* 注意:当前后处理按 source 去重,同文档多 chunk 可能被合并掉,只保留最高分 content。</p>
|
||||
* <p>一条 block 对应一个 chunk 级证据身份({@code evidenceKey})。
|
||||
* 同文档多个相关 chunk 可以同时存在,受 maxChunksPerDocument / return-n 约束。</p>
|
||||
*
|
||||
* <p>投影到 Agent 时由 {@code RagResultProjector} 转为 {@code RagEvidence}
|
||||
* (document_id / source / title / breadcrumb / excerpt)。</p>
|
||||
* <p>投影到 Agent 时由 {@code RagResultProjector} 转为 {@code RagEvidence},
|
||||
* {@code document_id} 优先使用 evidenceKey(chunk 级)。</p>
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class EvidenceBlock {
|
||||
|
||||
/** 来源标识,常见为文件路径、upload:docId 或 docId。 */
|
||||
/** 文档级 id;用于每文档上限统计。 */
|
||||
private String docId;
|
||||
|
||||
/** 切片序号。 */
|
||||
private Integer chunkIndex;
|
||||
|
||||
/**
|
||||
* 片段级身份;去重与投影 document_id 的首选。
|
||||
*/
|
||||
private String evidenceKey;
|
||||
|
||||
/** 来源标识,常见为文件路径、upload:docId 或 docId;可重复。 */
|
||||
private String source;
|
||||
|
||||
private String title;
|
||||
|
||||
@@ -9,8 +9,8 @@ import java.util.Map;
|
||||
/**
|
||||
* L1 向量命中后、后处理前的统一候选结构。
|
||||
*
|
||||
* <p>由 {@code KnowledgeDocumentRetriever} 从 {@code VectorSearchService.SearchResult} 映射而来。
|
||||
* 后处理会基于它做归一化、规则 boost、去重并生成 {@link EvidenceBlock}。</p>
|
||||
* <p>由检索适配器从 {@code KnowledgeSearchHit} / 向量结果映射而来。
|
||||
* 后处理会基于它做归一化、规则 boost、chunk 级去重并生成 {@link EvidenceBlock}。</p>
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
@@ -19,9 +19,24 @@ public class RetrievedEvidenceCandidate {
|
||||
/** 向量库记录 id。 */
|
||||
private String id;
|
||||
|
||||
/**
|
||||
* 文档级 id(metadata.docId 等)。
|
||||
* 用于每文档 chunk 上限;不等于 evidenceKey。
|
||||
*/
|
||||
private String docId;
|
||||
|
||||
/** 文档内切片序号;可能为空(老数据)。 */
|
||||
private Integer chunkIndex;
|
||||
|
||||
/**
|
||||
* 片段级去重/投影主键。
|
||||
* 通常为 docId#chunk-N,fallback 为 vector:{id}。
|
||||
*/
|
||||
private String evidenceKey;
|
||||
|
||||
/**
|
||||
* 来源标识(_source / source / filePath / docId 等)。
|
||||
* 当前后处理去重主要依赖该字段,粒度偏文档级。
|
||||
* 可与同文档其他 chunk 重复;不再作为唯一去重键。
|
||||
*/
|
||||
private String source;
|
||||
|
||||
@@ -54,7 +69,7 @@ public class RetrievedEvidenceCandidate {
|
||||
|
||||
/**
|
||||
* 扁平化 metadata(string map)。
|
||||
* 可能含 docId、chunkIndex、category、kb_scope 等;chunkIndex 尚未提升为一等字段。
|
||||
* 可能含 docId、chunkIndex、category、kb_scope 等。
|
||||
*/
|
||||
private Map<String, String> metadata;
|
||||
|
||||
|
||||
@@ -18,27 +18,9 @@ import java.util.Set;
|
||||
/**
|
||||
* 把 legacy {@code LookupResult} JSON 投影成冻结的 Agent 可见 RAG 契约。
|
||||
*
|
||||
* <h3>为什么需要投影</h3>
|
||||
* 内部检索结果字段较多(retrievalTrace、rerankTrace、score、hitReasons、contextPack…),
|
||||
* Agent / EvidenceGuard 只应看到受控、有界、可引用的子集。
|
||||
*
|
||||
* <h3>保留给 Agent 的字段</h3>
|
||||
* <ul>
|
||||
* <li>evidenceStatus / toolCallId / query</li>
|
||||
* <li>evidence[]:document_id, source, title, breadcrumb, excerpt</li>
|
||||
* <li>relevanceLevel、truncated、returned_count</li>
|
||||
* </ul>
|
||||
*
|
||||
* <h3>刻意丢弃</h3>
|
||||
* score、hitReasons、retrievalTrace、rerankTrace、contextPack 等内部可观测细节。
|
||||
*
|
||||
* <h3>去重与预算(读代码关键)</h3>
|
||||
* <ul>
|
||||
* <li>按 {@code document_id} 去重;若 block 无 document_id,则回退 source/title</li>
|
||||
* <li>因此同 source 的多个 chunk 在此也会被压成 1 条(与后处理文档级去重叠加)</li>
|
||||
* <li>条数上限 {@link ToolProjectionLimits#maxEvidence()},excerpt 字符上限,
|
||||
* 以及总 UTF-8 字节预算(超限从后往前删 evidence)</li>
|
||||
* </ul>
|
||||
* <h3>去重</h3>
|
||||
* 按 chunk 级证据身份去重(优先 evidenceKey / document_id),
|
||||
* <b>不再</b>因为 source 相同就丢弃第二条 chunk。
|
||||
*/
|
||||
public final class RagResultProjector {
|
||||
|
||||
@@ -50,11 +32,6 @@ public final class RagResultProjector {
|
||||
this.limits = limits;
|
||||
}
|
||||
|
||||
/**
|
||||
* @param request Agent 原始请求(用于回填/截断 query)
|
||||
* @param toolCallId 框架分配的本次工具调用 id,进入结果供证据引用
|
||||
* @param rawResponse legacy LookupResult 的 JSON 字符串
|
||||
*/
|
||||
public ProjectedToolResult project(RagToolRequest request, String toolCallId,
|
||||
String rawResponse) throws Exception {
|
||||
if (request == null || toolCallId == null || toolCallId.isBlank()) {
|
||||
@@ -69,8 +46,7 @@ public final class RagResultProjector {
|
||||
String query = bounded(request.query(), limits.maxQueryChars());
|
||||
truncated = !query.equals(request.query());
|
||||
List<RagEvidence> evidence = new ArrayList<>();
|
||||
// 文档级唯一集合:相同 documentId 只保留首次出现
|
||||
Set<String> documentIds = new HashSet<>();
|
||||
Set<String> evidenceIds = new HashSet<>();
|
||||
JsonNode blocks = root.has("evidenceBlocks") ? root.get("evidenceBlocks") : root.get("evidence_blocks");
|
||||
if (blocks != null && blocks.isArray()) {
|
||||
int ordinal = 0;
|
||||
@@ -89,11 +65,16 @@ public final class RagResultProjector {
|
||||
}
|
||||
String source = text(block, "source");
|
||||
String title = text(block, "title");
|
||||
// legacy EvidenceBlock 通常没有 document_id,实际常退化为 source
|
||||
String documentId = firstNonBlank(text(block, "document_id"), source, title,
|
||||
"legacy-document-" + ordinal);
|
||||
if (!documentIds.add(documentId)) {
|
||||
// 同 documentId 重复:丢弃后续条,并标记 truncated
|
||||
// Chunk-scoped identity first; do not collapse on source alone.
|
||||
String documentId = firstPresent(
|
||||
text(block, "evidenceKey"),
|
||||
text(block, "evidence_key"),
|
||||
text(block, "document_id"),
|
||||
text(block, "documentId"),
|
||||
composeChunkId(text(block, "docId"), text(block, "doc_id"),
|
||||
text(block, "chunkIndex"), text(block, "chunk_index")),
|
||||
"legacy-evidence-" + ordinal);
|
||||
if (!evidenceIds.add(documentId)) {
|
||||
truncated = true;
|
||||
continue;
|
||||
}
|
||||
@@ -118,12 +99,36 @@ public final class RagResultProjector {
|
||||
? relevanceLevel(root) : null;
|
||||
RagToolResult result = new RagToolResult(
|
||||
status, toolCallId, query, evidence, evidence.size(), relevanceLevel, truncated);
|
||||
// 总字节预算:仍超限则从尾部删 evidence,直到放得下或变 no_evidence
|
||||
result = fitBudget(result, truncated);
|
||||
return new ProjectedToolResult(objectMapper.writeValueAsString(result), result.evidenceStatus());
|
||||
}
|
||||
|
||||
/** 按 maxAgentUtf8Bytes 从后往前删 evidence,保证 Agent 侧 payload 有界。 */
|
||||
private static String composeChunkId(String docId, String docIdAlt, String chunkIndex, String chunkIndexAlt) {
|
||||
String id = firstPresentOrNull(docId, docIdAlt);
|
||||
String idx = firstPresentOrNull(chunkIndex, chunkIndexAlt);
|
||||
if (id == null || idx == null) {
|
||||
return null;
|
||||
}
|
||||
return id + "#chunk-" + idx;
|
||||
}
|
||||
|
||||
private static String firstPresent(String... values) {
|
||||
String found = firstPresentOrNull(values);
|
||||
return found == null ? "unknown-document" : found;
|
||||
}
|
||||
|
||||
private static String firstPresentOrNull(String... values) {
|
||||
if (values == null) {
|
||||
return null;
|
||||
}
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private RagToolResult fitBudget(RagToolResult result, boolean truncated) throws Exception {
|
||||
RagToolResult current = result;
|
||||
while (bytes(objectMapper.writeValueAsString(current)) > limits.maxAgentUtf8Bytes()
|
||||
@@ -146,15 +151,6 @@ public final class RagResultProjector {
|
||||
return value == null || value.isNull() ? "" : value.asText("");
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return "unknown-document";
|
||||
}
|
||||
|
||||
private static String nullable(String value) {
|
||||
return value == null || value.isBlank() ? null : value;
|
||||
}
|
||||
|
||||
@@ -1,40 +1,36 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import com.superbiz.agent.service.retrieval.KnowledgeSearchHit;
|
||||
import com.superbiz.agent.service.retrieval.KnowledgeSearchMode;
|
||||
import com.superbiz.agent.service.retrieval.KnowledgeSearchPort;
|
||||
import com.superbiz.agent.service.retrieval.KnowledgeSearchRequest;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Locale;
|
||||
|
||||
/**
|
||||
* L1 向量检索适配器。
|
||||
*
|
||||
* <p>职责是把 {@link VectorSearchService} 的原始命中,转成 pipeline 统一使用的
|
||||
* <p>通过 {@link KnowledgeSearchPort} 拉取候选,映射为 pipeline 统一的
|
||||
* {@link RetrievedEvidenceCandidate},并记录单次 attempt 的 trace。</p>
|
||||
*
|
||||
* <p>本类不做质量判断、不做 rerank、不做上下文打包;那些属于后处理阶段。</p>
|
||||
*
|
||||
* <h3>字段映射要点</h3>
|
||||
* <ul>
|
||||
* <li>{@code source}:优先 metadata._source / source / filePath / docId</li>
|
||||
* <li>{@code title/breadcrumb}:来自 chunk metadata,用于展示与规则 boost</li>
|
||||
* <li>{@code score}:兼容后的距离分(SDK 为 L2;Spring AI 路径会映射成兼容 L2)</li>
|
||||
* <li>metadata 中的 chunkIndex 目前只留在 map 里,未提升为一等字段</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeDocumentRetriever {
|
||||
|
||||
private final VectorSearchService vectorSearchService;
|
||||
private final ObjectMapper objectMapper;
|
||||
private final KnowledgeSearchPort knowledgeSearchPort;
|
||||
|
||||
public KnowledgeDocumentRetriever(VectorSearchService vectorSearchService, ObjectMapper objectMapper) {
|
||||
this.vectorSearchService = vectorSearchService;
|
||||
this.objectMapper = objectMapper;
|
||||
@Value("${retrieval.search.mode:dense}")
|
||||
private String searchMode = "dense";
|
||||
|
||||
public KnowledgeDocumentRetriever(KnowledgeSearchPort knowledgeSearchPort) {
|
||||
this.knowledgeSearchPort = knowledgeSearchPort;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -43,22 +39,24 @@ public class KnowledgeDocumentRetriever {
|
||||
* @param attemptName 写入 trace 的 attempt 名(FILTERED_VECTOR / UNFILTERED_VECTOR 等)
|
||||
* @param query 检索文本
|
||||
* @param categoryFilter 可选 category 元数据过滤;null 表示不过滤
|
||||
* @param topK 召回条数
|
||||
* @param topK 召回条数(retrieve-k)
|
||||
* @return attempt 元信息 + 候选列表;异常时 candidates 为空,error 记在 attempt 上
|
||||
*/
|
||||
public RetrievalAttemptResult retrieve(String attemptName, String query, String categoryFilter, int topK) {
|
||||
long start = System.currentTimeMillis();
|
||||
try {
|
||||
List<VectorSearchService.SearchResult> results =
|
||||
vectorSearchService.searchSimilarDocuments(query, topK, categoryFilter);
|
||||
List<RetrievedEvidenceCandidate> candidates = toCandidates(attemptName, results);
|
||||
KnowledgeSearchMode mode = "hybrid".equalsIgnoreCase(trim(searchMode))
|
||||
? KnowledgeSearchMode.HYBRID
|
||||
: KnowledgeSearchMode.DENSE;
|
||||
List<KnowledgeSearchHit> hits = knowledgeSearchPort.search(
|
||||
new KnowledgeSearchRequest(query, topK, categoryFilter, mode));
|
||||
List<RetrievedEvidenceCandidate> candidates = toCandidates(attemptName, hits);
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, candidates.size(), null,
|
||||
(int) (System.currentTimeMillis() - start), topScore(results)),
|
||||
(int) (System.currentTimeMillis() - start), topScore(hits)),
|
||||
candidates
|
||||
);
|
||||
} catch (Exception e) {
|
||||
// 检索失败不向上抛:由上层按“无候选 / 低质量”路径继续(例如 fallback retry)
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, 0, e.getMessage(),
|
||||
(int) (System.currentTimeMillis() - start), null),
|
||||
@@ -67,42 +65,29 @@ public class KnowledgeDocumentRetriever {
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 将向量库原始结果规范化为候选证据。
|
||||
* originalRank 从 1 开始,对应向量召回顺序(尚未规则 rerank)。
|
||||
*/
|
||||
private List<RetrievedEvidenceCandidate> toCandidates(String attemptName,
|
||||
List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
private List<RetrievedEvidenceCandidate> toCandidates(String attemptName, List<KnowledgeSearchHit> hits) {
|
||||
if (hits == null || hits.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<RetrievedEvidenceCandidate> candidates = new ArrayList<>();
|
||||
for (int i = 0; i < results.size(); i++) {
|
||||
VectorSearchService.SearchResult result = results.get(i);
|
||||
Map<String, String> metadata = parseMetadata(result.getMetadata());
|
||||
// source 是后续去重/展示的主标识;当前实现偏“文档级”,同文档多 chunk 可能共享 source
|
||||
String source = firstNonBlank(
|
||||
metadata.get("_source"),
|
||||
metadata.get("source"),
|
||||
metadata.get("filePath"),
|
||||
metadata.get("docId"),
|
||||
result.getMetadata(),
|
||||
result.getId()
|
||||
);
|
||||
List<RetrievedEvidenceCandidate> candidates = new ArrayList<>(hits.size());
|
||||
for (KnowledgeSearchHit hit : hits) {
|
||||
candidates.add(RetrievedEvidenceCandidate.builder()
|
||||
.id(result.getId())
|
||||
.source(source)
|
||||
.title(metadata.get("title"))
|
||||
.breadcrumb(metadata.get("breadcrumb"))
|
||||
.content(result.getContent())
|
||||
.id(hit.id())
|
||||
.docId(hit.docId())
|
||||
.chunkIndex(hit.chunkIndex())
|
||||
.evidenceKey(hit.evidenceKey())
|
||||
.source(hit.source())
|
||||
.title(hit.title())
|
||||
.breadcrumb(hit.breadcrumb())
|
||||
.content(hit.content())
|
||||
.retrievalLayer("L1")
|
||||
.retrievalAttempt(attemptName)
|
||||
.score((double) result.getScore())
|
||||
.rawScore(result.getRawScore())
|
||||
.scoreLabel(result.getScoreLabel())
|
||||
.originalRank(i + 1)
|
||||
.metadata(metadata)
|
||||
.hitReasons(List.of("semantic_rank:" + (i + 1), "attempt:" + attemptName))
|
||||
.score(hit.score())
|
||||
.rawScore(hit.rawScore())
|
||||
.scoreLabel(hit.scoreLabel())
|
||||
.originalRank(hit.originalRank())
|
||||
.metadata(hit.metadata() == null ? java.util.Map.of() : hit.metadata())
|
||||
.hitReasons(List.of("semantic_rank:" + hit.originalRank(), "attempt:" + attemptName))
|
||||
.build());
|
||||
}
|
||||
return candidates;
|
||||
@@ -120,7 +105,6 @@ public class KnowledgeDocumentRetriever {
|
||||
.query(query)
|
||||
.categoryFilter(categoryFilter)
|
||||
.candidateCount(candidateCount)
|
||||
// 此处 usable 只表示“有候选且无错误”;后处理还会用相似度阈值再收紧
|
||||
.usable(errorMessage == null && candidateCount > 0)
|
||||
.errorMessage(errorMessage)
|
||||
.durationMs(durationMs)
|
||||
@@ -128,46 +112,19 @@ public class KnowledgeDocumentRetriever {
|
||||
.build();
|
||||
}
|
||||
|
||||
private Double topScore(List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
private Double topScore(List<KnowledgeSearchHit> hits) {
|
||||
if (hits == null || hits.isEmpty() || hits.get(0).score() == null) {
|
||||
return null;
|
||||
}
|
||||
return (double) results.get(0).getScore();
|
||||
return hits.get(0).score();
|
||||
}
|
||||
|
||||
/** metadata 在向量库中多为 JSON 字符串,这里压成 string map 方便后处理读取。 */
|
||||
private Map<String, String> parseMetadata(String metadata) {
|
||||
if (metadata == null || metadata.isBlank()) {
|
||||
return Map.of();
|
||||
}
|
||||
try {
|
||||
Map<?, ?> raw = objectMapper.readValue(metadata, Map.class);
|
||||
Map<String, String> parsed = new LinkedHashMap<>();
|
||||
for (Map.Entry<?, ?> entry : raw.entrySet()) {
|
||||
if (entry.getKey() != null && entry.getValue() != null) {
|
||||
parsed.put(String.valueOf(entry.getKey()), String.valueOf(entry.getValue()));
|
||||
}
|
||||
}
|
||||
return parsed;
|
||||
} catch (Exception ignored) {
|
||||
return Map.of();
|
||||
}
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
private static String trim(String value) {
|
||||
return value == null ? "" : value.trim().toLowerCase(Locale.ROOT);
|
||||
}
|
||||
|
||||
/**
|
||||
* 单次检索 attempt 的结果包。
|
||||
*
|
||||
* @param attempt 可观测元数据(耗时、过滤条件、错误等)
|
||||
* @param candidates 规范化后的证据候选
|
||||
*/
|
||||
public record RetrievalAttemptResult(RetrievalTrace.Attempt attempt,
|
||||
List<RetrievedEvidenceCandidate> candidates) {
|
||||
|
||||
@@ -5,11 +5,13 @@ import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import com.superbiz.agent.service.retrieval.EvidenceIdentity;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
@@ -18,24 +20,18 @@ import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* 检索后处理:分数归一化、规则 rerank、证据块组装、相关等级判定。
|
||||
* 检索后处理:分数归一化、规则 rerank、chunk 级证据组装、相关等级判定。
|
||||
*
|
||||
* <h3>处理步骤</h3>
|
||||
* <ol>
|
||||
* <li>把候选 L2 距离归一成 0~1 的 baseScore</li>
|
||||
* <li>用 L0 hint(domain/entity/keyword)做规则加分,得到 finalScore</li>
|
||||
* <li>用 L0 hint 做规则加分,得到 finalScore</li>
|
||||
* <li>按 finalScore 降序排序</li>
|
||||
* <li>按 sourceKey 去重后生成 {@link EvidenceBlock}</li>
|
||||
* <li>按 evidenceKey 去重</li>
|
||||
* <li>按 maxChunksPerDocument 截断同文档 chunk</li>
|
||||
* <li>按 return-n 截断最终 evidence 条数</li>
|
||||
* <li>根据 top baseScore + hint 支撑计算 relevanceLevel</li>
|
||||
* </ol>
|
||||
*
|
||||
* <h3>重要行为(读代码时容易误解)</h3>
|
||||
* <ul>
|
||||
* <li>“rerank” 是规则加权,不是 cross-encoder / LLM rerank</li>
|
||||
* <li>去重 key 优先 source/title,粒度偏文档级:同文档多个 chunk 可能被压成一条,
|
||||
* 且 merge 时只合并 hitReasons,不拼接 content</li>
|
||||
* <li>content 在此截断到 800 字符;Agent 侧 projector 还可能再截断</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeEvidencePostProcessor {
|
||||
@@ -48,24 +44,28 @@ public class KnowledgeEvidencePostProcessor {
|
||||
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
|
||||
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
|
||||
|
||||
/** L2 距离上限,用于把距离映射到 [0,1] 相似度。 */
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
/** baseScore >= 该阈值,才可能判 HIGHLY_RELEVANT / PRECISE。 */
|
||||
@Value("${retrieval.normalization.highly-relevant-threshold:0.75}")
|
||||
private double highlyRelevantThreshold = 0.75;
|
||||
|
||||
/** baseScore >= 该阈值视为可用参考;低于则 isLowQuality=true,可能触发 unfiltered retry。 */
|
||||
@Value("${retrieval.normalization.reference-threshold:0.5}")
|
||||
private double referenceThreshold = 0.5;
|
||||
|
||||
/** 同一 docId 最多保留的 chunk 数。 */
|
||||
@Value("${rag.max-chunks-per-document:2}")
|
||||
private int maxChunksPerDocument = 2;
|
||||
|
||||
/**
|
||||
* 对一次 attempt 的候选做后处理,产出可交给打包/组装的证据结果。
|
||||
* 后处理后最多返回的 evidence 条数。
|
||||
* 0 或负数表示不在此层截断(仍可能被 projector 预算截断)。
|
||||
*/
|
||||
@Value("${rag.return-n:5}")
|
||||
private int returnN = 5;
|
||||
|
||||
public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) {
|
||||
List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates;
|
||||
// 先打分排序:baseScore 来自向量距离,finalScore = base + 规则 boost
|
||||
List<ScoredCandidate> ranked = safeCandidates.stream()
|
||||
.map(candidate -> score(query, candidate))
|
||||
.sorted(Comparator.comparingDouble(ScoredCandidate::finalScore).reversed())
|
||||
@@ -73,22 +73,34 @@ public class KnowledgeEvidencePostProcessor {
|
||||
|
||||
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
|
||||
List<RerankTrace.Item> traceItems = new ArrayList<>();
|
||||
Map<String, Integer> chunksPerDocument = new HashMap<>();
|
||||
int finalRank = 1;
|
||||
int accepted = 0;
|
||||
int effectiveReturnN = returnN > 0 ? returnN : Integer.MAX_VALUE;
|
||||
int effectiveMaxChunks = maxChunksPerDocument > 0 ? maxChunksPerDocument : Integer.MAX_VALUE;
|
||||
|
||||
for (ScoredCandidate scored : ranked) {
|
||||
if (accepted >= effectiveReturnN) {
|
||||
break;
|
||||
}
|
||||
RetrievedEvidenceCandidate candidate = scored.candidate();
|
||||
EvidenceBlock block = EvidenceBlock.builder()
|
||||
.source(candidate.getSource())
|
||||
.title(candidate.getTitle())
|
||||
.breadcrumb(candidate.getBreadcrumb())
|
||||
.retrievalLayer(candidate.getRetrievalLayer())
|
||||
.content(truncate(candidate.getContent(), 800))
|
||||
.score(candidate.getScore())
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons()))
|
||||
.build();
|
||||
// 注意:key 未使用 chunkIndex,同 source 的多个 chunk 会走 merge 分支
|
||||
String key = sourceKey(block, "candidate-" + candidate.getOriginalRank());
|
||||
if (!deduped.containsKey(key)) {
|
||||
deduped.put(key, block);
|
||||
String evidenceKey = resolveEvidenceKey(candidate);
|
||||
String docBucket = resolveDocBucket(candidate, evidenceKey);
|
||||
|
||||
if (deduped.containsKey(evidenceKey)) {
|
||||
mergeEvidence(deduped.get(evidenceKey), toBlock(candidate, evidenceKey, scored));
|
||||
continue;
|
||||
}
|
||||
|
||||
int used = chunksPerDocument.getOrDefault(docBucket, 0);
|
||||
if (used >= effectiveMaxChunks) {
|
||||
continue;
|
||||
}
|
||||
|
||||
EvidenceBlock block = toBlock(candidate, evidenceKey, scored);
|
||||
deduped.put(evidenceKey, block);
|
||||
chunksPerDocument.put(docBucket, used + 1);
|
||||
accepted++;
|
||||
traceItems.add(RerankTrace.Item.builder()
|
||||
.finalRank(finalRank++)
|
||||
.source(candidate.getSource())
|
||||
@@ -96,14 +108,9 @@ public class KnowledgeEvidencePostProcessor {
|
||||
.finalScore(scored.finalScore())
|
||||
.boostReasons(scored.boostReasons())
|
||||
.build());
|
||||
} else {
|
||||
// 重复 key:保留先放入的(更高分)content,只补充 reasons/breadcrumb
|
||||
mergeEvidence(deduped.get(key), block);
|
||||
}
|
||||
}
|
||||
|
||||
List<EvidenceBlock> blocks = new ArrayList<>(deduped.values());
|
||||
// topSimilarity 用排序后第一名的 baseScore(未含 boost),供质量阈值判断
|
||||
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).baseScore();
|
||||
RelevanceAssessment assessment = computeRelevance(query, ranked);
|
||||
return EvidencePostprocessResult.builder()
|
||||
@@ -117,10 +124,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
.build();
|
||||
}
|
||||
|
||||
/**
|
||||
* 是否低质量,用于触发 filtered -> unfiltered 降级。
|
||||
* 无可用证据,或 topSimilarity 低于 referenceThreshold,都视为低质量。
|
||||
*/
|
||||
public boolean isLowQuality(EvidencePostprocessResult result) {
|
||||
if (result == null || !result.hasUsableEvidence()) {
|
||||
return true;
|
||||
@@ -129,10 +132,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return topSimilarity == null || topSimilarity < referenceThreshold;
|
||||
}
|
||||
|
||||
/**
|
||||
* L2 距离 -> 相似度。
|
||||
* 距离越小越相似:similarity = 1 - min(l2, max) / max。
|
||||
*/
|
||||
public double normalizeL2(Double l2Score) {
|
||||
if (l2Score == null) {
|
||||
return 0.0;
|
||||
@@ -145,10 +144,43 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return referenceThreshold;
|
||||
}
|
||||
|
||||
/**
|
||||
* 规则打分:baseScore + domain/entity/keyword/source_type boost。
|
||||
* boost 只影响排序,不改变用于阈值判断的 baseScore。
|
||||
*/
|
||||
private EvidenceBlock toBlock(RetrievedEvidenceCandidate candidate,
|
||||
String evidenceKey,
|
||||
ScoredCandidate scored) {
|
||||
return EvidenceBlock.builder()
|
||||
.docId(candidate.getDocId())
|
||||
.chunkIndex(candidate.getChunkIndex())
|
||||
.evidenceKey(evidenceKey)
|
||||
.source(candidate.getSource())
|
||||
.title(candidate.getTitle())
|
||||
.breadcrumb(candidate.getBreadcrumb())
|
||||
.retrievalLayer(candidate.getRetrievalLayer())
|
||||
.content(truncate(candidate.getContent(), 800))
|
||||
.score(candidate.getScore())
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons()))
|
||||
.build();
|
||||
}
|
||||
|
||||
private String resolveEvidenceKey(RetrievedEvidenceCandidate candidate) {
|
||||
if (candidate.getEvidenceKey() != null && !candidate.getEvidenceKey().isBlank()) {
|
||||
return candidate.getEvidenceKey();
|
||||
}
|
||||
return EvidenceIdentity.evidenceKey(
|
||||
candidate.getDocId(),
|
||||
candidate.getChunkIndex(),
|
||||
candidate.getId(),
|
||||
candidate.getOriginalRank());
|
||||
}
|
||||
|
||||
private String resolveDocBucket(RetrievedEvidenceCandidate candidate, String evidenceKey) {
|
||||
String docId = EvidenceIdentity.trimToNull(candidate.getDocId());
|
||||
if (docId != null) {
|
||||
return docId;
|
||||
}
|
||||
// No docId: do not collapse unrelated fallback keys under one bucket.
|
||||
return evidenceKey;
|
||||
}
|
||||
|
||||
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate) {
|
||||
double baseScore = normalizeL2(candidate.getScore());
|
||||
double finalScore = baseScore;
|
||||
@@ -174,7 +206,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return new ScoredCandidate(candidate, baseScore, finalScore, boosts);
|
||||
}
|
||||
|
||||
/** 在 source/title/breadcrumb/content/metadata 拼接串上做子串匹配(大小写不敏感)。 */
|
||||
private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) {
|
||||
if (hints == null || hints.isEmpty()) {
|
||||
return false;
|
||||
@@ -194,7 +225,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return false;
|
||||
}
|
||||
|
||||
/** runbook / guide / case 类来源轻微加分。 */
|
||||
private boolean isPreferredSourceType(RetrievedEvidenceCandidate candidate) {
|
||||
Map<String, String> metadata = candidate.getMetadata();
|
||||
if (metadata == null || metadata.isEmpty()) {
|
||||
@@ -208,10 +238,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return normalized.contains("runbook") || normalized.contains("guide") || normalized.contains("case");
|
||||
}
|
||||
|
||||
/**
|
||||
* 相关等级只看 top1 的 baseScore(向量相似度),PRECISE 额外要求 L0 hint 有支撑。
|
||||
* 低于 referenceThreshold 时 level/hint 都为 null,表示不可用参考。
|
||||
*/
|
||||
private RelevanceAssessment computeRelevance(KnowledgeQuery query, List<ScoredCandidate> ranked) {
|
||||
if (ranked.isEmpty()) {
|
||||
return new RelevanceAssessment(null, null);
|
||||
@@ -235,9 +261,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
|| matchesAny(top.candidate(), query.getMatchedKeywords());
|
||||
}
|
||||
|
||||
/**
|
||||
* 同 key 合并策略:不覆盖已有 content(保留更高分的那条),只补 reasons 和空 breadcrumb。
|
||||
*/
|
||||
private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) {
|
||||
Set<String> reasons = new LinkedHashSet<>();
|
||||
if (existing.getHitReasons() != null) {
|
||||
@@ -264,14 +287,6 @@ public class KnowledgeEvidencePostProcessor {
|
||||
return new ArrayList<>(merged);
|
||||
}
|
||||
|
||||
/**
|
||||
* 当前去重 key:source -> title -> breadcrumb -> fallback。
|
||||
* 因此“同文档不同 chunk”若 source 相同,会被视为重复。
|
||||
*/
|
||||
private String sourceKey(EvidenceBlock block, String fallback) {
|
||||
return firstNonBlank(block.getSource(), block.getTitle(), block.getBreadcrumb(), fallback);
|
||||
}
|
||||
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
@@ -280,19 +295,13 @@ public class KnowledgeEvidencePostProcessor {
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
return EvidenceIdentity.firstNonBlank(values);
|
||||
}
|
||||
|
||||
private String nullToEmpty(String value) {
|
||||
return value == null ? "" : value;
|
||||
}
|
||||
|
||||
/** 内部打分结果:baseScore 用于阈值,finalScore 用于排序。 */
|
||||
private record ScoredCandidate(RetrievedEvidenceCandidate candidate,
|
||||
double baseScore,
|
||||
double finalScore,
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Chunk-level evidence identity helpers shared by retrieval mapping and post-processing.
|
||||
*/
|
||||
public final class EvidenceIdentity {
|
||||
|
||||
private EvidenceIdentity() {
|
||||
}
|
||||
|
||||
public static String evidenceKey(String docId, Integer chunkIndex, String vectorId, Integer originalRank) {
|
||||
String normalizedDocId = trimToNull(docId);
|
||||
if (normalizedDocId != null && chunkIndex != null) {
|
||||
return normalizedDocId + "#chunk-" + chunkIndex;
|
||||
}
|
||||
String normalizedVectorId = trimToNull(vectorId);
|
||||
if (normalizedVectorId != null) {
|
||||
return "vector:" + normalizedVectorId;
|
||||
}
|
||||
int rank = originalRank == null ? 0 : originalRank;
|
||||
return "rank:" + rank;
|
||||
}
|
||||
|
||||
public static String extractDocId(Map<String, String> metadata, String... fallbacks) {
|
||||
String fromMeta = firstNonBlank(
|
||||
metadataValue(metadata, "docId"),
|
||||
metadataValue(metadata, "doc_id"));
|
||||
if (fromMeta != null) {
|
||||
return fromMeta;
|
||||
}
|
||||
return firstNonBlank(fallbacks);
|
||||
}
|
||||
|
||||
public static Integer extractChunkIndex(Map<String, String> metadata) {
|
||||
String raw = firstNonBlank(
|
||||
metadataValue(metadata, "chunkIndex"),
|
||||
metadataValue(metadata, "chunk_index"));
|
||||
if (raw == null) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
return Integer.valueOf(raw.trim());
|
||||
} catch (NumberFormatException ignored) {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
public static String metadataValue(Map<String, String> metadata, String key) {
|
||||
if (metadata == null || key == null) {
|
||||
return null;
|
||||
}
|
||||
return trimToNull(metadata.get(key));
|
||||
}
|
||||
|
||||
public static String firstNonBlank(String... values) {
|
||||
if (values == null) {
|
||||
return null;
|
||||
}
|
||||
for (String value : values) {
|
||||
String trimmed = trimToNull(value);
|
||||
if (trimmed != null) {
|
||||
return trimmed;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
public static String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,24 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Normalized search hit returned by {@link KnowledgeSearchPort}.
|
||||
*/
|
||||
public record KnowledgeSearchHit(
|
||||
String id,
|
||||
String content,
|
||||
Double score,
|
||||
Double rawScore,
|
||||
String scoreLabel,
|
||||
String metadataJson,
|
||||
Map<String, String> metadata,
|
||||
String docId,
|
||||
Integer chunkIndex,
|
||||
String evidenceKey,
|
||||
String source,
|
||||
String title,
|
||||
String breadcrumb,
|
||||
int originalRank
|
||||
) {
|
||||
}
|
||||
@@ -0,0 +1,10 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
/**
|
||||
* Retrieval mode for {@link KnowledgeSearchPort}.
|
||||
* Delivery 1 only requires {@link #DENSE}; hybrid arrives in a later change.
|
||||
*/
|
||||
public enum KnowledgeSearchMode {
|
||||
DENSE,
|
||||
HYBRID
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Application boundary for knowledge semantic search.
|
||||
* Implementations may wrap VectorStore, hybrid engines, etc. without leaking SDK details upward.
|
||||
*/
|
||||
public interface KnowledgeSearchPort {
|
||||
|
||||
List<KnowledgeSearchHit> search(KnowledgeSearchRequest request);
|
||||
}
|
||||
@@ -0,0 +1,23 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
/**
|
||||
* Portable search request used by the knowledge pipeline.
|
||||
*/
|
||||
public record KnowledgeSearchRequest(
|
||||
String query,
|
||||
int topK,
|
||||
String categoryFilter,
|
||||
KnowledgeSearchMode mode
|
||||
) {
|
||||
public KnowledgeSearchRequest {
|
||||
if (topK <= 0) {
|
||||
throw new IllegalArgumentException("topK must be positive");
|
||||
}
|
||||
mode = mode == null ? KnowledgeSearchMode.DENSE : mode;
|
||||
query = query == null ? "" : query;
|
||||
}
|
||||
|
||||
public static KnowledgeSearchRequest dense(String query, int topK, String categoryFilter) {
|
||||
return new KnowledgeSearchRequest(query, topK, categoryFilter, KnowledgeSearchMode.DENSE);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,73 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Sparse-lite lexical ranking over already recalled candidates.
|
||||
* Not a substitute for inverted-index BM25; expands ordering signal only.
|
||||
*/
|
||||
public final class LexicalRanker {
|
||||
|
||||
private LexicalRanker() {
|
||||
}
|
||||
|
||||
public static List<KnowledgeSearchHit> rank(String query, List<KnowledgeSearchHit> candidates) {
|
||||
if (candidates == null || candidates.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
Set<String> terms = tokenize(query);
|
||||
if (terms.isEmpty()) {
|
||||
return List.copyOf(candidates);
|
||||
}
|
||||
List<ScoredHit> scored = new ArrayList<>(candidates.size());
|
||||
for (KnowledgeSearchHit hit : candidates) {
|
||||
String haystack = (nullToEmpty(hit.title()) + " "
|
||||
+ nullToEmpty(hit.breadcrumb()) + " "
|
||||
+ nullToEmpty(hit.content())).toLowerCase(Locale.ROOT);
|
||||
int hits = 0;
|
||||
for (String term : terms) {
|
||||
if (haystack.contains(term)) {
|
||||
hits++;
|
||||
}
|
||||
}
|
||||
double coverage = hits / (double) terms.size();
|
||||
scored.add(new ScoredHit(hit, coverage, hits));
|
||||
}
|
||||
scored.sort(Comparator
|
||||
.comparingDouble((ScoredHit s) -> s.coverage).reversed()
|
||||
.thenComparingInt((ScoredHit s) -> s.hits).reversed()
|
||||
.thenComparingInt(s -> s.hit.originalRank()));
|
||||
return scored.stream().map(s -> s.hit).toList();
|
||||
}
|
||||
|
||||
static Set<String> tokenize(String query) {
|
||||
if (query == null || query.isBlank()) {
|
||||
return Set.of();
|
||||
}
|
||||
String normalized = query.toLowerCase(Locale.ROOT);
|
||||
String[] parts = normalized.split("[^\\p{IsAlphabetic}\\p{IsDigit}]+");
|
||||
Set<String> terms = new LinkedHashSet<>();
|
||||
for (String part : parts) {
|
||||
if (part == null) {
|
||||
continue;
|
||||
}
|
||||
String term = part.trim();
|
||||
if (term.length() >= 2) {
|
||||
terms.add(term);
|
||||
}
|
||||
}
|
||||
return terms;
|
||||
}
|
||||
|
||||
private static String nullToEmpty(String value) {
|
||||
return value == null ? "" : value;
|
||||
}
|
||||
|
||||
private record ScoredHit(KnowledgeSearchHit hit, double coverage, int hits) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.HashMap;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Objects;
|
||||
import java.util.function.Function;
|
||||
|
||||
/**
|
||||
* Reciprocal Rank Fusion helpers.
|
||||
*
|
||||
* <pre>
|
||||
* RRF_w(d) = Σ w_i / (k + rank_i(d))
|
||||
* </pre>
|
||||
*/
|
||||
public final class RrfFusion {
|
||||
|
||||
private RrfFusion() {
|
||||
}
|
||||
|
||||
public static <T> List<Scored<T>> fuse(List<RankedPath<T>> paths,
|
||||
int rrfK,
|
||||
Function<T, String> identityFn) {
|
||||
if (paths == null || paths.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
int k = Math.max(1, rrfK);
|
||||
Map<String, Acc<T>> acc = new LinkedHashMap<>();
|
||||
for (RankedPath<T> path : paths) {
|
||||
if (path == null || path.items() == null || path.items().isEmpty()) {
|
||||
continue;
|
||||
}
|
||||
double weight = path.weight() <= 0 ? 1.0 : path.weight();
|
||||
List<T> items = path.items();
|
||||
for (int i = 0; i < items.size(); i++) {
|
||||
T item = items.get(i);
|
||||
if (item == null) {
|
||||
continue;
|
||||
}
|
||||
String id = identityFn.apply(item);
|
||||
if (id == null || id.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
int rank = i + 1;
|
||||
double contrib = weight / (k + rank);
|
||||
Acc<T> bucket = acc.computeIfAbsent(id, ignored -> new Acc<>(item));
|
||||
bucket.score += contrib;
|
||||
bucket.ranks.put(path.name(), rank);
|
||||
// Prefer first-seen item payload; callers should put preferred path first if needed.
|
||||
}
|
||||
}
|
||||
List<Scored<T>> scored = new ArrayList<>(acc.size());
|
||||
for (Map.Entry<String, Acc<T>> entry : acc.entrySet()) {
|
||||
Acc<T> value = entry.getValue();
|
||||
scored.add(new Scored<>(entry.getKey(), value.item, value.score, Map.copyOf(value.ranks)));
|
||||
}
|
||||
scored.sort(Comparator
|
||||
.comparingDouble((Scored<T> s) -> s.rrfScore()).reversed()
|
||||
.thenComparing(Scored::identity));
|
||||
return scored;
|
||||
}
|
||||
|
||||
public record RankedPath<T>(String name, List<T> items, double weight) {
|
||||
public RankedPath {
|
||||
Objects.requireNonNull(name, "name");
|
||||
items = items == null ? List.of() : List.copyOf(items);
|
||||
}
|
||||
}
|
||||
|
||||
public record Scored<T>(String identity, T item, double rrfScore, Map<String, Integer> ranks) {
|
||||
}
|
||||
|
||||
private static final class Acc<T> {
|
||||
private final T item;
|
||||
private double score;
|
||||
private final Map<String, Integer> ranks = new HashMap<>();
|
||||
|
||||
private Acc(T item) {
|
||||
this.item = item;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,211 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Component;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* {@link KnowledgeSearchPort} backed by the existing vector search facade.
|
||||
*
|
||||
* <ul>
|
||||
* <li>{@code DENSE}: single dense path (legacy behavior)</li>
|
||||
* <li>{@code HYBRID}: dense unfiltered + optional dense filtered + lexical rank, fused by RRF</li>
|
||||
* </ul>
|
||||
*/
|
||||
@Component
|
||||
public class VectorKnowledgeSearchAdapter implements KnowledgeSearchPort {
|
||||
|
||||
private final VectorSearchService vectorSearchService;
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
@Value("${retrieval.search.mode:dense}")
|
||||
private String configuredMode = "dense";
|
||||
|
||||
@Value("${retrieval.hybrid.rrf-k:60}")
|
||||
private int rrfK = 60;
|
||||
|
||||
@Value("${retrieval.hybrid.weight.dense-unfiltered:1.0}")
|
||||
private double weightDenseUnfiltered = 1.0;
|
||||
|
||||
@Value("${retrieval.hybrid.weight.dense-filtered:1.0}")
|
||||
private double weightDenseFiltered = 1.0;
|
||||
|
||||
@Value("${retrieval.hybrid.weight.lexical:1.0}")
|
||||
private double weightLexical = 1.0;
|
||||
|
||||
public VectorKnowledgeSearchAdapter(VectorSearchService vectorSearchService, ObjectMapper objectMapper) {
|
||||
this.vectorSearchService = vectorSearchService;
|
||||
this.objectMapper = objectMapper;
|
||||
}
|
||||
|
||||
@Override
|
||||
public List<KnowledgeSearchHit> search(KnowledgeSearchRequest request) {
|
||||
KnowledgeSearchMode mode = resolveMode(request.mode());
|
||||
if (mode == KnowledgeSearchMode.HYBRID) {
|
||||
return searchHybrid(request);
|
||||
}
|
||||
return searchDense(request.query(), request.topK(), request.categoryFilter());
|
||||
}
|
||||
|
||||
private List<KnowledgeSearchHit> searchDense(String query, int topK, String categoryFilter) {
|
||||
List<VectorSearchService.SearchResult> results = vectorSearchService.searchSimilarDocuments(
|
||||
query, topK, categoryFilter);
|
||||
return toHits(results);
|
||||
}
|
||||
|
||||
private List<KnowledgeSearchHit> searchHybrid(KnowledgeSearchRequest request) {
|
||||
String query = request.query();
|
||||
int topK = request.topK();
|
||||
String category = trimToNull(request.categoryFilter());
|
||||
|
||||
List<KnowledgeSearchHit> unfiltered = searchDense(query, topK, null);
|
||||
List<KnowledgeSearchHit> filtered = category == null
|
||||
? List.of()
|
||||
: searchDense(query, topK, category);
|
||||
|
||||
Map<String, KnowledgeSearchHit> unionByKey = new LinkedHashMap<>();
|
||||
for (KnowledgeSearchHit hit : unfiltered) {
|
||||
unionByKey.putIfAbsent(hit.evidenceKey(), hit);
|
||||
}
|
||||
for (KnowledgeSearchHit hit : filtered) {
|
||||
unionByKey.putIfAbsent(hit.evidenceKey(), hit);
|
||||
}
|
||||
List<KnowledgeSearchHit> union = new ArrayList<>(unionByKey.values());
|
||||
List<KnowledgeSearchHit> lexical = LexicalRanker.rank(query, union);
|
||||
|
||||
List<RrfFusion.RankedPath<KnowledgeSearchHit>> paths = new ArrayList<>();
|
||||
paths.add(new RrfFusion.RankedPath<>("dense_unfiltered", unfiltered, weightDenseUnfiltered));
|
||||
if (!filtered.isEmpty()) {
|
||||
paths.add(new RrfFusion.RankedPath<>("dense_filtered", filtered, weightDenseFiltered));
|
||||
}
|
||||
if (!lexical.isEmpty()) {
|
||||
paths.add(new RrfFusion.RankedPath<>("lexical", lexical, weightLexical));
|
||||
}
|
||||
|
||||
List<RrfFusion.Scored<KnowledgeSearchHit>> fused = RrfFusion.fuse(
|
||||
paths, rrfK, KnowledgeSearchHit::evidenceKey);
|
||||
|
||||
List<KnowledgeSearchHit> ordered = new ArrayList<>();
|
||||
int rank = 1;
|
||||
for (RrfFusion.Scored<KnowledgeSearchHit> scored : fused) {
|
||||
if (ordered.size() >= topK) {
|
||||
break;
|
||||
}
|
||||
KnowledgeSearchHit base = scored.item();
|
||||
Map<String, String> metadata = new LinkedHashMap<>(
|
||||
base.metadata() == null ? Map.of() : base.metadata());
|
||||
metadata.put("fusedScore", Double.toString(scored.rrfScore()));
|
||||
metadata.put("fusionRanks", scored.ranks().toString());
|
||||
metadata.put("fusionRank", Integer.toString(rank));
|
||||
ordered.add(new KnowledgeSearchHit(
|
||||
base.id(),
|
||||
base.content(),
|
||||
base.score(),
|
||||
base.rawScore(),
|
||||
base.scoreLabel(),
|
||||
base.metadataJson(),
|
||||
metadata,
|
||||
base.docId(),
|
||||
base.chunkIndex(),
|
||||
base.evidenceKey(),
|
||||
base.source(),
|
||||
base.title(),
|
||||
base.breadcrumb(),
|
||||
rank
|
||||
));
|
||||
rank++;
|
||||
}
|
||||
return ordered;
|
||||
}
|
||||
|
||||
private KnowledgeSearchMode resolveMode(KnowledgeSearchMode requestMode) {
|
||||
if (requestMode == KnowledgeSearchMode.HYBRID) {
|
||||
return KnowledgeSearchMode.HYBRID;
|
||||
}
|
||||
if (requestMode == KnowledgeSearchMode.DENSE) {
|
||||
// Allow global config to force hybrid even if caller passes DENSE default.
|
||||
String configured = configuredMode == null ? "dense" : configuredMode.trim().toLowerCase(Locale.ROOT);
|
||||
if ("hybrid".equals(configured)) {
|
||||
return KnowledgeSearchMode.HYBRID;
|
||||
}
|
||||
return KnowledgeSearchMode.DENSE;
|
||||
}
|
||||
String configured = configuredMode == null ? "dense" : configuredMode.trim().toLowerCase(Locale.ROOT);
|
||||
return "hybrid".equals(configured) ? KnowledgeSearchMode.HYBRID : KnowledgeSearchMode.DENSE;
|
||||
}
|
||||
|
||||
private List<KnowledgeSearchHit> toHits(List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<KnowledgeSearchHit> hits = new ArrayList<>(results.size());
|
||||
for (int i = 0; i < results.size(); i++) {
|
||||
hits.add(toHit(results.get(i), i + 1));
|
||||
}
|
||||
return hits;
|
||||
}
|
||||
|
||||
private KnowledgeSearchHit toHit(VectorSearchService.SearchResult result, int originalRank) {
|
||||
Map<String, String> metadata = parseMetadata(result.getMetadata());
|
||||
String docId = EvidenceIdentity.extractDocId(
|
||||
metadata,
|
||||
EvidenceIdentity.metadataValue(metadata, "_source"),
|
||||
EvidenceIdentity.metadataValue(metadata, "source"));
|
||||
Integer chunkIndex = EvidenceIdentity.extractChunkIndex(metadata);
|
||||
String evidenceKey = EvidenceIdentity.evidenceKey(docId, chunkIndex, result.getId(), originalRank);
|
||||
String source = EvidenceIdentity.firstNonBlank(
|
||||
EvidenceIdentity.metadataValue(metadata, "_source"),
|
||||
EvidenceIdentity.metadataValue(metadata, "source"),
|
||||
EvidenceIdentity.metadataValue(metadata, "filePath"),
|
||||
docId,
|
||||
result.getId());
|
||||
return new KnowledgeSearchHit(
|
||||
result.getId(),
|
||||
result.getContent(),
|
||||
(double) result.getScore(),
|
||||
result.getRawScore(),
|
||||
result.getScoreLabel(),
|
||||
result.getMetadata(),
|
||||
metadata,
|
||||
docId,
|
||||
chunkIndex,
|
||||
evidenceKey,
|
||||
source,
|
||||
EvidenceIdentity.metadataValue(metadata, "title"),
|
||||
EvidenceIdentity.metadataValue(metadata, "breadcrumb"),
|
||||
originalRank
|
||||
);
|
||||
}
|
||||
|
||||
private Map<String, String> parseMetadata(String metadata) {
|
||||
if (metadata == null || metadata.isBlank()) {
|
||||
return Map.of();
|
||||
}
|
||||
try {
|
||||
Map<?, ?> raw = objectMapper.readValue(metadata, Map.class);
|
||||
Map<String, String> parsed = new LinkedHashMap<>();
|
||||
for (Map.Entry<?, ?> entry : raw.entrySet()) {
|
||||
if (entry.getKey() != null && entry.getValue() != null) {
|
||||
parsed.put(String.valueOf(entry.getKey()), String.valueOf(entry.getValue()));
|
||||
}
|
||||
}
|
||||
return parsed;
|
||||
} catch (Exception ignored) {
|
||||
return Map.of();
|
||||
}
|
||||
}
|
||||
|
||||
private static String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
}
|
||||
@@ -10,6 +10,7 @@ import com.superbiz.agent.service.KnowledgeDocumentRetriever;
|
||||
import com.superbiz.agent.service.KnowledgeEvidencePostProcessor;
|
||||
import com.superbiz.agent.service.KnowledgeQueryTransformer;
|
||||
import com.superbiz.agent.service.LookupResultAssembler;
|
||||
import jakarta.annotation.PostConstruct;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
@@ -29,39 +30,37 @@ import java.util.Map;
|
||||
* <h3>主链路</h3>
|
||||
* <pre>
|
||||
* query
|
||||
* -> KnowledgeQueryTransformer // L0:domain/keyword hint,可选 category filter
|
||||
* -> KnowledgeDocumentRetriever // L1:向量召回 topK
|
||||
* -> KnowledgeEvidencePostProcessor // 归一化、规则 rerank、组装 evidenceBlocks
|
||||
* -> [可选] 去掉 category 后重试 // filtered 结果质量不足时
|
||||
* -> KnowledgeContextPacker // 按字符预算打包文本
|
||||
* -> LookupResultAssembler // 统一 LookupResult
|
||||
* -> KnowledgeQueryTransformer
|
||||
* -> KnowledgeDocumentRetriever (via KnowledgeSearchPort, retrieve-k)
|
||||
* -> KnowledgeEvidencePostProcessor (chunk dedup / caps / return-n)
|
||||
* -> [optional] unfiltered retry
|
||||
* -> KnowledgeContextPacker
|
||||
* -> LookupResultAssembler
|
||||
* </pre>
|
||||
*
|
||||
* <h3>降级策略</h3>
|
||||
* <ul>
|
||||
* <li>有 categoryFilter:先 FILTERED_VECTOR</li>
|
||||
* <li>若结果无证据或 topSimilarity < referenceThreshold:再 UNFILTERED_VECTOR_RETRY</li>
|
||||
* <li>无 categoryFilter:直接 UNFILTERED_VECTOR</li>
|
||||
* </ul>
|
||||
*
|
||||
* <p>注意:本类返回的是内部 {@link LookupResult},不是 Agent 最终看到的 JSON。
|
||||
* 会话级文档去重目前不在这里做。</p>
|
||||
*/
|
||||
@Slf4j
|
||||
@Component
|
||||
public class LookupKnowledgeTool {
|
||||
|
||||
/** 首次检索:带 L0 推导出的 category 过滤。 */
|
||||
private static final String ATTEMPT_FILTERED_VECTOR = "FILTERED_VECTOR";
|
||||
/** 首次检索:L0 未给出唯一 domain,不做 category 过滤。 */
|
||||
private static final String ATTEMPT_UNFILTERED_VECTOR = "UNFILTERED_VECTOR";
|
||||
/** 降级重试:去掉 category 过滤,用原始 query 再搜一次。 */
|
||||
private static final String ATTEMPT_UNFILTERED_VECTOR_RETRY = "UNFILTERED_VECTOR_RETRY";
|
||||
private static final String FALLBACK_NO_EVIDENCE = "filtered_vector_no_evidence";
|
||||
private static final String FALLBACK_LOW_QUALITY = "filtered_vector_low_quality";
|
||||
|
||||
@Value("${rag.top-k:3}")
|
||||
private int topK = 3;
|
||||
/**
|
||||
* Legacy single knob. Used only when retrieve-k / return-n are absent.
|
||||
*/
|
||||
@Value("${rag.top-k:0}")
|
||||
private int legacyTopK = 0;
|
||||
|
||||
@Value("${rag.retrieve-k:0}")
|
||||
private int retrieveKConfig = 0;
|
||||
|
||||
@Value("${rag.return-n:0}")
|
||||
private int returnNConfig = 0;
|
||||
|
||||
private int retrieveK = 20;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeQueryTransformer queryTransformer;
|
||||
@@ -78,19 +77,24 @@ public class LookupKnowledgeTool {
|
||||
@Autowired
|
||||
private LookupResultAssembler resultAssembler;
|
||||
|
||||
/**
|
||||
* 执行一次知识库检索,返回 evidence-first 的内部结果。
|
||||
*
|
||||
* @param query Agent / Harness 传入的检索语句(不是最终用户原话的完整上下文)
|
||||
* @return 含 evidenceBlocks、contextPack、retrievalTrace 的 LookupResult
|
||||
*/
|
||||
@PostConstruct
|
||||
void resolveRetrievalWidths() {
|
||||
int fallback = legacyTopK > 0 ? legacyTopK : 3;
|
||||
this.retrieveK = retrieveKConfig > 0 ? retrieveKConfig : fallback;
|
||||
// return-n is owned by post-processor config; keep field for observability only.
|
||||
if (returnNConfig <= 0 && legacyTopK > 0) {
|
||||
log.info("rag.return-n not set; post-processor will use its own default or rag.return-n binding");
|
||||
}
|
||||
log.info("lookup_knowledge widths: retrieveK={}, legacyTopK={}, returnNConfig={}",
|
||||
retrieveK, legacyTopK, returnNConfig);
|
||||
}
|
||||
|
||||
public LookupResult lookupKnowledge(String query) {
|
||||
log.info("========================================");
|
||||
log.info(">>> [工具调用] lookup_knowledge");
|
||||
log.info(">>> metadata: query_chars={}", query == null ? 0 : query.length());
|
||||
log.info(">>> metadata: query_chars={}, retrieveK={}", query == null ? 0 : query.length(), retrieveK);
|
||||
log.info("----------------------------------------");
|
||||
|
||||
// 1) Query understanding:L0 只产 hint/filter,不直接当事实证据
|
||||
KnowledgeQuery knowledgeQuery = queryTransformer.transform(query);
|
||||
log.info("[QueryTransformer] categoryFilter={}, domainHintCount={}, keywordCount={}",
|
||||
knowledgeQuery.getCategoryFilter(),
|
||||
@@ -100,7 +104,6 @@ public class LookupKnowledgeTool {
|
||||
List<RetrievalTrace.Attempt> attempts = new ArrayList<>();
|
||||
String fallbackReason = null;
|
||||
|
||||
// 2) 首次 L1 向量检索(有唯一 domain 则带 category filter)
|
||||
String firstAttemptName = knowledgeQuery.getCategoryFilter() == null
|
||||
? ATTEMPT_UNFILTERED_VECTOR
|
||||
: ATTEMPT_FILTERED_VECTOR;
|
||||
@@ -108,7 +111,7 @@ public class LookupKnowledgeTool {
|
||||
documentRetriever.retrieve(firstAttemptName,
|
||||
knowledgeQuery.getRewrittenQuery(),
|
||||
knowledgeQuery.getCategoryFilter(),
|
||||
topK);
|
||||
retrieveK);
|
||||
EvidencePostprocessResult selectedEvidence = evidencePostProcessor.process(
|
||||
knowledgeQuery,
|
||||
firstAttempt.candidates());
|
||||
@@ -116,8 +119,6 @@ public class LookupKnowledgeTool {
|
||||
attempts.add(firstAttempt.attempt());
|
||||
String selectedAttemptName = firstAttemptName;
|
||||
|
||||
// 3) filtered 路径质量不足时,去掉 category 用原始 query 重试一次
|
||||
// 重试结果会整体替换首次结果(不是与首次融合)
|
||||
if (knowledgeQuery.getCategoryFilter() != null && evidencePostProcessor.isLowQuality(selectedEvidence)) {
|
||||
fallbackReason = selectedEvidence.hasUsableEvidence()
|
||||
? FALLBACK_LOW_QUALITY
|
||||
@@ -128,7 +129,7 @@ public class LookupKnowledgeTool {
|
||||
documentRetriever.retrieve(ATTEMPT_UNFILTERED_VECTOR_RETRY,
|
||||
knowledgeQuery.getOriginalQuery(),
|
||||
null,
|
||||
topK);
|
||||
retrieveK);
|
||||
EvidencePostprocessResult retryEvidence = evidencePostProcessor.process(
|
||||
knowledgeQuery,
|
||||
retryAttempt.candidates());
|
||||
@@ -138,7 +139,6 @@ public class LookupKnowledgeTool {
|
||||
selectedAttemptName = ATTEMPT_UNFILTERED_VECTOR_RETRY;
|
||||
}
|
||||
|
||||
// 4) 打包 + 组装最终内部结果(供 projector / 审计消费)
|
||||
ContextPack contextPack = contextPacker.pack(selectedEvidence.getEvidenceBlocks());
|
||||
RetrievalTrace retrievalTrace = buildRetrievalTrace(knowledgeQuery, attempts, selectedAttemptName,
|
||||
fallbackReason, selectedEvidence);
|
||||
@@ -148,10 +148,6 @@ public class LookupKnowledgeTool {
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* 用后处理后的相似度回填 attempt 可观测字段。
|
||||
* usable 要求:有可用证据,且 topSimilarity 达到 reference 阈值。
|
||||
*/
|
||||
private void enrichAttempt(RetrievalTrace.Attempt attempt, EvidencePostprocessResult evidence) {
|
||||
attempt.setTopSimilarity(evidence.getTopSimilarity());
|
||||
attempt.setUsable(evidence.hasUsableEvidence()
|
||||
@@ -159,7 +155,6 @@ public class LookupKnowledgeTool {
|
||||
&& evidence.getTopSimilarity() >= evidencePostProcessor.getReferenceThreshold());
|
||||
}
|
||||
|
||||
/** 汇总本次检索的 query hint、attempt 列表与最终选用路径,便于 trace 回放。 */
|
||||
private RetrievalTrace buildRetrievalTrace(KnowledgeQuery query,
|
||||
List<RetrievalTrace.Attempt> attempts,
|
||||
String selectedAttempt,
|
||||
|
||||
+28
-7
@@ -16,14 +16,14 @@ class RagResultProjectorTest {
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
@Test
|
||||
void projectsOnlyDeduplicatedBoundedEvidence() throws Exception {
|
||||
void projectsDistinctChunkEvidenceEvenWhenSourceMatches() throws Exception {
|
||||
RagResultProjector projector = new RagResultProjector(objectMapper,
|
||||
new ToolProjectionLimits(2, 8, 2, 2, 20, 30, 4096));
|
||||
new ToolProjectionLimits(3, 8, 2, 2, 20, 30, 4096));
|
||||
String raw = """
|
||||
{"found":true,"evidenceBlocks":[
|
||||
{"source":"doc-1","title":"One","breadcrumb":"a","content":"123456789","score":0.99},
|
||||
{"source":"doc-1","title":"Duplicate","content":"ignored"},
|
||||
{"source":"doc-2","title":"Two","content":"second"}],
|
||||
{"source":"doc-1","evidenceKey":"doc-1#chunk-0","title":"One","breadcrumb":"a","content":"123456789","score":0.99},
|
||||
{"source":"doc-1","evidenceKey":"doc-1#chunk-1","title":"Two","content":"second-chunk"},
|
||||
{"source":"doc-2","evidenceKey":"doc-2#chunk-0","title":"Other","content":"third"}],
|
||||
"contextPack":{"packedText":"secret internal context"},"retrievalTrace":{"attempts":[]},"rerankTrace":{}}
|
||||
""";
|
||||
|
||||
@@ -32,7 +32,9 @@ class RagResultProjectorTest {
|
||||
|
||||
assertEquals(EvidenceStatus.EVIDENCE_FOUND, projected.evidenceStatus());
|
||||
assertEquals("framework-1", json.path("tool_call_id").asText());
|
||||
assertEquals(2, json.path("returned_count").asInt());
|
||||
assertEquals(3, json.path("returned_count").asInt());
|
||||
assertEquals("doc-1#chunk-0", json.path("evidence").get(0).path("document_id").asText());
|
||||
assertEquals("doc-1#chunk-1", json.path("evidence").get(1).path("document_id").asText());
|
||||
assertEquals("12345678", json.path("evidence").get(0).path("excerpt").asText());
|
||||
assertTrue(json.path("truncated").asBoolean());
|
||||
assertFalse(projected.agentResult().contains("contextPack"));
|
||||
@@ -40,12 +42,31 @@ class RagResultProjectorTest {
|
||||
assertFalse(projected.agentResult().contains("score"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void deduplicatesOnlySameEvidenceKey() throws Exception {
|
||||
RagResultProjector projector = new RagResultProjector(objectMapper, ToolProjectionLimits.defaults());
|
||||
String raw = """
|
||||
{"found":true,"evidenceBlocks":[
|
||||
{"source":"doc","evidenceKey":"doc#chunk-0","content":"first"},
|
||||
{"source":"doc","evidenceKey":"doc#chunk-0","content":"dup"},
|
||||
{"source":"doc","evidenceKey":"doc#chunk-1","content":"second"}]}
|
||||
""";
|
||||
|
||||
JsonNode json = objectMapper.readTree(
|
||||
projector.project(new RagToolRequest("q"), "call-1", raw).agentResult());
|
||||
|
||||
assertEquals(2, json.path("evidence").size());
|
||||
assertEquals("doc#chunk-0", json.path("evidence").get(0).path("document_id").asText());
|
||||
assertEquals("doc#chunk-1", json.path("evidence").get(1).path("document_id").asText());
|
||||
assertTrue(json.path("truncated").asBoolean());
|
||||
}
|
||||
|
||||
@Test
|
||||
void preservesCamelAndSnakeCaseReferenceLevel() throws Exception {
|
||||
RagResultProjector projector = new RagResultProjector(objectMapper, ToolProjectionLimits.defaults());
|
||||
String camel = """
|
||||
{"found":true,"relevanceLevel":"REFERENCE",
|
||||
"evidenceBlocks":[{"source":"doc","content":"generic guidance"}]}
|
||||
"evidenceBlocks":[{"source":"doc","evidenceKey":"doc#chunk-0","content":"generic guidance"}]}
|
||||
""";
|
||||
String snake = camel.replace("relevanceLevel", "relevance_level");
|
||||
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class KnowledgeEvidencePostProcessorTest {
|
||||
|
||||
private KnowledgeEvidencePostProcessor processor;
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
processor = new KnowledgeEvidencePostProcessor();
|
||||
ReflectionTestUtils.setField(processor, "maxChunksPerDocument", 2);
|
||||
ReflectionTestUtils.setField(processor, "returnN", 5);
|
||||
}
|
||||
|
||||
@Test
|
||||
void keepsDistinctChunksFromSameDocument() {
|
||||
EvidencePostprocessResult result = processor.process(query(), List.of(
|
||||
candidate("shared", 0, "shared#chunk-0", "c0", 0.2),
|
||||
candidate("shared", 1, "shared#chunk-1", "c1", 0.3)
|
||||
));
|
||||
|
||||
assertEquals(2, result.getEvidenceBlockCount());
|
||||
assertEquals("shared#chunk-0", result.getEvidenceBlocks().get(0).getEvidenceKey());
|
||||
assertEquals("shared#chunk-1", result.getEvidenceBlocks().get(1).getEvidenceKey());
|
||||
assertEquals("c0", result.getEvidenceBlocks().get(0).getContent());
|
||||
assertEquals("c1", result.getEvidenceBlocks().get(1).getContent());
|
||||
}
|
||||
|
||||
@Test
|
||||
void mergesTrueDuplicateEvidenceKeysWithoutReplacingContent() {
|
||||
EvidencePostprocessResult result = processor.process(query(), List.of(
|
||||
candidate("shared", 0, "shared#chunk-0", "keep-me", 0.2),
|
||||
RetrievedEvidenceCandidate.builder()
|
||||
.id("dup")
|
||||
.docId("shared")
|
||||
.chunkIndex(0)
|
||||
.evidenceKey("shared#chunk-0")
|
||||
.source("shared.md")
|
||||
.content("drop-me")
|
||||
.score(0.25)
|
||||
.originalRank(2)
|
||||
.hitReasons(List.of("extra"))
|
||||
.metadata(Map.of())
|
||||
.build()
|
||||
));
|
||||
|
||||
assertEquals(1, result.getEvidenceBlockCount());
|
||||
assertEquals("keep-me", result.getEvidenceBlocks().get(0).getContent());
|
||||
assertTrue(result.getEvidenceBlocks().get(0).getHitReasons().contains("extra"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void enforcesMaxChunksPerDocument() {
|
||||
ReflectionTestUtils.setField(processor, "maxChunksPerDocument", 2);
|
||||
EvidencePostprocessResult result = processor.process(query(), List.of(
|
||||
candidate("shared", 0, "shared#chunk-0", "c0", 0.1),
|
||||
candidate("shared", 1, "shared#chunk-1", "c1", 0.2),
|
||||
candidate("shared", 2, "shared#chunk-2", "c2", 0.3),
|
||||
candidate("other", 0, "other#chunk-0", "o0", 0.15)
|
||||
));
|
||||
|
||||
assertEquals(3, result.getEvidenceBlockCount());
|
||||
assertEquals(2, result.getEvidenceBlocks().stream()
|
||||
.filter(block -> "shared".equals(block.getDocId()))
|
||||
.count());
|
||||
assertTrue(result.getEvidenceBlocks().stream()
|
||||
.noneMatch(block -> "shared#chunk-2".equals(block.getEvidenceKey())));
|
||||
}
|
||||
|
||||
@Test
|
||||
void enforcesReturnN() {
|
||||
ReflectionTestUtils.setField(processor, "returnN", 1);
|
||||
EvidencePostprocessResult result = processor.process(query(), List.of(
|
||||
candidate("a", 0, "a#chunk-0", "a0", 0.1),
|
||||
candidate("b", 0, "b#chunk-0", "b0", 0.2)
|
||||
));
|
||||
|
||||
assertEquals(1, result.getEvidenceBlockCount());
|
||||
assertEquals("a#chunk-0", result.getEvidenceBlocks().get(0).getEvidenceKey());
|
||||
}
|
||||
|
||||
private static KnowledgeQuery query() {
|
||||
return KnowledgeQuery.builder()
|
||||
.originalQuery("q")
|
||||
.rewrittenQuery("q")
|
||||
.domainHints(List.of())
|
||||
.matchedKeywords(List.of())
|
||||
.entities(List.of())
|
||||
.build();
|
||||
}
|
||||
|
||||
private static RetrievedEvidenceCandidate candidate(String docId,
|
||||
int chunkIndex,
|
||||
String evidenceKey,
|
||||
String content,
|
||||
double score) {
|
||||
return RetrievedEvidenceCandidate.builder()
|
||||
.id(evidenceKey)
|
||||
.docId(docId)
|
||||
.chunkIndex(chunkIndex)
|
||||
.evidenceKey(evidenceKey)
|
||||
.source(docId + ".md")
|
||||
.content(content)
|
||||
.score(score)
|
||||
.originalRank(chunkIndex + 1)
|
||||
.hitReasons(List.of("base"))
|
||||
.metadata(Map.of("docId", docId, "chunkIndex", String.valueOf(chunkIndex)))
|
||||
.build();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class RrfFusionTest {
|
||||
|
||||
@Test
|
||||
void multiPathAgreementOutranksSinglePathHead() {
|
||||
List<String> dense = List.of("a", "b", "c");
|
||||
List<String> lexical = List.of("c", "b", "d");
|
||||
|
||||
List<RrfFusion.Scored<String>> fused = RrfFusion.fuse(
|
||||
List.of(
|
||||
new RrfFusion.RankedPath<>("dense", dense, 1.0),
|
||||
new RrfFusion.RankedPath<>("lexical", lexical, 1.0)
|
||||
),
|
||||
60,
|
||||
s -> s
|
||||
);
|
||||
|
||||
// c: dense#3 + lexical#1 ; b: dense#2 + lexical#2 ; a: dense#1 only
|
||||
// With k=60, c edges b slightly, and both beat single-path a.
|
||||
assertEquals("c", fused.get(0).identity());
|
||||
assertEquals("b", fused.get(1).identity());
|
||||
assertEquals("a", fused.get(2).identity());
|
||||
assertTrue(fused.get(0).rrfScore() > fused.get(2).rrfScore());
|
||||
}
|
||||
|
||||
@Test
|
||||
void pathWeightCanElevateSecondaryPath() {
|
||||
List<String> dense = List.of("a", "b");
|
||||
List<String> lexical = List.of("b", "a");
|
||||
|
||||
List<RrfFusion.Scored<String>> fused = RrfFusion.fuse(
|
||||
List.of(
|
||||
new RrfFusion.RankedPath<>("dense", dense, 1.0),
|
||||
new RrfFusion.RankedPath<>("lexical", lexical, 2.0)
|
||||
),
|
||||
60,
|
||||
s -> s
|
||||
);
|
||||
|
||||
assertEquals("b", fused.get(0).identity());
|
||||
}
|
||||
}
|
||||
+53
@@ -0,0 +1,53 @@
|
||||
package com.superbiz.agent.service.retrieval;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class VectorKnowledgeSearchAdapterHybridTest {
|
||||
|
||||
@Test
|
||||
void hybridFusesFilteredAndUnfilteredDensePaths() {
|
||||
VectorSearchService vectorSearchService = mock(VectorSearchService.class);
|
||||
when(vectorSearchService.searchSimilarDocuments("pool timeout", 3, null)).thenReturn(List.of(
|
||||
result("u1", "{\"_source\":\"a.md\",\"docId\":\"a\",\"chunkIndex\":0,\"title\":\"generic\"}", "generic pool", 0.4f),
|
||||
result("u2", "{\"_source\":\"b.md\",\"docId\":\"b\",\"chunkIndex\":0,\"title\":\"other\"}", "other", 0.5f)
|
||||
));
|
||||
when(vectorSearchService.searchSimilarDocuments("pool timeout", 3, "mysql")).thenReturn(List.of(
|
||||
result("f1", "{\"_source\":\"c.md\",\"docId\":\"c\",\"chunkIndex\":0,\"title\":\"mysql pool timeout\"}", "mysql pool timeout runbook", 0.35f)
|
||||
));
|
||||
|
||||
VectorKnowledgeSearchAdapter adapter = new VectorKnowledgeSearchAdapter(vectorSearchService, new ObjectMapper());
|
||||
ReflectionTestUtils.setField(adapter, "configuredMode", "hybrid");
|
||||
ReflectionTestUtils.setField(adapter, "rrfK", 60);
|
||||
ReflectionTestUtils.setField(adapter, "weightDenseUnfiltered", 1.0);
|
||||
ReflectionTestUtils.setField(adapter, "weightDenseFiltered", 1.0);
|
||||
ReflectionTestUtils.setField(adapter, "weightLexical", 1.0);
|
||||
|
||||
List<KnowledgeSearchHit> hits = adapter.search(
|
||||
new KnowledgeSearchRequest("pool timeout", 3, "mysql", KnowledgeSearchMode.HYBRID));
|
||||
|
||||
assertEquals(3, hits.size());
|
||||
assertTrue(hits.stream().anyMatch(hit -> "c#chunk-0".equals(hit.evidenceKey())));
|
||||
assertTrue(hits.get(0).metadata().containsKey("fusedScore"));
|
||||
}
|
||||
|
||||
private static VectorSearchService.SearchResult result(String id, String metadata, String content, float score) {
|
||||
VectorSearchService.SearchResult result = new VectorSearchService.SearchResult();
|
||||
result.setId(id);
|
||||
result.setMetadata(metadata);
|
||||
result.setContent(content);
|
||||
result.setScore(score);
|
||||
result.setRawScore((double) score);
|
||||
result.setScoreLabel("l2_distance");
|
||||
return result;
|
||||
}
|
||||
}
|
||||
@@ -10,6 +10,7 @@ import com.superbiz.agent.service.KnowledgeIndexService;
|
||||
import com.superbiz.agent.service.KnowledgeQueryTransformer;
|
||||
import com.superbiz.agent.service.LookupResultAssembler;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import com.superbiz.agent.service.retrieval.VectorKnowledgeSearchAdapter;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.mockito.Mock;
|
||||
@@ -45,15 +46,17 @@ class LookupKnowledgeToolTest {
|
||||
MockitoAnnotations.openMocks(this);
|
||||
|
||||
KnowledgeEvidencePostProcessor postProcessor = new KnowledgeEvidencePostProcessor();
|
||||
ReflectionTestUtils.setField(postProcessor, "returnN", 5);
|
||||
ReflectionTestUtils.setField(postProcessor, "maxChunksPerDocument", 2);
|
||||
KnowledgeContextPacker contextPacker = new KnowledgeContextPacker();
|
||||
tool = new LookupKnowledgeTool();
|
||||
ReflectionTestUtils.setField(tool, "queryTransformer", new KnowledgeQueryTransformer(knowledgeIndexService));
|
||||
ReflectionTestUtils.setField(tool, "documentRetriever",
|
||||
new KnowledgeDocumentRetriever(vectorSearchService, new ObjectMapper()));
|
||||
new KnowledgeDocumentRetriever(new VectorKnowledgeSearchAdapter(vectorSearchService, new ObjectMapper())));
|
||||
ReflectionTestUtils.setField(tool, "evidencePostProcessor", postProcessor);
|
||||
ReflectionTestUtils.setField(tool, "contextPacker", contextPacker);
|
||||
ReflectionTestUtils.setField(tool, "resultAssembler", new LookupResultAssembler());
|
||||
ReflectionTestUtils.setField(tool, "topK", 3);
|
||||
ReflectionTestUtils.setField(tool, "retrieveK", 3);
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -61,8 +64,7 @@ class LookupKnowledgeToolTest {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult result = searchResult(
|
||||
"vec-1",
|
||||
"db.md",
|
||||
"{\"_source\":\"db.md\",\"title\":\"Database Doc\",\"category\":\"database\"}",
|
||||
"{\"_source\":\"db.md\",\"docId\":\"db\",\"chunkIndex\":0,\"title\":\"Database Doc\",\"category\":\"database\"}",
|
||||
"mysql timeout runbook",
|
||||
0.2f);
|
||||
|
||||
@@ -75,6 +77,7 @@ class LookupKnowledgeToolTest {
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals(1, lookup.getEvidenceBlockCount());
|
||||
assertEquals("db.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
assertEquals("db#chunk-0", lookup.getEvidenceBlocks().get(0).getEvidenceKey());
|
||||
assertNotNull(lookup.getContextPack());
|
||||
assertTrue(lookup.getContextPack().getPackedText().contains("mysql timeout runbook"));
|
||||
assertEquals("FILTERED_VECTOR", lookup.getRetrievalTrace().getSelectedAttempt());
|
||||
@@ -89,14 +92,12 @@ class LookupKnowledgeToolTest {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult weak = searchResult(
|
||||
"weak",
|
||||
"weak.md",
|
||||
"{\"_source\":\"weak.md\",\"title\":\"Weak\"}",
|
||||
"{\"_source\":\"weak.md\",\"docId\":\"weak\",\"chunkIndex\":0,\"title\":\"Weak\"}",
|
||||
"weak candidate",
|
||||
1.4f);
|
||||
VectorSearchService.SearchResult strong = searchResult(
|
||||
"strong",
|
||||
"strong.md",
|
||||
"{\"_source\":\"strong.md\",\"title\":\"Strong\"}",
|
||||
"{\"_source\":\"strong.md\",\"docId\":\"strong\",\"chunkIndex\":0,\"title\":\"Strong\"}",
|
||||
"mysql timeout strong runbook",
|
||||
0.2f);
|
||||
|
||||
@@ -122,8 +123,7 @@ class LookupKnowledgeToolTest {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult strong = searchResult(
|
||||
"strong",
|
||||
"strong.md",
|
||||
"{\"_source\":\"strong.md\",\"title\":\"Strong\"}",
|
||||
"{\"_source\":\"strong.md\",\"docId\":\"strong\",\"chunkIndex\":0,\"title\":\"Strong\"}",
|
||||
"mysql timeout strong runbook",
|
||||
0.2f);
|
||||
|
||||
@@ -164,8 +164,7 @@ class LookupKnowledgeToolTest {
|
||||
void noL0HintUsesUnfilteredVectorSearch() {
|
||||
VectorSearchService.SearchResult result = searchResult(
|
||||
"vec-1",
|
||||
"perf.md",
|
||||
"{\"_source\":\"perf.md\",\"title\":\"Perf\"}",
|
||||
"{\"_source\":\"perf.md\",\"docId\":\"perf\",\"chunkIndex\":0,\"title\":\"Perf\"}",
|
||||
"performance tuning guide",
|
||||
0.3f);
|
||||
|
||||
@@ -187,14 +186,12 @@ class LookupKnowledgeToolTest {
|
||||
KnowledgeEntry entry = entry("payment.md", "Payment", "ERR_TIMEOUT", "payment");
|
||||
VectorSearchService.SearchResult first = searchResult(
|
||||
"a",
|
||||
"a.md",
|
||||
"{\"_source\":\"a.md\",\"title\":\"Generic\",\"category\":\"other\"}",
|
||||
"{\"_source\":\"a.md\",\"docId\":\"a\",\"chunkIndex\":0,\"title\":\"Generic\",\"category\":\"other\"}",
|
||||
"generic troubleshooting",
|
||||
0.4f);
|
||||
VectorSearchService.SearchResult second = searchResult(
|
||||
"b",
|
||||
"b.md",
|
||||
"{\"_source\":\"b.md\",\"title\":\"Payment ERR_TIMEOUT\",\"breadcrumb\":\"Payment > Timeout\",\"category\":\"payment\"}",
|
||||
"{\"_source\":\"b.md\",\"docId\":\"b\",\"chunkIndex\":0,\"title\":\"Payment ERR_TIMEOUT\",\"breadcrumb\":\"Payment > Timeout\",\"category\":\"payment\"}",
|
||||
"payment ERR_TIMEOUT timeout diagnosis",
|
||||
0.45f);
|
||||
|
||||
@@ -213,18 +210,16 @@ class LookupKnowledgeToolTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void deduplicatesEvidenceBlocksBySource() {
|
||||
void keepsDistinctChunksFromSameSource() {
|
||||
KnowledgeEntry entry = entry("shared.md", "Shared", "shared", "payment");
|
||||
VectorSearchService.SearchResult first = searchResult(
|
||||
"a",
|
||||
"shared.md",
|
||||
"{\"_source\":\"shared.md\",\"title\":\"Shared\"}",
|
||||
"{\"_source\":\"shared.md\",\"docId\":\"shared\",\"chunkIndex\":0,\"title\":\"Shared\"}",
|
||||
"shared content 1",
|
||||
0.2f);
|
||||
VectorSearchService.SearchResult second = searchResult(
|
||||
"b",
|
||||
"shared.md",
|
||||
"{\"_source\":\"shared.md\",\"title\":\"Shared\"}",
|
||||
"{\"_source\":\"shared.md\",\"docId\":\"shared\",\"chunkIndex\":1,\"title\":\"Shared\"}",
|
||||
"shared content 2",
|
||||
0.25f);
|
||||
|
||||
@@ -236,8 +231,11 @@ class LookupKnowledgeToolTest {
|
||||
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals(2, lookup.getEvidenceCandidateCount());
|
||||
assertEquals(1, lookup.getEvidenceBlockCount());
|
||||
assertEquals("shared.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
assertEquals(2, lookup.getEvidenceBlockCount());
|
||||
assertEquals("shared#chunk-0", lookup.getEvidenceBlocks().get(0).getEvidenceKey());
|
||||
assertEquals("shared#chunk-1", lookup.getEvidenceBlocks().get(1).getEvidenceKey());
|
||||
assertEquals("shared content 1", lookup.getEvidenceBlocks().get(0).getContent());
|
||||
assertEquals("shared content 2", lookup.getEvidenceBlocks().get(1).getContent());
|
||||
}
|
||||
|
||||
private KnowledgeEntry entry(String filePath, String title, String keyword, String category) {
|
||||
@@ -251,7 +249,6 @@ class LookupKnowledgeToolTest {
|
||||
}
|
||||
|
||||
private VectorSearchService.SearchResult searchResult(String id,
|
||||
String source,
|
||||
String metadata,
|
||||
String content,
|
||||
float score) {
|
||||
|
||||
Reference in New Issue
Block a user