feat(rag): modularize knowledge retrieval pipeline
This commit is contained in:
@@ -5,6 +5,7 @@
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
# Modular RAG Pipeline — Acceptance
|
||||
|
||||
## 验收状态
|
||||
|
||||
状态:通过,OpenSpec 已归档。
|
||||
|
||||
任务完成:
|
||||
|
||||
- OpenSpec tasks:31/31 完成。
|
||||
- Review 后新增去重边界修复和回归测试。
|
||||
- OpenSpec archive:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`。
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
```text
|
||||
Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
```text
|
||||
PASS
|
||||
```
|
||||
|
||||
说明:仅出现 Windows LF/CRLF warning,无 whitespace error。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
覆盖:
|
||||
|
||||
- filtered L1 成功不 retry。
|
||||
- filtered L1 低质量触发 raw unfiltered retry。
|
||||
- filtered L1 无 evidence 触发 raw unfiltered retry。
|
||||
- L0 hint 不作为 standalone fact evidence。
|
||||
- 无 L0 hint 时直接 unfiltered vector search。
|
||||
- rerank 使用 hint match 并记录 trace。
|
||||
- context pack 保留 source/title/breadcrumb/hit reasons。
|
||||
- evidence blocks 按 source 去重。
|
||||
- session dedup 不再返回可消费 evidence/context。
|
||||
- recorder 记录 retrieval trace、rerank trace、context pack summary 和 evidence summaries。
|
||||
|
||||
```powershell
|
||||
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
说明:
|
||||
|
||||
- `MilvusConnectionTest` 需要 `MILVUS_TOKEN` 环境变量,直接读 `System.getenv`,不会自动读 `application.yml`。
|
||||
- 注入该环境变量后完整测试通过。
|
||||
|
||||
## 实现验收
|
||||
|
||||
已验证行为:
|
||||
|
||||
- `LookupKnowledgeTool` 已变为 pipeline orchestrator。
|
||||
- `LookupResult` 新契约包含 `evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
|
||||
- 旧 `primary/supplement` 字段和 DTO 已删除。
|
||||
- `ToolInvocationRecorder` 不再依赖 `result.getPrimary()`。
|
||||
- filtered retrieval 失败时会记录 `filtered_vector_no_evidence` 或 `filtered_vector_low_quality`。
|
||||
- no-evidence 情况返回 `found=false` 且保留 retrieval trace。
|
||||
- session dedup 情况返回 `found=false` 且 evidence/context 为空。
|
||||
|
||||
## 未验证项
|
||||
|
||||
人工 Demo 未执行:
|
||||
|
||||
- 还没有通过真实 Chat/AIOps 会话观察 Agent 是否稳定按 `contextPack.packedText` 和 `evidenceBlocks` 引用证据。
|
||||
|
||||
风险:
|
||||
|
||||
- 工具 JSON 契约是 L4 breaking change,prompt 已更新,但真实对话行为仍建议做一次端到端 demo。
|
||||
|
||||
## 后续建议
|
||||
|
||||
- 增加一组 RAG eval cases,固定 query、期望 evidence source、期望 fallback path。
|
||||
- 将 `MilvusConnectionTest` 改成 Spring 配置驱动或 integration profile,避免配置源混用。
|
||||
- 后续可在评测数据足够后再考虑 model-based rerank 或 hybrid retrieval。
|
||||
@@ -0,0 +1,54 @@
|
||||
# Modular RAG Pipeline — Brief
|
||||
|
||||
## 背景
|
||||
|
||||
`lookup_knowledge` 已经能返回知识库证据,但实现集中在 `LookupKnowledgeTool` 内部:L0 查询分析、L1 向量召回、相关性归一化、证据组装、会话去重和 trace 入库耦合在一起。
|
||||
|
||||
旧返回契约 `primary/supplement` 也延续了“L0 是主结果、L1 是补充”的语义,和当前设计目标不一致。新的目标是让 L0 只作为 query understanding / filter / rerank / trace 信号,让 L1 向量检索成为事实证据来源。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将 `lookup_knowledge` 改造成模块化 RAG pipeline。
|
||||
- 保留显式 Agent tool 边界,不改工具名和 query 参数。
|
||||
- L0 只提供领域、关键词、实体、category filter 和 trace hint。
|
||||
- L1 filtered vector retrieval 失败或低质量时,降级为 raw query unfiltered L1 retry。
|
||||
- 输出 evidence-first contract:`evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
|
||||
- 保持 `tool_invocation` 表结构稳定,把新 trace 写入 `retrieval_details` JSON。
|
||||
|
||||
## 范围
|
||||
|
||||
已完成:
|
||||
|
||||
- 新增 pipeline DTO:`KnowledgeQuery`、`RetrievedEvidenceCandidate`、`ContextPack`、`RetrievalTrace`、`RerankTrace`、`EvidencePostprocessResult`。
|
||||
- 新增 pipeline service:`KnowledgeQueryTransformer`、`KnowledgeDocumentRetriever`、`KnowledgeEvidencePostProcessor`、`KnowledgeContextPacker`、`LookupResultAssembler`。
|
||||
- 重构 `LookupKnowledgeTool` 为薄 orchestration 层。
|
||||
- 迁移 `LookupResult`,删除 `primary/supplement` 字段和 `PrimaryResult` / `SupplementResult` 类。
|
||||
- 更新 `ToolInvocationRecorder`,记录 query transform、retrieval trace、context pack summary、rerank trace、fallback reason 和 evidence summaries。
|
||||
- 更新 executor prompt 和 RAG 架构文档。
|
||||
- 补充 lookup、recorder、fallback、rerank、context pack、session dedup 测试。
|
||||
|
||||
非目标:
|
||||
|
||||
- 不引入 implicit Advisor。
|
||||
- 不引入 cross-encoder、BM25、RRF、Elasticsearch、OpenSearch。
|
||||
- 不改文档上传、chunk、embedding 写入、Milvus schema。
|
||||
- 不改变 Agent 何时调用 `lookup_knowledge`。
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
- `openspec/changes/archive/2026-07-06-modular-rag-pipeline`
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4 breaking interface。
|
||||
|
||||
原因:
|
||||
|
||||
- 删除旧 `LookupResult.primary` / `LookupResult.supplement`。
|
||||
- `lookup_knowledge` tool JSON 输出形状变化。
|
||||
|
||||
缓解:
|
||||
|
||||
- 工具名和输入参数保持不变。
|
||||
- in-repo 消费方、测试和 prompt 同步迁移。
|
||||
- `tool_invocation` 表结构不变。
|
||||
@@ -0,0 +1,72 @@
|
||||
# Modular RAG Pipeline — Decisions
|
||||
|
||||
## D1: `lookup_knowledge` 保持显式工具
|
||||
|
||||
不把知识检索做成隐式 Advisor。Agent 仍显式调用 `lookup_knowledge(query)`,这样 trace、Verifier、Eval 都能看到工具调用边界。
|
||||
|
||||
## D2: L0 只做 query understanding
|
||||
|
||||
L0 产出:
|
||||
|
||||
- `domainHints`
|
||||
- `matchedKeywords`
|
||||
- `entities`
|
||||
- `categoryFilter`
|
||||
- `l0Titles`
|
||||
- `l0MatchCount`
|
||||
|
||||
L0 不再直接转成 fact evidence。L0 hint 可以影响 filter、rerank、trace,但不能在 L1 失败时冒充知识证据。
|
||||
|
||||
## D3: MVP 降级策略采用 unfiltered L1 retry
|
||||
|
||||
流程:
|
||||
|
||||
```text
|
||||
filtered L1 with L0 category filter
|
||||
-> empty / no final evidence / below reference threshold
|
||||
-> raw query unfiltered L1 retry
|
||||
-> still no evidence => no_evidence
|
||||
```
|
||||
|
||||
取舍:
|
||||
|
||||
- 简单、可解释、适合 MVP。
|
||||
- 避免引入 BM25/RRF/multi-query/cross-encoder 的复杂度。
|
||||
- 代价是低质量场景多一次向量查询,已通过 trace 记录 attempt duration。
|
||||
|
||||
## D4: 删除 `primary/supplement`
|
||||
|
||||
这是一次 L4 breaking interface change。
|
||||
|
||||
删除原因:
|
||||
|
||||
- `primary/supplement` 绑定旧语义:L0 primary、L1 supplement。
|
||||
- 新设计中事实证据来自 `evidenceBlocks/contextPack`。
|
||||
|
||||
迁移结果:
|
||||
|
||||
- `LookupResult` 暴露 evidence-first 字段。
|
||||
- `PrimaryResult` / `SupplementResult` 已删除。
|
||||
- 生产代码和测试不再引用 `getPrimary()` / `getSupplement()`。
|
||||
|
||||
## D5: Trace 表结构保持稳定
|
||||
|
||||
`tool_invocation` 表不新增列。新增信息写入 `retrieval_details` JSON:
|
||||
|
||||
- `query_transform`
|
||||
- `retrieval_trace`
|
||||
- `context_pack_summary`
|
||||
- `rerank_trace`
|
||||
- `fallback_reason`
|
||||
- `evidence_blocks`
|
||||
|
||||
原因:当前 trace、Verifier、Eval 已经以 `tool_invocation` 为证据入口,JSON details 足够承载 RAG 细节,避免 schema churn。
|
||||
|
||||
## D6: 会话去重不返回可消费证据
|
||||
|
||||
Review 后修正:
|
||||
|
||||
- dedup result 的 `found=false` 必须和 evidence/context 语义一致。
|
||||
- 返回消息说明文档已检索过。
|
||||
- 不再返回 `evidenceBlocks/contextPack`,避免 Agent 重复使用同一证据。
|
||||
- 保留 `retrievalTrace` 和 `retrievedDomainsThisSession` 便于可观测。
|
||||
@@ -0,0 +1,56 @@
|
||||
# Modular RAG Pipeline — Evidence
|
||||
|
||||
## 代码证据
|
||||
|
||||
关键入口:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
|
||||
新增模块:
|
||||
|
||||
- `KnowledgeQueryTransformer`:复用 `KnowledgeIndexService.analyzeQuery`,把 L0 转成 query hints 和可选 `categoryFilter`。
|
||||
- `KnowledgeDocumentRetriever`:封装 `VectorSearchService.searchSimilarDocuments(query, topK, category)`,统一 filtered / unfiltered attempt。
|
||||
- `KnowledgeEvidencePostProcessor`:归一化 L2、创建 evidence blocks、source dedup、规则 rerank、输出 `RerankTrace`。
|
||||
- `KnowledgeContextPacker`:按字符预算打包 evidence,保留 source/title/breadcrumb/hit reasons。
|
||||
- `LookupResultAssembler`:统一组装 evidence-first result、no-evidence result、session dedup result。
|
||||
|
||||
## 设计证据
|
||||
|
||||
已有文档约束:
|
||||
|
||||
- `mvp/architecture/rag-architecture.md`:RAG 应表达为可解释 pipeline,而不是一坨工具逻辑。
|
||||
- `mvp/architecture/retrieval-observability.md`:L0 是 hint/explainability 层,L1 是语义检索主路径。
|
||||
- `devflow/glossary/CONTEXT.md`:`lookup_knowledge` 是显式 Agent evidence tool,`tool_invocation` 是 trace / verifier / eval 的证据来源。
|
||||
|
||||
OpenSpec 对齐:
|
||||
|
||||
- `openspec/changes/modular-rag-pipeline/proposal.md`
|
||||
- `openspec/changes/modular-rag-pipeline/design.md`
|
||||
- `openspec/changes/modular-rag-pipeline/specs/rag-knowledge-retrieval/spec.md`
|
||||
- `openspec/changes/modular-rag-pipeline/tasks.md`
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 一次到位做模块化 RAG,而不是只做小补丁。
|
||||
- L0 不再作为事实证据兜底。
|
||||
- filtered L1 不准时,MVP 降级为 raw query unfiltered L1 retry。
|
||||
- 可以新增字段,并删除旧字段以换取后续流程清晰。
|
||||
|
||||
## Review 发现
|
||||
|
||||
Review 中发现一个非阻塞但应修复的问题:
|
||||
|
||||
- 会话去重命中时,返回 `found=false` 但仍带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费同一份证据。
|
||||
|
||||
修复:
|
||||
|
||||
- `LookupResultAssembler.deduped` 清空可消费 evidence/context,只保留 message、trace、relevance hint 和 session domain memory。
|
||||
- 新增 `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
|
||||
|
||||
## 非阻塞观察
|
||||
|
||||
- `MilvusConnectionTest` 仍直接依赖 `MILVUS_TOKEN` 环境变量;主配置中已有 token,但测试不读 Spring 配置。
|
||||
- 测试日志仍有 ANTLR 版本 warning,不影响测试通过。
|
||||
- 控制台在部分命令输出中仍会出现中文编码显示问题,但源码按 UTF-8 读取时关键用户提示文本正常。
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -17,6 +17,7 @@
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||
@@ -35,8 +36,9 @@ SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
7. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
8. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
9. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
10. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
11. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
# 模块化 RAG Pipeline 架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前已实现架构
|
||||
**关联 OpenSpec**:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文记录 `lookup_knowledge` 的当前模块化 RAG 实现。它是 [rag-architecture.md](rag-architecture.md) 的落地版,重点说明代码模块、数据契约、降级策略和可观测性边界。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 保留显式 Agent Tool:`lookup_knowledge(query)`。
|
||||
- L0 只作为 query understanding / filter / rerank / trace hint。
|
||||
- L1 向量检索作为事实证据来源。
|
||||
- filtered L1 低质量时,降级为 raw query unfiltered L1 retry。
|
||||
- 输出 evidence-first contract,替代旧 `primary/supplement`。
|
||||
|
||||
## 2. 当前链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Agent["Executor Agent"] --> Tool["LookupKnowledgeTool.lookupKnowledge(query)"]
|
||||
|
||||
Tool --> Transform["KnowledgeQueryTransformer"]
|
||||
Transform --> KQ["KnowledgeQuery"]
|
||||
|
||||
KQ --> Retriever["KnowledgeDocumentRetriever"]
|
||||
Retriever --> Attempt1["FILTERED_VECTOR or UNFILTERED_VECTOR"]
|
||||
Attempt1 --> Post1["KnowledgeEvidencePostProcessor"]
|
||||
Post1 --> Quality{"usable evidence?"}
|
||||
Quality -->|yes| Pack
|
||||
Quality -->|no and categoryFilter exists| Retry["UNFILTERED_VECTOR_RETRY"]
|
||||
Retry --> Post2["KnowledgeEvidencePostProcessor"]
|
||||
Post2 --> Pack["KnowledgeContextPacker"]
|
||||
|
||||
Pack --> Assemble["LookupResultAssembler"]
|
||||
Assemble --> Result["LookupResult"]
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Trace["tool_invocation.retrieval_details"]
|
||||
Result --> Agent
|
||||
```
|
||||
|
||||
对应代码:
|
||||
|
||||
| 阶段 | 类 | 职责 |
|
||||
|---|---|---|
|
||||
| Tool Boundary | `LookupKnowledgeTool` | 接收 Agent 工具调用,编排 pipeline,处理 session dedup 和 recorder |
|
||||
| Query Transformation | `KnowledgeQueryTransformer` | 复用 L0,输出 query hints 和可选 category filter |
|
||||
| Retrieval | `KnowledgeDocumentRetriever` | 调用 `VectorSearchService`,统一 filtered / unfiltered attempt |
|
||||
| Post-Retrieval | `KnowledgeEvidencePostProcessor` | L2 归一化、证据块构建、source dedup、规则 rerank |
|
||||
| Context Packing | `KnowledgeContextPacker` | 按字符预算打包 Agent 可消费 context |
|
||||
| Result Assembly | `LookupResultAssembler` | 统一 evidence result、no-evidence result、dedup result |
|
||||
| Observability | `ToolInvocationRecorder` | 写入 query transform、retrieval trace、rerank trace、context pack summary |
|
||||
|
||||
## 3. L0 与 L1 边界
|
||||
|
||||
L0 来源于 `KnowledgeIndexService.analyzeQuery`,输出进入 `KnowledgeQuery`:
|
||||
|
||||
```text
|
||||
originalQuery
|
||||
rewrittenQuery
|
||||
domainHints
|
||||
matchedKeywords
|
||||
entities
|
||||
categoryFilter
|
||||
l0Titles
|
||||
l0MatchCount
|
||||
```
|
||||
|
||||
L0 可以做:
|
||||
|
||||
- 给 L1 提供单一 category filter。
|
||||
- 给 rerank 提供 domain / keyword / entity boost 信号。
|
||||
- 给 trace 提供解释信息。
|
||||
|
||||
L0 不再做:
|
||||
|
||||
- 不因唯一命中直接返回文档正文。
|
||||
- 不在 L1 无结果时作为事实证据兜底。
|
||||
- 不进入 `evidenceBlocks`,除非未来明确引入新的 evidence source 规则。
|
||||
|
||||
L1 通过 `VectorSearchService.searchSimilarDocuments(query, topK, category)` 执行,内部仍保留 Spring AI VectorStore 优先和 Milvus SDK fallback。
|
||||
|
||||
## 4. 降级策略
|
||||
|
||||
MVP 降级策略保持简单:
|
||||
|
||||
```text
|
||||
if categoryFilter exists:
|
||||
run FILTERED_VECTOR
|
||||
if empty / no evidence / top similarity < referenceThreshold:
|
||||
run UNFILTERED_VECTOR_RETRY with original query
|
||||
else:
|
||||
run UNFILTERED_VECTOR
|
||||
```
|
||||
|
||||
fallback reason:
|
||||
|
||||
| reason | 含义 |
|
||||
|---|---|
|
||||
| `filtered_vector_no_evidence` | filtered L1 无候选或 post-processing 后无 evidence |
|
||||
| `filtered_vector_low_quality` | filtered L1 有候选,但 top normalized similarity 低于 `retrieval.normalization.reference-threshold` |
|
||||
|
||||
当 retry 后仍无证据:
|
||||
|
||||
- `found=false`
|
||||
- `evidenceStatus=no_evidence`
|
||||
- 保留 `retrievalTrace`
|
||||
- 不返回 L0 文档作为事实证据
|
||||
|
||||
## 5. Evidence-First Contract
|
||||
|
||||
`LookupResult` 当前核心字段:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
旧字段已删除:
|
||||
|
||||
```text
|
||||
primary
|
||||
supplement
|
||||
```
|
||||
|
||||
这是一项 L4 breaking interface change。项目内已同步迁移:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `ToolInvocationRecorder`
|
||||
- executor prompts
|
||||
- lookup / recorder tests
|
||||
- RAG architecture docs
|
||||
- OpenSpec 主 spec
|
||||
|
||||
## 6. Trace 结构
|
||||
|
||||
`retrieval_details` 保持 JSON 扩展,不改表结构。关键内容:
|
||||
|
||||
```json
|
||||
{
|
||||
"query_transform": {},
|
||||
"retrieval_trace": {
|
||||
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallback_reason": "filtered_vector_no_evidence",
|
||||
"attempts": []
|
||||
},
|
||||
"context_pack_summary": {},
|
||||
"rerank_trace": {},
|
||||
"evidence_blocks": []
|
||||
}
|
||||
```
|
||||
|
||||
这样 Trace API、Verifier、Eval 可以继续从 `tool_invocation` 读取证据链。
|
||||
|
||||
## 7. Review 修正
|
||||
|
||||
归档前 review 发现:session dedup 命中时返回 `found=false`,但仍携带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费证据。
|
||||
|
||||
当前行为已修正:
|
||||
|
||||
- dedup result 不再返回可消费 evidence/context。
|
||||
- 保留 message、retrieval trace、relevance hint 和 retrieved domains。
|
||||
- 测试覆盖:`LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
|
||||
|
||||
## 8. 验证
|
||||
|
||||
已执行:
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
|
||||
openspec validate --all --strict
|
||||
git diff --check
|
||||
```
|
||||
|
||||
结果:全部通过。
|
||||
|
||||
注意:
|
||||
|
||||
- `MilvusConnectionTest` 直接读 `MILVUS_TOKEN` 环境变量,不读 Spring 配置。
|
||||
- 完整测试需要在 Maven 进程里注入该环境变量。
|
||||
|
||||
## 9. 后续演进
|
||||
|
||||
建议后续按评测结果推进,而不是先堆复杂能力:
|
||||
|
||||
- 增加 RAG eval cases:固定 query、期望 source、期望 fallback path。
|
||||
- 引入更严格的 evidence grounding 检查。
|
||||
- 当规则 rerank 不足时,再考虑 model-based rerank。
|
||||
- 当召回覆盖率不足时,再考虑 BM25/RRF/hybrid retrieval。
|
||||
@@ -1,6 +1,6 @@
|
||||
# RAG 新架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
||||
|
||||
@@ -44,9 +44,13 @@ flowchart TD
|
||||
SdkFallback --> Results
|
||||
SdkOnly --> Results
|
||||
|
||||
Results --> Normalize["relevance normalization"]
|
||||
Normalize --> Dedup["session dedup: RetrievedDocTracker"]
|
||||
Dedup --> Output["LookupResult"]
|
||||
Results --> Retry{"filtered result usable?"}
|
||||
Retry -->|no| RetryL1["raw query unfiltered L1 retry"]
|
||||
Retry -->|yes| Post["post-retrieval processing"]
|
||||
RetryL1 --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Dedup["session dedup: RetrievedDocTracker"]
|
||||
Dedup --> Output["LookupResult: evidenceBlocks / contextPack / traces"]
|
||||
Output --> Record["tool_invocation record"]
|
||||
Output --> Agent
|
||||
```
|
||||
@@ -66,10 +70,13 @@ Agent Executor
|
||||
-> Spring AI VectorStore only
|
||||
-> mode=sdk
|
||||
-> Milvus SDK only
|
||||
-> result normalization
|
||||
-> post-retrieval processing
|
||||
-> relevanceLevel
|
||||
-> completenessHint
|
||||
-> score/rawScore/scoreLabel
|
||||
-> evidenceBlocks
|
||||
-> rerankTrace
|
||||
-> context packing
|
||||
-> contextPack
|
||||
-> session dedup
|
||||
-> RetrievedDocTracker
|
||||
-> tool_invocation record
|
||||
@@ -180,10 +187,11 @@ L0 负责:
|
||||
- metadata/category filter candidate
|
||||
- trace 中的 hit reason
|
||||
|
||||
L0 不再默认负责:
|
||||
L0 不再负责:
|
||||
|
||||
```text
|
||||
L0 unique hit -> 直接作为最终检索结果
|
||||
L1 no result -> 返回 L0 文档作为事实证据
|
||||
```
|
||||
|
||||
当前职责是:
|
||||
@@ -192,8 +200,10 @@ L0 unique hit -> 直接作为最终检索结果
|
||||
query / AIOps payload
|
||||
-> L0 matched keywords / domains / entities
|
||||
-> category filter candidate
|
||||
-> L1 semantic retrieval
|
||||
-> relevance normalization
|
||||
-> filtered L1 semantic retrieval
|
||||
-> low-quality? raw query unfiltered L1 retry
|
||||
-> post-retrieval processing
|
||||
-> context packing
|
||||
```
|
||||
|
||||
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
|
||||
@@ -319,33 +329,24 @@ AIOps payload
|
||||
|
||||
## 9. Evidence 与去重
|
||||
|
||||
当前 evidence 输出仍以 `LookupResult` 和工具返回文本为主,已经具备:
|
||||
当前 evidence 输出已从旧 `primary/supplement` 迁移为 evidence-first contract,核心字段包括:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层记录。
|
||||
- relevance level。
|
||||
- completeness hint。
|
||||
- session 级文档去重。
|
||||
- domain 行动记忆。
|
||||
- `tool_invocation` 明细记录。
|
||||
- `evidenceBlocks`
|
||||
- `contextPack`
|
||||
- `retrievalTrace`
|
||||
- `rerankTrace`
|
||||
- `relevanceLevel`
|
||||
- `completenessHint`
|
||||
- `retrievedDomainsThisSession`
|
||||
- `tool_invocation.retrieval_details`
|
||||
|
||||
后续更完整的 evidence block 目标:
|
||||
evidence block 结构:
|
||||
|
||||
```text
|
||||
source
|
||||
docId
|
||||
chunkIndex
|
||||
title
|
||||
breadcrumb
|
||||
score
|
||||
rawScore
|
||||
scoreLabel
|
||||
hitReason
|
||||
content
|
||||
expandedFrom
|
||||
source / title / breadcrumb / retrievalLayer / content / score / hitReasons
|
||||
```
|
||||
|
||||
这部分应作为下一阶段增强,而不是当前已完全完成能力。
|
||||
context pack 会按重排后的证据顺序生成 Agent 可消费的紧凑上下文,并保留 included/omitted sources 供 trace 检查。
|
||||
|
||||
## 10. 评测与验收
|
||||
|
||||
@@ -380,18 +381,18 @@ RAG 架构变更必须先过评测,再认为可合入主链路。
|
||||
- Markdown chunk 保留 `title` 和 `breadcrumb`。
|
||||
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
|
||||
- AIOps payload 生成推荐知识库 query。
|
||||
- `tool_invocation` 记录 relevance level 和 dedup reason。
|
||||
- `tool_invocation` 记录 relevance level、dedup reason、evidence summaries、retrieval trace、rerank trace 和 context pack summary。
|
||||
- `lookup_knowledge` 输出使用 evidence-first contract,不再暴露旧 `primary/supplement` 字段。
|
||||
- RAG offline baseline 和 live acceptance 脚本已补齐。
|
||||
|
||||
## 12. 后续演进
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 完整 evidence block 结构化输出。
|
||||
2. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||
3. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||
4. Query Transformer / MultiQuery 的可回退接入。
|
||||
5. VectorStore 写入路径评估。
|
||||
1. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||
3. Query Transformer / MultiQuery 的可回退接入。
|
||||
4. VectorStore 写入路径评估。
|
||||
|
||||
暂不优先:
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 检索与可观测性架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
|
||||
|
||||
@@ -13,7 +13,7 @@
|
||||
- 检索结果如何归一化、去重、记录。
|
||||
- 如何通过 trace 和 eval 判断检索质量。
|
||||
|
||||
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1,也不在 L1 失败时作为事实证据兜底。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||
|
||||
## 2. 检索总图
|
||||
|
||||
@@ -35,9 +35,13 @@ flowchart TD
|
||||
|
||||
Spring --> Candidates["L1 candidates"]
|
||||
SDK --> Candidates
|
||||
Candidates --> Normalize["relevance normalization"]
|
||||
L0Result --> Normalize
|
||||
Normalize --> Result["LookupResult"]
|
||||
Candidates --> Quality{"filtered L1 usable?"}
|
||||
Quality -->|no| Retry["raw query unfiltered L1 retry"]
|
||||
Quality -->|yes| Post["post-retrieval processing"]
|
||||
Retry --> Post
|
||||
L0Result --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Result["LookupResult evidenceBlocks/contextPack/traces"]
|
||||
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Final["final tool output"]
|
||||
@@ -69,6 +73,7 @@ singleDomainOrNull
|
||||
|
||||
```text
|
||||
matches=1 -> skip L1 -> 直接返回 L0 文档正文
|
||||
L1 无可用证据 -> 返回 L0 文档正文
|
||||
```
|
||||
|
||||
原因:
|
||||
@@ -123,12 +128,12 @@ SDK fallback 保留的价值:
|
||||
| `rawScore` | 底层检索实现原始分数 |
|
||||
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
|
||||
|
||||
工具层再把 L0/L1 情况归一为:
|
||||
post-retrieval 层再把检索候选归一为:
|
||||
|
||||
| relevanceLevel | 含义 |
|
||||
|---|---|
|
||||
| `PRECISE` | L0 单命中且 L1 相似度高 |
|
||||
| `HIGHLY_RELEVANT` | L1 相似度高,或 L0 多命中且 L1 支撑强 |
|
||||
| `PRECISE` | L1 相似度高且 query hint 与候选证据互相支撑 |
|
||||
| `HIGHLY_RELEVANT` | L1 相似度高 |
|
||||
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
|
||||
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
|
||||
|
||||
@@ -169,7 +174,7 @@ title + breadcrumb + content
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LookupResult["LookupResult"] --> Agent["Agent context"]
|
||||
LookupResult["LookupResult: evidenceBlocks/contextPack/traces"] --> Agent["Agent context"]
|
||||
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
@@ -195,10 +200,13 @@ success
|
||||
`retrieval_details` 承载更细信息,例如:
|
||||
|
||||
- L0 命中文档标题和路径。
|
||||
- L1 分数。
|
||||
- L1 attempts、fallback reason、分数和 similarity。
|
||||
- retrieved domains。
|
||||
- evidence status。
|
||||
- dedup reason。
|
||||
- evidence block summaries。
|
||||
- context pack summary。
|
||||
- rerank trace。
|
||||
|
||||
## 8. 去重与行动记忆
|
||||
|
||||
@@ -252,15 +260,15 @@ trace inspection
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 完整 evidence block 输出。
|
||||
2. 邻居 chunk / 同章节上下文扩展。
|
||||
3. metadata taxonomy 清理。
|
||||
4. Query Transformer / MultiQuery 可回退接入。
|
||||
5. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||
1. 邻居 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理。
|
||||
3. Query Transformer / MultiQuery 可回退接入。
|
||||
4. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 重新引入 L0 直接返回。
|
||||
- 重新引入 L0 文档作为 L1 失败时的事实证据兜底。
|
||||
- 一次性迁移所有写入路径。
|
||||
- 在没有评测收益前引入 rerank / RRF / BM25。
|
||||
- 在没有评测收益前引入模型 rerank / RRF / BM25。
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,396 @@
|
||||
# Modular RAG Pipeline Decisions
|
||||
|
||||
## Discover Summary
|
||||
|
||||
- Capability source: sm-flow Discover using local repository evidence and existing OpenSpec/devflow context.
|
||||
- Slug: `modular-rag-pipeline`.
|
||||
- Scale: standard.
|
||||
- Goal: turn `lookup_knowledge` into a modular RAG pipeline suitable for Agent engineering interview use, while keeping the explicit Agent tool and evidence trace.
|
||||
|
||||
## Context Evidence
|
||||
|
||||
### Existing Architecture
|
||||
|
||||
- `mvp/architecture/rag-architecture.md` documents the desired boundary: mature framework retrieval plus business-observable orchestration.
|
||||
- `mvp/architecture/retrieval-observability.md` says L0 is a hint/explainability layer and L1 semantic retrieval is the default recall path.
|
||||
- `VectorSearchService` is already the retrieval facade and supports Spring AI `VectorStore` with SDK fallback.
|
||||
- `LookupKnowledgeTool` currently still owns query analysis, L1 invocation, relevance normalization, result assembly, evidence block construction, session dedup, and recorder calls.
|
||||
|
||||
### Existing Evidence Blocks
|
||||
|
||||
- `EvidenceBlock` already exists.
|
||||
- `LookupResult` already has `evidenceBlocks`, `evidenceCandidateCount`, and `evidenceBlockCount`.
|
||||
- `LookupKnowledgeTool` currently builds evidence blocks internally.
|
||||
- `ToolInvocationRecorder` already persists compact evidence block summaries in `retrieval_details`.
|
||||
|
||||
Conclusion: evidence blocks are partially implemented, but the post-retrieval module boundary is not.
|
||||
|
||||
### Current Gaps
|
||||
|
||||
- Context packing is not a first-class module.
|
||||
- Rerank is not a first-class module; only original semantic rank and trace summary sorting exist.
|
||||
- The `primary` / `supplement` result model still encodes old L0/L1 semantics.
|
||||
- `lookup_knowledge` tool description still describes the old two-stage retrieval model.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Question | Mode | Status |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Q1 | Boundary | Should this be internal-only refactor or allow return contract changes? | user-interview | confirmed |
|
||||
| Q2 | Fallback | If filtered L1 fails, should L0 provide weak fallback evidence? | user-interview | confirmed |
|
||||
| Q3 | Interface | Can `LookupResult` add new fields and remove old `primary` / `supplement` if simpler? | user-interview | confirmed |
|
||||
| Q4 | Architecture | Does current repo already have evidence blocks, context packing, and rerank? | evidence-driven | resolved |
|
||||
| Q5 | Compatibility | Which in-repo consumers reference `primary` / `supplement`? | evidence-driven | resolved |
|
||||
| Q6 | Commit detail | What are the exact fields and thresholds for traces/context pack/low quality? | user-interview or specify | pending |
|
||||
|
||||
## Confirmed User Decisions
|
||||
|
||||
### D1: Prefer one-shot modular RAG refactor
|
||||
|
||||
User confirmed that the change can be done "一次到位" instead of only doing a compatibility-preserving internal refactor.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Create full pipeline modules now.
|
||||
- Do not leave `LookupKnowledgeTool` as a large procedural class.
|
||||
|
||||
### D2: L0 is not a normal evidence retrieval path
|
||||
|
||||
User challenged the first design because it made L0 participate in too many flows.
|
||||
|
||||
Confirmed direction:
|
||||
|
||||
```text
|
||||
L0 -> query understanding / category filter / domain/entity/keyword signal
|
||||
L1 -> main vector retrieval
|
||||
```
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Do not model L0 and L1 as equal retrievers in the normal path.
|
||||
- L0 output can influence filter, rerank, and trace.
|
||||
|
||||
### D3: MVP fallback is unfiltered L1 retry
|
||||
|
||||
User proposed a simpler MVP fallback:
|
||||
|
||||
```text
|
||||
Filtered L1 using L0 category filter
|
||||
-> if inaccurate or empty
|
||||
-> retry raw query with L1 and no L0 filter
|
||||
```
|
||||
|
||||
Confirmed direction:
|
||||
|
||||
- Use filtered vector retrieval first when L0 provides an unambiguous category.
|
||||
- If filtered retrieval is low-quality, retry unfiltered vector retrieval with raw query.
|
||||
- Do not return L0 documents as fact evidence fallback in this MVP design.
|
||||
|
||||
### D4: Result contract can change
|
||||
|
||||
User confirmed new fields can be added and old fields can be removed if the later flow becomes cleaner.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
|
||||
- Preferred contract becomes `evidenceBlocks + contextPack + retrievalTrace + rerankTrace`.
|
||||
- This is a breaking interface change and must be treated as L4.
|
||||
|
||||
## Evidence-Driven Findings
|
||||
|
||||
### E1: Existing spec conflict
|
||||
|
||||
`openspec/specs/rag-knowledge-retrieval/spec.md` currently says L1 no-result or failure should return an L0-based primary result. This conflicts with the confirmed design.
|
||||
|
||||
Required OpenSpec update:
|
||||
|
||||
- Replace L0 primary fallback with unfiltered vector retry.
|
||||
- Define no-evidence behavior when both filtered and unfiltered L1 fail.
|
||||
|
||||
### E2: Existing compatibility-field requirement conflict
|
||||
|
||||
`openspec/specs/rag-knowledge-retrieval/spec.md` currently requires `primary` and `supplement` compatibility fields to remain when evidence blocks exist.
|
||||
|
||||
Required OpenSpec update:
|
||||
|
||||
- Remove compatibility-field requirement.
|
||||
- Define evidence blocks and context pack as the preferred tool result contract.
|
||||
|
||||
### E3: Primary/supplement references are localized
|
||||
|
||||
Search found concrete Java references in:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeToolTest`
|
||||
- `ToolInvocationRecorderTest`
|
||||
- `LookupResult`
|
||||
- `PrimaryResult`
|
||||
- `SupplementResult`
|
||||
|
||||
No broad in-repo service usage was found beyond tool implementation, recorder, tests, prompts, and historical docs.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- One-shot migration is feasible if tests and prompts are updated in the same change.
|
||||
|
||||
### E4: Existing trace contract must be preserved
|
||||
|
||||
`tool_invocation` is used by diagnosis trace, verifier, and evaluation code. The database table does not need to change for this design if new details remain inside `retrieval_details`.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Keep table-level fields stable.
|
||||
- Enrich JSON `retrieval_details` with `retrieval_trace`, `rerank_trace`, `context_pack_summary`, and fallback reason.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L4 breaking interface.
|
||||
|
||||
Reason:
|
||||
|
||||
- Removes or changes old result fields consumed by current tests and possibly by Agent prompt behavior.
|
||||
- Changes `lookup_knowledge` tool JSON shape.
|
||||
|
||||
Mitigation:
|
||||
|
||||
- Keep tool name and input signature unchanged.
|
||||
- Update all in-repo consumers in the same change.
|
||||
- Keep `tool_invocation` table schema stable.
|
||||
- Add tests for the new result contract.
|
||||
- Update prompt text to teach Agent to use `contextPack` and `evidenceBlocks`.
|
||||
|
||||
## Proposed Implementation Shape
|
||||
|
||||
Pipeline classes:
|
||||
|
||||
- `KnowledgeQueryTransformer`
|
||||
- `KnowledgeDocumentRetriever`
|
||||
- `KnowledgeEvidencePostProcessor`
|
||||
- `KnowledgeContextPacker`
|
||||
- `LookupResultAssembler`
|
||||
|
||||
New or updated DTOs:
|
||||
|
||||
- `KnowledgeQuery`
|
||||
- `RetrievedEvidenceCandidate`
|
||||
- `EvidencePostprocessResult`
|
||||
- `ContextPack`
|
||||
- `RetrievalTrace`
|
||||
- `RerankTrace`
|
||||
- `LookupResult`
|
||||
|
||||
Policy:
|
||||
|
||||
- L0-derived filter is optional and only used when unambiguous.
|
||||
- Filtered L1 low-quality result triggers raw unfiltered L1 retry.
|
||||
- Rule-based rerank is sufficient for MVP.
|
||||
- Context packing uses character budget first, not exact token counting.
|
||||
|
||||
## Pending For Commit
|
||||
|
||||
- Specify exact fields for `ContextPack`, `RetrievalTrace`, and `RerankTrace`.
|
||||
- Specify low-quality trigger for unfiltered retry.
|
||||
- Decide whether `PrimaryResult` / `SupplementResult` classes are deleted or deprecated during the first apply.
|
||||
- Write OpenSpec `design.md`, specs, and executable `tasks.md`.
|
||||
|
||||
## Specify Results
|
||||
|
||||
Created committed-design artifacts:
|
||||
|
||||
- `design.md`
|
||||
- `specs/rag-knowledge-retrieval/spec.md`
|
||||
- `tasks.md`
|
||||
|
||||
Resolved pending items:
|
||||
|
||||
- `ContextPack` minimum fields: `packedText`, `strategy`, `charBudget`, `usedChars`, `includedSources`, `omittedSources`.
|
||||
- `RetrievalTrace` minimum behavior: record filtered attempt, unfiltered retry when used, fallback reason, and no-evidence paths.
|
||||
- `RerankTrace` minimum behavior: record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks.
|
||||
- Low-quality trigger: empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
|
||||
- `PrimaryResult` / `SupplementResult`: may be removed during apply if all compile-time usages are migrated.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| proposal goals/scope -> design decisions | aligned |
|
||||
| design module boundaries -> specs behavior | aligned |
|
||||
| specs observable behavior -> tasks | aligned |
|
||||
| interface impact -> design/tasks migration work | aligned |
|
||||
|
||||
No cross-artifact gaps remain for Commit.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
lookup_knowledge(query)
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
Risk assessment:
|
||||
|
||||
- The architecture keeps the explicit Agent tool boundary and does not move retrieval into an implicit Advisor.
|
||||
- Data ownership is clearer: query hints belong to transformer, vector candidates to retriever, evidence/context/traces to post-retrieval pipeline, persistence summaries to recorder.
|
||||
- The largest risk is the L4 result contract change; design and tasks require prompt/test/recorder migration in the same apply.
|
||||
- Database migration risk is low because `tool_invocation` table fields remain stable and new trace details stay in JSON.
|
||||
- Latency risk from unfiltered retry is accepted for MVP because retry only happens below reference quality.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
OpenSpec validation:
|
||||
|
||||
```text
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
File integrity:
|
||||
|
||||
- proposal exists and states problem, proposed change, scope, non-goals, risks, and interface impact.
|
||||
- design exists and records module boundaries, decisions, migration, rollback, and risks.
|
||||
- specs exist and define observable behavior for modular pipeline, unfiltered retry, evidence-first result, context pack, rerank, and L0 hint boundaries.
|
||||
- tasks exist and are executable vertical slices.
|
||||
|
||||
Consistency:
|
||||
|
||||
- proposal core concepts are represented in design.
|
||||
- design decisions are represented in specs and tasks.
|
||||
- tasks have verifiable implementation and test steps.
|
||||
- L4 interface impact is recorded and mapped to migration tasks.
|
||||
|
||||
Status: ready to mark as Committed OpenSpec.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, and tests.
|
||||
|
||||
Reference implementation and affected files inspected:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- `src/main/resources/prompts/executor-prompt.md`
|
||||
- `src/main/resources/application.yml`
|
||||
|
||||
Technical stack checklist:
|
||||
|
||||
- Request/response structure: `lookup_knowledge` returns a Java DTO serialized as Agent tool JSON.
|
||||
- Retrieval facade: `VectorSearchService.searchSimilarDocuments(query, topK, category)` is the stable retrieval API and already hides Spring AI VectorStore vs SDK fallback.
|
||||
- L0 query hints: `KnowledgeIndexService.analyzeQuery` returns `L0Hint(matches, matchedKeywords, domains, entities, titles)` and `singleDomainOrNull()`.
|
||||
- Session dedup: `RetrievedDocTracker` stores `sessionId -> domain -> filePath` and exposes backward-compatible `isAlreadyRetrieved`.
|
||||
- Trace persistence: `ToolInvocationRecorder.recordLookupKnowledge` writes stable table columns and JSON `retrieval_details`.
|
||||
- Prompt consumers: executor prompts still describe old L0/L1 and `primary.content`; these must be migrated.
|
||||
- Tests: `LookupKnowledgeToolTest` is the main consumer of old `primary` / `supplement` assertions; recorder tests verify retrieval details.
|
||||
|
||||
Implementation decision:
|
||||
|
||||
- Add pipeline DTOs under `com.superbiz.agent.dto`.
|
||||
- Add pipeline services under `com.superbiz.agent.service`.
|
||||
- Keep `VectorSearchService` unchanged.
|
||||
- Keep `lookup_knowledge` tool name and query argument unchanged.
|
||||
- Delete `PrimaryResult` / `SupplementResult` only after production and tests stop referencing them.
|
||||
|
||||
## Apply Results
|
||||
|
||||
Completed tasks: 31/31.
|
||||
|
||||
Implemented:
|
||||
|
||||
- Added modular RAG DTOs: `KnowledgeQuery`, `RetrievedEvidenceCandidate`, `ContextPack`, `RetrievalTrace`, `RerankTrace`, `EvidencePostprocessResult`.
|
||||
- Added pipeline services: `KnowledgeQueryTransformer`, `KnowledgeDocumentRetriever`, `KnowledgeEvidencePostProcessor`, `KnowledgeContextPacker`, `LookupResultAssembler`.
|
||||
- Refactored `LookupKnowledgeTool` into a thin orchestrator.
|
||||
- Migrated `LookupResult` to evidence-first fields and removed `primary` / `supplement`.
|
||||
- Deleted `PrimaryResult` and `SupplementResult`.
|
||||
- Updated `ToolInvocationRecorder` to persist query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries in `retrieval_details`.
|
||||
- Updated executor prompt text for `contextPack` / `evidenceBlocks`.
|
||||
- Updated RAG architecture docs.
|
||||
- Rewrote lookup and recorder tests for the new contract.
|
||||
|
||||
Verification:
|
||||
|
||||
```text
|
||||
mvn -q -DskipTests compile
|
||||
PASS
|
||||
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
PASS
|
||||
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
PASS: Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
Full suite attempt:
|
||||
|
||||
```text
|
||||
mvn -q test
|
||||
FAIL
|
||||
```
|
||||
|
||||
The full suite failed on pre-existing/environment-dependent tests:
|
||||
|
||||
- `MilvusConnectionTest.connect`: `MILVUS_TOKEN` not set.
|
||||
- `RedisSessionManagerTest`: Redis JSON contains legacy `messagePairCount`, not accepted by current `SessionContext`.
|
||||
- Spring context / repository tests attempted MySQL/Flyway and failed when database connectivity was unavailable in the first sandboxed run.
|
||||
|
||||
The full suite was retried outside the sandbox after approval. It still failed for the Milvus token and Redis serialization issues above, so these failures are not attributed to the modular RAG change.
|
||||
|
||||
Diff scope reviewed:
|
||||
|
||||
- RAG implementation: DTOs, pipeline services, `LookupKnowledgeTool`, `ToolInvocationRecorder`.
|
||||
- Contract cleanup: removed `PrimaryResult` / `SupplementResult`, updated `LookupResult`.
|
||||
- Tests: lookup tool and recorder tests.
|
||||
- Prompts: executor and chat executor guidance.
|
||||
- Docs/OpenSpec: modular RAG change files and architecture docs.
|
||||
|
||||
Known remaining risk:
|
||||
|
||||
- Runtime Agent prompt behavior should be demo-tested manually because the tool JSON contract changed from `primary/supplement` to `evidenceBlocks/contextPack/traces`.
|
||||
|
||||
## Review Results
|
||||
|
||||
Review finding:
|
||||
|
||||
- Session dedup returned `found=false` but still carried `evidenceBlocks` and `contextPack`, allowing the Agent to consume duplicate evidence despite the dedup message.
|
||||
|
||||
Fix:
|
||||
|
||||
- `LookupResultAssembler.deduped` now returns an empty evidence/context payload while preserving trace, relevance hint, and retrieved-domain memory.
|
||||
- Added `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`.
|
||||
|
||||
Additional verification:
|
||||
|
||||
```text
|
||||
mvn -q -DskipTests compile
|
||||
PASS
|
||||
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
PASS
|
||||
|
||||
MILVUS_TOKEN=<application.yml milvus.token> mvn -q test
|
||||
PASS
|
||||
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
PASS: Change 'modular-rag-pipeline' is valid
|
||||
|
||||
git diff --check
|
||||
PASS with LF/CRLF warnings only
|
||||
```
|
||||
|
||||
Updated full-suite note:
|
||||
|
||||
- `MilvusConnectionTest` passes when `MILVUS_TOKEN` is injected into the Maven process from `application.yml`.
|
||||
- `RedisSessionManagerTest` now passes after marking computed `SessionContext` getters as non-serialized JSON properties and ignoring unknown legacy Redis fields.
|
||||
@@ -0,0 +1,193 @@
|
||||
## Context
|
||||
|
||||
The current `lookup_knowledge` implementation is operational but still organized around a legacy L0/L1 result model. `LookupKnowledgeTool` currently performs query analysis, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording in one class. Evidence blocks already exist, but post-retrieval processing is not a first-class pipeline boundary.
|
||||
|
||||
The project direction is already documented as:
|
||||
|
||||
- keep `lookup_knowledge` as an explicit Agent evidence tool;
|
||||
- treat L0 as domain/entity/keyword hint data;
|
||||
- use L1 vector retrieval as the main evidence source;
|
||||
- keep Spring AI `VectorStore` behind `VectorSearchService`;
|
||||
- preserve `tool_invocation` as the evidence trace for Verifier, Eval, and trace APIs.
|
||||
|
||||
This change turns that architecture into code structure and updates the tool result contract so future Agent, Verifier, and Eval flows consume structured evidence instead of the old `primary` / `supplement` split.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Refactor `lookup_knowledge` into a modular RAG pipeline.
|
||||
- Keep L0 as query understanding and retrieval-control signal.
|
||||
- Keep L1 vector retrieval as the main document retrieval path.
|
||||
- Add a simple MVP fallback: retry raw query through unfiltered L1 when filtered L1 is low quality.
|
||||
- Move relevance normalization, evidence block creation, deduplication, and lightweight rerank into post-retrieval processing.
|
||||
- Add context packing as a first-class output.
|
||||
- Replace the old `primary` / `supplement` result contract with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
|
||||
- Keep `lookup_knowledge` tool name and input argument unchanged.
|
||||
- Keep `tool_invocation` table schema stable and enrich `retrieval_details`.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not replace `lookup_knowledge` with an implicit Advisor.
|
||||
- Do not add a model-based reranker, cross-encoder, BM25, RRF, Elasticsearch, or OpenSearch.
|
||||
- Do not migrate document upload, chunking, embedding write path, Milvus schema, or vector collection layout.
|
||||
- Do not change when the Agent decides to call `lookup_knowledge`.
|
||||
- Do not introduce multi-query expansion in this change.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision: Use explicit pipeline components
|
||||
|
||||
Create a local pipeline behind `LookupKnowledgeTool`:
|
||||
|
||||
```text
|
||||
LookupKnowledgeTool
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
Rationale: this keeps the Agent tool boundary stable while making the RAG flow easy to test and explain. It also maps cleanly to Spring AI modular RAG concepts without hiding business observability inside an Advisor.
|
||||
|
||||
Alternative considered: keep all logic in `LookupKnowledgeTool` and only add fields. Rejected because the class would continue to mix query transformation, retrieval, post-processing, and trace responsibilities.
|
||||
|
||||
### Decision: L0 is a query transformer signal, not a main retriever
|
||||
|
||||
`KnowledgeQueryTransformer` will wrap the existing `KnowledgeIndexService.analyzeQuery` behavior and produce a transformed query object containing:
|
||||
|
||||
- original query
|
||||
- rewritten query, initially equal to the raw query unless a future rule rewrites it
|
||||
- domain hints
|
||||
- matched keywords
|
||||
- entities
|
||||
- optional category filter
|
||||
- L0 titles for trace only
|
||||
|
||||
L0 matches must not be converted into normal evidence candidates in the main path.
|
||||
|
||||
Rationale: L0 keyword/frontmatter matching is useful for controlling retrieval, but it is not reliable enough to be treated as fact evidence when L1 cannot support it.
|
||||
|
||||
Alternative considered: combine L0 and L1 into one candidate list. Rejected because it makes L0 an equal retrieval layer again and conflicts with the desired architecture.
|
||||
|
||||
### Decision: Use filtered L1 first, then unfiltered L1 retry
|
||||
|
||||
`KnowledgeDocumentRetriever` will perform:
|
||||
|
||||
```text
|
||||
attempt 1: vector search with L0-derived category filter, when unambiguous
|
||||
attempt 2: raw query vector search without the L0-derived filter, when attempt 1 is low quality
|
||||
```
|
||||
|
||||
Filtered retrieval is low quality when any of the following is true:
|
||||
|
||||
- no candidates are returned;
|
||||
- post-processing would produce zero evidence blocks;
|
||||
- top candidate normalized similarity is below `retrieval.normalization.reference-threshold`.
|
||||
|
||||
Rationale: the most likely MVP failure mode is an over-strict or wrong metadata filter. A raw unfiltered vector retry addresses that without adding a complex multi-stage fallback policy.
|
||||
|
||||
Alternative considered: return L0 documents as weak fallback evidence. Rejected for this MVP because it can let keyword hints masquerade as factual evidence.
|
||||
|
||||
### Decision: Keep rerank rule-based
|
||||
|
||||
`KnowledgeEvidencePostProcessor` will rerank with deterministic signals:
|
||||
|
||||
- vector score / normalized similarity;
|
||||
- domain match with query hints;
|
||||
- entity match;
|
||||
- keyword match;
|
||||
- source type or metadata priority when available;
|
||||
- retrieval attempt, with filtered hits not automatically preferred over stronger unfiltered hits.
|
||||
|
||||
The rerank trace should explain major score contributions per final evidence block.
|
||||
|
||||
Rationale: a rule-based reranker is explainable, cheap, testable, and enough for the interview-oriented MVP. It also avoids introducing model latency and new dependencies.
|
||||
|
||||
Alternative considered: model-based rerank. Rejected as out of scope until evaluation shows a need.
|
||||
|
||||
### Decision: Context pack becomes the Agent-facing content
|
||||
|
||||
`KnowledgeContextPacker` will turn final evidence blocks into a compact context package:
|
||||
|
||||
```text
|
||||
packedText
|
||||
strategy
|
||||
charBudget
|
||||
usedChars
|
||||
includedSources
|
||||
omittedSources
|
||||
```
|
||||
|
||||
MVP packing uses a character budget rather than exact token counting. The packer preserves source/title/breadcrumb/hit reasons and truncates content only after preserving metadata.
|
||||
|
||||
Rationale: the Agent should consume curated evidence context instead of inferring semantics from `primary` and `supplement`.
|
||||
|
||||
Alternative considered: keep `primary` and `supplement` as the Agent-facing fields. Rejected because those names preserve the outdated L0 primary / L1 supplement model.
|
||||
|
||||
### Decision: Break the old result contract deliberately
|
||||
|
||||
`LookupResult` will move to the preferred contract:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
`primary` and `supplement` may be removed in this change, and `PrimaryResult` / `SupplementResult` may be deleted if no production code remains after migration.
|
||||
|
||||
Rationale: this is an MVP project intended to demonstrate clean modular RAG. Keeping obsolete fields would force later code to preserve misleading semantics.
|
||||
|
||||
Alternative considered: add new fields while keeping old fields deprecated. Rejected because the user explicitly accepted removing old fields if the later flow becomes cleaner.
|
||||
|
||||
### Decision: Preserve persistence compatibility at table level
|
||||
|
||||
`ToolInvocationRecorder` will stop depending on `result.getPrimary()` for output preview. Preview should come from `contextPack.packedText` or the top evidence block content. `retrieval_details` will include compact summaries:
|
||||
|
||||
- query transform summary;
|
||||
- retrieval trace with attempts and fallback reason;
|
||||
- rerank trace summary;
|
||||
- context pack summary;
|
||||
- evidence block summaries.
|
||||
|
||||
Rationale: trace consumers already read `tool_invocation` rows. The database schema can remain stable while JSON details evolve.
|
||||
|
||||
Alternative considered: add columns for each new trace object. Rejected because the current trace model already stores retrieval-specific details in JSON and does not need schema churn for this change.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Breaking tool output shape can affect Agent prompt behavior. -> Mitigation: update tool description and executor prompt to prefer `contextPack` and `evidenceBlocks`.
|
||||
- [Risk] Existing tests assert `primary` / `supplement`. -> Mitigation: migrate tests to evidence/context/traces in the same change.
|
||||
- [Risk] Rule-based rerank can reorder evidence unexpectedly. -> Mitigation: persist `rerankTrace` and cover ordering behavior with tests.
|
||||
- [Risk] Context packing can omit useful evidence under a small budget. -> Mitigation: record included and omitted sources and keep the character budget configurable.
|
||||
- [Risk] Existing OpenSpec requirements conflict with the new fallback policy. -> Mitigation: update `rag-knowledge-retrieval` delta before apply and validate the change.
|
||||
- [Risk] Unfiltered retry may increase latency. -> Mitigation: retry only when filtered result is empty or below reference quality, and record attempt counts/duration in trace.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add new internal DTOs and pipeline components while keeping `LookupKnowledgeTool` as the public tool bean.
|
||||
2. Migrate `LookupKnowledgeTool` orchestration to the pipeline.
|
||||
3. Update `LookupResult` to the new result contract and remove `primary` / `supplement` usages.
|
||||
4. Update `ToolInvocationRecorder` preview and retrieval details to use context/evidence/traces.
|
||||
5. Update prompt text and tool description.
|
||||
6. Update tests for the new contract.
|
||||
7. Run targeted unit tests and OpenSpec validation.
|
||||
|
||||
Rollback strategy:
|
||||
|
||||
- Revert the change as one unit if Agent tool behavior regresses.
|
||||
- Because the table schema remains stable and the tool input is unchanged, rollback does not require database migration.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- The exact default context pack character budget should be finalized during implementation; recommended MVP default is 3000 to 5000 characters.
|
||||
- Whether to delete `PrimaryResult` and `SupplementResult` immediately depends on final compile-time references after migration.
|
||||
@@ -0,0 +1,172 @@
|
||||
# Modular RAG Pipeline Proposal
|
||||
|
||||
## Problem
|
||||
|
||||
`lookup_knowledge` already exposes structured retrieval evidence, but the runtime flow is still concentrated inside `LookupKnowledgeTool`. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.
|
||||
|
||||
The current result contract also still carries the old `primary` / `supplement` model, where `primary` means L0 exact match and `supplement` means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.
|
||||
|
||||
## Proposed Change
|
||||
|
||||
Refactor `lookup_knowledge` into a modular RAG pipeline while preserving the explicit Agent tool boundary and `tool_invocation` evidence trace.
|
||||
|
||||
Target pipeline:
|
||||
|
||||
```text
|
||||
LookupKnowledgeTool
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
### Query Transformation
|
||||
|
||||
Introduce a query transformer around the current L0 analysis.
|
||||
|
||||
L0 SHALL provide:
|
||||
|
||||
- domain/category hints
|
||||
- matched keywords
|
||||
- extracted entities
|
||||
- optional metadata filter candidate
|
||||
- traceable query understanding data
|
||||
|
||||
L0 SHALL NOT act as a main evidence retrieval path in the normal flow.
|
||||
|
||||
### Retrieval
|
||||
|
||||
L1 vector retrieval remains the main document retrieval path through `VectorSearchService`, which already supports Spring AI `VectorStore` as the preferred path and Milvus SDK fallback.
|
||||
|
||||
MVP fallback strategy:
|
||||
|
||||
```text
|
||||
1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
|
||||
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
|
||||
3. If unfiltered retrieval also fails, return no_evidence.
|
||||
```
|
||||
|
||||
The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.
|
||||
|
||||
### Post-Retrieval Processing
|
||||
|
||||
Move evidence post-processing out of `LookupKnowledgeTool`.
|
||||
|
||||
The post-processor SHALL handle:
|
||||
|
||||
- relevance normalization
|
||||
- evidence block creation
|
||||
- source-level deduplication
|
||||
- rule-based lightweight rerank
|
||||
- retrieval trace and rerank trace generation
|
||||
|
||||
The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.
|
||||
|
||||
### Context Packing
|
||||
|
||||
Add a context packing step that converts final evidence blocks into an Agent-facing context package.
|
||||
|
||||
The packer SHALL:
|
||||
|
||||
- keep source/title/breadcrumb visible
|
||||
- obey a configurable character budget in the MVP
|
||||
- prioritize reranked evidence order
|
||||
- avoid duplicate source content
|
||||
- produce a compact summary of included and omitted evidence
|
||||
|
||||
### Result Contract
|
||||
|
||||
This change intentionally updates the `lookup_knowledge` return contract.
|
||||
|
||||
New preferred contract:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
The old `primary` and `supplement` fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- Refactor `LookupKnowledgeTool` into a thin tool boundary and pipeline orchestrator.
|
||||
- Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
|
||||
- Update `LookupResult` to prefer `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
|
||||
- Remove or deprecate `primary` / `supplement` according to the final spec.
|
||||
- Update `ToolInvocationRecorder` to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace.
|
||||
- Update tool description / prompt wording so Agent behavior matches the new contract.
|
||||
- Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
|
||||
- Update `rag-knowledge-retrieval` spec to remove L0 primary fallback and old compatibility-field requirements.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Replacing the explicit `lookup_knowledge` tool with an implicit Advisor.
|
||||
- Introducing a model-based reranker or cross-encoder.
|
||||
- Introducing BM25, RRF, Elasticsearch, or OpenSearch.
|
||||
- Migrating document upload, chunking, embedding writes, or Milvus schema.
|
||||
- Changing the Agent decision of when to call `lookup_knowledge`.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Impact level: L4 breaking interface change.
|
||||
|
||||
Reason:
|
||||
|
||||
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
|
||||
- The JSON returned by the `lookup_knowledge` Agent tool changes shape.
|
||||
- Tests and internal consumers that read `primary` / `supplement` must migrate to `evidenceBlocks` and `contextPack`.
|
||||
|
||||
Known affected areas:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `LookupResult`
|
||||
- `PrimaryResult` / `SupplementResult`
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeToolTest`
|
||||
- `ToolInvocationRecorderTest`
|
||||
- Agent tool prompt / executor prompt references
|
||||
- `rag-knowledge-retrieval` OpenSpec requirements
|
||||
- RAG docs under `mvp/architecture/`
|
||||
|
||||
Migration approach:
|
||||
|
||||
- Update all in-repo consumers in the same change.
|
||||
- Keep `lookup_knowledge` tool name and input argument unchanged.
|
||||
- Keep `tool_invocation` persistence compatible at table level while enriching `retrieval_details`.
|
||||
- Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.
|
||||
|
||||
## Context Constraints
|
||||
|
||||
Relevant project decisions:
|
||||
|
||||
- `lookup_knowledge` remains an explicit Agent evidence tool.
|
||||
- L0 is already documented as a hint layer, not a final decision layer.
|
||||
- Spring AI `VectorStore` is already the preferred retrieval path behind `VectorSearchService`.
|
||||
- Milvus SDK fallback remains valuable for MVP resilience.
|
||||
- `tool_invocation` is the stable evidence trace used by trace inspection, Verifier, and Eval.
|
||||
- Existing `rag-knowledge-retrieval` spec still contains legacy fallback and compatibility-field requirements that must be changed.
|
||||
|
||||
## Risks
|
||||
|
||||
- Breaking result contract may affect prompt behavior because the tool JSON changes.
|
||||
- Removing `primary` / `supplement` requires updating tests and recorder preview logic.
|
||||
- Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
|
||||
- Context packing can hide useful evidence if the budget is too small.
|
||||
- Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- What exact minimum fields should `ContextPack`, `RetrievalTrace`, and `RerankTrace` expose in the committed spec?
|
||||
- Should `PrimaryResult` and `SupplementResult` classes be deleted immediately or left deprecated for one change cycle?
|
||||
- What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?
|
||||
+121
@@ -0,0 +1,121 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
|
||||
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
|
||||
|
||||
#### Scenario: Pipeline components execute in order
|
||||
- **WHEN** `lookup_knowledge` receives a query
|
||||
- **THEN** the system SHALL transform the query before retrieval
|
||||
- **AND** it SHALL retrieve vector candidates before post-processing
|
||||
- **AND** it SHALL build evidence blocks before context packing
|
||||
- **AND** it SHALL record trace details after result assembly
|
||||
|
||||
#### Scenario: Tool boundary remains explicit
|
||||
- **WHEN** the modular pipeline is used
|
||||
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
|
||||
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
|
||||
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
|
||||
|
||||
#### Scenario: Filtered retrieval succeeds
|
||||
- **WHEN** L0 provides an unambiguous category filter
|
||||
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
|
||||
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
|
||||
|
||||
#### Scenario: Filtered retrieval returns no evidence
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
|
||||
|
||||
#### Scenario: Filtered retrieval is below reference quality
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
|
||||
|
||||
#### Scenario: Both retrieval attempts fail
|
||||
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
|
||||
- **THEN** the tool SHALL return `found=false`
|
||||
- **AND** the tool SHALL set evidence status to `no_evidence`
|
||||
- **AND** the tool SHALL NOT return L0 documents as fact evidence
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
|
||||
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
|
||||
|
||||
#### Scenario: Evidence result contains context and traces
|
||||
- **WHEN** `lookup_knowledge` returns usable evidence
|
||||
- **THEN** the result SHALL include `evidenceBlocks`
|
||||
- **AND** it SHALL include `contextPack`
|
||||
- **AND** it SHALL include `retrievalTrace`
|
||||
- **AND** it SHALL include `rerankTrace`
|
||||
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
|
||||
|
||||
#### Scenario: No-evidence result keeps traceability
|
||||
- **WHEN** `lookup_knowledge` returns no usable evidence
|
||||
- **THEN** the result SHALL include `found=false`
|
||||
- **AND** it SHALL include a message explaining that no knowledge evidence was found
|
||||
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
|
||||
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
|
||||
|
||||
#### Scenario: Context pack preserves source metadata
|
||||
- **WHEN** evidence blocks are packed
|
||||
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
|
||||
|
||||
#### Scenario: Context pack respects budget
|
||||
- **WHEN** final evidence content exceeds the configured context budget
|
||||
- **THEN** the packer SHALL truncate content rather than source metadata
|
||||
- **AND** it SHALL record included and omitted sources in the context pack summary
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
|
||||
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
|
||||
|
||||
#### Scenario: Rerank trace records score contributions
|
||||
- **WHEN** candidates are reranked
|
||||
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
|
||||
|
||||
#### Scenario: Query hints influence rerank without becoming evidence
|
||||
- **WHEN** L0 query hints match candidate metadata or content
|
||||
- **THEN** the reranker MAY boost the candidate
|
||||
- **AND** the evidence block SHALL record the hint as a hit reason
|
||||
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
|
||||
|
||||
#### Scenario: L0 produces traceable hint data
|
||||
- **WHEN** L0 matches one or more indexed knowledge entries
|
||||
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
|
||||
|
||||
#### Scenario: L0 does not bypass semantic retrieval
|
||||
- **WHEN** L0 returns exactly one match
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
|
||||
|
||||
#### Scenario: L0 hints do not become normal evidence
|
||||
- **WHEN** L1 retrieval returns no usable evidence
|
||||
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
|
||||
- **AND** L0 hint data MAY still be recorded in retrieval trace details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
|
||||
|
||||
#### Scenario: Evidence block contains source metadata
|
||||
- **WHEN** a `lookup_knowledge` call returns evidence
|
||||
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
|
||||
|
||||
#### Scenario: Evidence blocks are the primary evidence contract
|
||||
- **WHEN** evidence blocks are returned
|
||||
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
|
||||
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
|
||||
|
||||
## REMOVED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
|
||||
**Reason**: The new MVP fallback strategy retries unfiltered L1 retrieval when L0-derived filtering appears to suppress usable semantic evidence. Returning L0 documents as fallback fact evidence can let keyword hints masquerade as verified knowledge.
|
||||
|
||||
**Migration**: Use the new requirement "Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality". If filtered and unfiltered L1 both fail, return `no_evidence` while preserving L0 hint data in trace details only.
|
||||
@@ -0,0 +1,51 @@
|
||||
## 1. Pipeline Models
|
||||
|
||||
- [x] 1.1 Add query transformation model for original query, rewritten query, domain hints, matched keywords, entities, category filter, and trace-only L0 titles.
|
||||
- [x] 1.2 Add retrieved evidence candidate model that normalizes vector result metadata, score semantics, source, title, breadcrumb, content, retrieval attempt, and original rank.
|
||||
- [x] 1.3 Add context pack, retrieval trace, and rerank trace DTOs for the new `LookupResult` contract.
|
||||
- [x] 1.4 Update `LookupResult` to expose evidence blocks, context pack, retrieval trace, rerank trace, relevance level, completeness hint, retrieved domains, and message.
|
||||
|
||||
## 2. Query And Retrieval Pipeline
|
||||
|
||||
- [x] 2.1 Implement `KnowledgeQueryTransformer` by reusing `KnowledgeIndexService.analyzeQuery` and mapping L0 output to query hints and optional category filter.
|
||||
- [x] 2.2 Implement `KnowledgeDocumentRetriever` as a wrapper around `VectorSearchService` for filtered and unfiltered vector retrieval attempts.
|
||||
- [x] 2.3 Implement low-quality detection using empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
|
||||
- [x] 2.4 Implement unfiltered raw-query retry when filtered retrieval is low quality and record fallback reason in retrieval trace.
|
||||
|
||||
## 3. Post-Retrieval Processing
|
||||
|
||||
- [x] 3.1 Move relevance normalization out of `LookupKnowledgeTool` into `KnowledgeEvidencePostProcessor`.
|
||||
- [x] 3.2 Move evidence block creation and source-level deduplication out of `LookupKnowledgeTool` into the post-processor.
|
||||
- [x] 3.3 Implement rule-based lightweight rerank using vector similarity, domain match, entity match, keyword match, and metadata/source-type signals.
|
||||
- [x] 3.4 Ensure L0 hints influence filter/rerank/trace only and are not returned as standalone fact evidence when L1 has no usable evidence.
|
||||
|
||||
## 4. Context Packing And Result Assembly
|
||||
|
||||
- [x] 4.1 Implement `KnowledgeContextPacker` with a configurable MVP character budget.
|
||||
- [x] 4.2 Pack evidence blocks while preserving source, title, breadcrumb, and hit reasons before truncating content.
|
||||
- [x] 4.3 Implement `LookupResultAssembler` to build evidence-first results for usable evidence, no-evidence, and session dedup cases.
|
||||
- [x] 4.4 Remove or migrate all `primary` and `supplement` result usage from production code.
|
||||
|
||||
## 5. Tool Boundary And Trace Recording
|
||||
|
||||
- [x] 5.1 Refactor `LookupKnowledgeTool` into a thin orchestrator that invokes the pipeline and handles tool boundary concerns.
|
||||
- [x] 5.2 Update `ToolInvocationRecorder.LookupKnowledgeRecord` to summarize context pack, retrieval trace, rerank trace, fallback reason, and evidence blocks without relying on `result.getPrimary()`.
|
||||
- [x] 5.3 Preserve stable `tool_invocation` table fields and store new retrieval details in JSON.
|
||||
- [x] 5.4 Update `@Tool` description and relevant executor prompt text to describe evidence blocks, context pack, and no-realtime-data boundaries.
|
||||
|
||||
## 6. Tests And Evaluation
|
||||
|
||||
- [x] 6.1 Update `LookupKnowledgeToolTest` for evidence-first result contract and removal of `primary` / `supplement`.
|
||||
- [x] 6.2 Add tests for filtered L1 success without retry.
|
||||
- [x] 6.3 Add tests for filtered low-quality retrieval triggering raw unfiltered L1 retry.
|
||||
- [x] 6.4 Add tests proving L0 hint data does not become standalone fact evidence when L1 fails.
|
||||
- [x] 6.5 Add tests for rerank ordering, rerank trace, context pack budget behavior, and source metadata preservation.
|
||||
- [x] 6.6 Update `ToolInvocationRecorderTest` for new retrieval detail summaries and output preview source.
|
||||
- [x] 6.7 Run targeted Java tests for lookup knowledge and recorder changes.
|
||||
- [x] 6.8 Run OpenSpec validation for `modular-rag-pipeline`.
|
||||
|
||||
## 7. Documentation Cleanup
|
||||
|
||||
- [x] 7.1 Update RAG architecture docs to reflect modular pipeline, unfiltered vector retry, and evidence-first contract.
|
||||
- [x] 7.2 Update retrieval observability docs to remove L0 primary fallback and `primary` / `supplement` compatibility language.
|
||||
- [x] 7.3 Review git diff to confirm only expected RAG, prompt, test, and spec files changed.
|
||||
@@ -4,15 +4,20 @@
|
||||
Define the runtime contract for the explicit `lookup_knowledge` Agent tool, including how L0 keyword/frontmatter hints cooperate with L1 semantic retrieval while preserving metadata filters, fallback evidence, and traceable retrieval details.
|
||||
## Requirements
|
||||
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it as domain, entity, and explainability hint data rather than as the sole final retrieval decision.
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
|
||||
|
||||
#### Scenario: L0 produces traceable hint data
|
||||
- **WHEN** L0 matches one or more indexed knowledge entries
|
||||
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
|
||||
|
||||
#### Scenario: L0 does not bypass semantic retrieval by default
|
||||
#### Scenario: L0 does not bypass semantic retrieval
|
||||
- **WHEN** L0 returns exactly one match
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is unavailable or explicitly disabled by configuration
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
|
||||
|
||||
#### Scenario: L0 hints do not become normal evidence
|
||||
- **WHEN** L1 retrieval returns no usable evidence
|
||||
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
|
||||
- **AND** L0 hint data MAY still be recorded in retrieval trace details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use L0 domain as optional L1 filter
|
||||
The retrieval flow SHALL use L0 domain/category information as an optional metadata filter for L1 retrieval when the domain is unambiguous.
|
||||
@@ -25,19 +30,6 @@ The retrieval flow SHALL use L0 domain/category information as an optional metad
|
||||
- **WHEN** L0 hint data contains zero domains or multiple domains
|
||||
- **THEN** the L1 retrieval request SHALL run without an L0-derived category filter
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
|
||||
The retrieval flow SHALL still return useful L0 evidence when L1 produces no usable result.
|
||||
|
||||
#### Scenario: L1 has no results
|
||||
- **WHEN** L0 has at least one match and L1 returns no candidates
|
||||
- **THEN** the tool SHALL return an L0-based primary result
|
||||
- **AND** the relevance assessment SHALL not claim semantic support from L1
|
||||
|
||||
#### Scenario: L1 fails
|
||||
- **WHEN** L0 has at least one match and L1 retrieval throws or fails
|
||||
- **THEN** the tool SHALL return an L0-based primary result
|
||||
- **AND** the tool invocation record SHALL preserve the L0 hint details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL persist L0 hints
|
||||
The system SHALL persist L0 hint details in `tool_invocation.retrieval_details` for `lookup_knowledge` calls.
|
||||
|
||||
@@ -50,15 +42,16 @@ The system SHALL persist L0 hint details in `tool_invocation.retrieval_details`
|
||||
- **THEN** the recorded retrieval layer SHALL be `L0+L1`
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks in addition to the existing compatibility fields.
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
|
||||
|
||||
#### Scenario: Evidence block contains source metadata
|
||||
- **WHEN** a `lookup_knowledge` call returns evidence
|
||||
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
|
||||
|
||||
#### Scenario: Compatibility fields remain available
|
||||
#### Scenario: Evidence blocks are the primary evidence contract
|
||||
- **WHEN** evidence blocks are returned
|
||||
- **THEN** the existing `primary` and `supplement` result fields SHALL remain available when their source evidence exists
|
||||
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
|
||||
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
|
||||
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
|
||||
@@ -151,3 +144,86 @@ Spring AI Milvus integration SHALL be configured to use the existing collection
|
||||
- **WHEN** Spring AI Milvus VectorStore is configured
|
||||
- **THEN** it SHALL use the existing id, content, vector, and metadata field names
|
||||
- **AND** it SHALL use the configured embedding dimension and metric type compatible with existing vectors
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
|
||||
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
|
||||
|
||||
#### Scenario: Pipeline components execute in order
|
||||
- **WHEN** `lookup_knowledge` receives a query
|
||||
- **THEN** the system SHALL transform the query before retrieval
|
||||
- **AND** it SHALL retrieve vector candidates before post-processing
|
||||
- **AND** it SHALL build evidence blocks before context packing
|
||||
- **AND** it SHALL record trace details after result assembly
|
||||
|
||||
#### Scenario: Tool boundary remains explicit
|
||||
- **WHEN** the modular pipeline is used
|
||||
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
|
||||
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
|
||||
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
|
||||
|
||||
#### Scenario: Filtered retrieval succeeds
|
||||
- **WHEN** L0 provides an unambiguous category filter
|
||||
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
|
||||
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
|
||||
|
||||
#### Scenario: Filtered retrieval returns no evidence
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
|
||||
|
||||
#### Scenario: Filtered retrieval is below reference quality
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
|
||||
|
||||
#### Scenario: Both retrieval attempts fail
|
||||
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
|
||||
- **THEN** the tool SHALL return `found=false`
|
||||
- **AND** the tool SHALL set evidence status to `no_evidence`
|
||||
- **AND** the tool SHALL NOT return L0 documents as fact evidence
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
|
||||
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
|
||||
|
||||
#### Scenario: Evidence result contains context and traces
|
||||
- **WHEN** `lookup_knowledge` returns usable evidence
|
||||
- **THEN** the result SHALL include `evidenceBlocks`
|
||||
- **AND** it SHALL include `contextPack`
|
||||
- **AND** it SHALL include `retrievalTrace`
|
||||
- **AND** it SHALL include `rerankTrace`
|
||||
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
|
||||
|
||||
#### Scenario: No-evidence result keeps traceability
|
||||
- **WHEN** `lookup_knowledge` returns no usable evidence
|
||||
- **THEN** the result SHALL include `found=false`
|
||||
- **AND** it SHALL include a message explaining that no knowledge evidence was found
|
||||
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
|
||||
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
|
||||
|
||||
#### Scenario: Context pack preserves source metadata
|
||||
- **WHEN** evidence blocks are packed
|
||||
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
|
||||
|
||||
#### Scenario: Context pack respects budget
|
||||
- **WHEN** final evidence content exceeds the configured context budget
|
||||
- **THEN** the packer SHALL truncate content rather than source metadata
|
||||
- **AND** it SHALL record included and omitted sources in the context pack summary
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
|
||||
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
|
||||
|
||||
#### Scenario: Rerank trace records score contributions
|
||||
- **WHEN** candidates are reranked
|
||||
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
|
||||
|
||||
#### Scenario: Query hints influence rerank without becoming evidence
|
||||
- **WHEN** L0 query hints match candidate metadata or content
|
||||
- **THEN** the reranker MAY boost the candidate
|
||||
- **AND** the evidence block SHALL record the hint as a hit reason
|
||||
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
package com.superbiz.agent.domain.model;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonIgnore;
|
||||
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
@@ -20,6 +22,7 @@ import java.util.Map;
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public class SessionContext implements Serializable {
|
||||
|
||||
private static final long serialVersionUID = 1L;
|
||||
@@ -118,6 +121,7 @@ public class SessionContext implements Serializable {
|
||||
/**
|
||||
* 获取聊天历史副本,避免调用方直接修改内部列表。
|
||||
*/
|
||||
@JsonIgnore
|
||||
public List<Map<String, String>> getMessageHistorySnapshot() {
|
||||
if (this.messageHistory == null || this.messageHistory.isEmpty()) {
|
||||
return new ArrayList<>();
|
||||
@@ -141,6 +145,7 @@ public class SessionContext implements Serializable {
|
||||
this.lastActiveAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
@JsonIgnore
|
||||
public int getMessagePairCount() {
|
||||
return this.messageHistory == null ? 0 : this.messageHistory.size() / 2;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Agent-facing packed context assembled from final evidence blocks.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class ContextPack {
|
||||
|
||||
private String packedText;
|
||||
|
||||
private String strategy;
|
||||
|
||||
private Integer charBudget;
|
||||
|
||||
private Integer usedChars;
|
||||
|
||||
private List<String> includedSources;
|
||||
|
||||
private List<String> omittedSources;
|
||||
}
|
||||
@@ -0,0 +1,32 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Output of post-retrieval processing before context packing.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class EvidencePostprocessResult {
|
||||
|
||||
private Integer candidateCount;
|
||||
|
||||
private Integer evidenceBlockCount;
|
||||
|
||||
private List<EvidenceBlock> evidenceBlocks;
|
||||
|
||||
private String relevanceLevel;
|
||||
|
||||
private String completenessHint;
|
||||
|
||||
private RerankTrace rerankTrace;
|
||||
|
||||
private Double topSimilarity;
|
||||
|
||||
public boolean hasUsableEvidence() {
|
||||
return evidenceBlocks != null && !evidenceBlocks.isEmpty();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Query understanding output used by the knowledge retrieval pipeline.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class KnowledgeQuery {
|
||||
|
||||
private String originalQuery;
|
||||
|
||||
private String rewrittenQuery;
|
||||
|
||||
private List<String> domainHints;
|
||||
|
||||
private List<String> matchedKeywords;
|
||||
|
||||
private List<String> entities;
|
||||
|
||||
private String categoryFilter;
|
||||
|
||||
private List<String> l0Titles;
|
||||
|
||||
private Integer l0MatchCount;
|
||||
}
|
||||
@@ -17,21 +17,26 @@ public class LookupResult {
|
||||
*/
|
||||
private boolean found;
|
||||
|
||||
/**
|
||||
* 主要结果(L0 精确匹配)
|
||||
*/
|
||||
private PrimaryResult primary;
|
||||
|
||||
/**
|
||||
* 补充结果(L1 语义检索)
|
||||
*/
|
||||
private SupplementResult supplement;
|
||||
|
||||
/**
|
||||
* Structured evidence blocks after retrieval post-processing.
|
||||
*/
|
||||
private List<EvidenceBlock> evidenceBlocks;
|
||||
|
||||
/**
|
||||
* Packed Agent-facing context assembled from evidence blocks.
|
||||
*/
|
||||
private ContextPack contextPack;
|
||||
|
||||
/**
|
||||
* Retrieval attempts and fallback trace.
|
||||
*/
|
||||
private RetrievalTrace retrievalTrace;
|
||||
|
||||
/**
|
||||
* Rule-based rerank explanation.
|
||||
*/
|
||||
private RerankTrace rerankTrace;
|
||||
|
||||
/**
|
||||
* Candidate count before evidence deduplication.
|
||||
*/
|
||||
|
||||
@@ -1,39 +0,0 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* L0 精确匹配结果
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class PrimaryResult {
|
||||
|
||||
/**
|
||||
* 文档内容(前 2000 字符)
|
||||
*/
|
||||
private String content;
|
||||
|
||||
/**
|
||||
* 文档来源路径
|
||||
*/
|
||||
private String source;
|
||||
|
||||
/**
|
||||
* 匹配类型(exact_L0)
|
||||
*/
|
||||
private String matchType;
|
||||
|
||||
/**
|
||||
* 置信度(high / low)
|
||||
*/
|
||||
private String confidence;
|
||||
|
||||
/**
|
||||
* 可用的章节列表(预留字段,MVP 返回 null)
|
||||
*/
|
||||
private List<String> availableSections;
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Rule-based rerank explanation for final evidence blocks.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RerankTrace {
|
||||
|
||||
private List<Item> items;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
public static class Item {
|
||||
private Integer finalRank;
|
||||
private String source;
|
||||
private Double baseScore;
|
||||
private Double finalScore;
|
||||
private List<String> boostReasons;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,45 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Trace of retrieval attempts used by lookup_knowledge.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RetrievalTrace {
|
||||
|
||||
private String originalQuery;
|
||||
|
||||
private String rewrittenQuery;
|
||||
|
||||
private String categoryFilter;
|
||||
|
||||
private String selectedAttempt;
|
||||
|
||||
private String fallbackReason;
|
||||
|
||||
private String evidenceStatus;
|
||||
|
||||
private Map<String, Object> queryHints;
|
||||
|
||||
private List<Attempt> attempts;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
public static class Attempt {
|
||||
private String name;
|
||||
private String query;
|
||||
private String categoryFilter;
|
||||
private Integer candidateCount;
|
||||
private Boolean usable;
|
||||
private String errorMessage;
|
||||
private Integer durationMs;
|
||||
private Double topScore;
|
||||
private Double topSimilarity;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Normalized vector retrieval candidate before evidence post-processing.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RetrievedEvidenceCandidate {
|
||||
|
||||
private String id;
|
||||
|
||||
private String source;
|
||||
|
||||
private String title;
|
||||
|
||||
private String breadcrumb;
|
||||
|
||||
private String content;
|
||||
|
||||
private String retrievalLayer;
|
||||
|
||||
private String retrievalAttempt;
|
||||
|
||||
private Double score;
|
||||
|
||||
private Double rawScore;
|
||||
|
||||
private String scoreLabel;
|
||||
|
||||
private Integer originalRank;
|
||||
|
||||
private Map<String, String> metadata;
|
||||
|
||||
private List<String> hitReasons;
|
||||
}
|
||||
@@ -1,27 +0,0 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
/**
|
||||
* L1 语义检索补充结果
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class SupplementResult {
|
||||
|
||||
/**
|
||||
* 文档内容片段
|
||||
*/
|
||||
private String content;
|
||||
|
||||
/**
|
||||
* 文档来源
|
||||
*/
|
||||
private String source;
|
||||
|
||||
/**
|
||||
* 匹配类型(semantic_L1)
|
||||
*/
|
||||
private String matchType;
|
||||
}
|
||||
@@ -0,0 +1,70 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Packs final evidence blocks into compact Agent-facing context.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeContextPacker {
|
||||
|
||||
@Value("${rag.context-pack.char-budget:4000}")
|
||||
private int charBudget = 4000;
|
||||
|
||||
public ContextPack pack(List<EvidenceBlock> blocks) {
|
||||
List<EvidenceBlock> safeBlocks = blocks == null ? List.of() : blocks;
|
||||
StringBuilder packed = new StringBuilder();
|
||||
List<String> included = new ArrayList<>();
|
||||
List<String> omitted = new ArrayList<>();
|
||||
|
||||
int rank = 1;
|
||||
for (EvidenceBlock block : safeBlocks) {
|
||||
String header = buildHeader(rank, block);
|
||||
String content = block.getContent() == null ? "" : block.getContent();
|
||||
int remaining = charBudget - packed.length() - header.length();
|
||||
if (remaining <= 0) {
|
||||
omitted.add(block.getSource());
|
||||
continue;
|
||||
}
|
||||
String body = content.length() <= remaining ? content : content.substring(0, Math.max(0, remaining)) + "...";
|
||||
packed.append(header).append(body).append("\n\n");
|
||||
included.add(block.getSource());
|
||||
rank++;
|
||||
}
|
||||
|
||||
return ContextPack.builder()
|
||||
.packedText(packed.toString().trim())
|
||||
.strategy("ranked_evidence_char_budget")
|
||||
.charBudget(charBudget)
|
||||
.usedChars(packed.length())
|
||||
.includedSources(included)
|
||||
.omittedSources(omitted)
|
||||
.build();
|
||||
}
|
||||
|
||||
private String buildHeader(int rank, EvidenceBlock block) {
|
||||
StringBuilder header = new StringBuilder();
|
||||
header.append("[Evidence ").append(rank).append("]\n");
|
||||
appendLine(header, "source", block.getSource());
|
||||
appendLine(header, "title", block.getTitle());
|
||||
appendLine(header, "breadcrumb", block.getBreadcrumb());
|
||||
appendLine(header, "layer", block.getRetrievalLayer());
|
||||
if (block.getHitReasons() != null && !block.getHitReasons().isEmpty()) {
|
||||
appendLine(header, "reasons", String.join(", ", block.getHitReasons()));
|
||||
}
|
||||
header.append("content:\n");
|
||||
return header.toString();
|
||||
}
|
||||
|
||||
private void appendLine(StringBuilder builder, String key, String value) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
builder.append(key).append(": ").append(value).append("\n");
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,139 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Vector retrieval adapter for the modular knowledge pipeline.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeDocumentRetriever {
|
||||
|
||||
private final VectorSearchService vectorSearchService;
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
public KnowledgeDocumentRetriever(VectorSearchService vectorSearchService, ObjectMapper objectMapper) {
|
||||
this.vectorSearchService = vectorSearchService;
|
||||
this.objectMapper = objectMapper;
|
||||
}
|
||||
|
||||
public RetrievalAttemptResult retrieve(String attemptName, String query, String categoryFilter, int topK) {
|
||||
long start = System.currentTimeMillis();
|
||||
try {
|
||||
List<VectorSearchService.SearchResult> results =
|
||||
vectorSearchService.searchSimilarDocuments(query, topK, categoryFilter);
|
||||
List<RetrievedEvidenceCandidate> candidates = toCandidates(attemptName, results);
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, candidates.size(), null,
|
||||
(int) (System.currentTimeMillis() - start), topScore(results)),
|
||||
candidates
|
||||
);
|
||||
} catch (Exception e) {
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, 0, e.getMessage(),
|
||||
(int) (System.currentTimeMillis() - start), null),
|
||||
List.of()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
private List<RetrievedEvidenceCandidate> toCandidates(String attemptName,
|
||||
List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<RetrievedEvidenceCandidate> candidates = new ArrayList<>();
|
||||
for (int i = 0; i < results.size(); i++) {
|
||||
VectorSearchService.SearchResult result = results.get(i);
|
||||
Map<String, String> metadata = parseMetadata(result.getMetadata());
|
||||
String source = firstNonBlank(
|
||||
metadata.get("_source"),
|
||||
metadata.get("source"),
|
||||
metadata.get("filePath"),
|
||||
metadata.get("docId"),
|
||||
result.getMetadata(),
|
||||
result.getId()
|
||||
);
|
||||
candidates.add(RetrievedEvidenceCandidate.builder()
|
||||
.id(result.getId())
|
||||
.source(source)
|
||||
.title(metadata.get("title"))
|
||||
.breadcrumb(metadata.get("breadcrumb"))
|
||||
.content(result.getContent())
|
||||
.retrievalLayer("L1")
|
||||
.retrievalAttempt(attemptName)
|
||||
.score((double) result.getScore())
|
||||
.rawScore(result.getRawScore())
|
||||
.scoreLabel(result.getScoreLabel())
|
||||
.originalRank(i + 1)
|
||||
.metadata(metadata)
|
||||
.hitReasons(List.of("semantic_rank:" + (i + 1), "attempt:" + attemptName))
|
||||
.build());
|
||||
}
|
||||
return candidates;
|
||||
}
|
||||
|
||||
private RetrievalTrace.Attempt attempt(String name,
|
||||
String query,
|
||||
String categoryFilter,
|
||||
int candidateCount,
|
||||
String errorMessage,
|
||||
int durationMs,
|
||||
Double topScore) {
|
||||
return RetrievalTrace.Attempt.builder()
|
||||
.name(name)
|
||||
.query(query)
|
||||
.categoryFilter(categoryFilter)
|
||||
.candidateCount(candidateCount)
|
||||
.usable(errorMessage == null && candidateCount > 0)
|
||||
.errorMessage(errorMessage)
|
||||
.durationMs(durationMs)
|
||||
.topScore(topScore)
|
||||
.build();
|
||||
}
|
||||
|
||||
private Double topScore(List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
return (double) results.get(0).getScore();
|
||||
}
|
||||
|
||||
private Map<String, String> parseMetadata(String metadata) {
|
||||
if (metadata == null || metadata.isBlank()) {
|
||||
return Map.of();
|
||||
}
|
||||
try {
|
||||
Map<?, ?> raw = objectMapper.readValue(metadata, Map.class);
|
||||
Map<String, String> parsed = new LinkedHashMap<>();
|
||||
for (Map.Entry<?, ?> entry : raw.entrySet()) {
|
||||
if (entry.getKey() != null && entry.getValue() != null) {
|
||||
parsed.put(String.valueOf(entry.getKey()), String.valueOf(entry.getValue()));
|
||||
}
|
||||
}
|
||||
return parsed;
|
||||
} catch (Exception ignored) {
|
||||
return Map.of();
|
||||
}
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
public record RetrievalAttemptResult(RetrievalTrace.Attempt attempt,
|
||||
List<RetrievedEvidenceCandidate> candidates) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,251 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Post-retrieval evidence normalization, rerank, and evidence block assembly.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeEvidencePostProcessor {
|
||||
|
||||
private static final String LEVEL_PRECISE = "PRECISE";
|
||||
private static final String LEVEL_HIGHLY_RELEVANT = "HIGHLY_RELEVANT";
|
||||
private static final String LEVEL_REFERENCE = "REFERENCE";
|
||||
|
||||
private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档";
|
||||
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
|
||||
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
|
||||
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
@Value("${retrieval.normalization.highly-relevant-threshold:0.75}")
|
||||
private double highlyRelevantThreshold = 0.75;
|
||||
|
||||
@Value("${retrieval.normalization.reference-threshold:0.5}")
|
||||
private double referenceThreshold = 0.5;
|
||||
|
||||
public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) {
|
||||
List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates;
|
||||
List<ScoredCandidate> ranked = safeCandidates.stream()
|
||||
.map(candidate -> score(query, candidate))
|
||||
.sorted(Comparator.comparingDouble(ScoredCandidate::finalScore).reversed())
|
||||
.toList();
|
||||
|
||||
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
|
||||
List<RerankTrace.Item> traceItems = new ArrayList<>();
|
||||
int finalRank = 1;
|
||||
for (ScoredCandidate scored : ranked) {
|
||||
RetrievedEvidenceCandidate candidate = scored.candidate();
|
||||
EvidenceBlock block = EvidenceBlock.builder()
|
||||
.source(candidate.getSource())
|
||||
.title(candidate.getTitle())
|
||||
.breadcrumb(candidate.getBreadcrumb())
|
||||
.retrievalLayer(candidate.getRetrievalLayer())
|
||||
.content(truncate(candidate.getContent(), 800))
|
||||
.score(candidate.getScore())
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons()))
|
||||
.build();
|
||||
String key = sourceKey(block, "candidate-" + candidate.getOriginalRank());
|
||||
if (!deduped.containsKey(key)) {
|
||||
deduped.put(key, block);
|
||||
traceItems.add(RerankTrace.Item.builder()
|
||||
.finalRank(finalRank++)
|
||||
.source(candidate.getSource())
|
||||
.baseScore(scored.baseScore())
|
||||
.finalScore(scored.finalScore())
|
||||
.boostReasons(scored.boostReasons())
|
||||
.build());
|
||||
} else {
|
||||
mergeEvidence(deduped.get(key), block);
|
||||
}
|
||||
}
|
||||
|
||||
List<EvidenceBlock> blocks = new ArrayList<>(deduped.values());
|
||||
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).baseScore();
|
||||
RelevanceAssessment assessment = computeRelevance(query, ranked);
|
||||
return EvidencePostprocessResult.builder()
|
||||
.candidateCount(safeCandidates.size())
|
||||
.evidenceBlockCount(blocks.size())
|
||||
.evidenceBlocks(blocks)
|
||||
.relevanceLevel(assessment.level())
|
||||
.completenessHint(assessment.hint())
|
||||
.rerankTrace(RerankTrace.builder().items(traceItems).build())
|
||||
.topSimilarity(topSimilarity)
|
||||
.build();
|
||||
}
|
||||
|
||||
public boolean isLowQuality(EvidencePostprocessResult result) {
|
||||
if (result == null || !result.hasUsableEvidence()) {
|
||||
return true;
|
||||
}
|
||||
Double topSimilarity = result.getTopSimilarity();
|
||||
return topSimilarity == null || topSimilarity < referenceThreshold;
|
||||
}
|
||||
|
||||
public double normalizeL2(Double l2Score) {
|
||||
if (l2Score == null) {
|
||||
return 0.0;
|
||||
}
|
||||
double clamped = Math.min(l2Score, maxL2Distance);
|
||||
return Math.max(0.0, 1.0 - clamped / maxL2Distance);
|
||||
}
|
||||
|
||||
public double getReferenceThreshold() {
|
||||
return referenceThreshold;
|
||||
}
|
||||
|
||||
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate) {
|
||||
double baseScore = normalizeL2(candidate.getScore());
|
||||
double finalScore = baseScore;
|
||||
List<String> boosts = new ArrayList<>();
|
||||
|
||||
if (matchesAny(candidate, query.getDomainHints())) {
|
||||
finalScore += 0.15;
|
||||
boosts.add("domain_match:+0.15");
|
||||
}
|
||||
if (matchesAny(candidate, query.getEntities())) {
|
||||
finalScore += 0.20;
|
||||
boosts.add("entity_match:+0.20");
|
||||
}
|
||||
if (matchesAny(candidate, query.getMatchedKeywords())) {
|
||||
finalScore += 0.10;
|
||||
boosts.add("keyword_match:+0.10");
|
||||
}
|
||||
if (isPreferredSourceType(candidate)) {
|
||||
finalScore += 0.05;
|
||||
boosts.add("source_type:+0.05");
|
||||
}
|
||||
|
||||
return new ScoredCandidate(candidate, baseScore, finalScore, boosts);
|
||||
}
|
||||
|
||||
private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) {
|
||||
if (hints == null || hints.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
String haystack = String.join(" ",
|
||||
nullToEmpty(candidate.getSource()),
|
||||
nullToEmpty(candidate.getTitle()),
|
||||
nullToEmpty(candidate.getBreadcrumb()),
|
||||
nullToEmpty(candidate.getContent()),
|
||||
candidate.getMetadata() == null ? "" : candidate.getMetadata().toString()
|
||||
).toLowerCase(Locale.ROOT);
|
||||
for (String hint : hints) {
|
||||
if (hint != null && !hint.isBlank() && haystack.contains(hint.toLowerCase(Locale.ROOT))) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
private boolean isPreferredSourceType(RetrievedEvidenceCandidate candidate) {
|
||||
Map<String, String> metadata = candidate.getMetadata();
|
||||
if (metadata == null || metadata.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
String type = firstNonBlank(metadata.get("source_type"), metadata.get("documentType"), metadata.get("type"));
|
||||
if (type == null) {
|
||||
return false;
|
||||
}
|
||||
String normalized = type.toLowerCase(Locale.ROOT);
|
||||
return normalized.contains("runbook") || normalized.contains("guide") || normalized.contains("case");
|
||||
}
|
||||
|
||||
private RelevanceAssessment computeRelevance(KnowledgeQuery query, List<ScoredCandidate> ranked) {
|
||||
if (ranked.isEmpty()) {
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
ScoredCandidate top = ranked.get(0);
|
||||
if (top.baseScore() >= highlyRelevantThreshold && hasHintSupport(query, top)) {
|
||||
return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE);
|
||||
}
|
||||
if (top.baseScore() >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
|
||||
}
|
||||
if (top.baseScore() >= referenceThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
|
||||
private boolean hasHintSupport(KnowledgeQuery query, ScoredCandidate top) {
|
||||
return matchesAny(top.candidate(), query.getDomainHints())
|
||||
|| matchesAny(top.candidate(), query.getEntities())
|
||||
|| matchesAny(top.candidate(), query.getMatchedKeywords());
|
||||
}
|
||||
|
||||
private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) {
|
||||
Set<String> reasons = new LinkedHashSet<>();
|
||||
if (existing.getHitReasons() != null) {
|
||||
reasons.addAll(existing.getHitReasons());
|
||||
}
|
||||
if (incoming.getHitReasons() != null) {
|
||||
reasons.addAll(incoming.getHitReasons());
|
||||
}
|
||||
existing.setHitReasons(new ArrayList<>(reasons));
|
||||
if ((existing.getBreadcrumb() == null || existing.getBreadcrumb().isBlank())
|
||||
&& incoming.getBreadcrumb() != null) {
|
||||
existing.setBreadcrumb(incoming.getBreadcrumb());
|
||||
}
|
||||
}
|
||||
|
||||
private List<String> mergeReasons(List<String> base, List<String> boosts) {
|
||||
Set<String> merged = new LinkedHashSet<>();
|
||||
if (base != null) {
|
||||
merged.addAll(base);
|
||||
}
|
||||
if (boosts != null) {
|
||||
merged.addAll(boosts);
|
||||
}
|
||||
return new ArrayList<>(merged);
|
||||
}
|
||||
|
||||
private String sourceKey(EvidenceBlock block, String fallback) {
|
||||
return firstNonBlank(block.getSource(), block.getTitle(), block.getBreadcrumb(), fallback);
|
||||
}
|
||||
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String nullToEmpty(String value) {
|
||||
return value == null ? "" : value;
|
||||
}
|
||||
|
||||
private record ScoredCandidate(RetrievedEvidenceCandidate candidate,
|
||||
double baseScore,
|
||||
double finalScore,
|
||||
List<String> boostReasons) {
|
||||
}
|
||||
|
||||
private record RelevanceAssessment(String level, String hint) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,38 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Converts a raw Agent query into retrieval-control hints.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeQueryTransformer {
|
||||
|
||||
private final KnowledgeIndexService knowledgeIndexService;
|
||||
|
||||
public KnowledgeQueryTransformer(KnowledgeIndexService knowledgeIndexService) {
|
||||
this.knowledgeIndexService = knowledgeIndexService;
|
||||
}
|
||||
|
||||
public KnowledgeQuery transform(String rawQuery) {
|
||||
String normalized = rawQuery == null ? "" : rawQuery.trim();
|
||||
KnowledgeIndexService.L0Hint hint = knowledgeIndexService.analyzeQuery(normalized);
|
||||
return KnowledgeQuery.builder()
|
||||
.originalQuery(normalized)
|
||||
.rewrittenQuery(normalized)
|
||||
.domainHints(safeList(hint.domains()))
|
||||
.matchedKeywords(safeList(hint.matchedKeywords()))
|
||||
.entities(safeList(hint.entities()))
|
||||
.categoryFilter(hint.singleDomainOrNull())
|
||||
.l0Titles(safeList(hint.titles()))
|
||||
.l0MatchCount(hint.matches() == null ? 0 : hint.matches().size())
|
||||
.build();
|
||||
}
|
||||
|
||||
private List<String> safeList(List<String> values) {
|
||||
return values == null ? List.of() : values;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Assembles evidence-first LookupResult instances.
|
||||
*/
|
||||
@Service
|
||||
public class LookupResultAssembler {
|
||||
|
||||
public LookupResult assemble(EvidencePostprocessResult evidence,
|
||||
ContextPack contextPack,
|
||||
RetrievalTrace retrievalTrace) {
|
||||
boolean found = evidence != null && evidence.hasUsableEvidence();
|
||||
return LookupResult.builder()
|
||||
.found(found)
|
||||
.evidenceBlocks(evidence != null ? evidence.getEvidenceBlocks() : List.of())
|
||||
.evidenceCandidateCount(evidence != null ? evidence.getCandidateCount() : 0)
|
||||
.evidenceBlockCount(evidence != null ? evidence.getEvidenceBlockCount() : 0)
|
||||
.contextPack(contextPack)
|
||||
.retrievalTrace(retrievalTrace)
|
||||
.rerankTrace(evidence != null ? evidence.getRerankTrace() : null)
|
||||
.relevanceLevel(evidence != null ? evidence.getRelevanceLevel() : null)
|
||||
.completenessHint(evidence != null ? evidence.getCompletenessHint() : null)
|
||||
.message(found ? null : "知识库未检索到可用证据,请结合日志、指标、告警继续排查")
|
||||
.build();
|
||||
}
|
||||
|
||||
public LookupResult deduped(LookupResult original, List<String> retrievedDomains, String docKey) {
|
||||
return LookupResult.builder()
|
||||
.found(false)
|
||||
.message("文档已在本会话中检索过,无需重复召回: " + docKey)
|
||||
.evidenceBlocks(List.of())
|
||||
.evidenceCandidateCount(0)
|
||||
.evidenceBlockCount(0)
|
||||
.retrievalTrace(original.getRetrievalTrace())
|
||||
.rerankTrace(original.getRerankTrace())
|
||||
.relevanceLevel(original.getRelevanceLevel())
|
||||
.completenessHint(original.getCompletenessHint())
|
||||
.retrievedDomainsThisSession(retrievedDomains)
|
||||
.build();
|
||||
}
|
||||
}
|
||||
@@ -3,10 +3,13 @@ package com.superbiz.agent.service;
|
||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import lombok.Builder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
@@ -138,6 +141,21 @@ public class ToolInvocationRecorder {
|
||||
if (record.evidenceBlocks() != null && !record.evidenceBlocks().isEmpty()) {
|
||||
details.put("evidence_blocks", record.evidenceBlocks());
|
||||
}
|
||||
if (record.queryTransform() != null && !record.queryTransform().isEmpty()) {
|
||||
details.put("query_transform", record.queryTransform());
|
||||
}
|
||||
if (record.retrievalTrace() != null && !record.retrievalTrace().isEmpty()) {
|
||||
details.put("retrieval_trace", record.retrievalTrace());
|
||||
}
|
||||
if (record.contextPack() != null && !record.contextPack().isEmpty()) {
|
||||
details.put("context_pack_summary", record.contextPack());
|
||||
}
|
||||
if (record.rerankTrace() != null && !record.rerankTrace().isEmpty()) {
|
||||
details.put("rerank_trace", record.rerankTrace());
|
||||
}
|
||||
if (record.fallbackReason() != null && !record.fallbackReason().isBlank()) {
|
||||
details.put("fallback_reason", record.fallbackReason());
|
||||
}
|
||||
if (record.relevanceLevel() != null) {
|
||||
details.put("relevance_level", record.relevanceLevel());
|
||||
}
|
||||
@@ -222,40 +240,33 @@ public class ToolInvocationRecorder {
|
||||
List<Double> l1Scores,
|
||||
Integer evidenceCandidateCount,
|
||||
Integer evidenceBlockCount,
|
||||
List<Map<String, Object>> evidenceBlocks
|
||||
List<Map<String, Object>> evidenceBlocks,
|
||||
Map<String, Object> queryTransform,
|
||||
Map<String, Object> retrievalTrace,
|
||||
Map<String, Object> contextPack,
|
||||
Map<String, Object> rerankTrace,
|
||||
String fallbackReason
|
||||
) {
|
||||
public static LookupKnowledgeRecord from(String query,
|
||||
KnowledgeIndexService.L0Hint l0Hint,
|
||||
List<VectorSearchService.SearchResult> l1Results,
|
||||
boolean highConfidence,
|
||||
public static LookupKnowledgeRecord from(KnowledgeQuery query,
|
||||
LookupResult result,
|
||||
String domain,
|
||||
String dedupReason,
|
||||
int durationMs,
|
||||
double l1TopSimilarity) {
|
||||
List<KnowledgeEntry> l0Matches = l0Hint != null ? l0Hint.matches() : List.of();
|
||||
boolean hasL0 = l0Matches != null && !l0Matches.isEmpty();
|
||||
boolean hasL1 = l1Results != null && !l1Results.isEmpty();
|
||||
String layer;
|
||||
if (hasL0 && !highConfidence) {
|
||||
layer = "L0+L1";
|
||||
} else if (hasL0) {
|
||||
layer = "L0";
|
||||
} else if (hasL1) {
|
||||
layer = "L1";
|
||||
} else {
|
||||
layer = null;
|
||||
}
|
||||
int durationMs) {
|
||||
RetrievalTrace trace = result != null ? result.getRetrievalTrace() : null;
|
||||
String layer = trace != null ? trace.getSelectedAttempt() : null;
|
||||
|
||||
String outputPreview = null;
|
||||
int outputLength = 0;
|
||||
boolean truncated = false;
|
||||
if (result != null && result.getPrimary() != null && result.getPrimary().getContent() != null) {
|
||||
outputPreview = result.getPrimary().getContent();
|
||||
if (result != null && result.getContextPack() != null
|
||||
&& result.getContextPack().getPackedText() != null) {
|
||||
outputPreview = result.getContextPack().getPackedText();
|
||||
outputLength = outputPreview.length();
|
||||
truncated = outputLength > OUTPUT_PREVIEW_LIMIT;
|
||||
} else if (hasL1 && l1Results.get(0).getContent() != null) {
|
||||
outputPreview = l1Results.get(0).getContent();
|
||||
} else if (result != null && result.getEvidenceBlocks() != null
|
||||
&& !result.getEvidenceBlocks().isEmpty()
|
||||
&& result.getEvidenceBlocks().get(0).getContent() != null) {
|
||||
outputPreview = result.getEvidenceBlocks().get(0).getContent();
|
||||
outputLength = outputPreview.length();
|
||||
truncated = outputLength > OUTPUT_PREVIEW_LIMIT;
|
||||
}
|
||||
@@ -265,29 +276,19 @@ public class ToolInvocationRecorder {
|
||||
evidenceStatus = EVIDENCE_STATUS_DEDUPED;
|
||||
} else if (result == null || !result.isFound()) {
|
||||
evidenceStatus = EVIDENCE_STATUS_NO_EVIDENCE;
|
||||
} else if (trace != null && trace.getEvidenceStatus() != null) {
|
||||
evidenceStatus = trace.getEvidenceStatus();
|
||||
}
|
||||
|
||||
List<String> l0Titles = new ArrayList<>();
|
||||
if (hasL0) {
|
||||
for (int i = 0; i < Math.min(3, l0Matches.size()); i++) {
|
||||
l0Titles.add(l0Matches.get(i).getTitle());
|
||||
}
|
||||
}
|
||||
|
||||
List<Double> l1Scores = new ArrayList<>();
|
||||
if (hasL1) {
|
||||
for (int i = 0; i < Math.min(3, l1Results.size()); i++) {
|
||||
l1Scores.add((double) l1Results.get(i).getScore());
|
||||
}
|
||||
}
|
||||
List<Double> l1Scores = collectAttemptScores(trace);
|
||||
|
||||
return LookupKnowledgeRecord.builder()
|
||||
.query(query)
|
||||
.query(query != null ? query.getOriginalQuery() : null)
|
||||
.outputPreview(outputPreview)
|
||||
.outputLength(outputLength)
|
||||
.retrievalLayer(layer)
|
||||
.l0MatchCount(hasL0 ? l0Matches.size() : null)
|
||||
.l1MatchCount(hasL1 ? l1Results.size() : null)
|
||||
.l0MatchCount(query != null ? query.getL0MatchCount() : null)
|
||||
.l1MatchCount(totalCandidateCount(trace))
|
||||
.truncated(truncated)
|
||||
.relevanceLevel(result != null ? result.getRelevanceLevel() : null)
|
||||
.completenessHint(result != null ? result.getCompletenessHint() : null)
|
||||
@@ -296,16 +297,21 @@ public class ToolInvocationRecorder {
|
||||
.durationMs(durationMs)
|
||||
.success(true)
|
||||
.evidenceStatus(evidenceStatus)
|
||||
.l0Titles(l0Titles)
|
||||
.l0MatchedKeywords(l0Hint != null ? l0Hint.matchedKeywords() : List.of())
|
||||
.l0Domains(l0Hint != null ? l0Hint.domains() : List.of())
|
||||
.l0Entities(l0Hint != null ? l0Hint.entities() : List.of())
|
||||
.l1TopScore(hasL1 ? (double) l1Results.get(0).getScore() : null)
|
||||
.l1TopSimilarity(hasL1 ? l1TopSimilarity : null)
|
||||
.l0Titles(query != null ? query.getL0Titles() : List.of())
|
||||
.l0MatchedKeywords(query != null ? query.getMatchedKeywords() : List.of())
|
||||
.l0Domains(query != null ? query.getDomainHints() : List.of())
|
||||
.l0Entities(query != null ? query.getEntities() : List.of())
|
||||
.l1TopScore(firstAttemptScore(trace))
|
||||
.l1TopSimilarity(firstAttemptSimilarity(trace))
|
||||
.l1Scores(l1Scores)
|
||||
.evidenceCandidateCount(result != null ? result.getEvidenceCandidateCount() : null)
|
||||
.evidenceBlockCount(result != null ? result.getEvidenceBlockCount() : null)
|
||||
.evidenceBlocks(result != null ? summarizeEvidenceBlocks(result.getEvidenceBlocks()) : List.of())
|
||||
.queryTransform(summarizeQueryTransform(query))
|
||||
.retrievalTrace(summarizeRetrievalTrace(trace))
|
||||
.contextPack(summarizeContextPack(result != null ? result.getContextPack() : null))
|
||||
.rerankTrace(summarizeRerankTrace(result != null ? result.getRerankTrace() : null))
|
||||
.fallbackReason(trace != null ? trace.getFallbackReason() : null)
|
||||
.build();
|
||||
}
|
||||
|
||||
@@ -331,5 +337,130 @@ public class ToolInvocationRecorder {
|
||||
}
|
||||
return summaries;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeQueryTransform(KnowledgeQuery query) {
|
||||
if (query == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("original_query", query.getOriginalQuery());
|
||||
summary.put("rewritten_query", query.getRewrittenQuery());
|
||||
summary.put("category_filter", query.getCategoryFilter());
|
||||
summary.put("domain_hints", query.getDomainHints());
|
||||
summary.put("matched_keywords", query.getMatchedKeywords());
|
||||
summary.put("entities", query.getEntities());
|
||||
summary.put("l0_titles", query.getL0Titles());
|
||||
summary.put("l0_match_count", query.getL0MatchCount());
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeRetrievalTrace(RetrievalTrace trace) {
|
||||
if (trace == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("selected_attempt", trace.getSelectedAttempt());
|
||||
summary.put("fallback_reason", trace.getFallbackReason());
|
||||
summary.put("evidence_status", trace.getEvidenceStatus());
|
||||
summary.put("category_filter", trace.getCategoryFilter());
|
||||
List<Map<String, Object>> attempts = new ArrayList<>();
|
||||
if (trace.getAttempts() != null) {
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
Map<String, Object> item = new LinkedHashMap<>();
|
||||
item.put("name", attempt.getName());
|
||||
item.put("category_filter", attempt.getCategoryFilter());
|
||||
item.put("candidate_count", attempt.getCandidateCount());
|
||||
item.put("usable", attempt.getUsable());
|
||||
item.put("duration_ms", attempt.getDurationMs());
|
||||
item.put("top_score", attempt.getTopScore());
|
||||
item.put("top_similarity", attempt.getTopSimilarity());
|
||||
item.put("error_message", attempt.getErrorMessage());
|
||||
attempts.add(item);
|
||||
}
|
||||
}
|
||||
summary.put("attempts", attempts);
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeContextPack(ContextPack contextPack) {
|
||||
if (contextPack == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("strategy", contextPack.getStrategy());
|
||||
summary.put("char_budget", contextPack.getCharBudget());
|
||||
summary.put("used_chars", contextPack.getUsedChars());
|
||||
summary.put("included_sources", contextPack.getIncludedSources());
|
||||
summary.put("omitted_sources", contextPack.getOmittedSources());
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeRerankTrace(RerankTrace trace) {
|
||||
if (trace == null || trace.getItems() == null || trace.getItems().isEmpty()) {
|
||||
return Map.of();
|
||||
}
|
||||
List<Map<String, Object>> items = new ArrayList<>();
|
||||
for (int i = 0; i < Math.min(5, trace.getItems().size()); i++) {
|
||||
RerankTrace.Item item = trace.getItems().get(i);
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("final_rank", item.getFinalRank());
|
||||
summary.put("source", item.getSource());
|
||||
summary.put("base_score", item.getBaseScore());
|
||||
summary.put("final_score", item.getFinalScore());
|
||||
summary.put("boost_reasons", item.getBoostReasons());
|
||||
items.add(summary);
|
||||
}
|
||||
return Map.of("items", items);
|
||||
}
|
||||
|
||||
private static Integer totalCandidateCount(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
int total = 0;
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getCandidateCount() != null) {
|
||||
total += attempt.getCandidateCount();
|
||||
}
|
||||
}
|
||||
return total;
|
||||
}
|
||||
|
||||
private static Double firstAttemptScore(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopScore() != null) {
|
||||
return attempt.getTopScore();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static Double firstAttemptSimilarity(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopSimilarity() != null) {
|
||||
return attempt.getTopSimilarity();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static List<Double> collectAttemptScores(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<Double> scores = new ArrayList<>();
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopScore() != null) {
|
||||
scores.add(attempt.getTopScore());
|
||||
}
|
||||
}
|
||||
return scores;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,11 +1,17 @@
|
||||
package com.superbiz.agent.tool;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.*;
|
||||
import com.superbiz.agent.service.KnowledgeIndexService;
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.service.KnowledgeContextPacker;
|
||||
import com.superbiz.agent.service.KnowledgeDocumentRetriever;
|
||||
import com.superbiz.agent.service.KnowledgeEvidencePostProcessor;
|
||||
import com.superbiz.agent.service.KnowledgeQueryTransformer;
|
||||
import com.superbiz.agent.service.LookupResultAssembler;
|
||||
import com.superbiz.agent.service.ToolInvocationRecorder;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.tool.annotation.Tool;
|
||||
@@ -16,41 +22,39 @@ import org.springframework.stereotype.Component;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.stream.Collectors;
|
||||
import java.util.UUID;
|
||||
|
||||
/**
|
||||
* 知识库查询工具
|
||||
* 提供给 Agent 的混合检索工具(L0 + L1)
|
||||
* 内置归一化层:将 L0 匹配数 + L1 L2 距离归一化为统一质量等级
|
||||
* Agent-facing explicit knowledge retrieval tool.
|
||||
*/
|
||||
@Slf4j
|
||||
@Component
|
||||
public class LookupKnowledgeTool {
|
||||
|
||||
private static final String LEVEL_PRECISE = "PRECISE";
|
||||
private static final String LEVEL_HIGHLY_RELEVANT = "HIGHLY_RELEVANT";
|
||||
private static final String LEVEL_REFERENCE = "REFERENCE";
|
||||
private static final String ATTEMPT_FILTERED_VECTOR = "FILTERED_VECTOR";
|
||||
private static final String ATTEMPT_UNFILTERED_VECTOR = "UNFILTERED_VECTOR";
|
||||
private static final String ATTEMPT_UNFILTERED_VECTOR_RETRY = "UNFILTERED_VECTOR_RETRY";
|
||||
private static final String FALLBACK_NO_EVIDENCE = "filtered_vector_no_evidence";
|
||||
private static final String FALLBACK_LOW_QUALITY = "filtered_vector_low_quality";
|
||||
|
||||
private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档";
|
||||
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
|
||||
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
|
||||
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
@Value("${retrieval.normalization.highly-relevant-threshold:0.75}")
|
||||
private double highlyRelevantThreshold = 0.75;
|
||||
|
||||
@Value("${retrieval.normalization.reference-threshold:0.5}")
|
||||
private double referenceThreshold = 0.5;
|
||||
@Value("${rag.top-k:3}")
|
||||
private int topK = 3;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeIndexService knowledgeIndexService;
|
||||
private KnowledgeQueryTransformer queryTransformer;
|
||||
|
||||
@Autowired
|
||||
private VectorSearchService vectorSearchService;
|
||||
private KnowledgeDocumentRetriever documentRetriever;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeEvidencePostProcessor evidencePostProcessor;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeContextPacker contextPacker;
|
||||
|
||||
@Autowired
|
||||
private LookupResultAssembler resultAssembler;
|
||||
|
||||
@Autowired
|
||||
private ToolInvocationRecorder toolInvocationRecorder;
|
||||
@@ -58,26 +62,20 @@ public class LookupKnowledgeTool {
|
||||
@Autowired
|
||||
private RetrievedDocTracker retrievedDocTracker;
|
||||
|
||||
@Autowired
|
||||
private ObjectMapper objectMapper;
|
||||
|
||||
/**
|
||||
* 查询知识库文档
|
||||
* 查询知识库文档。
|
||||
*
|
||||
* @param query 查询关键词
|
||||
* @return 查询结果
|
||||
*/
|
||||
@Tool(description = "查询内部知识库文档,获取错误码定义、接口文档、排障步骤、配置说明等背景信息。" +
|
||||
"采用两阶段检索:L0 精确匹配关键词(< 10ms),L1 语义检索补充(200-500ms)。" +
|
||||
"IMPORTANT: 遇到错误码、接口名、配置项、排障问题时,优先使用此工具。" +
|
||||
"支持的查询场景:" +
|
||||
"1) 错误码定义 - 查询错误码的含义和处理方法,例如 'ERR_TIMEOUT'、'ERR_CONNECTION_REFUSED';" +
|
||||
"2) 接口文档 - 查询 API 接口定义、参数说明、返回格式,例如 'payment-gateway'、'/api/v1/orders';" +
|
||||
"3) 排障步骤 - 查询故障诊断流程、最佳实践,例如 '支付超时排查'、'数据库连接池配置';" +
|
||||
"4) 配置说明 - 查询系统配置、中间件参数,例如 'HikariCP'、'Redis 集群配置'。" +
|
||||
@Tool(description = "查询内部知识库文档,获取错误码定义、接口文档、排障步骤、配置说明和历史案例等背景知识。" +
|
||||
"工具会执行 query understanding、向量检索、证据去重、轻量重排和上下文打包,并返回 evidenceBlocks、contextPack、retrievalTrace、rerankTrace。" +
|
||||
"L0 仅用于领域/关键词/实体 hint 和可选 metadata filter,不作为事实证据兜底。" +
|
||||
"当带 filter 的向量检索低质量或无结果时,会跳过 L0 filter,用原始 query 再做一次无过滤语义检索。" +
|
||||
"不要用此工具查询实时运行状态;日志、指标、告警等实时事实应使用 query_logs、query_metrics 或告警工具。" +
|
||||
"参数 query: 查询关键词或描述")
|
||||
public LookupResult lookupKnowledge(String query) {
|
||||
String requestId = java.util.UUID.randomUUID().toString().substring(0, 8);
|
||||
String requestId = UUID.randomUUID().toString().substring(0, 8);
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
log.info("========================================");
|
||||
@@ -86,566 +84,179 @@ public class LookupKnowledgeTool {
|
||||
log.info(">>> RequestId: {}", requestId);
|
||||
log.info("----------------------------------------");
|
||||
|
||||
// Step 1: L0 hint 分析
|
||||
long l0Start = System.currentTimeMillis();
|
||||
KnowledgeIndexService.L0Hint l0Hint = knowledgeIndexService.analyzeQuery(query);
|
||||
List<KnowledgeEntry> l0Matches = l0Hint.matches();
|
||||
long l0Time = System.currentTimeMillis() - l0Start;
|
||||
log.info("[L0 Hint] 完成: matches={}, domains={}, keywords={}, time={}ms",
|
||||
l0Matches.size(), l0Hint.domains(), l0Hint.matchedKeywords(), l0Time);
|
||||
if (!l0Matches.isEmpty()) {
|
||||
log.info("[L0 Hint] 找到文档:");
|
||||
for (int i = 0; i < Math.min(3, l0Matches.size()); i++) {
|
||||
KnowledgeEntry entry = l0Matches.get(i);
|
||||
log.info(" - [{}] 标题: {}, 路径: {}, 域: {}", i+1, entry.getTitle(), entry.getFilePath(), entry.getCategory());
|
||||
}
|
||||
KnowledgeQuery knowledgeQuery = queryTransformer.transform(query);
|
||||
log.info("[QueryTransformer] rewrittenQuery={}, categoryFilter={}, domains={}, keywords={}",
|
||||
knowledgeQuery.getRewrittenQuery(),
|
||||
knowledgeQuery.getCategoryFilter(),
|
||||
knowledgeQuery.getDomainHints(),
|
||||
knowledgeQuery.getMatchedKeywords());
|
||||
|
||||
List<RetrievalTrace.Attempt> attempts = new ArrayList<>();
|
||||
String fallbackReason = null;
|
||||
|
||||
String firstAttemptName = knowledgeQuery.getCategoryFilter() == null
|
||||
? ATTEMPT_UNFILTERED_VECTOR
|
||||
: ATTEMPT_FILTERED_VECTOR;
|
||||
KnowledgeDocumentRetriever.RetrievalAttemptResult firstAttempt =
|
||||
documentRetriever.retrieve(firstAttemptName,
|
||||
knowledgeQuery.getRewrittenQuery(),
|
||||
knowledgeQuery.getCategoryFilter(),
|
||||
topK);
|
||||
EvidencePostprocessResult selectedEvidence = evidencePostProcessor.process(
|
||||
knowledgeQuery,
|
||||
firstAttempt.candidates());
|
||||
enrichAttempt(firstAttempt.attempt(), selectedEvidence);
|
||||
attempts.add(firstAttempt.attempt());
|
||||
String selectedAttemptName = firstAttemptName;
|
||||
|
||||
if (knowledgeQuery.getCategoryFilter() != null && evidencePostProcessor.isLowQuality(selectedEvidence)) {
|
||||
fallbackReason = selectedEvidence.hasUsableEvidence()
|
||||
? FALLBACK_LOW_QUALITY
|
||||
: FALLBACK_NO_EVIDENCE;
|
||||
log.info("[RetrievalFallback] {} -> retry without category filter", fallbackReason);
|
||||
|
||||
KnowledgeDocumentRetriever.RetrievalAttemptResult retryAttempt =
|
||||
documentRetriever.retrieve(ATTEMPT_UNFILTERED_VECTOR_RETRY,
|
||||
knowledgeQuery.getOriginalQuery(),
|
||||
null,
|
||||
topK);
|
||||
EvidencePostprocessResult retryEvidence = evidencePostProcessor.process(
|
||||
knowledgeQuery,
|
||||
retryAttempt.candidates());
|
||||
enrichAttempt(retryAttempt.attempt(), retryEvidence);
|
||||
attempts.add(retryAttempt.attempt());
|
||||
selectedEvidence = retryEvidence;
|
||||
selectedAttemptName = ATTEMPT_UNFILTERED_VECTOR_RETRY;
|
||||
}
|
||||
|
||||
// Step 2: L1 默认调用;L0 只提供可解释 hint 和可选 category filter
|
||||
List<VectorSearchService.SearchResult> l1Results = List.of();
|
||||
String l0CategoryFilter = l0Hint.singleDomainOrNull();
|
||||
try {
|
||||
log.info("[L1 语义检索] 触发L1语义检索, categoryFilter={}", l0CategoryFilter);
|
||||
long l1Start = System.currentTimeMillis();
|
||||
l1Results = vectorSearchService.searchSimilarDocuments(query, 3, l0CategoryFilter);
|
||||
long l1Time = System.currentTimeMillis() - l1Start;
|
||||
log.info("[L1 语义检索] 完成: matches={}, time={}ms",
|
||||
l1Results != null ? l1Results.size() : 0, l1Time);
|
||||
if (l1Results != null && !l1Results.isEmpty()) {
|
||||
log.info("[L1 语义检索] 找到文档:");
|
||||
for (int i = 0; i < Math.min(3, l1Results.size()); i++) {
|
||||
VectorSearchService.SearchResult result = l1Results.get(i);
|
||||
log.info(" - [{}] 文档ID: {}, L2距离: {}", i+1, result.getId(), String.format("%.4f", result.getScore()));
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.warn("[L1 语义检索] 调用失败,保留L0 fallback: {}", e.getMessage());
|
||||
l1Results = List.of();
|
||||
}
|
||||
ContextPack contextPack = contextPacker.pack(selectedEvidence.getEvidenceBlocks());
|
||||
RetrievalTrace retrievalTrace = buildRetrievalTrace(knowledgeQuery, attempts, selectedAttemptName,
|
||||
fallbackReason, selectedEvidence);
|
||||
LookupResult result = resultAssembler.assemble(selectedEvidence, contextPack, retrievalTrace);
|
||||
|
||||
// Step 3: 归一化质量等级判定
|
||||
float l1TopScore = (l1Results != null && !l1Results.isEmpty()) ? l1Results.get(0).getScore() : Float.MAX_VALUE;
|
||||
RelevanceAssessment assessment = computeRelevance(l0Matches.size(), l1TopScore);
|
||||
boolean highConfidence = isHighConfidence(l0Matches.size(), l1TopScore);
|
||||
log.info("[归一化] relevanceLevel={}, completenessHint={}", assessment.level, assessment.hint);
|
||||
if (l1TopScore != Float.MAX_VALUE) {
|
||||
double similarity = normalizeL2(l1TopScore);
|
||||
log.info("[归一化] L2距离={}, similarity={}", String.format("%.4f", l1TopScore), String.format("%.4f", similarity));
|
||||
}
|
||||
|
||||
// Step 4: 组装结果
|
||||
LookupResult result = buildResult(l0Matches, l1Results, highConfidence);
|
||||
result.setRelevanceLevel(assessment.level);
|
||||
result.setCompletenessHint(assessment.hint);
|
||||
|
||||
// Step 5: session 级去重过滤 + 域级行动记忆
|
||||
String sessionId = SessionContextHolder.getSessionId();
|
||||
String domain = extractDomain(l0Matches, l1Results);
|
||||
|
||||
String domain = extractDomain(knowledgeQuery);
|
||||
if (sessionId != null && result.isFound()) {
|
||||
String docKey = extractDocKey(result);
|
||||
if (docKey != null && retrievedDocTracker.isAlreadyRetrieved(sessionId, docKey)) {
|
||||
log.info("[去重] 文档已在本会话中检索过,跳过: {}", docKey);
|
||||
List<String> retrievedDomains = retrievedDocTracker.getRetrievedDomains(sessionId);
|
||||
saveToolInvocation(query, l0Hint, l1Results, highConfidence, startTime, result, domain, "doc_retrieved");
|
||||
return LookupResult.builder()
|
||||
.found(false)
|
||||
.message("文档已在本会话中检索过,无需重复召回: " + docKey)
|
||||
.relevanceLevel(assessment.level)
|
||||
.completenessHint(assessment.hint)
|
||||
.retrievedDomainsThisSession(retrievedDomains)
|
||||
.build();
|
||||
saveToolInvocation(knowledgeQuery, startTime, result, domain, "doc_retrieved");
|
||||
LookupResult deduped = resultAssembler.deduped(
|
||||
result,
|
||||
retrievedDocTracker.getRetrievedDomains(sessionId),
|
||||
docKey);
|
||||
logReturn(deduped, startTime);
|
||||
return deduped;
|
||||
}
|
||||
if (docKey != null) {
|
||||
retrievedDocTracker.markRetrieved(sessionId, domain, docKey);
|
||||
}
|
||||
}
|
||||
|
||||
// 附加行动记忆
|
||||
if (sessionId != null) {
|
||||
result.setRetrievedDomainsThisSession(retrievedDocTracker.getRetrievedDomains(sessionId));
|
||||
}
|
||||
|
||||
// 记录结构化结果摘要
|
||||
long totalTime = System.currentTimeMillis() - startTime;
|
||||
log.info("----------------------------------------");
|
||||
log.info("<<< [工具返回] lookup_knowledge");
|
||||
log.info("<<< 结果: found={}, relevanceLevel={}, 耗时: {}ms",
|
||||
result.isFound(), result.getRelevanceLevel(), totalTime);
|
||||
log.info("<<< 行动记忆: retrievedDomainsThisSession={}", result.getRetrievedDomainsThisSession());
|
||||
|
||||
if (!l0Matches.isEmpty()) {
|
||||
KnowledgeEntry top = l0Matches.get(0);
|
||||
log.info("<<< [L0 主结果] 标题: {}", top.getTitle());
|
||||
log.info("<<< [L0 主结果] 来源: {}", top.getFilePath());
|
||||
log.info("<<< [L0 主结果] 域: {}", top.getCategory());
|
||||
if (top.getSummary() != null) {
|
||||
log.info("<<< [L0 主结果] 摘要: {}", top.getSummary());
|
||||
}
|
||||
String content = result.getPrimary() != null ? result.getPrimary().getContent() : null;
|
||||
if (content != null) {
|
||||
int headingCount = countMdHeadings(content);
|
||||
log.info("<<< [L0 主结果] 内容: {} 字符, {} 个章节", content.length(), headingCount);
|
||||
}
|
||||
}
|
||||
|
||||
if (l1Results != null && !l1Results.isEmpty()) {
|
||||
VectorSearchService.SearchResult topL1 = l1Results.get(0);
|
||||
log.info("<<< [L1 补充] 来源: {}", topL1.getMetadata() != null ? topL1.getMetadata() : topL1.getId());
|
||||
log.info("<<< [L1 补充] L2距离: {}, similarity: {}",
|
||||
String.format("%.4f", topL1.getScore()),
|
||||
String.format("%.4f", normalizeL2(topL1.getScore())));
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
|
||||
// 记录 tool_invocation
|
||||
saveToolInvocation(query, l0Hint, l1Results, highConfidence, startTime, result, domain, null);
|
||||
|
||||
saveToolInvocation(knowledgeQuery, startTime, result, domain, null);
|
||||
logReturn(result, startTime);
|
||||
return result;
|
||||
}
|
||||
|
||||
// ==================== 归一化层 ====================
|
||||
|
||||
/**
|
||||
* L2 距离 Min-Max 归一化到 [0,1] similarity
|
||||
* BGE-M3 输出 L2 归一化单位向量,L2 距离硬上界 = 2.0
|
||||
* similarity = 1 - min(score, maxL2Distance) / maxL2Distance
|
||||
* score=0 → 1.0(完全相同),score=2.0 → 0.0(完全相反)
|
||||
*/
|
||||
double normalizeL2(float l2Score) {
|
||||
double clamped = Math.min(l2Score, maxL2Distance);
|
||||
return 1.0 - clamped / maxL2Distance;
|
||||
private void enrichAttempt(RetrievalTrace.Attempt attempt, EvidencePostprocessResult evidence) {
|
||||
attempt.setTopSimilarity(evidence.getTopSimilarity());
|
||||
attempt.setUsable(evidence.hasUsableEvidence()
|
||||
&& evidence.getTopSimilarity() != null
|
||||
&& evidence.getTopSimilarity() >= evidencePostProcessor.getReferenceThreshold());
|
||||
}
|
||||
|
||||
/**
|
||||
* 归一化质量等级判定
|
||||
*
|
||||
* @param l0MatchCount L0 匹配数
|
||||
* @param l1TopScore L1 最高分(L2 距离),无 L1 结果时传 Float.MAX_VALUE
|
||||
* @return RelevanceAssessment(level + hint)
|
||||
*/
|
||||
RelevanceAssessment computeRelevance(int l0MatchCount, float l1TopScore) {
|
||||
double l1Similarity = (l1TopScore != Float.MAX_VALUE) ? normalizeL2(l1TopScore) : 0.0;
|
||||
private RetrievalTrace buildRetrievalTrace(KnowledgeQuery query,
|
||||
List<RetrievalTrace.Attempt> attempts,
|
||||
String selectedAttempt,
|
||||
String fallbackReason,
|
||||
EvidencePostprocessResult evidence) {
|
||||
Map<String, Object> queryHints = new LinkedHashMap<>();
|
||||
queryHints.put("domains", query.getDomainHints());
|
||||
queryHints.put("matched_keywords", query.getMatchedKeywords());
|
||||
queryHints.put("entities", query.getEntities());
|
||||
queryHints.put("l0_titles", query.getL0Titles());
|
||||
queryHints.put("l0_match_count", query.getL0MatchCount());
|
||||
|
||||
// L0 唯一匹配 + L1 高分 → PRECISE
|
||||
if (l0MatchCount == 1 && l1Similarity >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE);
|
||||
return RetrievalTrace.builder()
|
||||
.originalQuery(query.getOriginalQuery())
|
||||
.rewrittenQuery(query.getRewrittenQuery())
|
||||
.categoryFilter(query.getCategoryFilter())
|
||||
.selectedAttempt(selectedAttempt)
|
||||
.fallbackReason(fallbackReason)
|
||||
.evidenceStatus(evidence.hasUsableEvidence()
|
||||
? ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED
|
||||
: ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE)
|
||||
.queryHints(queryHints)
|
||||
.attempts(attempts)
|
||||
.build();
|
||||
}
|
||||
|
||||
// L0 唯一匹配但缺少 L1 支持 → REFERENCE
|
||||
if (l0MatchCount == 1) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
private String extractDomain(KnowledgeQuery query) {
|
||||
if (query.getCategoryFilter() != null && !query.getCategoryFilter().isBlank()) {
|
||||
return query.getCategoryFilter();
|
||||
}
|
||||
|
||||
// L0 命中 + L1 高分 → HIGHLY_RELEVANT
|
||||
if (l0MatchCount > 1 && l1Similarity >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
|
||||
if (query.getDomainHints() != null && !query.getDomainHints().isEmpty()) {
|
||||
return query.getDomainHints().get(0);
|
||||
}
|
||||
|
||||
// 仅 L1 高分 → HIGHLY_RELEVANT
|
||||
if (l0MatchCount == 0 && l1Similarity >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
|
||||
}
|
||||
|
||||
// L0 多匹配 + L1 中分 → REFERENCE
|
||||
if (l0MatchCount > 1 && l1Similarity >= referenceThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
|
||||
// 仅 L1 中分 → REFERENCE
|
||||
if (l0MatchCount == 0 && l1Similarity >= referenceThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
|
||||
// L0 多匹配 + 无 L1 / L1 低分 → REFERENCE(L0 命中本身有价值)
|
||||
if (l0MatchCount > 1) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
|
||||
// 无有效结果
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
|
||||
boolean isHighConfidence(int l0MatchCount, float l1TopScore) {
|
||||
if (l0MatchCount != 1) {
|
||||
return false;
|
||||
}
|
||||
if (l1TopScore == Float.MAX_VALUE) {
|
||||
return true;
|
||||
}
|
||||
return normalizeL2(l1TopScore) >= highlyRelevantThreshold;
|
||||
}
|
||||
|
||||
/**
|
||||
* 归一化评估结果
|
||||
*/
|
||||
record RelevanceAssessment(String level, String hint) {}
|
||||
|
||||
// ==================== 域提取 ====================
|
||||
|
||||
/**
|
||||
* 从检索结果中提取域信息
|
||||
* 优先使用 L0 的 category,兜底从 L1 metadata 解析
|
||||
*/
|
||||
private String extractDomain(List<KnowledgeEntry> l0Matches, List<VectorSearchService.SearchResult> l1Results) {
|
||||
// 优先 L0
|
||||
if (l0Matches != null && !l0Matches.isEmpty()) {
|
||||
String category = l0Matches.get(0).getCategory();
|
||||
if (category != null && !category.isBlank()) {
|
||||
return category;
|
||||
}
|
||||
}
|
||||
|
||||
// 兜底 L1:从 metadata JSON 中解析 category
|
||||
if (l1Results != null && !l1Results.isEmpty()) {
|
||||
try {
|
||||
String metadata = l1Results.get(0).getMetadata();
|
||||
if (metadata != null && metadata.contains("category")) {
|
||||
var node = objectMapper.readTree(metadata);
|
||||
if (node.has("category")) {
|
||||
return node.get("category").asText();
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("L1 metadata 解析 category 失败: {}", e.getMessage());
|
||||
}
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
// ==================== 入库 ====================
|
||||
private String extractDocKey(LookupResult result) {
|
||||
if (result.getEvidenceBlocks() == null || result.getEvidenceBlocks().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
EvidenceBlock first = result.getEvidenceBlocks().get(0);
|
||||
if (first.getSource() != null && !first.getSource().isBlank()) {
|
||||
return first.getSource();
|
||||
}
|
||||
if (first.getTitle() != null && !first.getTitle().isBlank()) {
|
||||
return first.getTitle();
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* 保存工具调用明细到 tool_invocation 表
|
||||
*/
|
||||
private void saveToolInvocation(String query, KnowledgeIndexService.L0Hint l0Hint,
|
||||
List<VectorSearchService.SearchResult> l1Results,
|
||||
boolean highConfidence, long startTime,
|
||||
LookupResult result, String domain, String dedupReason) {
|
||||
private void saveToolInvocation(KnowledgeQuery query,
|
||||
long startTime,
|
||||
LookupResult result,
|
||||
String domain,
|
||||
String dedupReason) {
|
||||
try {
|
||||
String sessionId = SessionContextHolder.getSessionId();
|
||||
if (sessionId == null) return;
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
double l1TopSimilarity = (l1Results != null && !l1Results.isEmpty())
|
||||
? normalizeL2(l1Results.get(0).getScore())
|
||||
: -1;
|
||||
|
||||
ToolInvocationRecorder.LookupKnowledgeRecord record = ToolInvocationRecorder.LookupKnowledgeRecord.from(
|
||||
if (SessionContextHolder.getSessionId() == null) {
|
||||
return;
|
||||
}
|
||||
ToolInvocationRecorder.LookupKnowledgeRecord record =
|
||||
ToolInvocationRecorder.LookupKnowledgeRecord.from(
|
||||
query,
|
||||
l0Hint,
|
||||
l1Results,
|
||||
highConfidence,
|
||||
result,
|
||||
domain,
|
||||
dedupReason,
|
||||
(int) duration,
|
||||
l1TopSimilarity
|
||||
(int) Math.max(0, System.currentTimeMillis() - startTime)
|
||||
);
|
||||
toolInvocationRecorder.recordLookupKnowledge(record);
|
||||
log.debug("tool_invocation 已保存: sessionId={}, layer={}, relevanceLevel={}, duration={}ms",
|
||||
sessionId, record.retrievalLayer(), record.relevanceLevel(), duration);
|
||||
} catch (Exception e) {
|
||||
log.error("保存 tool_invocation 失败", e);
|
||||
}
|
||||
}
|
||||
|
||||
// ==================== 结果组装 ====================
|
||||
|
||||
private LookupResult buildResult(
|
||||
List<KnowledgeEntry> l0Matches,
|
||||
List<VectorSearchService.SearchResult> l1Results,
|
||||
boolean highConfidence
|
||||
) {
|
||||
LookupResult.LookupResultBuilder builder = LookupResult.builder();
|
||||
|
||||
// 构建 primary(L0 结果)
|
||||
PrimaryResult primary = null;
|
||||
if (l0Matches != null && !l0Matches.isEmpty()) {
|
||||
KnowledgeEntry first = l0Matches.get(0);
|
||||
boolean hasL1 = l1Results != null && !l1Results.isEmpty();
|
||||
boolean needFullContent = highConfidence || !hasL1;
|
||||
String content = needFullContent
|
||||
? buildCompactSummary(first)
|
||||
: buildMetadataOnlySummary(first);
|
||||
|
||||
if (content != null) {
|
||||
primary = PrimaryResult.builder()
|
||||
.content(content)
|
||||
.source(first.getFilePath())
|
||||
.matchType("exact_L0")
|
||||
.confidence(highConfidence ? "high" : "low")
|
||||
.availableSections(null)
|
||||
.build();
|
||||
log.debug("L0结果已构建: source={}, contentLength={}", first.getFilePath(), content.length());
|
||||
} else {
|
||||
log.warn("L0匹配但文件读取失败: {}", first.getFilePath());
|
||||
}
|
||||
}
|
||||
builder.primary(primary);
|
||||
|
||||
// 构建 supplement(L1 结果)
|
||||
SupplementResult supplement = null;
|
||||
boolean hasL1 = l1Results != null && !l1Results.isEmpty();
|
||||
if (hasL1) {
|
||||
VectorSearchService.SearchResult firstL1 = l1Results.get(0);
|
||||
supplement = SupplementResult.builder()
|
||||
.content(firstL1.getContent())
|
||||
.source(firstL1.getMetadata())
|
||||
.matchType("semantic_L1")
|
||||
.build();
|
||||
log.debug("L1结果已构建: source={}, score={}", firstL1.getMetadata(), firstL1.getScore());
|
||||
}
|
||||
builder.supplement(supplement);
|
||||
|
||||
EvidencePostprocessResult evidence = buildEvidenceBlocks(l0Matches, l1Results);
|
||||
builder.evidenceBlocks(evidence.blocks());
|
||||
builder.evidenceCandidateCount(evidence.candidateCount());
|
||||
builder.evidenceBlockCount(evidence.blocks().size());
|
||||
|
||||
boolean found = (primary != null) || (supplement != null);
|
||||
builder.found(found);
|
||||
|
||||
return builder.build();
|
||||
}
|
||||
|
||||
private EvidencePostprocessResult buildEvidenceBlocks(
|
||||
List<KnowledgeEntry> l0Matches,
|
||||
List<VectorSearchService.SearchResult> l1Results) {
|
||||
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
|
||||
int candidateCount = 0;
|
||||
|
||||
if (l0Matches != null) {
|
||||
for (int i = 0; i < l0Matches.size(); i++) {
|
||||
KnowledgeEntry entry = l0Matches.get(i);
|
||||
candidateCount++;
|
||||
EvidenceBlock block = EvidenceBlock.builder()
|
||||
.source(entry.getFilePath())
|
||||
.title(entry.getTitle())
|
||||
.breadcrumb(null)
|
||||
.retrievalLayer("L0")
|
||||
.content(buildMetadataOnlySummary(entry))
|
||||
.score(null)
|
||||
.hitReasons(buildL0HitReasons(entry, i + 1))
|
||||
.build();
|
||||
mergeEvidence(deduped, sourceKey(block, "l0-" + i), block);
|
||||
}
|
||||
}
|
||||
|
||||
if (l1Results != null) {
|
||||
for (int i = 0; i < l1Results.size(); i++) {
|
||||
VectorSearchService.SearchResult result = l1Results.get(i);
|
||||
candidateCount++;
|
||||
Map<String, String> metadata = parseMetadata(result.getMetadata());
|
||||
String source = firstNonBlank(
|
||||
metadata.get("_source"),
|
||||
metadata.get("docId"),
|
||||
result.getMetadata(),
|
||||
result.getId()
|
||||
);
|
||||
EvidenceBlock block = EvidenceBlock.builder()
|
||||
.source(source)
|
||||
.title(metadata.get("title"))
|
||||
.breadcrumb(metadata.get("breadcrumb"))
|
||||
.retrievalLayer("L1")
|
||||
.content(truncate(result.getContent(), 800))
|
||||
.score((double) result.getScore())
|
||||
.hitReasons(List.of("semantic_rank:" + (i + 1)))
|
||||
.build();
|
||||
mergeEvidence(deduped, sourceKey(block, "l1-" + i), block);
|
||||
}
|
||||
}
|
||||
|
||||
return new EvidencePostprocessResult(candidateCount, new ArrayList<>(deduped.values()));
|
||||
}
|
||||
|
||||
private void mergeEvidence(Map<String, EvidenceBlock> deduped, String key, EvidenceBlock incoming) {
|
||||
EvidenceBlock existing = deduped.get(key);
|
||||
if (existing == null) {
|
||||
deduped.put(key, incoming);
|
||||
return;
|
||||
}
|
||||
|
||||
List<String> mergedReasons = new ArrayList<>();
|
||||
if (existing.getHitReasons() != null) {
|
||||
mergedReasons.addAll(existing.getHitReasons());
|
||||
}
|
||||
if (incoming.getHitReasons() != null) {
|
||||
for (String reason : incoming.getHitReasons()) {
|
||||
if (!mergedReasons.contains(reason)) {
|
||||
mergedReasons.add(reason);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
String mergedLayer = existing.getRetrievalLayer();
|
||||
if (incoming.getRetrievalLayer() != null && !incoming.getRetrievalLayer().equals(mergedLayer)) {
|
||||
mergedLayer = "L0+L1";
|
||||
}
|
||||
|
||||
existing.setRetrievalLayer(mergedLayer);
|
||||
existing.setHitReasons(mergedReasons);
|
||||
if (existing.getScore() == null && incoming.getScore() != null) {
|
||||
existing.setScore(incoming.getScore());
|
||||
}
|
||||
if ((existing.getBreadcrumb() == null || existing.getBreadcrumb().isBlank())
|
||||
&& incoming.getBreadcrumb() != null) {
|
||||
existing.setBreadcrumb(incoming.getBreadcrumb());
|
||||
}
|
||||
}
|
||||
|
||||
private List<String> buildL0HitReasons(KnowledgeEntry entry, int rank) {
|
||||
List<String> reasons = new ArrayList<>();
|
||||
reasons.add("l0_rank:" + rank);
|
||||
if (entry.getKeywords() != null && !entry.getKeywords().isEmpty()) {
|
||||
reasons.add("l0_keywords:" + String.join(",", entry.getKeywords()));
|
||||
}
|
||||
if (entry.getCategory() != null && !entry.getCategory().isBlank()) {
|
||||
reasons.add("domain:" + entry.getCategory());
|
||||
}
|
||||
return reasons;
|
||||
}
|
||||
|
||||
private String sourceKey(EvidenceBlock block, String fallback) {
|
||||
return firstNonBlank(block.getSource(), block.getTitle(), block.getBreadcrumb(), fallback);
|
||||
}
|
||||
|
||||
private Map<String, String> parseMetadata(String metadata) {
|
||||
if (metadata == null || metadata.isBlank()) {
|
||||
return Map.of();
|
||||
}
|
||||
try {
|
||||
Map<?, ?> raw = objectMapper.readValue(metadata, Map.class);
|
||||
Map<String, String> result = new LinkedHashMap<>();
|
||||
for (Map.Entry<?, ?> entry : raw.entrySet()) {
|
||||
if (entry.getKey() != null && entry.getValue() != null) {
|
||||
result.put(String.valueOf(entry.getKey()), String.valueOf(entry.getValue()));
|
||||
}
|
||||
}
|
||||
return result;
|
||||
} catch (Exception e) {
|
||||
return Map.of();
|
||||
}
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private record EvidencePostprocessResult(int candidateCount, List<EvidenceBlock> blocks) {}
|
||||
|
||||
private int countMdHeadings(String content) {
|
||||
if (content == null) return 0;
|
||||
return (int) content.lines()
|
||||
.filter(l -> l.trim().startsWith("##"))
|
||||
.count();
|
||||
}
|
||||
|
||||
private String buildCompactSummary(KnowledgeEntry entry) {
|
||||
String rawContent = knowledgeIndexService.readDocument(entry.getFilePath(), 2000);
|
||||
if (rawContent == null) return null;
|
||||
|
||||
String body = rawContent;
|
||||
if (body.startsWith("---")) {
|
||||
int end = body.indexOf("---", 3);
|
||||
if (end != -1) {
|
||||
body = body.substring(end + 3).trim();
|
||||
}
|
||||
}
|
||||
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append("文档: ").append(entry.getTitle()).append("\n");
|
||||
if (entry.getSummary() != null) {
|
||||
sb.append("摘要: ").append(entry.getSummary()).append("\n");
|
||||
}
|
||||
|
||||
String headings = body.lines()
|
||||
.filter(l -> l.trim().startsWith("##"))
|
||||
.map(l -> " - " + l.trim().replaceAll("^#+\\s*", ""))
|
||||
.collect(Collectors.joining("\n"));
|
||||
if (!headings.isEmpty()) {
|
||||
sb.append("章节:\n").append(headings).append("\n");
|
||||
}
|
||||
sb.append("---\n");
|
||||
|
||||
String textContent = body.lines()
|
||||
.filter(l -> !l.trim().startsWith("#") && !l.trim().isEmpty())
|
||||
.collect(Collectors.joining("\n"))
|
||||
.trim();
|
||||
|
||||
int maxBodyChars = body.length() < 500 ? 800 : 500;
|
||||
if (textContent.length() > maxBodyChars) {
|
||||
sb.append(textContent, 0, maxBodyChars).append("...");
|
||||
} else {
|
||||
sb.append(textContent);
|
||||
}
|
||||
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
private String buildMetadataOnlySummary(KnowledgeEntry entry) {
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append("文档: ").append(entry.getTitle()).append("\n");
|
||||
if (entry.getSummary() != null) {
|
||||
sb.append("摘要: ").append(entry.getSummary()).append("\n");
|
||||
}
|
||||
if (entry.getKeywords() != null && !entry.getKeywords().isEmpty()) {
|
||||
sb.append("关键词: ").append(String.join(", ", entry.getKeywords())).append("\n");
|
||||
}
|
||||
sb.append("来源: ").append(entry.getFilePath()).append("\n");
|
||||
return sb.toString();
|
||||
}
|
||||
|
||||
private String extractFirstMeaningfulLine(String content, int maxLen) {
|
||||
if (content == null || content.isBlank()) return "(空)";
|
||||
|
||||
String text = content.trim();
|
||||
if (text.startsWith("---")) {
|
||||
int end = text.indexOf("---", 3);
|
||||
if (end != -1) {
|
||||
text = text.substring(end + 3);
|
||||
}
|
||||
}
|
||||
|
||||
String[] lines = text.split("\n");
|
||||
for (String line : lines) {
|
||||
String tl = line.trim();
|
||||
if (!tl.isEmpty() && !tl.startsWith("#")) {
|
||||
return tl.length() <= maxLen ? tl : tl.substring(0, maxLen) + "...";
|
||||
}
|
||||
}
|
||||
|
||||
for (String line : lines) {
|
||||
if (!line.trim().isEmpty()) {
|
||||
String tl = line.trim();
|
||||
return tl.length() <= maxLen ? tl : tl.substring(0, maxLen) + "...";
|
||||
}
|
||||
}
|
||||
|
||||
return "(无有效内容)";
|
||||
}
|
||||
|
||||
private String extractDocKey(LookupResult result) {
|
||||
if (result.getPrimary() != null && result.getPrimary().getSource() != null) {
|
||||
return result.getPrimary().getSource();
|
||||
}
|
||||
if (result.getSupplement() != null && result.getSupplement().getSource() != null) {
|
||||
return result.getSupplement().getSource();
|
||||
}
|
||||
return null;
|
||||
private void logReturn(LookupResult result, long startTime) {
|
||||
long totalTime = System.currentTimeMillis() - startTime;
|
||||
log.info("----------------------------------------");
|
||||
log.info("<<< [工具返回] lookup_knowledge");
|
||||
log.info("<<< 结果: found={}, relevanceLevel={}, evidenceBlocks={}, 耗时: {}ms",
|
||||
result.isFound(),
|
||||
result.getRelevanceLevel(),
|
||||
result.getEvidenceBlockCount(),
|
||||
totalTime);
|
||||
log.info("<<< 行动记忆: retrievedDomainsThisSession={}", result.getRetrievedDomainsThisSession());
|
||||
if (result.getRetrievalTrace() != null) {
|
||||
log.info("<<< 检索路径: selectedAttempt={}, fallbackReason={}",
|
||||
result.getRetrievalTrace().getSelectedAttempt(),
|
||||
result.getRetrievalTrace().getFallbackReason());
|
||||
}
|
||||
log.info("========================================");
|
||||
}
|
||||
}
|
||||
|
||||
@@ -37,3 +37,5 @@
|
||||
- relevanceLevel=HIGHLY_RELEVANT + 域已在 retrievedDomainsThisSession → 禁止再次调用
|
||||
- relevanceLevel=REFERENCE → 先指出缺什么维度,再定向补充一次
|
||||
- completenessHint 是知识库给你的天花板信号,信任它
|
||||
- lookup_knowledge 的事实证据以 evidenceBlocks 和 contextPack.packedText 为准,不要假设 L0 hint 本身就是事实证据
|
||||
- retrievalTrace 只用于理解检索路径和降级原因,不能单独作为诊断事实
|
||||
|
||||
@@ -35,9 +35,9 @@
|
||||
| `query` | 查询关键词或描述。例如:`ERR_TIMEOUT`、`payment-gateway`、`支付为什么失败` |
|
||||
|
||||
**内部机制**:
|
||||
工具内部自动执行「先精确匹配(L0),未命中则语义检索(L1)」的两阶段检索逻辑,你无需关心哪一层。
|
||||
工具内部会先做 query understanding,使用领域/关键词/实体 hint 控制向量检索;如果带 filter 的向量检索低质量,会用原始 query 再执行一次无过滤语义检索。
|
||||
|
||||
**返回结果**:包含 `found`(是否找到)、`primary.content`(文档内容)、`primary.match_type`(来源标记:`exact_L0` 或 `semantic_L1`)等字段。
|
||||
**返回结果**:包含 `found`(是否找到)、`evidenceBlocks`(结构化证据)、`contextPack.packedText`(可直接引用的证据上下文)、`retrievalTrace`(检索路径)和 `rerankTrace`(重排解释)等字段。
|
||||
|
||||
**使用规则**:
|
||||
- 当你查到了错误码、接口名、服务名时:**必须**调用此工具
|
||||
|
||||
@@ -2,7 +2,10 @@ package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -129,23 +132,51 @@ class ToolInvocationRecorderTest {
|
||||
.evidenceCandidateCount(3)
|
||||
.evidenceBlockCount(1)
|
||||
.evidenceBlocks(List.of(block))
|
||||
.contextPack(ContextPack.builder()
|
||||
.packedText("packed evidence")
|
||||
.strategy("ranked_evidence_char_budget")
|
||||
.charBudget(4000)
|
||||
.usedChars(15)
|
||||
.includedSources(List.of("doc.md"))
|
||||
.omittedSources(List.of())
|
||||
.build())
|
||||
.retrievalTrace(RetrievalTrace.builder()
|
||||
.originalQuery("query")
|
||||
.rewrittenQuery("query")
|
||||
.selectedAttempt("UNFILTERED_VECTOR")
|
||||
.evidenceStatus(ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED)
|
||||
.attempts(List.of(RetrievalTrace.Attempt.builder()
|
||||
.name("UNFILTERED_VECTOR")
|
||||
.candidateCount(3)
|
||||
.topScore(0.42)
|
||||
.topSimilarity(0.79)
|
||||
.usable(true)
|
||||
.build()))
|
||||
.build())
|
||||
.build();
|
||||
KnowledgeQuery query = KnowledgeQuery.builder()
|
||||
.originalQuery("query")
|
||||
.rewrittenQuery("query")
|
||||
.domainHints(List.of())
|
||||
.matchedKeywords(List.of())
|
||||
.entities(List.of())
|
||||
.l0Titles(List.of())
|
||||
.l0MatchCount(0)
|
||||
.build();
|
||||
|
||||
ToolInvocationRecorder.LookupKnowledgeRecord record = ToolInvocationRecorder.LookupKnowledgeRecord.from(
|
||||
"query",
|
||||
KnowledgeIndexService.L0Hint.empty(),
|
||||
List.of(),
|
||||
false,
|
||||
query,
|
||||
result,
|
||||
null,
|
||||
null,
|
||||
10,
|
||||
-1
|
||||
10
|
||||
);
|
||||
|
||||
assertEquals(3, record.evidenceCandidateCount());
|
||||
assertEquals(1, record.evidenceBlockCount());
|
||||
assertEquals(1, record.evidenceBlocks().size());
|
||||
assertTrue(String.valueOf(record.evidenceBlocks().get(0).get("content_preview")).endsWith("..."));
|
||||
assertEquals("UNFILTERED_VECTOR", record.retrievalLayer());
|
||||
assertEquals(3, record.l1MatchCount());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,27 +1,36 @@
|
||||
package com.superbiz.agent.tool;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.service.KnowledgeContextPacker;
|
||||
import com.superbiz.agent.service.KnowledgeDocumentRetriever;
|
||||
import com.superbiz.agent.service.KnowledgeEvidencePostProcessor;
|
||||
import com.superbiz.agent.service.KnowledgeIndexService;
|
||||
import com.superbiz.agent.service.KnowledgeQueryTransformer;
|
||||
import com.superbiz.agent.service.LookupResultAssembler;
|
||||
import com.superbiz.agent.service.ToolInvocationRecorder;
|
||||
import com.superbiz.agent.service.VectorSearchService;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.mockito.InjectMocks;
|
||||
import org.mockito.Mock;
|
||||
import org.mockito.MockitoAnnotations;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.util.Collections;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
import static org.mockito.ArgumentMatchers.*;
|
||||
import static org.mockito.Mockito.*;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.Mockito.never;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
/**
|
||||
* LookupKnowledgeTool 单元测试
|
||||
* LookupKnowledgeTool evidence-first contract tests.
|
||||
*/
|
||||
class LookupKnowledgeToolTest {
|
||||
|
||||
@@ -37,305 +46,271 @@ class LookupKnowledgeToolTest {
|
||||
@Mock
|
||||
private RetrievedDocTracker retrievedDocTracker;
|
||||
|
||||
@Mock
|
||||
private ObjectMapper objectMapper;
|
||||
|
||||
@InjectMocks
|
||||
private LookupKnowledgeTool tool;
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
MockitoAnnotations.openMocks(this);
|
||||
|
||||
KnowledgeEvidencePostProcessor postProcessor = new KnowledgeEvidencePostProcessor();
|
||||
KnowledgeContextPacker contextPacker = new KnowledgeContextPacker();
|
||||
tool = new LookupKnowledgeTool();
|
||||
ReflectionTestUtils.setField(tool, "queryTransformer", new KnowledgeQueryTransformer(knowledgeIndexService));
|
||||
ReflectionTestUtils.setField(tool, "documentRetriever",
|
||||
new KnowledgeDocumentRetriever(vectorSearchService, new ObjectMapper()));
|
||||
ReflectionTestUtils.setField(tool, "evidencePostProcessor", postProcessor);
|
||||
ReflectionTestUtils.setField(tool, "contextPacker", contextPacker);
|
||||
ReflectionTestUtils.setField(tool, "resultAssembler", new LookupResultAssembler());
|
||||
ReflectionTestUtils.setField(tool, "toolInvocationRecorder", toolInvocationRecorder);
|
||||
ReflectionTestUtils.setField(tool, "retrievedDocTracker", retrievedDocTracker);
|
||||
ReflectionTestUtils.setField(tool, "topK", 3);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_uniqueMatch_usesL0HintAndL1() {
|
||||
// 准备 L0 唯一匹配
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("test.md")
|
||||
.title("Test Doc")
|
||||
.keywords(List.of("ERR_TIMEOUT"))
|
||||
.summary("Test summary")
|
||||
.build();
|
||||
void filteredL1SuccessDoesNotRetry() {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult result = searchResult(
|
||||
"vec-1",
|
||||
"db.md",
|
||||
"{\"_source\":\"db.md\",\"title\":\"Database Doc\",\"category\":\"database\"}",
|
||||
"mysql timeout runbook",
|
||||
0.2f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("ERR_TIMEOUT"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("test.md", 2000))
|
||||
.thenReturn("Test content * 用于构建紧凑摘要 * keyword2");
|
||||
VectorSearchService.SearchResult l1Result = new VectorSearchService.SearchResult();
|
||||
l1Result.setContent("L1 supporting content");
|
||||
l1Result.setMetadata("l1-source");
|
||||
l1Result.setScore(0.4f);
|
||||
when(vectorSearchService.searchSimilarDocuments("ERR_TIMEOUT", 3, null))
|
||||
.thenReturn(List.of(l1Result));
|
||||
when(knowledgeIndexService.analyzeQuery("mysql timeout")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, "database"))
|
||||
.thenReturn(List.of(result));
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("ERR_TIMEOUT");
|
||||
LookupResult lookup = tool.lookupKnowledge("mysql timeout");
|
||||
|
||||
// 验证结果
|
||||
assertTrue(result.isFound());
|
||||
assertNotNull(result.getPrimary());
|
||||
assertEquals("high", result.getPrimary().getConfidence());
|
||||
assertEquals("exact_L0", result.getPrimary().getMatchType());
|
||||
// 唯一匹配 → buildCompactSummary(),内容为结构化摘要
|
||||
String content = result.getPrimary().getContent();
|
||||
assertTrue(content.contains("文档: Test Doc"));
|
||||
assertTrue(content.contains("摘要: Test summary"));
|
||||
assertTrue(content.contains("Test content"));
|
||||
assertNotNull(result.getSupplement()); // L0 唯一命中仍调用 L1
|
||||
assertEquals("PRECISE", result.getRelevanceLevel());
|
||||
assertNotNull(result.getEvidenceBlocks());
|
||||
assertEquals(2, result.getEvidenceCandidateCount());
|
||||
assertEquals(2, result.getEvidenceBlockCount());
|
||||
assertEquals("L0", result.getEvidenceBlocks().get(0).getRetrievalLayer());
|
||||
assertTrue(result.getEvidenceBlocks().get(0).getHitReasons().stream()
|
||||
.anyMatch(reason -> reason.contains("ERR_TIMEOUT")));
|
||||
|
||||
verify(vectorSearchService).searchSimilarDocuments("ERR_TIMEOUT", 3, null);
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals(1, lookup.getEvidenceBlockCount());
|
||||
assertEquals("db.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
assertNotNull(lookup.getContextPack());
|
||||
assertTrue(lookup.getContextPack().getPackedText().contains("mysql timeout runbook"));
|
||||
assertEquals("FILTERED_VECTOR", lookup.getRetrievalTrace().getSelectedAttempt());
|
||||
assertEquals(1, lookup.getRetrievalTrace().getAttempts().size());
|
||||
assertEquals("PRECISE", lookup.getRelevanceLevel());
|
||||
verify(vectorSearchService).searchSimilarDocuments("mysql timeout", 3, "database");
|
||||
verify(vectorSearchService, never()).searchSimilarDocuments("mysql timeout", 3, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_multipleMatches_lowConfidence() {
|
||||
// 准备 L0 多个匹配
|
||||
KnowledgeEntry entry1 = KnowledgeEntry.builder()
|
||||
.filePath("doc1.md")
|
||||
.title("测试文档1")
|
||||
.keywords(List.of("超时"))
|
||||
.summary("这是一个测试文档")
|
||||
.build();
|
||||
void filteredLowQualityTriggersRawUnfilteredRetry() {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult weak = searchResult(
|
||||
"weak",
|
||||
"weak.md",
|
||||
"{\"_source\":\"weak.md\",\"title\":\"Weak\"}",
|
||||
"weak candidate",
|
||||
1.4f);
|
||||
VectorSearchService.SearchResult strong = searchResult(
|
||||
"strong",
|
||||
"strong.md",
|
||||
"{\"_source\":\"strong.md\",\"title\":\"Strong\"}",
|
||||
"mysql timeout strong runbook",
|
||||
0.2f);
|
||||
|
||||
KnowledgeEntry entry2 = KnowledgeEntry.builder()
|
||||
.filePath("doc2.md")
|
||||
.keywords(List.of("超时"))
|
||||
.build();
|
||||
when(knowledgeIndexService.analyzeQuery("mysql timeout")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, "database"))
|
||||
.thenReturn(List.of(weak));
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, null))
|
||||
.thenReturn(List.of(strong));
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("超时"))
|
||||
.thenReturn(hint(entry1, entry2));
|
||||
// 多匹配 + L1 有结果 → buildMetadataOnlySummary(),不读文件,不调用 readDocument
|
||||
LookupResult lookup = tool.lookupKnowledge("mysql timeout");
|
||||
|
||||
// 准备 L1 结果
|
||||
VectorSearchService.SearchResult l1Result = new VectorSearchService.SearchResult();
|
||||
l1Result.setContent("L1 content");
|
||||
l1Result.setMetadata("l1-source");
|
||||
|
||||
when(vectorSearchService.searchSimilarDocuments("超时", 3, null))
|
||||
.thenReturn(List.of(l1Result));
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("超时");
|
||||
|
||||
// 验证结果
|
||||
assertTrue(result.isFound());
|
||||
assertNotNull(result.getPrimary());
|
||||
assertEquals("low", result.getPrimary().getConfidence()); // 多个匹配 = 低置信度
|
||||
// 多匹配 + L1 有结果 → 仅元数据摘要
|
||||
String content = result.getPrimary().getContent();
|
||||
assertTrue(content.contains("文档: 测试文档1"));
|
||||
assertTrue(content.contains("摘要: 这是一个测试文档"));
|
||||
assertTrue(content.contains("关键词: 超时"));
|
||||
assertTrue(content.contains("来源: doc1.md"));
|
||||
|
||||
assertNotNull(result.getSupplement()); // 低置信度调用 L1
|
||||
assertEquals("L1 content", result.getSupplement().getContent());
|
||||
assertEquals("semantic_L1", result.getSupplement().getMatchType());
|
||||
|
||||
// 验证 L1 被调用,readDocument 未被调用(多匹配不走 buildCompactSummary)
|
||||
verify(vectorSearchService).searchSimilarDocuments("超时", 3, null);
|
||||
verify(knowledgeIndexService, never()).readDocument(anyString(), anyInt());
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals("UNFILTERED_VECTOR_RETRY", lookup.getRetrievalTrace().getSelectedAttempt());
|
||||
assertEquals("filtered_vector_low_quality", lookup.getRetrievalTrace().getFallbackReason());
|
||||
assertEquals(2, lookup.getRetrievalTrace().getAttempts().size());
|
||||
assertEquals("strong.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
verify(vectorSearchService).searchSimilarDocuments("mysql timeout", 3, "database");
|
||||
verify(vectorSearchService).searchSimilarDocuments("mysql timeout", 3, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_noL0Match_onlyL1() {
|
||||
// L0 未匹配
|
||||
void filteredNoEvidenceTriggersRawUnfilteredRetry() {
|
||||
KnowledgeEntry entry = entry("db.md", "Database Doc", "mysql", "database");
|
||||
VectorSearchService.SearchResult strong = searchResult(
|
||||
"strong",
|
||||
"strong.md",
|
||||
"{\"_source\":\"strong.md\",\"title\":\"Strong\"}",
|
||||
"mysql timeout strong runbook",
|
||||
0.2f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("mysql timeout")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, "database"))
|
||||
.thenReturn(Collections.emptyList());
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, null))
|
||||
.thenReturn(List.of(strong));
|
||||
|
||||
LookupResult lookup = tool.lookupKnowledge("mysql timeout");
|
||||
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals("UNFILTERED_VECTOR_RETRY", lookup.getRetrievalTrace().getSelectedAttempt());
|
||||
assertEquals("filtered_vector_no_evidence", lookup.getRetrievalTrace().getFallbackReason());
|
||||
assertEquals("strong.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
}
|
||||
|
||||
@Test
|
||||
void l0HintsDoNotBecomeStandaloneEvidenceWhenL1Fails() {
|
||||
KnowledgeEntry entry = entry("fallback.md", "Fallback Doc", "fallback", "database");
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("fallback")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("fallback", 3, "database"))
|
||||
.thenReturn(Collections.emptyList());
|
||||
when(vectorSearchService.searchSimilarDocuments("fallback", 3, null))
|
||||
.thenReturn(Collections.emptyList());
|
||||
|
||||
LookupResult lookup = tool.lookupKnowledge("fallback");
|
||||
|
||||
assertFalse(lookup.isFound());
|
||||
assertEquals(0, lookup.getEvidenceBlockCount());
|
||||
assertTrue(lookup.getEvidenceBlocks().isEmpty());
|
||||
assertEquals("no_evidence", lookup.getRetrievalTrace().getEvidenceStatus());
|
||||
assertTrue(String.valueOf(lookup.getRetrievalTrace().getQueryHints()).contains("Fallback Doc"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void noL0HintUsesUnfilteredVectorSearch() {
|
||||
VectorSearchService.SearchResult result = searchResult(
|
||||
"vec-1",
|
||||
"perf.md",
|
||||
"{\"_source\":\"perf.md\",\"title\":\"Perf\"}",
|
||||
"performance tuning guide",
|
||||
0.3f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("性能优化"))
|
||||
.thenReturn(KnowledgeIndexService.L0Hint.empty());
|
||||
|
||||
// 准备 L1 结果
|
||||
VectorSearchService.SearchResult l1Result = new VectorSearchService.SearchResult();
|
||||
l1Result.setContent("L1 semantic result");
|
||||
l1Result.setMetadata("l1-doc");
|
||||
|
||||
when(vectorSearchService.searchSimilarDocuments("性能优化", 3, null))
|
||||
.thenReturn(List.of(l1Result));
|
||||
.thenReturn(List.of(result));
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("性能优化");
|
||||
|
||||
// 验证结果
|
||||
assertTrue(result.isFound());
|
||||
assertNull(result.getPrimary()); // L0 未命中
|
||||
assertNotNull(result.getSupplement()); // 只有 L1 结果
|
||||
assertEquals("L1 semantic result", result.getSupplement().getContent());
|
||||
LookupResult lookup = tool.lookupKnowledge("性能优化");
|
||||
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals("UNFILTERED_VECTOR", lookup.getRetrievalTrace().getSelectedAttempt());
|
||||
assertEquals("perf.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
verify(vectorSearchService).searchSimilarDocuments("性能优化", 3, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_noMatch() {
|
||||
// L0 和 L1 都未匹配
|
||||
when(knowledgeIndexService.analyzeQuery("不存在的内容"))
|
||||
.thenReturn(KnowledgeIndexService.L0Hint.empty());
|
||||
when(vectorSearchService.searchSimilarDocuments("不存在的内容", 3, null))
|
||||
.thenReturn(Collections.emptyList());
|
||||
void rerankUsesHintMatchesAndContextPackPreservesMetadata() {
|
||||
KnowledgeEntry entry = entry("payment.md", "Payment", "ERR_TIMEOUT", "payment");
|
||||
VectorSearchService.SearchResult first = searchResult(
|
||||
"a",
|
||||
"a.md",
|
||||
"{\"_source\":\"a.md\",\"title\":\"Generic\",\"category\":\"other\"}",
|
||||
"generic troubleshooting",
|
||||
0.4f);
|
||||
VectorSearchService.SearchResult second = searchResult(
|
||||
"b",
|
||||
"b.md",
|
||||
"{\"_source\":\"b.md\",\"title\":\"Payment ERR_TIMEOUT\",\"breadcrumb\":\"Payment > Timeout\",\"category\":\"payment\"}",
|
||||
"payment ERR_TIMEOUT timeout diagnosis",
|
||||
0.45f);
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("不存在的内容");
|
||||
when(knowledgeIndexService.analyzeQuery("ERR_TIMEOUT")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("ERR_TIMEOUT", 3, "payment"))
|
||||
.thenReturn(List.of(first, second));
|
||||
|
||||
// 验证结果
|
||||
assertFalse(result.isFound());
|
||||
assertNull(result.getPrimary());
|
||||
assertNull(result.getSupplement());
|
||||
LookupResult lookup = tool.lookupKnowledge("ERR_TIMEOUT");
|
||||
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals("b.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
assertTrue(lookup.getRerankTrace().getItems().get(0).getBoostReasons().stream()
|
||||
.anyMatch(reason -> reason.startsWith("domain_match")));
|
||||
assertTrue(lookup.getContextPack().getPackedText().contains("Payment > Timeout"));
|
||||
assertTrue(lookup.getContextPack().getPackedText().contains("reasons:"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_l0MatchButReadFails() {
|
||||
// L0 匹配但文件读取失败
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("nonexistent.md")
|
||||
.keywords(List.of("test"))
|
||||
void deduplicatesEvidenceBlocksBySource() {
|
||||
KnowledgeEntry entry = entry("shared.md", "Shared", "shared", "payment");
|
||||
VectorSearchService.SearchResult first = searchResult(
|
||||
"a",
|
||||
"shared.md",
|
||||
"{\"_source\":\"shared.md\",\"title\":\"Shared\"}",
|
||||
"shared content 1",
|
||||
0.2f);
|
||||
VectorSearchService.SearchResult second = searchResult(
|
||||
"b",
|
||||
"shared.md",
|
||||
"{\"_source\":\"shared.md\",\"title\":\"Shared\"}",
|
||||
"shared content 2",
|
||||
0.25f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("shared")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("shared", 3, "payment"))
|
||||
.thenReturn(List.of(first, second));
|
||||
|
||||
LookupResult lookup = tool.lookupKnowledge("shared");
|
||||
|
||||
assertTrue(lookup.isFound());
|
||||
assertEquals(2, lookup.getEvidenceCandidateCount());
|
||||
assertEquals(1, lookup.getEvidenceBlockCount());
|
||||
assertEquals("shared.md", lookup.getEvidenceBlocks().get(0).getSource());
|
||||
}
|
||||
|
||||
@Test
|
||||
void sessionDedupDoesNotReturnConsumableEvidenceAgain() {
|
||||
KnowledgeEntry entry = entry("payment.md", "Payment", "ERR_TIMEOUT", "payment");
|
||||
VectorSearchService.SearchResult result = searchResult(
|
||||
"vec-1",
|
||||
"payment.md",
|
||||
"{\"_source\":\"payment.md\",\"title\":\"Payment\",\"category\":\"payment\"}",
|
||||
"payment timeout runbook",
|
||||
0.2f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("ERR_TIMEOUT")).thenReturn(hint(entry));
|
||||
when(vectorSearchService.searchSimilarDocuments("ERR_TIMEOUT", 3, "payment"))
|
||||
.thenReturn(List.of(result));
|
||||
when(retrievedDocTracker.isAlreadyRetrieved("session-1", "payment.md")).thenReturn(true);
|
||||
when(retrievedDocTracker.getRetrievedDomains("session-1")).thenReturn(List.of("payment"));
|
||||
|
||||
SessionContextHolder.setSessionId("session-1");
|
||||
try {
|
||||
LookupResult lookup = tool.lookupKnowledge("ERR_TIMEOUT");
|
||||
|
||||
assertFalse(lookup.isFound());
|
||||
assertEquals(0, lookup.getEvidenceBlockCount());
|
||||
assertTrue(lookup.getEvidenceBlocks().isEmpty());
|
||||
assertTrue(lookup.getMessage().contains("无需重复召回"));
|
||||
assertEquals(List.of("payment"), lookup.getRetrievedDomainsThisSession());
|
||||
} finally {
|
||||
SessionContextHolder.clear();
|
||||
}
|
||||
}
|
||||
|
||||
private KnowledgeEntry entry(String filePath, String title, String keyword, String category) {
|
||||
return KnowledgeEntry.builder()
|
||||
.filePath(filePath)
|
||||
.title(title)
|
||||
.keywords(List.of(keyword))
|
||||
.summary(title + " summary")
|
||||
.category(category)
|
||||
.build();
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("test"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("nonexistent.md", 2000))
|
||||
.thenReturn(null); // 读取失败
|
||||
when(vectorSearchService.searchSimilarDocuments("test", 3, null))
|
||||
.thenReturn(Collections.emptyList());
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("test");
|
||||
|
||||
// 验证:found 为 false,因为无法读取内容且无 L1 补充
|
||||
assertFalse(result.isFound());
|
||||
assertNull(result.getPrimary());
|
||||
assertNull(result.getSupplement());
|
||||
|
||||
verify(vectorSearchService).searchSimilarDocuments("test", 3, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_l1ReturnsNull() {
|
||||
// L0 未匹配,L1 返回 null
|
||||
when(knowledgeIndexService.analyzeQuery("query"))
|
||||
.thenReturn(KnowledgeIndexService.L0Hint.empty());
|
||||
when(vectorSearchService.searchSimilarDocuments("query", 3, null))
|
||||
.thenReturn(null);
|
||||
|
||||
// 执行查询
|
||||
LookupResult result = tool.lookupKnowledge("query");
|
||||
|
||||
// 验证
|
||||
assertFalse(result.isFound());
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_availableSectionsIsNull() {
|
||||
// 验证 availableSections 字段为 null(MVP 预留字段)
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("test.md")
|
||||
.keywords(List.of("test"))
|
||||
.build();
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("test"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("test.md", 2000))
|
||||
.thenReturn("Content");
|
||||
when(vectorSearchService.searchSimilarDocuments("test", 3, null))
|
||||
.thenReturn(Collections.emptyList());
|
||||
|
||||
LookupResult result = tool.lookupKnowledge("test");
|
||||
|
||||
assertNotNull(result.getPrimary());
|
||||
assertNull(result.getPrimary().getAvailableSections()); // MVP 返回 null
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_appliesSingleL0DomainAsL1Filter() {
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("db.md")
|
||||
.title("Database Doc")
|
||||
.keywords(List.of("mysql"))
|
||||
.summary("Database summary")
|
||||
.category("database")
|
||||
.build();
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("mysql timeout"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("db.md", 2000))
|
||||
.thenReturn("Database content");
|
||||
when(vectorSearchService.searchSimilarDocuments("mysql timeout", 3, "database"))
|
||||
.thenReturn(Collections.emptyList());
|
||||
|
||||
LookupResult result = tool.lookupKnowledge("mysql timeout");
|
||||
|
||||
assertTrue(result.isFound());
|
||||
verify(vectorSearchService).searchSimilarDocuments("mysql timeout", 3, "database");
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_l1FailureKeepsL0Fallback() {
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("fallback.md")
|
||||
.title("Fallback Doc")
|
||||
.keywords(List.of("fallback"))
|
||||
.summary("Fallback summary")
|
||||
.build();
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("fallback"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("fallback.md", 2000))
|
||||
.thenReturn("Fallback content");
|
||||
when(vectorSearchService.searchSimilarDocuments("fallback", 3, null))
|
||||
.thenThrow(new RuntimeException("milvus unavailable"));
|
||||
|
||||
LookupResult result = tool.lookupKnowledge("fallback");
|
||||
|
||||
assertTrue(result.isFound());
|
||||
assertNotNull(result.getPrimary());
|
||||
assertNull(result.getSupplement());
|
||||
assertEquals("REFERENCE", result.getRelevanceLevel());
|
||||
verify(vectorSearchService).searchSimilarDocuments("fallback", 3, null);
|
||||
}
|
||||
|
||||
@Test
|
||||
void testLookup_deduplicatesEvidenceBlocksBySource() throws Exception {
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath("shared.md")
|
||||
.title("Shared Doc")
|
||||
.keywords(List.of("shared"))
|
||||
.summary("Shared summary")
|
||||
.build();
|
||||
VectorSearchService.SearchResult l1Result = new VectorSearchService.SearchResult();
|
||||
l1Result.setContent("Shared semantic content");
|
||||
String metadata = "{\"docId\":\"doc-1\",\"_source\":\"shared.md\",\"title\":\"Shared Doc\"}";
|
||||
l1Result.setMetadata(metadata);
|
||||
l1Result.setScore(0.2f);
|
||||
|
||||
when(knowledgeIndexService.analyzeQuery("shared"))
|
||||
.thenReturn(hint(entry));
|
||||
when(knowledgeIndexService.readDocument("shared.md", 2000))
|
||||
.thenReturn("Shared content");
|
||||
when(vectorSearchService.searchSimilarDocuments("shared", 3, null))
|
||||
.thenReturn(List.of(l1Result));
|
||||
when(objectMapper.readValue(metadata, Map.class))
|
||||
.thenReturn(Map.of("docId", "doc-1", "_source", "shared.md", "title", "Shared Doc"));
|
||||
|
||||
LookupResult result = tool.lookupKnowledge("shared");
|
||||
|
||||
assertTrue(result.isFound());
|
||||
assertEquals(2, result.getEvidenceCandidateCount());
|
||||
assertEquals(1, result.getEvidenceBlockCount());
|
||||
assertEquals("shared.md", result.getEvidenceBlocks().get(0).getSource());
|
||||
assertEquals("L0+L1", result.getEvidenceBlocks().get(0).getRetrievalLayer());
|
||||
assertTrue(result.getEvidenceBlocks().get(0).getHitReasons().contains("semantic_rank:1"));
|
||||
assertTrue(result.getEvidenceBlocks().get(0).getHitReasons().stream()
|
||||
.anyMatch(reason -> reason.startsWith("l0_keywords:")));
|
||||
private VectorSearchService.SearchResult searchResult(String id,
|
||||
String source,
|
||||
String metadata,
|
||||
String content,
|
||||
float score) {
|
||||
VectorSearchService.SearchResult result = new VectorSearchService.SearchResult();
|
||||
result.setId(id);
|
||||
result.setMetadata(metadata);
|
||||
result.setContent(content);
|
||||
result.setScore(score);
|
||||
result.setRawScore((double) score);
|
||||
result.setScoreLabel("l2_distance");
|
||||
return result;
|
||||
}
|
||||
|
||||
private KnowledgeIndexService.L0Hint hint(KnowledgeEntry... entries) {
|
||||
List<KnowledgeEntry> matches = List.of(entries);
|
||||
List<String> keywords = matches.stream()
|
||||
.flatMap(entry -> entry.getKeywords() == null ? java.util.stream.Stream.empty() : entry.getKeywords().stream())
|
||||
.flatMap(entry -> entry.getKeywords() == null
|
||||
? java.util.stream.Stream.empty()
|
||||
: entry.getKeywords().stream())
|
||||
.distinct()
|
||||
.toList();
|
||||
List<String> domains = matches.stream()
|
||||
|
||||
Reference in New Issue
Block a user