diff --git a/mvp/issues/ISS-008-executor-narrow-scope-overreach.md b/mvp/issues/ISS-008-executor-narrow-scope-overreach.md new file mode 100644 index 0000000..6a73a8d --- /dev/null +++ b/mvp/issues/ISS-008-executor-narrow-scope-overreach.md @@ -0,0 +1,308 @@ +# ISS-008 Executor 窄范围查询越界 + +**严重程度**:中 +**状态**:已修复 +**发现时间**:2026-07-08 +**关联**: +- `ISS-007-verifier-evidence-summary-fidelity` +- `executor-structured-output-v2` +- `executor-evidence-attribution-hallucination` + +--- + +## 背景 + +当前 Chat 诊断链路已经演进为: + +```text +Planner + -> Executor + -> VerifierInputHook / Gatekeeper + -> Verifier + -> Composer +``` + +其中 Executor 的定位已经从“生成最终诊断答案”收敛为: + +```text +证据收集 + 微观事实提炼 +``` + +但在窄范围问题中,Executor 仍可能把用户只要求确认的一件事扩展成多条 claim,例如用户只问 `HighCPUUsage`,Executor 可能顺手输出内存、连接池、数据库或修复建议相关内容。 + +这类问题不一定是证据伪造。很多时候工具返回里确实有其它信息,但它们不属于当前用户问题的范围。Gatekeeper 只能校验证据引用真假,不能完整承担“用户意图范围控制”;Verifier 虽然可以降级,但会增加链路负担。 + +因此本 issue 采用低成本的 Prompt-first 修复:先收紧 Executor prompt,不改 Planner,不引入 `scope_contract`。 + +--- + +## 问题类型 + +### 1. 窄范围查询越界 + +用户问题只要求确认一个服务、告警、日志、订单或时间窗口,但 Executor 输出了用户未要求的 claim。 + +示例: + +```text +只确认 payment-service 是否存在 HighCPUUsage,不要分析订单、OOM、数据库慢查询、连接池或 user-service。 +``` + +错误输出包括: + +- `HighMemoryUsage` +- `SlowResponse` +- `order-123` +- `HikariCP` +- `DB / database` +- `user-service` + +### 2. Observation 变成 Diagnosis + +Executor 本应输出观察事实,却输出根因、风险、修复建议或经验推断。 + +错误输出包括: + +- “CPU 过高是请求超时的根因” +- “建议扩容” +- “通常这种情况是数据库慢查询导致” +- “存在内存泄漏风险” + +### 3. Runbook 通用知识变成当前事实 + +Runbook、Skill、知识库可以指导要查什么,但不能直接变成本次环境已发生的事实。 + +错误输出包括: + +```text +Runbook 中说 HighCPUUsage 常见原因是流量突增,所以当前环境发生了流量突增。 +``` + +--- + +## 修复决策 + +本期只修 Executor prompt。 + +### 本期做 + +1. 强化 Executor 单一职责:证据收集 + 微观事实提炼。 +2. 增加 `角色边界 HARD-GATE`。 +3. 增加 `窄范围确认任务 HARD-GATE`。 +4. 增加工具使用边界,避免为补全故事而扩展检索。 +5. 增加输出前自检,要求输出 JSON 前删除越界 claim。 + +### 本期不做 + +1. 不改 Planner。 +2. 不新增 `scope_contract`。 +3. 不解析 Planner 输出中的 scope。 +4. 不做 Gatekeeper scope 校验。 +5. 不改多 Agent 编排。 + +--- + +## 设计原则 + +### Claim 要少,Evidence 可以多 + +窄范围任务下,Executor 应输出最少必要 claim,通常 1 条,最多 2 条。 + +但 claim 数量限制不限制 `evidence_bindings` 数量。一条核心 claim 可以绑定多条直接相关证据。 + +```text +正确: +1 条 claim + 多条 evidence_bindings + +错误: +为了展示多条证据,把同一个观察事实拆成多条 claim +``` + +### 只输出当前问题范围内的 Observation + +窄范围任务下,`claims` 只能使用: + +- `observation` +- `negative_observation` + +禁止使用: + +- `root_cause` +- `risk` +- `recommendation` +- 其它建议类或诊断类 claim + +### 证据不足时不要补故事 + +如果工具没有返回可被精确引用的证据: + +```text +source_invocation_id + raw_path + evidence_excerpt +``` + +Executor 不应生成 confirmed claim,应写入 `missing_info`。 + +--- + +## Prompt 修复点 + +已更新: + +- `src/main/resources/prompts/chat-executor-prompt.md` + +核心新增约束: + +1. `角色边界 HARD-GATE` +2. `窄范围确认任务 HARD-GATE` +3. `工具使用边界` +4. `输出前自检` +5. 条目级 `raw_path` 强约束:同一条工具数组项只能绑定一次,禁止输出 `$.alerts[0].alert_name`、`$.alerts[0].state` 等字段级子路径。 + +--- + +## 验证结果 + +### 2026-07-08 E2E 验证 + +输入: + +```text +只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。 +``` + +第一次验证发现: + +- Executor 已经只输出 `payment-service + HighCPUUsage` 相关 observation,没有输出越界 claim。 +- 但 Executor 额外生成了字段级 `raw_path`: + - `$.alerts[0].alert_name` + - `$.alerts[0].state` +- 当前 Gatekeeper 只支持条目级路径 `$.alerts[i]` / `$.logs[i]` / `$.evidence_blocks[i]`,因此判定为 `REJECT`。 + +已追加 prompt 约束: + +```text +同一条工具数组项只能绑定一次。 +不要为了引用其中多个字段而拆成多个 evidence_bindings。 +raw_path 禁止指向字段级子路径。 +``` + +第二次验证结果: + +```text +sessionId: iss008-narrow-highcpu-rerun-20260708-215510 +verdict: PASS +groundedness_score: 1.0 +gatekeeper_result.status: pass +gatekeeper_result.severity: none +claim_count: 1 +claim_type: observation +forbidden_hits: none +``` + +Executor claim: + +```text +payment-service 当前存在 HighCPUUsage 告警,CPU 使用率持续超过 80%,当前值为 92%,告警状态为 firing,已持续 25 分钟。 +``` + +最终答案未出现以下排除项: + +- `HighMemoryUsage` +- `SlowResponse` +- `order-123` +- `订单123` +- `OOM` +- `DB / database` +- `HikariCP` +- `connection pool / 连接池` +- `user-service` + +--- + +## 验收标准 + +### 1. HighCPUUsage 窄范围 + +输入: + +```text +只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。 +``` + +期望: + +- `claims` 只围绕 `payment-service + HighCPUUsage`。 +- 不出现 `HighMemoryUsage`。 +- 不出现 `SlowResponse`。 +- 不出现 `order-123`。 +- 不出现 `OOM`。 +- 不出现 `DB / database`。 +- 不出现 `HikariCP / connection pool`。 +- 不出现 `user-service`。 + +### 2. HighMemoryUsage 窄范围 + +输入: + +```text +只确认 order-service 是否存在 HighMemoryUsage。 +``` + +期望: + +- 可以输出内存使用率、告警状态、持续时间等观察事实。 +- 不输出“内存泄漏已确认”。 +- 不输出扩容、重启、修改 JVM 参数等修复建议。 + +### 3. SlowResponse 窄范围 + +输入: + +```text +只确认 user-service 是否存在 SlowResponse 告警和慢请求日志。 +``` + +期望: + +- 可以绑定 alert 和 logs 多条证据。 +- 不推断数据库连接池耗尽。 +- 不推断下游服务故障。 + +### 4. 用户明确排除项 + +输入: + +```text +只看 order-service 支付失败日志,不要分析 HikariCP。 +``` + +期望: + +- `claim_text` 不出现 HikariCP 确认结论。 +- 最终答案不出现 HikariCP 确认结论。 + +### 5. 证据不足 + +输入: + +```text +只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。 +``` + +期望: + +- 如果工具返回 `logs=[]`,Executor 不编造 positive claim。 +- 输出 `negative_observation` 或 `missing_info`。 +- 不返回 `generic-service` 占位事实。 + +--- + +## 后续增强 + +如果 Prompt-first 后仍不稳定,再考虑: + +1. Planner 输出 `scope_contract`。 +2. Gatekeeper 增加 scope 校验。 +3. eval fixture 增加 forbidden claim 自动断言。 + +本期暂不进入这些改造。 diff --git a/mvp/issues/ISS-009-negative-observation-no-evidence-reference.md b/mvp/issues/ISS-009-negative-observation-no-evidence-reference.md new file mode 100644 index 0000000..cdf3dad --- /dev/null +++ b/mvp/issues/ISS-009-negative-observation-no-evidence-reference.md @@ -0,0 +1,239 @@ +# ISS-009 negative_observation 精确引用 no-evidence 结果 + +**严重程度**:中 +**状态**:已修复 +**发现时间**:2026-07-08 +**关联**: +- `ISS-007-verifier-evidence-summary-fidelity` +- `ISS-008-executor-narrow-scope-overreach` +- `executor-structured-output-v2` + +--- + +## 背景 + +ISS-007 已经把正向证据引用收敛为: + +```text +source_invocation_id + raw_path + evidence_excerpt +``` + +Gatekeeper 通过 `tool_invocation.retrieval_details.evidence_refs` 校验 Executor 引用是否真实存在。 + +但负向观察存在一个缺口:当工具明确返回“没查到”时,结果通常是空数组: + +```json +{ + "logs": [], + "total": 0, + "message": "未找到匹配的日志" +} +``` + +这时没有 `$.logs[0]`、`$.alerts[0]` 或 `$.evidence_blocks[0]` 可以引用。Executor 如果输出 `negative_observation`,Gatekeeper 无法稳定验证它引用的“无证据结果”,容易降级为 `LOW_CONFID` 或 `REJECT`。 + +--- + +## 问题 + +用户问: + +```text +只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。 +``` + +工具返回: + +```json +{ + "success": false, + "logs": [], + "total": 0, + "message": "未找到匹配的日志" +} +``` + +合理 claim 是: + +```json +{ + "claim_type": "negative_observation", + "claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。" +} +``` + +但旧设计只支持正向数组项: + +```text +$.alerts[i] +$.logs[i] +$.evidence_blocks[i] +``` + +因此负向观察缺少可回溯的精确引用点。 + +--- + +## 修复决策 + +给 no-hit / no-evidence 结果增加一等证据引用: + +```json +{ + "raw_path": "$.no_evidence", + "text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志" +} +``` + +Executor 可以引用: + +```json +{ + "tool_name": "query_logs", + "source_invocation_id": 123, + "raw_path": "$.no_evidence", + "evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence" +} +``` + +### 语义边界 + +`$.no_evidence` 只表示: + +```text +该工具对当前查询返回无匹配证据。 +``` + +它不表示: + +- 问题绝对不存在。 +- 根因被排除。 +- 系统已经健康。 +- 没有必要继续排查。 + +--- + +## 实施范围 + +### 已修改 + +- `ToolInvocationRecorder` + - `query_logs` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。 + - `query_metrics` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。 + - `lookup_knowledge` no-hit 且无 evidence blocks 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。 +- `ExecutorGatekeeperService` + - 复用既有 `evidence_refs` 校验逻辑,无需新增特殊分支。 + - `$.no_evidence` 和普通 raw_path 一样必须存在于 `retrieval_details.evidence_refs`。 +- `chat-executor-prompt.md` + - 明确 `negative_observation` 必须引用 `$.no_evidence`。 + - 明确没有实际工具调用时禁止使用 `$.no_evidence`。 + +### 未修改 + +- 不新增数据库表。 +- 不新增复杂 metadata。 +- 不改变 Planner。 +- 不改变 Agent 编排。 + +--- + +## 验收标准 + +1. `query_logs` 返回 `logs=[] / total=0 / evidence_status=no_evidence` 时,`tool_invocation.retrieval_details.evidence_refs` 包含: + +```json +{ + "raw_path": "$.no_evidence" +} +``` + +2. Executor 输出 `negative_observation` 并引用 `$.no_evidence` 时,Gatekeeper 可以校验通过。 + +3. Executor 如果用 `$.no_evidence` 搭配正向证据文本,例如 `HikariCP active=50/50`,Gatekeeper 必须拒绝。 + +4. `$.no_evidence` 不得被解释为“问题绝对不存在”,只能表达“当前查询未检索到匹配证据”。 + +5. HikariCP negative E2E: + +```text +只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。 +``` + +期望: + +- 工具返回 no-hit。 +- 不返回 `generic-service`。 +- Executor 输出 `negative_observation`。 +- `raw_path="$.no_evidence"`。 +- Gatekeeper `pass/none`。 +- Verifier 不误判为正向 HikariCP 证据。 + +--- + +## 验证记录 + +### 2026-07-08 单元测试 + +命令: + +```text +mvn '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,QueryLogsToolsTest' test +``` + +结果: + +```text +Tests run: 23, Failures: 0, Errors: 0, Skipped: 0 +BUILD SUCCESS +``` + +覆盖点: + +- `query_logs` no-hit 生成 `$.no_evidence`。 +- `query_metrics` no-hit 生成 `$.no_evidence`。 +- `lookup_knowledge` no-hit 生成 `$.no_evidence`。 +- Gatekeeper 可以校验 `$.no_evidence`。 +- 多个 no-evidence 调用存在时,Gatekeeper 可按 `tool_name + raw_path + evidence_excerpt` 唯一回填 `source_invocation_id`。 +- `negative_observation` 混绑正向 `$.logs[i]` 会被拒绝。 +- HikariCP negative mock 不返回 `generic-service`。 + +### 2026-07-08 E2E 验证 + +输入: + +```text +只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。 +``` + +最终通过 session: + +```text +sessionId: iss009-hikari-negative-latest-20260708-232428 +verdict: PASS +groundedness_score: 1.0 +gatekeeper_result.status: pass +gatekeeper_result.severity: none +claim_count: 1 +claim_type: negative_observation +raw_path: $.no_evidence +generic_service_hit: false +overstate_hit: false +``` + +Executor claim: + +```text +当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。 +``` + +最终答案: + +```text +本次查询在 inventory-service 中未发现 HikariCP 连接池耗尽的日志记录,检索结果未匹配到相关证据。 +``` + +说明: + +- Executor 仍可能输出多个 `$.no_evidence` binding。 +- 如果 `source_invocation_id` 缺失,Gatekeeper 会按 `tool_name + raw_path + evidence_excerpt` 唯一匹配真实 invocation 并写入 warning。 +- 最终答案不使用“排除”“确认没有”“不存在该问题”等过度表达。 diff --git a/mvp/issues/README.md b/mvp/issues/README.md index 4019314..f337c1c 100644 --- a/mvp/issues/README.md +++ b/mvp/issues/README.md @@ -9,6 +9,8 @@ | ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) | | ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) | | ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) | +| ISS-008 | Executor 窄范围查询越界 | 中 | 已修复 | [ISS-008-executor-narrow-scope-overreach.md](ISS-008-executor-narrow-scope-overreach.md) | +| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 中 | 已修复 | [ISS-009-negative-observation-no-evidence-reference.md](ISS-009-negative-observation-no-evidence-reference.md) | | executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) | | executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) | | expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) | diff --git a/src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java b/src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java index ccbe637..0f2144a 100644 --- a/src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java +++ b/src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java @@ -156,7 +156,8 @@ public class ExecutorGatekeeperService { SEVERITY_LOW_CONFID); continue; } - validateEvidenceBinding(binding, validInvocations, target, claim.get("claim_id"), result); + validateEvidenceBinding(binding, validInvocations, target, + claim.get("claim_id"), claim.get("claim_type"), result); } } @@ -181,7 +182,7 @@ public class ExecutorGatekeeperService { SEVERITY_LOW_CONFID); continue; } - validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), result); + validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), null, result); } } } @@ -190,6 +191,7 @@ public class ExecutorGatekeeperService { Map validInvocations, String target, Object ownerId, + Object ownerType, GatekeeperResult result) { Map checked = new LinkedHashMap<>(); checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId)); @@ -197,18 +199,6 @@ public class ExecutorGatekeeperService { checked.put("source_invocation_id", binding.get("source_invocation_id")); checked.put("raw_path", stringValue(binding.get("raw_path"))); - Long id = singleInvocationId(binding); - if (id == null) { - checked.put("status", STATUS_FAIL); - checked.put("rule", RULE_INVOCATION_REF); - checked.put("message", "source_invocation_id is required"); - result.checked(checked); - result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id", - "source_invocation_id is required", SEVERITY_LOW_CONFID); - return; - } - checked.put("source_invocation_id", id); - String claimedToolName = stringValue(binding.get("tool_name")); if (claimedToolName.isBlank()) { checked.put("status", STATUS_FAIL); @@ -219,6 +209,46 @@ public class ExecutorGatekeeperService { return; } + String rawPath = stringValue(binding.get("raw_path")); + if ("negative_observation".equals(stringValue(ownerType)) && !rawPath.isBlank() + && !"$.no_evidence".equals(rawPath)) { + checked.put("status", STATUS_FAIL); + checked.put("rule", RULE_RAW_PATH); + checked.put("message", "negative_observation must only bind $.no_evidence references"); + result.checked(checked); + result.fail(RULE_RAW_PATH, target + ".raw_path", + "negative_observation must only bind $.no_evidence references", SEVERITY_REJECT); + return; + } + + Long id = singleInvocationId(binding); + if (id == null) { + id = uniqueInvocationIdByToolRawPathAndExcerpt(validInvocations, claimedToolName, rawPath, + stringValue(binding.get("evidence_excerpt"))); + if (id == null) { + id = uniqueInvocationIdByToolAndRawPath(validInvocations, claimedToolName, rawPath); + } + if (id != null) { + result.warn(Map.of( + "rule", "evidence.invocation_auto_backfill_by_raw_path", + "message", "source_invocation_id was auto-filled from the unique evidence reference candidate", + "tool_name", claimedToolName, + "raw_path", rawPath, + "source_invocation_id", id + )); + } + } + if (id == null) { + checked.put("status", STATUS_FAIL); + checked.put("rule", RULE_INVOCATION_REF); + checked.put("message", "source_invocation_id is required"); + result.checked(checked); + result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id", + "source_invocation_id is required", SEVERITY_LOW_CONFID); + return; + } + checked.put("source_invocation_id", id); + ToolInvocation invocation = validInvocations.get(id); if (invocation == null) { checked.put("status", STATUS_FAIL); @@ -240,7 +270,6 @@ public class ExecutorGatekeeperService { return; } - String rawPath = stringValue(binding.get("raw_path")); if (rawPath.isBlank()) { checked.put("status", STATUS_FAIL); checked.put("rule", RULE_RAW_PATH); @@ -298,6 +327,82 @@ public class ExecutorGatekeeperService { result.checked(checked); } + private Long uniqueInvocationIdByToolAndRawPath(Map validInvocations, + String toolName, + String rawPath) { + if (toolName == null || toolName.isBlank() || rawPath == null || rawPath.isBlank()) { + return null; + } + Long matchedId = null; + for (Map.Entry entry : validInvocations.entrySet()) { + ToolInvocation invocation = entry.getValue(); + if (!Objects.equals(toolName, invocation.getToolName())) { + continue; + } + if (!evidenceRefsByRawPath(invocation.getRetrievalDetails()).containsKey(rawPath)) { + continue; + } + if (matchedId != null) { + return null; + } + matchedId = entry.getKey(); + } + return matchedId; + } + + private Long uniqueInvocationIdByToolRawPathAndExcerpt(Map validInvocations, + String toolName, + String rawPath, + String excerpt) { + if (toolName == null || toolName.isBlank() + || rawPath == null || rawPath.isBlank() + || excerpt == null || excerpt.isBlank()) { + return null; + } + Long matchedId = null; + for (Map.Entry entry : validInvocations.entrySet()) { + ToolInvocation invocation = entry.getValue(); + if (!Objects.equals(toolName, invocation.getToolName())) { + continue; + } + String matchedText = evidenceRefsByRawPath(invocation.getRetrievalDetails()).get(rawPath); + if (matchedText == null || !isBackfillCandidateSupported(rawPath, excerpt, matchedText)) { + continue; + } + if (matchedId != null) { + return null; + } + matchedId = entry.getKey(); + } + return matchedId; + } + + private boolean isBackfillCandidateSupported(String rawPath, String excerpt, String matchedText) { + if ("$.no_evidence".equals(rawPath)) { + String excerptQuery = semicolonField(excerpt, "query"); + String matchedQuery = semicolonField(matchedText, "query"); + if (!excerptQuery.isBlank() && !matchedQuery.isBlank() + && !normalized(excerptQuery).equals(normalized(matchedQuery))) { + return false; + } + } + return isExcerptSupported(excerpt, matchedText); + } + + private String semicolonField(String text, String field) { + if (text == null || text.isBlank() || field == null || field.isBlank()) { + return ""; + } + String prefix = field + "="; + for (String part : text.split(";")) { + String trimmed = part.trim(); + if (trimmed.regionMatches(true, 0, prefix, 0, prefix.length())) { + return trimmed.substring(prefix.length()).trim(); + } + } + return ""; + } + private Long singleInvocationId(Map binding) { Long singular = asLong(binding.get("source_invocation_id")); if (singular != null) { diff --git a/src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java b/src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java index 4f6c2ca..d0d03ed 100644 --- a/src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java +++ b/src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java @@ -88,7 +88,7 @@ public class ToolInvocationRecorder { if (extraDetails != null && !extraDetails.isEmpty()) { details.putAll(extraDetails); } - List> evidenceRefs = extractEvidenceRefs(toolName, output, details); + List> evidenceRefs = extractEvidenceRefs(toolName, inputParams, output, details); if (!evidenceRefs.isEmpty()) { details.put("evidence_refs", evidenceRefs); } @@ -145,6 +145,12 @@ public class ToolInvocationRecorder { } details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks()); List> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks()); + if (evidenceRefs.isEmpty() + && EVIDENCE_STATUS_NO_EVIDENCE.equals(normalizeEvidenceStatus(record.success(), record.evidenceStatus()))) { + Map input = new LinkedHashMap<>(); + input.put("query", record.query()); + evidenceRefs = List.of(noEvidenceRef("lookup_knowledge", input, null, details)); + } if (!evidenceRefs.isEmpty()) { details.put("evidence_refs", evidenceRefs); } @@ -195,7 +201,10 @@ public class ToolInvocationRecorder { return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED; } - private List> extractEvidenceRefs(String toolName, String output, Map details) { + private List> extractEvidenceRefs(String toolName, + Map inputParams, + String output, + Map details) { if (output == null || output.isBlank()) { return List.of(); } @@ -205,11 +214,14 @@ public class ToolInvocationRecorder { try { JsonNode root = objectMapper.readTree(output); + boolean noEvidence = EVIDENCE_STATUS_NO_EVIDENCE.equals(stringValue(details.get("evidence_status"))); if ("query_metrics".equals(toolName)) { - return evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText); + List> refs = evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText); + return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs; } if ("query_logs".equals(toolName)) { - return evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText); + List> refs = evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText); + return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs; } } catch (Exception e) { log.debug("extract evidence_refs failed for tool={}", toolName, e); @@ -217,6 +229,42 @@ public class ToolInvocationRecorder { return List.of(); } + private Map noEvidenceRef(String toolName, + Map inputParams, + JsonNode root, + Map details) { + List parts = new ArrayList<>(); + addPart(parts, toolName + " returned no evidence"); + addPart(parts, "evidence_status=" + stringValue(details.get("evidence_status"))); + String query = firstNonBlank( + root == null ? null : textField(root, "query"), + inputParams == null ? null : inputParams.get("query") + ); + if (!query.isBlank()) { + addPart(parts, "query=" + query); + } + String topic = firstNonBlank( + root == null ? null : textField(root, "log_topic"), + inputParams == null ? null : inputParams.get("log_topic"), + details == null ? null : details.get("log_topic"), + details == null ? null : details.get("metric_family") + ); + if (!topic.isBlank()) { + addPart(parts, "topic=" + topic); + } + if (root != null && root.has("total")) { + addPart(parts, "total=" + root.path("total").asText()); + } + String message = root == null ? "" : textField(root, "message"); + if (!message.isBlank()) { + addPart(parts, "message=" + message); + } + return Map.of( + "raw_path", "$.no_evidence", + "text", bounded(String.join("; ", parts), 500) + ); + } + private List> evidenceRefsFromArray(JsonNode arrayNode, String pathPrefix, java.util.function.Function textExtractor) { @@ -314,6 +362,10 @@ public class ToolInvocationRecorder { return ""; } + private String stringValue(Object value) { + return value == null ? "" : String.valueOf(value); + } + private String bounded(String value, int limit) { if (value == null) { return ""; diff --git a/src/main/resources/prompts/chat-composer-prompt.md b/src/main/resources/prompts/chat-composer-prompt.md index 20fb1dc..8812799 100644 --- a/src/main/resources/prompts/chat-composer-prompt.md +++ b/src/main/resources/prompts/chat-composer-prompt.md @@ -8,6 +8,8 @@ - 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。 - 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。 - 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。 +- 当 `allowed_claims` 中存在 `claim_type=negative_observation`,或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。 +- 对 `negative_observation` / `$.no_evidence`,禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。 ## 输入字段 @@ -26,6 +28,7 @@ - 可以表达确认结论。 - 只能使用 `allowed_claims` 和 `recommended_actions`。 - 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。 +- 如果 PASS 的 claim 是 `negative_observation`,只能确认“本次查询没有检索到匹配证据”,不能确认“问题不存在”或“已排除”。 ### LOW_CONFID diff --git a/src/main/resources/prompts/chat-executor-prompt.md b/src/main/resources/prompts/chat-executor-prompt.md index 6aa3d20..4748907 100644 --- a/src/main/resources/prompts/chat-executor-prompt.md +++ b/src/main/resources/prompts/chat-executor-prompt.md @@ -7,15 +7,114 @@ - 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。 - 你是证据收集与微观事实提炼器,不是最终答复生成器。 +## 角色边界 HARD-GATE + +你只负责证据收集与微观事实提炼,只能输出“当前工具证据可以直接支持的观察事实”。 + +你不是: +- 根因诊断器。 +- 修复方案生成器。 +- Runbook 转述器。 +- 经验推断器。 +- 最终用户答复生成器。 + +除非本轮工具返回中存在直接证据,否则禁止输出: +- 根因确认。 +- 修复建议。 +- 扩展排查方向。 +- 历史经验。 +- 通用知识。 +- 与用户问题无关的服务、指标、订单、错误码、组件。 + ## 规则 - 按顺序执行,不可跳过步骤。 - 所有事实性结论必须来自本轮 evidence tools 的返回。 -- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。 +- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导,不能直接写成本次事故的已确认事实。 - 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。 -- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。 -- 对窄范围确认问题,只输出与用户问题直接相关的 observation / negative_observation。通常 1 条 claim,最多 2 条 claim;不要限制 evidence_bindings 数量。 +- 不要使用“通常情况下”“根据经验”“很可能已经发生”“可能是”“推测”“理论上”等无证据推断词来伪装事实。 - 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。 +## 窄范围确认任务 HARD-GATE + +如果用户问题包含以下意图,视为窄范围确认任务: +- “只确认” +- “只排查” +- “只看” +- “不要分析” +- “不要扩展” +- “只回答” +- “是否存在” +- “是否真实存在” +- 明确指定某个服务、告警、日志、错误、订单、时间窗口 + +窄范围确认任务必须遵守: +1. `claims` 只能输出 `observation` 或 `negative_observation`。 +2. claim 数量必须是最少必要数量,通常 1 条,最多 2 条。 +3. claim 数量限制不限制 `evidence_bindings` 数量;一条 claim 可以绑定多条直接相关证据。 +4. 不得把同一观察事实拆成多条 claim。 +5. 只能围绕用户明确要求的目标对象和主题输出 claim。 +6. 用户明确排除的对象、服务、告警、订单、数据库、连接池、下游依赖,禁止出现在 claim 中。 +7. 禁止输出根因类、风险类或建议类 claim,例如 `root_cause`、`risk`、`recommendation`。 +8. 如果证据不足,不要补合理化解释;优先写入 `missing_info`。 +9. 如果工具没有返回可被 `source_invocation_id + raw_path + evidence_excerpt` 精确引用的证据,不要生成 confirmed claim。 +10. 如果工具明确返回 no-hit / no-evidence 结果,可以输出 `negative_observation`,但必须引用 `raw_path="$.no_evidence"`。 +11. 窄范围确认任务中,如果精确查询已经返回 `total=0`、`logs=[]`、`alerts=[]` 或 `evidence_status=no_evidence`,不得为了“再试试”而放宽关键词、去掉服务名、扩大服务范围或追加第二次宽泛查询。 + +窄范围任务的理想输出是: +- 1 条核心 claim。 +- 多条直接相关 `evidence_bindings`。 +- 必要的 `missing_info`。 + +同一条工具数组项只能绑定一次。不要为了引用其中多个字段而拆成多个 `evidence_bindings`。 + +正确: +```json +{ + "raw_path": "$.alerts[0]", + "evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m" +} +``` + +错误: +```json +{ "raw_path": "$.alerts[0].alert_name", "evidence_excerpt": "HighCPUUsage" } +{ "raw_path": "$.alerts[0].state", "evidence_excerpt": "firing" } +``` + +负向观察示例: +```json +{ + "claim_type": "negative_observation", + "claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。", + "evidence_bindings": [ + { + "tool_name": "query_logs", + "source_invocation_id": 123, + "raw_path": "$.no_evidence", + "evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence" + } + ] +} +``` + +`$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”,不能表示“问题不存在”或“根因被排除”。没有实际调用工具时,禁止使用 `$.no_evidence`。 + +`negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`。禁止把其它服务的正向日志或告警绑定到同一个 `negative_observation`,即使这些日志可以说明“不是当前服务”。 + +输出 `negative_observation` 或基于 `$.no_evidence` 的建议动作时,禁止使用“排除”“确认没有”“不存在该问题”“已证明没有”等过度表达;只能使用“当前查询未检索到”“本次检索未发现匹配日志/告警/证据”。 + +## 工具使用边界 + +你只能调用回答当前用户问题所必需的工具。 + +- 问告警状态:优先使用 `query_metrics`。 +- 问日志现象:优先使用 `query_logs`。 +- 问知识解释或排查步骤:才使用 `lookup_knowledge`。 +- Runbook / Skill / 知识库只能帮助决定“查什么”,不能直接作为“当前环境发生了什么”的证据。 +- 如果当前工具结果已经足以回答用户问题,不要继续扩展检索。 +- 不要为了补全故事而查询用户没有要求的服务、组件或故障类型。 +- 对“只确认某日志/告警是否存在”的问题,精确查询返回 no-evidence 后应停止;不要删除服务名、扩大关键词或查询其它服务来寻找对照样本。 + ## 检索约束 ### 1. 判断重复:基于已检索上下文 @@ -67,6 +166,19 @@ ### missing_info `missing_info` 用来列出无法确认结论所缺少的具体证据。 +## 输出前自检 + +在输出 JSON 前,逐项检查: + +1. 每条 claim 是否直接回答了用户当前问题? +2. 每条 claim 是否都有真实 `evidence_bindings`? +3. 每个 `evidence_excerpt` 是否来自工具返回原文? +4. 是否出现了用户没有要求的服务、告警、订单、数据库、连接池或下游组件? +5. 是否把 Runbook / Skill / 知识库通用内容写成了当前事实? +6. 是否输出了根因、修复动作、风险判断或经验推断? + +只要任一项不通过,删除对应 claim,不要解释。 + ## 最终输出格式(严格契约) 你必须输出且只能输出一个 JSON 对象,不要输出 Markdown,不要输出代码块,不要输出 JSON 之外的解释文字。 @@ -87,7 +199,7 @@ "source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空", "tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool", "source_invocation_id": null, - "raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0]", + "raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0] / $.no_evidence", "evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据" } ] @@ -121,6 +233,8 @@ - `claims[*].evidence_bindings` 不能为空。 - `evidence_excerpt` 必须来自工具返回,不允许编造。 - `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。 +- 当且仅当工具明确返回 no-hit / no-evidence 结果时,允许使用 `$.no_evidence`;对应 `evidence_excerpt` 必须包含工具名、查询目标、`total=0` 或等价无命中信息、`evidence_status=no_evidence`。 +- `raw_path` 禁止指向字段级子路径,例如 `$.alerts[0].alert_name`、`$.alerts[0].state`、`$.logs[0].message` 都是非法路径。需要引用多个字段时,仍然只使用对应数组条目的 `raw_path`,并把必要字段合并进同一个 `evidence_excerpt`。 - `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。 - 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。 - 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。 diff --git a/src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java b/src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java index 5b14d59..3a9050d 100644 --- a/src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java +++ b/src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java @@ -33,6 +33,111 @@ class ExecutorGatekeeperServiceTest { assertTrue(((List) result.get("failed_rules")).isEmpty()); } + @Test + void validatePassesForNoEvidenceReference() { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of( + invocation(101L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志") + )); + ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository); + + Map result = service.validate("session-1", + validOutput(101L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"), + Map.of("status", "valid")); + + assertEquals("pass", result.get("status")); + assertEquals("none", result.get("severity")); + assertTrue(((List) result.get("failed_rules")).isEmpty()); + } + + @Test + void validateBackfillsNoEvidenceInvocationByRawPathWhenToolHasMultipleCalls() { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of( + invocation(101L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志"), + invocation(102L, "query_logs", "$.logs[0]", + "order-service HikariCP active=50/50 waiting=32") + )); + ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository); + + Map result = service.validate("session-1", + validOutput(null, "query_logs", "$.no_evidence", + "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence", + "negative_observation"), + Map.of("status", "valid")); + + assertEquals("pass", result.get("status")); + assertEquals("none", result.get("severity")); + assertTrue(((List) result.get("failed_rules")).isEmpty()); + assertEquals("evidence.invocation_auto_backfill_by_raw_path", + ((Map) ((List) result.get("warnings")).get(0)).get("rule")); + } + + @Test + void validateBackfillsNoEvidenceInvocationByExcerptWhenRawPathIsRepeated() { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of( + invocation(101L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; evidence_status=no_evidence; query=service:inventory-service AND HikariCP; total=0; message=未找到匹配的日志"), + invocation(102L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service; total=0; message=未找到匹配的日志") + )); + ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository); + + Map result = service.validate("session-1", + validOutput(null, "query_logs", "$.no_evidence", + "query_logs returned no evidence; query=service:inventory-service AND HikariCP; total=0; evidence_status=no_evidence", + "negative_observation"), + Map.of("status", "valid")); + + assertEquals("pass", result.get("status")); + assertEquals("none", result.get("severity")); + assertEquals(101L, + ((Map) ((List) result.get("checked_bindings")).get(0)).get("source_invocation_id")); + } + + @Test + void validateRejectsNoEvidenceReferenceWhenExcerptClaimsPositiveEvidence() { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of( + invocation(101L, "query_logs", "$.no_evidence", + "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志") + )); + ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository); + + Map result = service.validate("session-1", + validOutput(101L, "query_logs", "$.no_evidence", + "HikariCP active=50/50 waiting=32"), + Map.of("status", "valid")); + + assertEquals("fail", result.get("status")); + assertEquals("reject", result.get("severity")); + assertTrue(((List) result.get("failed_rules")).contains("evidence.excerpt_mismatch")); + } + + @Test + void validateRejectsPositiveBindingOnNegativeObservation() { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of( + invocation(101L, "query_logs", "$.logs[0]", + "order-service HikariCP active=50/50 waiting=32") + )); + ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository); + + Map result = service.validate("session-1", + validOutput(101L, "query_logs", "$.logs[0]", + "order-service HikariCP active=50/50 waiting=32", + "negative_observation"), + Map.of("status", "valid")); + + assertEquals("fail", result.get("status")); + assertEquals("reject", result.get("severity")); + assertTrue(((List) result.get("failed_rules")).contains("evidence.raw_path")); + } + @Test void validateFailsWhenRemovedFieldsArePresent() { ToolInvocationRepository repository = mock(ToolInvocationRepository.class); @@ -191,11 +296,21 @@ class ExecutorGatekeeperServiceTest { } private Map validOutput(Long invocationId, String toolName, String rawPath, String excerpt) { + return validOutput(invocationId, toolName, rawPath, excerpt, "symptom"); + } + + private Map validOutput(Long invocationId, + String toolName, + String rawPath, + String excerpt, + String claimType) { Map binding = new java.util.LinkedHashMap<>(); binding.put("source_type", "tool_trace"); binding.put("source_id", "trace-1"); binding.put("tool_name", toolName); - binding.put("source_invocation_id", invocationId); + if (invocationId != null) { + binding.put("source_invocation_id", invocationId); + } if (rawPath != null) { binding.put("raw_path", rawPath); } @@ -204,7 +319,7 @@ class ExecutorGatekeeperServiceTest { "answer_version", "executor_evidence_v2", "claims", List.of(Map.of( "claim_id", "claim-1", - "claim_type", "symptom", + "claim_type", claimType, "claim_text", "连接池 active 达到上限", "support_level", "direct", "evidence_bindings", List.of(binding) diff --git a/src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java b/src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java index cdb1a51..b595e32 100644 --- a/src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java +++ b/src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java @@ -27,7 +27,7 @@ class ToolInvocationRecorderTest { private final ObjectMapper objectMapper = new ObjectMapper(); @Test - void recordEvidenceToolPreservesNoEvidenceSemantics() { + void recordEvidenceToolPreservesNoEvidenceSemantics() throws Exception { ToolInvocationRepository repository = mock(ToolInvocationRepository.class); when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0)); ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper()); @@ -36,8 +36,8 @@ class ToolInvocationRecorderTest { try { recorder.recordEvidenceTool( "query_logs", - Map.of("query", "timeout"), - "{\"success\":false,\"message\":\"未找到匹配的日志\"}", + Map.of("query", "inventory-service HikariCP", "log_topic", "application-logs"), + "{\"success\":false,\"query\":\"inventory-service HikariCP\",\"log_topic\":\"application-logs\",\"logs\":[],\"total\":0,\"message\":\"未找到匹配的日志\"}", true, System.currentTimeMillis() - 10, null, @@ -57,6 +57,14 @@ class ToolInvocationRecorderTest { assertEquals(Boolean.TRUE, saved.getSuccess()); assertTrue(saved.getRetrievalDetails().contains("\"evidence_status\":\"no_evidence\"")); assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]")); + + JsonNode details = objectMapper.readTree(saved.getRetrievalDetails()); + assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText()); + String text = details.path("evidence_refs").get(0).path("text").asText(); + assertTrue(text.contains("query_logs returned no evidence")); + assertTrue(text.contains("inventory-service HikariCP")); + assertTrue(text.contains("total=0")); + assertTrue(text.contains("evidence_status=no_evidence")); } @Test @@ -125,6 +133,40 @@ class ToolInvocationRecorderTest { assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage")); } + @Test + void recordEvidenceToolExtractsMetricNoEvidenceRef() throws Exception { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0)); + ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper()); + SessionContextHolder.setSessionId("metric-no-evidence-session"); + + try { + recorder.recordEvidenceTool( + "query_metrics", + Map.of("query", "active_prometheus_alerts"), + """ + {"success":true,"alerts":[],"message":"成功检索到 0 个活动告警"} + """, + true, + System.currentTimeMillis() - 10, + null, + "prometheus_alerts", + ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE, + Map.of("metric_family", "prometheus_alerts") + ); + } finally { + SessionContextHolder.clear(); + } + + ArgumentCaptor captor = ArgumentCaptor.forClass(ToolInvocation.class); + verify(repository).save(captor.capture()); + JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails()); + + assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText()); + assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("query_metrics returned no evidence")); + assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("prometheus_alerts")); + } + @Test void recordLookupKnowledgePreservesRetrievalSpecificFields() { ToolInvocationRepository repository = mock(ToolInvocationRepository.class); @@ -190,6 +232,38 @@ class ToolInvocationRecorderTest { assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\"")); } + @Test + void recordLookupKnowledgeAddsNoEvidenceRefWhenNoBlocks() throws Exception { + ToolInvocationRepository repository = mock(ToolInvocationRepository.class); + when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0)); + ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper()); + SessionContextHolder.setSessionId("lookup-no-evidence-session"); + + ToolInvocationRecorder.LookupKnowledgeRecord record = ToolInvocationRecorder.LookupKnowledgeRecord.builder() + .query("inventory-service HikariCP") + .outputPreview("") + .outputLength(0) + .durationMs(12) + .success(true) + .evidenceStatus(ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE) + .evidenceBlocks(List.of()) + .build(); + + try { + recorder.recordLookupKnowledge(record); + } finally { + SessionContextHolder.clear(); + } + + ArgumentCaptor captor = ArgumentCaptor.forClass(ToolInvocation.class); + verify(repository).save(captor.capture()); + JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails()); + + assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText()); + assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("lookup_knowledge returned no evidence")); + assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("inventory-service HikariCP")); + } + @Test void lookupKnowledgeRecordFromSummarizesEvidenceBlocks() { EvidenceBlock block = EvidenceBlock.builder()