feat(agent): support no-evidence references
This commit is contained in:
@@ -0,0 +1,308 @@
|
|||||||
|
# ISS-008 Executor 窄范围查询越界
|
||||||
|
|
||||||
|
**严重程度**:中
|
||||||
|
**状态**:已修复
|
||||||
|
**发现时间**:2026-07-08
|
||||||
|
**关联**:
|
||||||
|
- `ISS-007-verifier-evidence-summary-fidelity`
|
||||||
|
- `executor-structured-output-v2`
|
||||||
|
- `executor-evidence-attribution-hallucination`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
当前 Chat 诊断链路已经演进为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Planner
|
||||||
|
-> Executor
|
||||||
|
-> VerifierInputHook / Gatekeeper
|
||||||
|
-> Verifier
|
||||||
|
-> Composer
|
||||||
|
```
|
||||||
|
|
||||||
|
其中 Executor 的定位已经从“生成最终诊断答案”收敛为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
证据收集 + 微观事实提炼
|
||||||
|
```
|
||||||
|
|
||||||
|
但在窄范围问题中,Executor 仍可能把用户只要求确认的一件事扩展成多条 claim,例如用户只问 `HighCPUUsage`,Executor 可能顺手输出内存、连接池、数据库或修复建议相关内容。
|
||||||
|
|
||||||
|
这类问题不一定是证据伪造。很多时候工具返回里确实有其它信息,但它们不属于当前用户问题的范围。Gatekeeper 只能校验证据引用真假,不能完整承担“用户意图范围控制”;Verifier 虽然可以降级,但会增加链路负担。
|
||||||
|
|
||||||
|
因此本 issue 采用低成本的 Prompt-first 修复:先收紧 Executor prompt,不改 Planner,不引入 `scope_contract`。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 问题类型
|
||||||
|
|
||||||
|
### 1. 窄范围查询越界
|
||||||
|
|
||||||
|
用户问题只要求确认一个服务、告警、日志、订单或时间窗口,但 Executor 输出了用户未要求的 claim。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单、OOM、数据库慢查询、连接池或 user-service。
|
||||||
|
```
|
||||||
|
|
||||||
|
错误输出包括:
|
||||||
|
|
||||||
|
- `HighMemoryUsage`
|
||||||
|
- `SlowResponse`
|
||||||
|
- `order-123`
|
||||||
|
- `HikariCP`
|
||||||
|
- `DB / database`
|
||||||
|
- `user-service`
|
||||||
|
|
||||||
|
### 2. Observation 变成 Diagnosis
|
||||||
|
|
||||||
|
Executor 本应输出观察事实,却输出根因、风险、修复建议或经验推断。
|
||||||
|
|
||||||
|
错误输出包括:
|
||||||
|
|
||||||
|
- “CPU 过高是请求超时的根因”
|
||||||
|
- “建议扩容”
|
||||||
|
- “通常这种情况是数据库慢查询导致”
|
||||||
|
- “存在内存泄漏风险”
|
||||||
|
|
||||||
|
### 3. Runbook 通用知识变成当前事实
|
||||||
|
|
||||||
|
Runbook、Skill、知识库可以指导要查什么,但不能直接变成本次环境已发生的事实。
|
||||||
|
|
||||||
|
错误输出包括:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Runbook 中说 HighCPUUsage 常见原因是流量突增,所以当前环境发生了流量突增。
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 修复决策
|
||||||
|
|
||||||
|
本期只修 Executor prompt。
|
||||||
|
|
||||||
|
### 本期做
|
||||||
|
|
||||||
|
1. 强化 Executor 单一职责:证据收集 + 微观事实提炼。
|
||||||
|
2. 增加 `角色边界 HARD-GATE`。
|
||||||
|
3. 增加 `窄范围确认任务 HARD-GATE`。
|
||||||
|
4. 增加工具使用边界,避免为补全故事而扩展检索。
|
||||||
|
5. 增加输出前自检,要求输出 JSON 前删除越界 claim。
|
||||||
|
|
||||||
|
### 本期不做
|
||||||
|
|
||||||
|
1. 不改 Planner。
|
||||||
|
2. 不新增 `scope_contract`。
|
||||||
|
3. 不解析 Planner 输出中的 scope。
|
||||||
|
4. 不做 Gatekeeper scope 校验。
|
||||||
|
5. 不改多 Agent 编排。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 设计原则
|
||||||
|
|
||||||
|
### Claim 要少,Evidence 可以多
|
||||||
|
|
||||||
|
窄范围任务下,Executor 应输出最少必要 claim,通常 1 条,最多 2 条。
|
||||||
|
|
||||||
|
但 claim 数量限制不限制 `evidence_bindings` 数量。一条核心 claim 可以绑定多条直接相关证据。
|
||||||
|
|
||||||
|
```text
|
||||||
|
正确:
|
||||||
|
1 条 claim + 多条 evidence_bindings
|
||||||
|
|
||||||
|
错误:
|
||||||
|
为了展示多条证据,把同一个观察事实拆成多条 claim
|
||||||
|
```
|
||||||
|
|
||||||
|
### 只输出当前问题范围内的 Observation
|
||||||
|
|
||||||
|
窄范围任务下,`claims` 只能使用:
|
||||||
|
|
||||||
|
- `observation`
|
||||||
|
- `negative_observation`
|
||||||
|
|
||||||
|
禁止使用:
|
||||||
|
|
||||||
|
- `root_cause`
|
||||||
|
- `risk`
|
||||||
|
- `recommendation`
|
||||||
|
- 其它建议类或诊断类 claim
|
||||||
|
|
||||||
|
### 证据不足时不要补故事
|
||||||
|
|
||||||
|
如果工具没有返回可被精确引用的证据:
|
||||||
|
|
||||||
|
```text
|
||||||
|
source_invocation_id + raw_path + evidence_excerpt
|
||||||
|
```
|
||||||
|
|
||||||
|
Executor 不应生成 confirmed claim,应写入 `missing_info`。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Prompt 修复点
|
||||||
|
|
||||||
|
已更新:
|
||||||
|
|
||||||
|
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||||
|
|
||||||
|
核心新增约束:
|
||||||
|
|
||||||
|
1. `角色边界 HARD-GATE`
|
||||||
|
2. `窄范围确认任务 HARD-GATE`
|
||||||
|
3. `工具使用边界`
|
||||||
|
4. `输出前自检`
|
||||||
|
5. 条目级 `raw_path` 强约束:同一条工具数组项只能绑定一次,禁止输出 `$.alerts[0].alert_name`、`$.alerts[0].state` 等字段级子路径。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 验证结果
|
||||||
|
|
||||||
|
### 2026-07-08 E2E 验证
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
|
||||||
|
```
|
||||||
|
|
||||||
|
第一次验证发现:
|
||||||
|
|
||||||
|
- Executor 已经只输出 `payment-service + HighCPUUsage` 相关 observation,没有输出越界 claim。
|
||||||
|
- 但 Executor 额外生成了字段级 `raw_path`:
|
||||||
|
- `$.alerts[0].alert_name`
|
||||||
|
- `$.alerts[0].state`
|
||||||
|
- 当前 Gatekeeper 只支持条目级路径 `$.alerts[i]` / `$.logs[i]` / `$.evidence_blocks[i]`,因此判定为 `REJECT`。
|
||||||
|
|
||||||
|
已追加 prompt 约束:
|
||||||
|
|
||||||
|
```text
|
||||||
|
同一条工具数组项只能绑定一次。
|
||||||
|
不要为了引用其中多个字段而拆成多个 evidence_bindings。
|
||||||
|
raw_path 禁止指向字段级子路径。
|
||||||
|
```
|
||||||
|
|
||||||
|
第二次验证结果:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId: iss008-narrow-highcpu-rerun-20260708-215510
|
||||||
|
verdict: PASS
|
||||||
|
groundedness_score: 1.0
|
||||||
|
gatekeeper_result.status: pass
|
||||||
|
gatekeeper_result.severity: none
|
||||||
|
claim_count: 1
|
||||||
|
claim_type: observation
|
||||||
|
forbidden_hits: none
|
||||||
|
```
|
||||||
|
|
||||||
|
Executor claim:
|
||||||
|
|
||||||
|
```text
|
||||||
|
payment-service 当前存在 HighCPUUsage 告警,CPU 使用率持续超过 80%,当前值为 92%,告警状态为 firing,已持续 25 分钟。
|
||||||
|
```
|
||||||
|
|
||||||
|
最终答案未出现以下排除项:
|
||||||
|
|
||||||
|
- `HighMemoryUsage`
|
||||||
|
- `SlowResponse`
|
||||||
|
- `order-123`
|
||||||
|
- `订单123`
|
||||||
|
- `OOM`
|
||||||
|
- `DB / database`
|
||||||
|
- `HikariCP`
|
||||||
|
- `connection pool / 连接池`
|
||||||
|
- `user-service`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 验收标准
|
||||||
|
|
||||||
|
### 1. HighCPUUsage 窄范围
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- `claims` 只围绕 `payment-service + HighCPUUsage`。
|
||||||
|
- 不出现 `HighMemoryUsage`。
|
||||||
|
- 不出现 `SlowResponse`。
|
||||||
|
- 不出现 `order-123`。
|
||||||
|
- 不出现 `OOM`。
|
||||||
|
- 不出现 `DB / database`。
|
||||||
|
- 不出现 `HikariCP / connection pool`。
|
||||||
|
- 不出现 `user-service`。
|
||||||
|
|
||||||
|
### 2. HighMemoryUsage 窄范围
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 order-service 是否存在 HighMemoryUsage。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- 可以输出内存使用率、告警状态、持续时间等观察事实。
|
||||||
|
- 不输出“内存泄漏已确认”。
|
||||||
|
- 不输出扩容、重启、修改 JVM 参数等修复建议。
|
||||||
|
|
||||||
|
### 3. SlowResponse 窄范围
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 user-service 是否存在 SlowResponse 告警和慢请求日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- 可以绑定 alert 和 logs 多条证据。
|
||||||
|
- 不推断数据库连接池耗尽。
|
||||||
|
- 不推断下游服务故障。
|
||||||
|
|
||||||
|
### 4. 用户明确排除项
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只看 order-service 支付失败日志,不要分析 HikariCP。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- `claim_text` 不出现 HikariCP 确认结论。
|
||||||
|
- 最终答案不出现 HikariCP 确认结论。
|
||||||
|
|
||||||
|
### 5. 证据不足
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- 如果工具返回 `logs=[]`,Executor 不编造 positive claim。
|
||||||
|
- 输出 `negative_observation` 或 `missing_info`。
|
||||||
|
- 不返回 `generic-service` 占位事实。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 后续增强
|
||||||
|
|
||||||
|
如果 Prompt-first 后仍不稳定,再考虑:
|
||||||
|
|
||||||
|
1. Planner 输出 `scope_contract`。
|
||||||
|
2. Gatekeeper 增加 scope 校验。
|
||||||
|
3. eval fixture 增加 forbidden claim 自动断言。
|
||||||
|
|
||||||
|
本期暂不进入这些改造。
|
||||||
@@ -0,0 +1,239 @@
|
|||||||
|
# ISS-009 negative_observation 精确引用 no-evidence 结果
|
||||||
|
|
||||||
|
**严重程度**:中
|
||||||
|
**状态**:已修复
|
||||||
|
**发现时间**:2026-07-08
|
||||||
|
**关联**:
|
||||||
|
- `ISS-007-verifier-evidence-summary-fidelity`
|
||||||
|
- `ISS-008-executor-narrow-scope-overreach`
|
||||||
|
- `executor-structured-output-v2`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
ISS-007 已经把正向证据引用收敛为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
source_invocation_id + raw_path + evidence_excerpt
|
||||||
|
```
|
||||||
|
|
||||||
|
Gatekeeper 通过 `tool_invocation.retrieval_details.evidence_refs` 校验 Executor 引用是否真实存在。
|
||||||
|
|
||||||
|
但负向观察存在一个缺口:当工具明确返回“没查到”时,结果通常是空数组:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"logs": [],
|
||||||
|
"total": 0,
|
||||||
|
"message": "未找到匹配的日志"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
这时没有 `$.logs[0]`、`$.alerts[0]` 或 `$.evidence_blocks[0]` 可以引用。Executor 如果输出 `negative_observation`,Gatekeeper 无法稳定验证它引用的“无证据结果”,容易降级为 `LOW_CONFID` 或 `REJECT`。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 问题
|
||||||
|
|
||||||
|
用户问:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
工具返回:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"success": false,
|
||||||
|
"logs": [],
|
||||||
|
"total": 0,
|
||||||
|
"message": "未找到匹配的日志"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
合理 claim 是:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"claim_type": "negative_observation",
|
||||||
|
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
但旧设计只支持正向数组项:
|
||||||
|
|
||||||
|
```text
|
||||||
|
$.alerts[i]
|
||||||
|
$.logs[i]
|
||||||
|
$.evidence_blocks[i]
|
||||||
|
```
|
||||||
|
|
||||||
|
因此负向观察缺少可回溯的精确引用点。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 修复决策
|
||||||
|
|
||||||
|
给 no-hit / no-evidence 结果增加一等证据引用:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"raw_path": "$.no_evidence",
|
||||||
|
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Executor 可以引用:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"source_invocation_id": 123,
|
||||||
|
"raw_path": "$.no_evidence",
|
||||||
|
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 语义边界
|
||||||
|
|
||||||
|
`$.no_evidence` 只表示:
|
||||||
|
|
||||||
|
```text
|
||||||
|
该工具对当前查询返回无匹配证据。
|
||||||
|
```
|
||||||
|
|
||||||
|
它不表示:
|
||||||
|
|
||||||
|
- 问题绝对不存在。
|
||||||
|
- 根因被排除。
|
||||||
|
- 系统已经健康。
|
||||||
|
- 没有必要继续排查。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 实施范围
|
||||||
|
|
||||||
|
### 已修改
|
||||||
|
|
||||||
|
- `ToolInvocationRecorder`
|
||||||
|
- `query_logs` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||||
|
- `query_metrics` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||||
|
- `lookup_knowledge` no-hit 且无 evidence blocks 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||||
|
- `ExecutorGatekeeperService`
|
||||||
|
- 复用既有 `evidence_refs` 校验逻辑,无需新增特殊分支。
|
||||||
|
- `$.no_evidence` 和普通 raw_path 一样必须存在于 `retrieval_details.evidence_refs`。
|
||||||
|
- `chat-executor-prompt.md`
|
||||||
|
- 明确 `negative_observation` 必须引用 `$.no_evidence`。
|
||||||
|
- 明确没有实际工具调用时禁止使用 `$.no_evidence`。
|
||||||
|
|
||||||
|
### 未修改
|
||||||
|
|
||||||
|
- 不新增数据库表。
|
||||||
|
- 不新增复杂 metadata。
|
||||||
|
- 不改变 Planner。
|
||||||
|
- 不改变 Agent 编排。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 验收标准
|
||||||
|
|
||||||
|
1. `query_logs` 返回 `logs=[] / total=0 / evidence_status=no_evidence` 时,`tool_invocation.retrieval_details.evidence_refs` 包含:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"raw_path": "$.no_evidence"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
2. Executor 输出 `negative_observation` 并引用 `$.no_evidence` 时,Gatekeeper 可以校验通过。
|
||||||
|
|
||||||
|
3. Executor 如果用 `$.no_evidence` 搭配正向证据文本,例如 `HikariCP active=50/50`,Gatekeeper 必须拒绝。
|
||||||
|
|
||||||
|
4. `$.no_evidence` 不得被解释为“问题绝对不存在”,只能表达“当前查询未检索到匹配证据”。
|
||||||
|
|
||||||
|
5. HikariCP negative E2E:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
期望:
|
||||||
|
|
||||||
|
- 工具返回 no-hit。
|
||||||
|
- 不返回 `generic-service`。
|
||||||
|
- Executor 输出 `negative_observation`。
|
||||||
|
- `raw_path="$.no_evidence"`。
|
||||||
|
- Gatekeeper `pass/none`。
|
||||||
|
- Verifier 不误判为正向 HikariCP 证据。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 验证记录
|
||||||
|
|
||||||
|
### 2026-07-08 单元测试
|
||||||
|
|
||||||
|
命令:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvn '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,QueryLogsToolsTest' test
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Tests run: 23, Failures: 0, Errors: 0, Skipped: 0
|
||||||
|
BUILD SUCCESS
|
||||||
|
```
|
||||||
|
|
||||||
|
覆盖点:
|
||||||
|
|
||||||
|
- `query_logs` no-hit 生成 `$.no_evidence`。
|
||||||
|
- `query_metrics` no-hit 生成 `$.no_evidence`。
|
||||||
|
- `lookup_knowledge` no-hit 生成 `$.no_evidence`。
|
||||||
|
- Gatekeeper 可以校验 `$.no_evidence`。
|
||||||
|
- 多个 no-evidence 调用存在时,Gatekeeper 可按 `tool_name + raw_path + evidence_excerpt` 唯一回填 `source_invocation_id`。
|
||||||
|
- `negative_observation` 混绑正向 `$.logs[i]` 会被拒绝。
|
||||||
|
- HikariCP negative mock 不返回 `generic-service`。
|
||||||
|
|
||||||
|
### 2026-07-08 E2E 验证
|
||||||
|
|
||||||
|
输入:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
最终通过 session:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId: iss009-hikari-negative-latest-20260708-232428
|
||||||
|
verdict: PASS
|
||||||
|
groundedness_score: 1.0
|
||||||
|
gatekeeper_result.status: pass
|
||||||
|
gatekeeper_result.severity: none
|
||||||
|
claim_count: 1
|
||||||
|
claim_type: negative_observation
|
||||||
|
raw_path: $.no_evidence
|
||||||
|
generic_service_hit: false
|
||||||
|
overstate_hit: false
|
||||||
|
```
|
||||||
|
|
||||||
|
Executor claim:
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。
|
||||||
|
```
|
||||||
|
|
||||||
|
最终答案:
|
||||||
|
|
||||||
|
```text
|
||||||
|
本次查询在 inventory-service 中未发现 HikariCP 连接池耗尽的日志记录,检索结果未匹配到相关证据。
|
||||||
|
```
|
||||||
|
|
||||||
|
说明:
|
||||||
|
|
||||||
|
- Executor 仍可能输出多个 `$.no_evidence` binding。
|
||||||
|
- 如果 `source_invocation_id` 缺失,Gatekeeper 会按 `tool_name + raw_path + evidence_excerpt` 唯一匹配真实 invocation 并写入 warning。
|
||||||
|
- 最终答案不使用“排除”“确认没有”“不存在该问题”等过度表达。
|
||||||
@@ -9,6 +9,8 @@
|
|||||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
|
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||||
|
| ISS-008 | Executor 窄范围查询越界 | 中 | 已修复 | [ISS-008-executor-narrow-scope-overreach.md](ISS-008-executor-narrow-scope-overreach.md) |
|
||||||
|
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 中 | 已修复 | [ISS-009-negative-observation-no-evidence-reference.md](ISS-009-negative-observation-no-evidence-reference.md) |
|
||||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
||||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
||||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||||
|
|||||||
@@ -156,7 +156,8 @@ public class ExecutorGatekeeperService {
|
|||||||
SEVERITY_LOW_CONFID);
|
SEVERITY_LOW_CONFID);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
validateEvidenceBinding(binding, validInvocations, target, claim.get("claim_id"), result);
|
validateEvidenceBinding(binding, validInvocations, target,
|
||||||
|
claim.get("claim_id"), claim.get("claim_type"), result);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -181,7 +182,7 @@ public class ExecutorGatekeeperService {
|
|||||||
SEVERITY_LOW_CONFID);
|
SEVERITY_LOW_CONFID);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), result);
|
validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), null, result);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -190,6 +191,7 @@ public class ExecutorGatekeeperService {
|
|||||||
Map<Long, ToolInvocation> validInvocations,
|
Map<Long, ToolInvocation> validInvocations,
|
||||||
String target,
|
String target,
|
||||||
Object ownerId,
|
Object ownerId,
|
||||||
|
Object ownerType,
|
||||||
GatekeeperResult result) {
|
GatekeeperResult result) {
|
||||||
Map<String, Object> checked = new LinkedHashMap<>();
|
Map<String, Object> checked = new LinkedHashMap<>();
|
||||||
checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId));
|
checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId));
|
||||||
@@ -197,18 +199,6 @@ public class ExecutorGatekeeperService {
|
|||||||
checked.put("source_invocation_id", binding.get("source_invocation_id"));
|
checked.put("source_invocation_id", binding.get("source_invocation_id"));
|
||||||
checked.put("raw_path", stringValue(binding.get("raw_path")));
|
checked.put("raw_path", stringValue(binding.get("raw_path")));
|
||||||
|
|
||||||
Long id = singleInvocationId(binding);
|
|
||||||
if (id == null) {
|
|
||||||
checked.put("status", STATUS_FAIL);
|
|
||||||
checked.put("rule", RULE_INVOCATION_REF);
|
|
||||||
checked.put("message", "source_invocation_id is required");
|
|
||||||
result.checked(checked);
|
|
||||||
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
|
|
||||||
"source_invocation_id is required", SEVERITY_LOW_CONFID);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
checked.put("source_invocation_id", id);
|
|
||||||
|
|
||||||
String claimedToolName = stringValue(binding.get("tool_name"));
|
String claimedToolName = stringValue(binding.get("tool_name"));
|
||||||
if (claimedToolName.isBlank()) {
|
if (claimedToolName.isBlank()) {
|
||||||
checked.put("status", STATUS_FAIL);
|
checked.put("status", STATUS_FAIL);
|
||||||
@@ -219,6 +209,46 @@ public class ExecutorGatekeeperService {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
String rawPath = stringValue(binding.get("raw_path"));
|
||||||
|
if ("negative_observation".equals(stringValue(ownerType)) && !rawPath.isBlank()
|
||||||
|
&& !"$.no_evidence".equals(rawPath)) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_RAW_PATH);
|
||||||
|
checked.put("message", "negative_observation must only bind $.no_evidence references");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_RAW_PATH, target + ".raw_path",
|
||||||
|
"negative_observation must only bind $.no_evidence references", SEVERITY_REJECT);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
Long id = singleInvocationId(binding);
|
||||||
|
if (id == null) {
|
||||||
|
id = uniqueInvocationIdByToolRawPathAndExcerpt(validInvocations, claimedToolName, rawPath,
|
||||||
|
stringValue(binding.get("evidence_excerpt")));
|
||||||
|
if (id == null) {
|
||||||
|
id = uniqueInvocationIdByToolAndRawPath(validInvocations, claimedToolName, rawPath);
|
||||||
|
}
|
||||||
|
if (id != null) {
|
||||||
|
result.warn(Map.of(
|
||||||
|
"rule", "evidence.invocation_auto_backfill_by_raw_path",
|
||||||
|
"message", "source_invocation_id was auto-filled from the unique evidence reference candidate",
|
||||||
|
"tool_name", claimedToolName,
|
||||||
|
"raw_path", rawPath,
|
||||||
|
"source_invocation_id", id
|
||||||
|
));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (id == null) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_INVOCATION_REF);
|
||||||
|
checked.put("message", "source_invocation_id is required");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
|
||||||
|
"source_invocation_id is required", SEVERITY_LOW_CONFID);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
checked.put("source_invocation_id", id);
|
||||||
|
|
||||||
ToolInvocation invocation = validInvocations.get(id);
|
ToolInvocation invocation = validInvocations.get(id);
|
||||||
if (invocation == null) {
|
if (invocation == null) {
|
||||||
checked.put("status", STATUS_FAIL);
|
checked.put("status", STATUS_FAIL);
|
||||||
@@ -240,7 +270,6 @@ public class ExecutorGatekeeperService {
|
|||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
String rawPath = stringValue(binding.get("raw_path"));
|
|
||||||
if (rawPath.isBlank()) {
|
if (rawPath.isBlank()) {
|
||||||
checked.put("status", STATUS_FAIL);
|
checked.put("status", STATUS_FAIL);
|
||||||
checked.put("rule", RULE_RAW_PATH);
|
checked.put("rule", RULE_RAW_PATH);
|
||||||
@@ -298,6 +327,82 @@ public class ExecutorGatekeeperService {
|
|||||||
result.checked(checked);
|
result.checked(checked);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private Long uniqueInvocationIdByToolAndRawPath(Map<Long, ToolInvocation> validInvocations,
|
||||||
|
String toolName,
|
||||||
|
String rawPath) {
|
||||||
|
if (toolName == null || toolName.isBlank() || rawPath == null || rawPath.isBlank()) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
Long matchedId = null;
|
||||||
|
for (Map.Entry<Long, ToolInvocation> entry : validInvocations.entrySet()) {
|
||||||
|
ToolInvocation invocation = entry.getValue();
|
||||||
|
if (!Objects.equals(toolName, invocation.getToolName())) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (!evidenceRefsByRawPath(invocation.getRetrievalDetails()).containsKey(rawPath)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (matchedId != null) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
matchedId = entry.getKey();
|
||||||
|
}
|
||||||
|
return matchedId;
|
||||||
|
}
|
||||||
|
|
||||||
|
private Long uniqueInvocationIdByToolRawPathAndExcerpt(Map<Long, ToolInvocation> validInvocations,
|
||||||
|
String toolName,
|
||||||
|
String rawPath,
|
||||||
|
String excerpt) {
|
||||||
|
if (toolName == null || toolName.isBlank()
|
||||||
|
|| rawPath == null || rawPath.isBlank()
|
||||||
|
|| excerpt == null || excerpt.isBlank()) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
Long matchedId = null;
|
||||||
|
for (Map.Entry<Long, ToolInvocation> entry : validInvocations.entrySet()) {
|
||||||
|
ToolInvocation invocation = entry.getValue();
|
||||||
|
if (!Objects.equals(toolName, invocation.getToolName())) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
String matchedText = evidenceRefsByRawPath(invocation.getRetrievalDetails()).get(rawPath);
|
||||||
|
if (matchedText == null || !isBackfillCandidateSupported(rawPath, excerpt, matchedText)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (matchedId != null) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
matchedId = entry.getKey();
|
||||||
|
}
|
||||||
|
return matchedId;
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean isBackfillCandidateSupported(String rawPath, String excerpt, String matchedText) {
|
||||||
|
if ("$.no_evidence".equals(rawPath)) {
|
||||||
|
String excerptQuery = semicolonField(excerpt, "query");
|
||||||
|
String matchedQuery = semicolonField(matchedText, "query");
|
||||||
|
if (!excerptQuery.isBlank() && !matchedQuery.isBlank()
|
||||||
|
&& !normalized(excerptQuery).equals(normalized(matchedQuery))) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return isExcerptSupported(excerpt, matchedText);
|
||||||
|
}
|
||||||
|
|
||||||
|
private String semicolonField(String text, String field) {
|
||||||
|
if (text == null || text.isBlank() || field == null || field.isBlank()) {
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
String prefix = field + "=";
|
||||||
|
for (String part : text.split(";")) {
|
||||||
|
String trimmed = part.trim();
|
||||||
|
if (trimmed.regionMatches(true, 0, prefix, 0, prefix.length())) {
|
||||||
|
return trimmed.substring(prefix.length()).trim();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
|
||||||
private Long singleInvocationId(Map<?, ?> binding) {
|
private Long singleInvocationId(Map<?, ?> binding) {
|
||||||
Long singular = asLong(binding.get("source_invocation_id"));
|
Long singular = asLong(binding.get("source_invocation_id"));
|
||||||
if (singular != null) {
|
if (singular != null) {
|
||||||
|
|||||||
@@ -88,7 +88,7 @@ public class ToolInvocationRecorder {
|
|||||||
if (extraDetails != null && !extraDetails.isEmpty()) {
|
if (extraDetails != null && !extraDetails.isEmpty()) {
|
||||||
details.putAll(extraDetails);
|
details.putAll(extraDetails);
|
||||||
}
|
}
|
||||||
List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, output, details);
|
List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, inputParams, output, details);
|
||||||
if (!evidenceRefs.isEmpty()) {
|
if (!evidenceRefs.isEmpty()) {
|
||||||
details.put("evidence_refs", evidenceRefs);
|
details.put("evidence_refs", evidenceRefs);
|
||||||
}
|
}
|
||||||
@@ -145,6 +145,12 @@ public class ToolInvocationRecorder {
|
|||||||
}
|
}
|
||||||
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
|
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
|
||||||
List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks());
|
List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks());
|
||||||
|
if (evidenceRefs.isEmpty()
|
||||||
|
&& EVIDENCE_STATUS_NO_EVIDENCE.equals(normalizeEvidenceStatus(record.success(), record.evidenceStatus()))) {
|
||||||
|
Map<String, Object> input = new LinkedHashMap<>();
|
||||||
|
input.put("query", record.query());
|
||||||
|
evidenceRefs = List.of(noEvidenceRef("lookup_knowledge", input, null, details));
|
||||||
|
}
|
||||||
if (!evidenceRefs.isEmpty()) {
|
if (!evidenceRefs.isEmpty()) {
|
||||||
details.put("evidence_refs", evidenceRefs);
|
details.put("evidence_refs", evidenceRefs);
|
||||||
}
|
}
|
||||||
@@ -195,7 +201,10 @@ public class ToolInvocationRecorder {
|
|||||||
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
|
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
|
||||||
}
|
}
|
||||||
|
|
||||||
private List<Map<String, Object>> extractEvidenceRefs(String toolName, String output, Map<String, Object> details) {
|
private List<Map<String, Object>> extractEvidenceRefs(String toolName,
|
||||||
|
Map<String, Object> inputParams,
|
||||||
|
String output,
|
||||||
|
Map<String, Object> details) {
|
||||||
if (output == null || output.isBlank()) {
|
if (output == null || output.isBlank()) {
|
||||||
return List.of();
|
return List.of();
|
||||||
}
|
}
|
||||||
@@ -205,11 +214,14 @@ public class ToolInvocationRecorder {
|
|||||||
|
|
||||||
try {
|
try {
|
||||||
JsonNode root = objectMapper.readTree(output);
|
JsonNode root = objectMapper.readTree(output);
|
||||||
|
boolean noEvidence = EVIDENCE_STATUS_NO_EVIDENCE.equals(stringValue(details.get("evidence_status")));
|
||||||
if ("query_metrics".equals(toolName)) {
|
if ("query_metrics".equals(toolName)) {
|
||||||
return evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText);
|
List<Map<String, Object>> refs = evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText);
|
||||||
|
return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs;
|
||||||
}
|
}
|
||||||
if ("query_logs".equals(toolName)) {
|
if ("query_logs".equals(toolName)) {
|
||||||
return evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText);
|
List<Map<String, Object>> refs = evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText);
|
||||||
|
return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs;
|
||||||
}
|
}
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
log.debug("extract evidence_refs failed for tool={}", toolName, e);
|
log.debug("extract evidence_refs failed for tool={}", toolName, e);
|
||||||
@@ -217,6 +229,42 @@ public class ToolInvocationRecorder {
|
|||||||
return List.of();
|
return List.of();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private Map<String, Object> noEvidenceRef(String toolName,
|
||||||
|
Map<String, Object> inputParams,
|
||||||
|
JsonNode root,
|
||||||
|
Map<String, Object> details) {
|
||||||
|
List<String> parts = new ArrayList<>();
|
||||||
|
addPart(parts, toolName + " returned no evidence");
|
||||||
|
addPart(parts, "evidence_status=" + stringValue(details.get("evidence_status")));
|
||||||
|
String query = firstNonBlank(
|
||||||
|
root == null ? null : textField(root, "query"),
|
||||||
|
inputParams == null ? null : inputParams.get("query")
|
||||||
|
);
|
||||||
|
if (!query.isBlank()) {
|
||||||
|
addPart(parts, "query=" + query);
|
||||||
|
}
|
||||||
|
String topic = firstNonBlank(
|
||||||
|
root == null ? null : textField(root, "log_topic"),
|
||||||
|
inputParams == null ? null : inputParams.get("log_topic"),
|
||||||
|
details == null ? null : details.get("log_topic"),
|
||||||
|
details == null ? null : details.get("metric_family")
|
||||||
|
);
|
||||||
|
if (!topic.isBlank()) {
|
||||||
|
addPart(parts, "topic=" + topic);
|
||||||
|
}
|
||||||
|
if (root != null && root.has("total")) {
|
||||||
|
addPart(parts, "total=" + root.path("total").asText());
|
||||||
|
}
|
||||||
|
String message = root == null ? "" : textField(root, "message");
|
||||||
|
if (!message.isBlank()) {
|
||||||
|
addPart(parts, "message=" + message);
|
||||||
|
}
|
||||||
|
return Map.of(
|
||||||
|
"raw_path", "$.no_evidence",
|
||||||
|
"text", bounded(String.join("; ", parts), 500)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode,
|
private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode,
|
||||||
String pathPrefix,
|
String pathPrefix,
|
||||||
java.util.function.Function<JsonNode, String> textExtractor) {
|
java.util.function.Function<JsonNode, String> textExtractor) {
|
||||||
@@ -314,6 +362,10 @@ public class ToolInvocationRecorder {
|
|||||||
return "";
|
return "";
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private String stringValue(Object value) {
|
||||||
|
return value == null ? "" : String.valueOf(value);
|
||||||
|
}
|
||||||
|
|
||||||
private String bounded(String value, int limit) {
|
private String bounded(String value, int limit) {
|
||||||
if (value == null) {
|
if (value == null) {
|
||||||
return "";
|
return "";
|
||||||
|
|||||||
@@ -8,6 +8,8 @@
|
|||||||
- 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。
|
- 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。
|
||||||
- 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。
|
- 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。
|
||||||
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。
|
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。
|
||||||
|
- 当 `allowed_claims` 中存在 `claim_type=negative_observation`,或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||||||
|
- 对 `negative_observation` / `$.no_evidence`,禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||||||
|
|
||||||
## 输入字段
|
## 输入字段
|
||||||
|
|
||||||
@@ -26,6 +28,7 @@
|
|||||||
- 可以表达确认结论。
|
- 可以表达确认结论。
|
||||||
- 只能使用 `allowed_claims` 和 `recommended_actions`。
|
- 只能使用 `allowed_claims` 和 `recommended_actions`。
|
||||||
- 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。
|
- 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。
|
||||||
|
- 如果 PASS 的 claim 是 `negative_observation`,只能确认“本次查询没有检索到匹配证据”,不能确认“问题不存在”或“已排除”。
|
||||||
|
|
||||||
### LOW_CONFID
|
### LOW_CONFID
|
||||||
|
|
||||||
|
|||||||
@@ -7,15 +7,114 @@
|
|||||||
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
|
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
|
||||||
- 你是证据收集与微观事实提炼器,不是最终答复生成器。
|
- 你是证据收集与微观事实提炼器,不是最终答复生成器。
|
||||||
|
|
||||||
|
## 角色边界 HARD-GATE
|
||||||
|
|
||||||
|
你只负责证据收集与微观事实提炼,只能输出“当前工具证据可以直接支持的观察事实”。
|
||||||
|
|
||||||
|
你不是:
|
||||||
|
- 根因诊断器。
|
||||||
|
- 修复方案生成器。
|
||||||
|
- Runbook 转述器。
|
||||||
|
- 经验推断器。
|
||||||
|
- 最终用户答复生成器。
|
||||||
|
|
||||||
|
除非本轮工具返回中存在直接证据,否则禁止输出:
|
||||||
|
- 根因确认。
|
||||||
|
- 修复建议。
|
||||||
|
- 扩展排查方向。
|
||||||
|
- 历史经验。
|
||||||
|
- 通用知识。
|
||||||
|
- 与用户问题无关的服务、指标、订单、错误码、组件。
|
||||||
|
|
||||||
## 规则
|
## 规则
|
||||||
- 按顺序执行,不可跳过步骤。
|
- 按顺序执行,不可跳过步骤。
|
||||||
- 所有事实性结论必须来自本轮 evidence tools 的返回。
|
- 所有事实性结论必须来自本轮 evidence tools 的返回。
|
||||||
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。
|
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导,不能直接写成本次事故的已确认事实。
|
||||||
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
|
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
|
||||||
- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。
|
- 不要使用“通常情况下”“根据经验”“很可能已经发生”“可能是”“推测”“理论上”等无证据推断词来伪装事实。
|
||||||
- 对窄范围确认问题,只输出与用户问题直接相关的 observation / negative_observation。通常 1 条 claim,最多 2 条 claim;不要限制 evidence_bindings 数量。
|
|
||||||
- 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。
|
- 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。
|
||||||
|
|
||||||
|
## 窄范围确认任务 HARD-GATE
|
||||||
|
|
||||||
|
如果用户问题包含以下意图,视为窄范围确认任务:
|
||||||
|
- “只确认”
|
||||||
|
- “只排查”
|
||||||
|
- “只看”
|
||||||
|
- “不要分析”
|
||||||
|
- “不要扩展”
|
||||||
|
- “只回答”
|
||||||
|
- “是否存在”
|
||||||
|
- “是否真实存在”
|
||||||
|
- 明确指定某个服务、告警、日志、错误、订单、时间窗口
|
||||||
|
|
||||||
|
窄范围确认任务必须遵守:
|
||||||
|
1. `claims` 只能输出 `observation` 或 `negative_observation`。
|
||||||
|
2. claim 数量必须是最少必要数量,通常 1 条,最多 2 条。
|
||||||
|
3. claim 数量限制不限制 `evidence_bindings` 数量;一条 claim 可以绑定多条直接相关证据。
|
||||||
|
4. 不得把同一观察事实拆成多条 claim。
|
||||||
|
5. 只能围绕用户明确要求的目标对象和主题输出 claim。
|
||||||
|
6. 用户明确排除的对象、服务、告警、订单、数据库、连接池、下游依赖,禁止出现在 claim 中。
|
||||||
|
7. 禁止输出根因类、风险类或建议类 claim,例如 `root_cause`、`risk`、`recommendation`。
|
||||||
|
8. 如果证据不足,不要补合理化解释;优先写入 `missing_info`。
|
||||||
|
9. 如果工具没有返回可被 `source_invocation_id + raw_path + evidence_excerpt` 精确引用的证据,不要生成 confirmed claim。
|
||||||
|
10. 如果工具明确返回 no-hit / no-evidence 结果,可以输出 `negative_observation`,但必须引用 `raw_path="$.no_evidence"`。
|
||||||
|
11. 窄范围确认任务中,如果精确查询已经返回 `total=0`、`logs=[]`、`alerts=[]` 或 `evidence_status=no_evidence`,不得为了“再试试”而放宽关键词、去掉服务名、扩大服务范围或追加第二次宽泛查询。
|
||||||
|
|
||||||
|
窄范围任务的理想输出是:
|
||||||
|
- 1 条核心 claim。
|
||||||
|
- 多条直接相关 `evidence_bindings`。
|
||||||
|
- 必要的 `missing_info`。
|
||||||
|
|
||||||
|
同一条工具数组项只能绑定一次。不要为了引用其中多个字段而拆成多个 `evidence_bindings`。
|
||||||
|
|
||||||
|
正确:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
错误:
|
||||||
|
```json
|
||||||
|
{ "raw_path": "$.alerts[0].alert_name", "evidence_excerpt": "HighCPUUsage" }
|
||||||
|
{ "raw_path": "$.alerts[0].state", "evidence_excerpt": "firing" }
|
||||||
|
```
|
||||||
|
|
||||||
|
负向观察示例:
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"claim_type": "negative_observation",
|
||||||
|
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"source_invocation_id": 123,
|
||||||
|
"raw_path": "$.no_evidence",
|
||||||
|
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”,不能表示“问题不存在”或“根因被排除”。没有实际调用工具时,禁止使用 `$.no_evidence`。
|
||||||
|
|
||||||
|
`negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`。禁止把其它服务的正向日志或告警绑定到同一个 `negative_observation`,即使这些日志可以说明“不是当前服务”。
|
||||||
|
|
||||||
|
输出 `negative_observation` 或基于 `$.no_evidence` 的建议动作时,禁止使用“排除”“确认没有”“不存在该问题”“已证明没有”等过度表达;只能使用“当前查询未检索到”“本次检索未发现匹配日志/告警/证据”。
|
||||||
|
|
||||||
|
## 工具使用边界
|
||||||
|
|
||||||
|
你只能调用回答当前用户问题所必需的工具。
|
||||||
|
|
||||||
|
- 问告警状态:优先使用 `query_metrics`。
|
||||||
|
- 问日志现象:优先使用 `query_logs`。
|
||||||
|
- 问知识解释或排查步骤:才使用 `lookup_knowledge`。
|
||||||
|
- Runbook / Skill / 知识库只能帮助决定“查什么”,不能直接作为“当前环境发生了什么”的证据。
|
||||||
|
- 如果当前工具结果已经足以回答用户问题,不要继续扩展检索。
|
||||||
|
- 不要为了补全故事而查询用户没有要求的服务、组件或故障类型。
|
||||||
|
- 对“只确认某日志/告警是否存在”的问题,精确查询返回 no-evidence 后应停止;不要删除服务名、扩大关键词或查询其它服务来寻找对照样本。
|
||||||
|
|
||||||
## 检索约束
|
## 检索约束
|
||||||
|
|
||||||
### 1. 判断重复:基于已检索上下文
|
### 1. 判断重复:基于已检索上下文
|
||||||
@@ -67,6 +166,19 @@
|
|||||||
### missing_info
|
### missing_info
|
||||||
`missing_info` 用来列出无法确认结论所缺少的具体证据。
|
`missing_info` 用来列出无法确认结论所缺少的具体证据。
|
||||||
|
|
||||||
|
## 输出前自检
|
||||||
|
|
||||||
|
在输出 JSON 前,逐项检查:
|
||||||
|
|
||||||
|
1. 每条 claim 是否直接回答了用户当前问题?
|
||||||
|
2. 每条 claim 是否都有真实 `evidence_bindings`?
|
||||||
|
3. 每个 `evidence_excerpt` 是否来自工具返回原文?
|
||||||
|
4. 是否出现了用户没有要求的服务、告警、订单、数据库、连接池或下游组件?
|
||||||
|
5. 是否把 Runbook / Skill / 知识库通用内容写成了当前事实?
|
||||||
|
6. 是否输出了根因、修复动作、风险判断或经验推断?
|
||||||
|
|
||||||
|
只要任一项不通过,删除对应 claim,不要解释。
|
||||||
|
|
||||||
## 最终输出格式(严格契约)
|
## 最终输出格式(严格契约)
|
||||||
|
|
||||||
你必须输出且只能输出一个 JSON 对象,不要输出 Markdown,不要输出代码块,不要输出 JSON 之外的解释文字。
|
你必须输出且只能输出一个 JSON 对象,不要输出 Markdown,不要输出代码块,不要输出 JSON 之外的解释文字。
|
||||||
@@ -87,7 +199,7 @@
|
|||||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空",
|
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空",
|
||||||
"tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool",
|
"tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool",
|
||||||
"source_invocation_id": null,
|
"source_invocation_id": null,
|
||||||
"raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0]",
|
"raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0] / $.no_evidence",
|
||||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -121,6 +233,8 @@
|
|||||||
- `claims[*].evidence_bindings` 不能为空。
|
- `claims[*].evidence_bindings` 不能为空。
|
||||||
- `evidence_excerpt` 必须来自工具返回,不允许编造。
|
- `evidence_excerpt` 必须来自工具返回,不允许编造。
|
||||||
- `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。
|
- `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。
|
||||||
|
- 当且仅当工具明确返回 no-hit / no-evidence 结果时,允许使用 `$.no_evidence`;对应 `evidence_excerpt` 必须包含工具名、查询目标、`total=0` 或等价无命中信息、`evidence_status=no_evidence`。
|
||||||
|
- `raw_path` 禁止指向字段级子路径,例如 `$.alerts[0].alert_name`、`$.alerts[0].state`、`$.logs[0].message` 都是非法路径。需要引用多个字段时,仍然只使用对应数组条目的 `raw_path`,并把必要字段合并进同一个 `evidence_excerpt`。
|
||||||
- `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。
|
- `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。
|
||||||
- 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。
|
- 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。
|
||||||
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
|
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
|
||||||
|
|||||||
@@ -33,6 +33,111 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validatePassesForNoEvidenceReference() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("pass", result.get("status"));
|
||||||
|
assertEquals("none", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateBackfillsNoEvidenceInvocationByRawPathWhenToolHasMultipleCalls() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志"),
|
||||||
|
invocation(102L, "query_logs", "$.logs[0]",
|
||||||
|
"order-service HikariCP active=50/50 waiting=32")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(null, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence",
|
||||||
|
"negative_observation"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("pass", result.get("status"));
|
||||||
|
assertEquals("none", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
||||||
|
assertEquals("evidence.invocation_auto_backfill_by_raw_path",
|
||||||
|
((Map<?, ?>) ((List<?>) result.get("warnings")).get(0)).get("rule"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateBackfillsNoEvidenceInvocationByExcerptWhenRawPathIsRepeated() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; evidence_status=no_evidence; query=service:inventory-service AND HikariCP; total=0; message=未找到匹配的日志"),
|
||||||
|
invocation(102L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service; total=0; message=未找到匹配的日志")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(null, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; query=service:inventory-service AND HikariCP; total=0; evidence_status=no_evidence",
|
||||||
|
"negative_observation"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("pass", result.get("status"));
|
||||||
|
assertEquals("none", result.get("severity"));
|
||||||
|
assertEquals(101L,
|
||||||
|
((Map<?, ?>) ((List<?>) result.get("checked_bindings")).get(0)).get("source_invocation_id"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateRejectsNoEvidenceReferenceWhenExcerptClaimsPositiveEvidence() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_logs", "$.no_evidence",
|
||||||
|
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_logs", "$.no_evidence",
|
||||||
|
"HikariCP active=50/50 waiting=32"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.excerpt_mismatch"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateRejectsPositiveBindingOnNegativeObservation() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_logs", "$.logs[0]",
|
||||||
|
"order-service HikariCP active=50/50 waiting=32")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_logs", "$.logs[0]",
|
||||||
|
"order-service HikariCP active=50/50 waiting=32",
|
||||||
|
"negative_observation"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void validateFailsWhenRemovedFieldsArePresent() {
|
void validateFailsWhenRemovedFieldsArePresent() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
@@ -191,11 +296,21 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
}
|
}
|
||||||
|
|
||||||
private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) {
|
private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) {
|
||||||
|
return validOutput(invocationId, toolName, rawPath, excerpt, "symptom");
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, Object> validOutput(Long invocationId,
|
||||||
|
String toolName,
|
||||||
|
String rawPath,
|
||||||
|
String excerpt,
|
||||||
|
String claimType) {
|
||||||
Map<String, Object> binding = new java.util.LinkedHashMap<>();
|
Map<String, Object> binding = new java.util.LinkedHashMap<>();
|
||||||
binding.put("source_type", "tool_trace");
|
binding.put("source_type", "tool_trace");
|
||||||
binding.put("source_id", "trace-1");
|
binding.put("source_id", "trace-1");
|
||||||
binding.put("tool_name", toolName);
|
binding.put("tool_name", toolName);
|
||||||
|
if (invocationId != null) {
|
||||||
binding.put("source_invocation_id", invocationId);
|
binding.put("source_invocation_id", invocationId);
|
||||||
|
}
|
||||||
if (rawPath != null) {
|
if (rawPath != null) {
|
||||||
binding.put("raw_path", rawPath);
|
binding.put("raw_path", rawPath);
|
||||||
}
|
}
|
||||||
@@ -204,7 +319,7 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
"answer_version", "executor_evidence_v2",
|
"answer_version", "executor_evidence_v2",
|
||||||
"claims", List.of(Map.of(
|
"claims", List.of(Map.of(
|
||||||
"claim_id", "claim-1",
|
"claim_id", "claim-1",
|
||||||
"claim_type", "symptom",
|
"claim_type", claimType,
|
||||||
"claim_text", "连接池 active 达到上限",
|
"claim_text", "连接池 active 达到上限",
|
||||||
"support_level", "direct",
|
"support_level", "direct",
|
||||||
"evidence_bindings", List.of(binding)
|
"evidence_bindings", List.of(binding)
|
||||||
|
|||||||
@@ -27,7 +27,7 @@ class ToolInvocationRecorderTest {
|
|||||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void recordEvidenceToolPreservesNoEvidenceSemantics() {
|
void recordEvidenceToolPreservesNoEvidenceSemantics() throws Exception {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||||
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||||
@@ -36,8 +36,8 @@ class ToolInvocationRecorderTest {
|
|||||||
try {
|
try {
|
||||||
recorder.recordEvidenceTool(
|
recorder.recordEvidenceTool(
|
||||||
"query_logs",
|
"query_logs",
|
||||||
Map.of("query", "timeout"),
|
Map.of("query", "inventory-service HikariCP", "log_topic", "application-logs"),
|
||||||
"{\"success\":false,\"message\":\"未找到匹配的日志\"}",
|
"{\"success\":false,\"query\":\"inventory-service HikariCP\",\"log_topic\":\"application-logs\",\"logs\":[],\"total\":0,\"message\":\"未找到匹配的日志\"}",
|
||||||
true,
|
true,
|
||||||
System.currentTimeMillis() - 10,
|
System.currentTimeMillis() - 10,
|
||||||
null,
|
null,
|
||||||
@@ -57,6 +57,14 @@ class ToolInvocationRecorderTest {
|
|||||||
assertEquals(Boolean.TRUE, saved.getSuccess());
|
assertEquals(Boolean.TRUE, saved.getSuccess());
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"evidence_status\":\"no_evidence\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"evidence_status\":\"no_evidence\""));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
|
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
|
||||||
|
|
||||||
|
JsonNode details = objectMapper.readTree(saved.getRetrievalDetails());
|
||||||
|
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
|
||||||
|
String text = details.path("evidence_refs").get(0).path("text").asText();
|
||||||
|
assertTrue(text.contains("query_logs returned no evidence"));
|
||||||
|
assertTrue(text.contains("inventory-service HikariCP"));
|
||||||
|
assertTrue(text.contains("total=0"));
|
||||||
|
assertTrue(text.contains("evidence_status=no_evidence"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -125,6 +133,40 @@ class ToolInvocationRecorderTest {
|
|||||||
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage"));
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void recordEvidenceToolExtractsMetricNoEvidenceRef() throws Exception {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||||
|
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||||
|
SessionContextHolder.setSessionId("metric-no-evidence-session");
|
||||||
|
|
||||||
|
try {
|
||||||
|
recorder.recordEvidenceTool(
|
||||||
|
"query_metrics",
|
||||||
|
Map.of("query", "active_prometheus_alerts"),
|
||||||
|
"""
|
||||||
|
{"success":true,"alerts":[],"message":"成功检索到 0 个活动告警"}
|
||||||
|
""",
|
||||||
|
true,
|
||||||
|
System.currentTimeMillis() - 10,
|
||||||
|
null,
|
||||||
|
"prometheus_alerts",
|
||||||
|
ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE,
|
||||||
|
Map.of("metric_family", "prometheus_alerts")
|
||||||
|
);
|
||||||
|
} finally {
|
||||||
|
SessionContextHolder.clear();
|
||||||
|
}
|
||||||
|
|
||||||
|
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
|
||||||
|
verify(repository).save(captor.capture());
|
||||||
|
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
|
||||||
|
|
||||||
|
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("query_metrics returned no evidence"));
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("prometheus_alerts"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void recordLookupKnowledgePreservesRetrievalSpecificFields() {
|
void recordLookupKnowledgePreservesRetrievalSpecificFields() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
@@ -190,6 +232,38 @@ class ToolInvocationRecorderTest {
|
|||||||
assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\""));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void recordLookupKnowledgeAddsNoEvidenceRefWhenNoBlocks() throws Exception {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||||
|
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||||
|
SessionContextHolder.setSessionId("lookup-no-evidence-session");
|
||||||
|
|
||||||
|
ToolInvocationRecorder.LookupKnowledgeRecord record = ToolInvocationRecorder.LookupKnowledgeRecord.builder()
|
||||||
|
.query("inventory-service HikariCP")
|
||||||
|
.outputPreview("")
|
||||||
|
.outputLength(0)
|
||||||
|
.durationMs(12)
|
||||||
|
.success(true)
|
||||||
|
.evidenceStatus(ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE)
|
||||||
|
.evidenceBlocks(List.of())
|
||||||
|
.build();
|
||||||
|
|
||||||
|
try {
|
||||||
|
recorder.recordLookupKnowledge(record);
|
||||||
|
} finally {
|
||||||
|
SessionContextHolder.clear();
|
||||||
|
}
|
||||||
|
|
||||||
|
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
|
||||||
|
verify(repository).save(captor.capture());
|
||||||
|
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
|
||||||
|
|
||||||
|
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("lookup_knowledge returned no evidence"));
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("inventory-service HikariCP"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void lookupKnowledgeRecordFromSummarizesEvidenceBlocks() {
|
void lookupKnowledgeRecordFromSummarizesEvidenceBlocks() {
|
||||||
EvidenceBlock block = EvidenceBlock.builder()
|
EvidenceBlock block = EvidenceBlock.builder()
|
||||||
|
|||||||
Reference in New Issue
Block a user