feat(agent): add executor gatekeeper hook

This commit is contained in:
aruo
2026-07-08 02:01:49 +08:00
parent 050cbc8fee
commit c5e496e715
23 changed files with 1343 additions and 30 deletions
+1
View File
@@ -7,6 +7,7 @@
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
@@ -0,0 +1,48 @@
# Acceptance: executor-gatekeeper-hook
## Implementation Result
Completed stage two of Executor Structured Output V2.
- Added `ExecutorGatekeeperService`.
- Added initial `schema.executor_v2` and `evidence.invocation_ref` rules.
- Added `gatekeeper_result` to Verifier payload.
- Stored `gatekeeper_result` in `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated Verifier prompt so Gatekeeper fail must not produce PASS.
- Added focused tests for schema failure, valid pass, fabricated invocation ids, tool name mismatch, hook payload, and persistence.
## Static Verification
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
- Coverage: OpenSpec change validity.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests for Gatekeeper service, hook integration, and ChatService persistence.
## Browser / Manual Verification
Not run. This stage changes backend validation and audit behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
- Excerpt similarity or phrase/utilization rules.
Reason: This phase intentionally covers deterministic schema and invocation-reference validation. Full live verification is better after Verifier V2 and Composer are implemented.
## Remaining Work
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer final-answer generation.
- Phase five: eval fixtures and full audit closure.
## Archive Status
Devflow archive files created for stage two. OpenSpec archive is expected before moving to stage three.
@@ -0,0 +1,44 @@
# Brief: executor-gatekeeper-hook
## Background
Stage one of Executor Structured Output V2 changed Chat Executor output to `executor_evidence_v2`, removing final-expression fields from Executor. That made the output structured, but it did not yet prevent deterministic evidence attribution failures such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names.
## Goal
Add a deterministic Gatekeeper between Executor output parsing and Verifier model execution.
The Gatekeeper should:
- Validate the initial Executor V2 schema.
- Validate `claims[].evidence_bindings[].source_invocation_ids` against current-session `tool_invocation` rows.
- Validate evidence binding `tool_name` against the persisted invocation tool name.
- Expose a small `gatekeeper_result` to Verifier and audit persistence.
## Scope
Included:
- New Gatekeeper validation service.
- `schema.executor_v2` initial rule.
- `evidence.invocation_ref` initial rule.
- `VerifierInputHook` payload integration.
- `VerifierContextHolder` storage.
- `ChatService` verifier evaluation persistence.
- Minimal verifier prompt update.
- Focused tests for Gatekeeper, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Excluded:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## OpenSpec
- Change: `openspec/changes/executor-gatekeeper-hook`
- Parent stage: `openspec/changes/archive/2026-07-07-executor-v2-output-contract`
@@ -0,0 +1,51 @@
# Decisions: executor-gatekeeper-hook
## Key Decisions
### Gatekeeper stays in VerifierInputHook
Decision: Gatekeeper is integrated inside `VerifierInputHook`, after Executor output parsing and before Verifier model execution.
Reason: The user explicitly chose to keep this version in the Verifier hook and not move validation into Executor hook. This preserves the current workflow orchestration.
### No retry in this phase
Decision: Gatekeeper failure does not trigger automatic Executor retry.
Reason: Retry behavior is intentionally deferred. This phase only validates, exposes, and audits deterministic failures.
### Initial rule set is intentionally small
Decision: Stage two implements only `schema.executor_v2` and `evidence.invocation_ref` as hard checks.
Reason: These rules catch the highest-confidence physical failures with low implementation risk. Excerpt similarity, hallucination phrases, and evidence utilization remain later enhancements.
### No new database schema
Decision: Persist `gatekeeper_result` in existing `diagnosis_session.self_evaluation.verifier_evaluation`.
Reason: The user asked to keep database fields minimal. Existing JSON audit storage is enough for this phase.
### Internal interface impact
Decision: This is an L2 internal interface extension.
Impact:
- Verifier payload gains `gatekeeper_result`.
- `VerifierContextHolder` gains Gatekeeper result storage.
- `self_evaluation.verifier_evaluation` gains `gatekeeper_result`.
- No external API, DTO, database table, or schema migration changes.
## Deferred Decisions
- Whether Gatekeeper should later trigger Executor retry.
- Whether `evidence.excerpt_similarity` should be hard fail or warn-only.
- Whether hallucination phrase and evidence utilization rules should be config-driven from metadata files.
- How Verifier V2 `claim_checks` should enforce Gatekeeper failures in code, beyond prompt instruction.
## Remaining Risks
- Verifier prompt compliance is not a deterministic guarantee; stage three should make Gatekeeper fail incompatible with PASS in Verifier V2 behavior.
- Excerpt authenticity is not checked in this phase, so real invocation ids can still be paired with misleading excerpt text until a later rule is implemented.
@@ -0,0 +1,54 @@
# Evidence: executor-gatekeeper-hook
## Context Evidence
- `executor-v2-output-contract` established `executor_evidence_v2` and removed Executor final-expression fields.
- `VerifierInputHook` is the existing integration point for explicit Verifier payload construction.
- `ChatService.persistVerifierEvaluation(...)` is the existing persistence path for verifier audit snapshots.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` provides the current-session invocation pool used by Gatekeeper.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- Implements `schema.executor_v2`.
- Implements `evidence.invocation_ref`.
- Returns `status`, `failed_rules`, `warnings`, and `errors`.
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Runs Gatekeeper after parsing Executor output and building trace summary.
- Adds `gatekeeper_result` to Verifier payload.
- Stores `gatekeeper_result` in `VerifierContextHolder`.
- `src/main/java/com/superbiz/agent/util/VerifierContextHolder.java`
- Stores per-request Gatekeeper result for later persistence.
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Wires `ExecutorGatekeeperService` into verifier hook construction.
- Persists `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- `src/main/resources/prompts/chat-verifier-prompt.md`
- Documents `gatekeeper_result` as an input.
- States Gatekeeper fail must not produce PASS.
## Test Evidence
- `src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java`
- Covers schema failure and valid pass behavior.
- Covers fabricated invocation ids and tool name mismatch.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`
- Covers Verifier payload containing `gatekeeper_result`.
- Covers hook behavior for fabricated invocation ids.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- Covers persistence of `gatekeeper_result` into verifier evaluation.
## Validation Evidence
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests.
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
+247 -28
View File
@@ -1,9 +1,10 @@
# Executor Structured Output V2 实施设计
# Executor Structured Output V2 实施 Issue
**状态**:待规划
**状态**:分阶段实施中
**严重程度**:高
**日期**:2026-07-07
**文档类型**:实施设计文档
**创建日期**:2026-07-07
**最后更新**:2026-07-08
**文档类型**:可执行 issue / 分阶段实施说明
**范围**:Chat Executor 输出结构、Gatekeeper、Verifier、Composer 数据契约调整
**关联问题**:
@@ -14,13 +15,15 @@
---
## 0. 给实现 Agent 的阅读入口
## 0. 给后续实现 Agent 的执行入口
这份文档用于指导 `Chat` 复杂诊断链路的一次增量改造。实现时不要先从字段表开始,而应按以下顺序阅读:
这份文档不是单纯的数据结构草案,而是 `Executor Structured Output V2` 的分阶段实施 issue。后续 agent 接手时,应先理解问题背景和阶段边界,再进入具体字段定义。
1. 先读 `0.1 - 0.7`,理解为什么改、当前链路是什么、目标链路是什么、哪些地方不能改。
2. 再读 `20 - 25`,理解代码影响面、实施阶段、回滚策略和完成标准。
3. 最后按需查阅 `2 - 19` 的详细数据契约、Gatekeeper 规则、Verifier/Composer 输入输出。
推荐阅读顺序:
1. 先读 `0.1 - 0.10`:理解为什么要改、当前问题在哪里、目标链路是什么、哪些事情明确不做,以及如何按 OpenSpec/devflow 分阶段落地。
2. 再读 `20 - 25`:理解代码影响面、阶段拆分、每阶段验收、回滚策略和最终完成标准。
3. 最后按需查阅 `2 - 19`:这些是实现时使用的详细数据契约、Gatekeeper 规则、Verifier/Composer 输入输出。
一句话目标:
@@ -31,18 +34,36 @@
最终用户答案交给 Composer 生成。
```
执行约束:
- 必须按阶段实施,不允许把五个阶段揉成一个大改动。
- 每个阶段都要先有 OpenSpec change,再实现、验证、归档到 devflow,并单独提交。
- 阶段未归档、未提交前,不进入下一阶段。
- 验收失败如果是代码问题,agent 自行修复;如果是设计决策不明确,停下来问用户。
- 本 issue 不要求一次性完成所有阶段;后续 agent 应从当前 git/OpenSpec/devflow 状态继续推进。
---
## 0.1 背景
当前 Chat 复杂诊断链路中,Executor 的输出既包含结构化证据归因,也包含最终面向用户的自然语言答案:
近期 Chat 复杂诊断链路中,多条诊断会话被 Verifier 判为 `LOW_CONFID`。这些会话并不是没有调用工具;Executor 通常已经调用了 `lookup_knowledge`、`query_metrics`、`query_logs` 等 evidence tools。
真正的问题是:Executor 在综合输出阶段把三类内容混在一起写成“当前事故结论”:
1. 本轮工具真实返回的事实。
2. runbook、skill 或知识库里的通用模式。
3. 模型基于经验补全的推断。
当前 Executor 输出中还包含最终面向用户的自然语言字段:
```text
diagnosis_summary
user_facing_answer
```
这导致一个核心问题:Executor 在证据还不充分时,容易提前把“可能方向”写成“确认结论”。后续 Verifier 虽然会校验,但它面对的是自然语言答案和结构化字段混在一起的输出,容易出现以下风险:
这会诱导 Executor 提前进入“诊断报告 / 用户表达”模式,在证据不充分时把“可能方向”写成“确认结论”。后续 Verifier 虽然会拦截,但它面对的是自然语言答案和结构化字段混在一起的输出,校验边界不稳定。
典型风险包括:
- Executor 编造不存在的 `source_invocation_ids`。
- Executor 使用真实 invocation id,但 `evidence_excerpt` 与真实工具输出不一致。
@@ -50,11 +71,38 @@ user_facing_answer
- Executor 在 `user_facing_answer` 里夹带 claims 中没有的根因、错误码、指标值或修复建议。
- Verifier 被迫从自然语言里逐字抽事实,校验边界不稳定。
因此,本次改造不是为了增加 Agent 数量,而是为了收紧职责边界和证据链路。
因此,本次改造不是为了“多加几个 Agent 显得完整”,而是为了收紧证据归因链路:Executor 只产出可校验材料,Verifier 只判断材料是否被证据支撑,最终表达交给 Composer。
---
## 0.2 当前实现
## 0.2 当前问题定义
本 issue 要解决的是:
```text
Executor 证据归因幻觉
```
更具体地说,Executor 已经调用了工具,但在最终输出时没有严格区分:
| 类型 | 定义 | 应进入哪里 |
|---|---|---|
| direct evidence | 本轮工具直接观测到的事实 | `claims` |
| reasonable inference | 能由证据合理推出但不是逐字出现的判断 | `claims`,但需要标注 `indirect`,并由 Verifier 判断是否说重 |
| hypothesis | 值得排查但尚未被当前证据确认的方向 | `hypotheses` |
| missing evidence | 还缺少哪些证据才能确认 | `missing_info` |
| user expression | 面向用户的自然语言答案 | Composer 输出,不由 Executor 输出 |
如果不拆开这些边界,Verifier 会长期承担两个不同任务:
1. 从自然语言中抽事实。
2. 判断事实是否有证据支撑。
这会让 Verifier 的职责过重,也让最终答案容易夹带未经校验的内容。
---
## 0.3 当前实现
当前代码中的复杂 Chat 链路是三 Agent 顺序执行:
@@ -90,7 +138,7 @@ Verifier 既要校验 structured output,又要扫描自然语言答案
---
## 0.3 目标设计
## 0.4 目标设计
目标链路仍保持单条顺序链路,不新增 Controller,不改 Planner:
@@ -124,23 +172,42 @@ Executor raw JSON
-> Composer 生成 user_facing_answer
```
核心原则:
- 单智能体优先:不为了抽象而新增 Controller;只有已经明确的职责边界才拆 Agent。
- Planner 不改:本 issue 的问题不在计划拆解,而在 Executor 输出和 Verifier 校验边界。
- Gatekeeper 不做诊断:只做确定性校验,拦截伪造 ID、字段退化、明显张冠李戴。
- Verifier 不再逐字扫描最终答案:主校验对象是结构化 `claims`。
- Composer 不补事实:只把 Verifier 允许的材料组织成中文答案。
---
## 0.4 分阶段实施总览
## 0.5 分阶段实施总览
| 阶段 | 目标 | 主要改动 | 验收重点 |
|---|---|---|---|
| 阶段一 | Executor V2 输出契约 | 改 `chat-executor-prompt.md`,移除 `diagnosis_summary` / `user_facing_answer` | Executor 只输出结构化诊断材料 |
| 阶段二 | Gatekeeper 接入 | 在 `VerifierInputHook` 中加入 Gatekeeper,写入 payload 和审计 | 伪造 id、工具名不匹配、schema 退化可被拦截 |
| 阶段三 | Verifier V2 | 改 `chat-verifier-prompt.md`,从 `facts_checked` 转向 `claim_checks` | Verifier 判断可推导性,不再逐字抽自然语言事实 |
| 阶段四 | Composer | 新增 Composer prompt/调用,由 ChatService 过滤输入 | 最终答案只使用 Verifier 允许的材料 |
| 阶段五 | 回归与审计 | 扩测试和 eval fixtures | 证明幻觉拦截、生效路径、审计回溯都可验证 |
| 阶段 | 名称 | 目标 | 状态 | 退出条件 |
|---|---|---|---|---|
| 1 | Executor V2 输出契约 | 移除 Executor 最终表达字段,只输出结构化诊断材料 | 已完成并归档 | `executor_evidence_v2` 可运行,最终用户不再看到 raw JSON |
| 2 | Gatekeeper 接入 VerifierInputHook | 在 Verifier 前加入确定性拦截与审计 | 进行中 | `gatekeeper_result` 进入 Verifier payload 和 `self_evaluation` |
| 3 | Verifier V2 可推导性校验 | 从 `facts_checked` 转向 `claim_checks`,判断 claims 是否可由证据推出 | 未开始 | Verifier 不再从自然语言答案抽取额外事实 |
| 4 | Composer 最终表达 | 由 Composer 根据 Verifier 允许材料生成最终用户答案 | 未开始 | PASS/LOW_CONFID/REJECT 都不读取 Executor `user_facing_answer` |
| 5 | 回归评测与审计闭环 | 用测试和 eval fixtures 证明幻觉拦截链路有效 | 未开始 | 伪造 ID、张冠李戴、hypothesis 写成事实等场景都有回归覆盖 |
更详细的实施拆解见 `21. 实施阶段`。
---
## 0.5 非目标
## 0.6 范围与非目标
### In scope
- 调整 Executor 输出契约,从 `executor_evidence_v1` 演进到 `executor_evidence_v2`。
- 在 `VerifierInputHook` 中接入 Gatekeeper。
- 将 Gatekeeper 结果写入 Verifier payload 和 `diagnosis_session.self_evaluation`。
- 将 Verifier 主输出从 `facts_checked` 迁移到 `claim_checks`,兼容期保留旧字段。
- 新增 Composer 表达层,最终用户答案只来自 Verifier 允许材料。
- 补齐关键回归测试和 eval fixtures。
### Out of scope
本 issue 第一版明确不做以下事情:
@@ -154,7 +221,7 @@ Executor raw JSON
---
## 0.6 关键设计决策
## 0.7 关键设计决策
| 决策 | 结论 | 原因 |
|---|---|---|
@@ -168,7 +235,7 @@ Executor raw JSON
---
## 0.7 实现约束
## 0.8 实现约束
实现时必须遵守:
@@ -181,7 +248,40 @@ Executor raw JSON
---
## 0.8 推荐实现切片
## 0.9 OpenSpec / devflow 落地要求
这个 issue 后续按 OpenSpec-first 的方式实施。每个阶段都必须留下可审计痕迹:
```text
OpenSpec proposal/design/spec/tasks
-> 实现代码
-> 最小验证
-> archive 到 openspec/changes/archive
-> 回填 devflow/projects
-> git commit
```
阶段交付物:
| 交付物 | 要求 |
|---|---|
| OpenSpec change | 每个阶段一个独立 change,不能复用上个阶段的 active change |
| tests | 至少覆盖该阶段的新增失败路径和主成功路径 |
| devflow | 记录 brief、decisions、evidence、acceptance |
| commit | 每个阶段单独提交,提交前确认 diff 只包含本阶段内容 |
推荐验证命令:
```powershell
mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate --specs
```
如阶段新增专门测试,应把测试类加入 Maven `-Dtest` 列表。
---
## 0.10 推荐实现切片
建议不要把所有改动塞进一个大 PR。推荐拆成以下可独立验收的切片:
@@ -1305,8 +1405,18 @@ Composer 输出严格 JSON。
## 21. 实施阶段
每个阶段都必须作为独立 OpenSpec change 落地。后续 agent 不应直接从本节复制代码实现,而应把本节转成该阶段的 `proposal.md`、`design.md`、`spec.md` 和 `tasks.md`。
### 阶段一:Executor V2 输出契约
状态:已完成并归档。
已完成记录:
- OpenSpec archive:`openspec/changes/archive/2026-07-07-executor-v2-output-contract`
- devflow:`devflow/projects/2026-07-07-executor-v2-output-contract`
- commit:`050cbc8 feat(agent): add executor evidence v2 contract`
目标:先让 Executor 不再输出最终诊断话术,只输出结构化诊断材料。
改动范围:
@@ -1323,17 +1433,40 @@ Composer 输出严格 JSON。
- `claims[].evidence_bindings` 仍要求非空。
- 现有流程即使尚未接入 Composer,也不会把 Executor raw JSON 直接当最终答案泄露给用户。
已知阶段性债务:
- `ChatService` 仍有临时 V2 renderer,用于 Composer 上线前避免 raw JSON 外泄。
- Verifier 仍以 `facts_checked` 为主,`claim_checks` 等待阶段三。
### 阶段二:Gatekeeper 接入 VerifierInputHook
状态:进行中。
当前 OpenSpec change:
```text
openspec/changes/executor-gatekeeper-hook
```
目标:在 Verifier 之前用确定性规则拦截物理级幻觉。
改动范围:
- 新增 Gatekeeper 规则接口和校验服务。
- 新增规则索引与元数据配置。
- 新增 Gatekeeper 校验服务,第一版先覆盖必要规则。
- 在 `VerifierInputHook` 中调用 Gatekeeper。
- 将 `gatekeeper_result` 写入 Verifier payload。
- 将 `gatekeeper_result` 持久化到 `self_evaluation.verifier_evaluation`。
- 更新 `chat-verifier-prompt.md`,明确 Gatekeeper fail 时不得 PASS。
第一版规则范围:
| 规则 | 必须实现 | 说明 |
|---|---:|---|
| `schema.executor_v2` | 是 | 校验 V2 基础字段、移除字段、数组字段 |
| `evidence.invocation_ref` | 是 | 校验 invocation id 存在于当前 session 且工具名匹配 |
| `evidence.excerpt_similarity` | 否 | 可作为后续阶段或 warn-only 增强 |
| `claim.hallucination_phrases` | 否 | 可作为后续配置化规则增强 |
| `claim.evidence_utilization` | 否 | 第一版不建议硬启用,避免中文分词误杀 |
验收标准:
@@ -1342,9 +1475,32 @@ Composer 输出严格 JSON。
- `tool_name` 与真实 invocation 不匹配时,Gatekeeper fail。
- 缺省 `hypotheses`、`recommended_actions`、`missing_info` 时,传给 Verifier 前会 normalize 为 `[]`。
- 旧的 `VerifierInputHook` 结构校验不再和 Gatekeeper 重复维护。
- `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` 可回溯到本轮校验结果。
建议测试:
- `ExecutorGatekeeperServiceTest`
- `VerifierInputHookTest`
- `ChatServiceSequentialAgentTest`
建议验证命令:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-gatekeeper-hook
cmd /c openspec validate --specs
```
阶段退出条件:
- OpenSpec change 已 archive。
- devflow 已回填 brief、evidence、decisions、acceptance。
- 本阶段代码和文档已单独 commit。
### 阶段三:Verifier V2 可推导性校验
状态:未开始。
目标:Verifier 不再逐字扫描自然语言,而是校验 claim 是否能由证据合理推出。
改动范围:
@@ -1355,6 +1511,13 @@ Composer 输出严格 JSON。
- 兼容期继续输出或由代码生成 `facts_checked`。
- `parseVerifierDecision(...)` 能处理 `claim_checks` 和旧 `facts_checked`。
核心设计:
- Verifier 主校验对象是 `executor_structured_output.claims`。
- `executor_raw_output` / `executor_final_answer` 只作为 debug/fallback 字段,不作为新增事实来源。
- Verifier 不要求 claim 与证据逐字一致,而是判断是否可以合理推出。
- Gatekeeper fail 时,Verifier 不允许输出 `PASS`。
验收标准:
- `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted` 均有测试覆盖。
@@ -1362,8 +1525,23 @@ Composer 输出严格 JSON。
- malformed/missing Executor 输出不回退自然语言抽事实,直接进入 `LOW_CONFID`。
- `facts_checked` 兼容字段能继续支撑现有低置信模板、retry_context 和评测用例。
建议测试重点:
- `parseVerifierDecision(...)` 能解析 `claim_checks`。
- `facts_checked` 能从 `claim_checks` 兼容映射。
- `gatekeeper_result.status=fail` + Verifier 返回 PASS 时,代码侧应降级或测试 prompt 禁止该行为。
- `unsupported` / `overstated` 能进入低置信模板需要的 missing evidence 语义。
阶段退出条件:
- Verifier prompt 已不再要求逐字扫描最终自然语言答案。
- 审计中同时可见 `claim_checks` 和兼容 `facts_checked`。
- 现有低置信重试链路不回归。
### 阶段四:Composer 输出最终答案
状态:未开始。
目标:把最终用户表达从 Executor 中移出,由 Composer 基于 Verifier 允许的材料生成。
改动范围:
@@ -1373,6 +1551,14 @@ Composer 输出严格 JSON。
- Composer 只接收 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。
- PASS/LOW_CONFID/REJECT 均不再读取 Executor 的 `user_facing_answer`。
核心设计:
- Composer 是表达层,不是诊断层。
- Composer 不接收 raw tool output。
- Composer 不接收未经筛选的完整 Executor output。
- ChatService 负责根据 `claim_checks` 过滤出 `allowed_claims`。
- `overstated` claim 不得作为确认事实输出,可降级为 hypothesis 或 missing_info。
验收标准:
- Composer 输出严格 JSON,包含 `answer_summary`、`recommended_actions`、`user_facing_answer`。
@@ -1380,9 +1566,26 @@ Composer 输出严格 JSON。
- `REJECT` 时 `allowed_hypotheses=[]`,最终答案不出现根因结论。
- `LOW_CONFID` 时必须区分已确认信息和可能方向。
- `PASS` 时只有存在 root cause 类型 allowed claim,才允许表达“根因已确认”。
- `ChatService` 不再依赖临时 V2 renderer 生成 PASS 用户答案。
建议测试重点:
- Composer 输入不包含 raw tool output。
- 未通过 Verifier 的 claim 不进入最终答案。
- `LOW_CONFID` 能输出已确认信息、缺口、下一步建议。
- `REJECT` 不输出 root cause 结论。
- Composer malformed 输出时有降级策略,不向用户泄露 raw JSON。
阶段退出条件:
- 最终用户答案唯一来源是 Composer 或固定降级模板。
- Executor 的 `user_facing_answer` 不再参与任何最终答案路径。
- Stage 1 的临时 renderer 被移除或明确只作为关闭 Composer 时的安全降级,不恢复 Executor 表达。
### 阶段五:回归评测与审计闭环
状态:未开始。
目标:确认新链路真的降低证据归因幻觉,而不是只改变字段名。
改动范围:
@@ -1400,6 +1603,22 @@ Composer 输出严格 JSON。
- 最终答案中不再出现未通过 Verifier 的 claim。
- 审计链路能从 final answer 回溯到 Composer 输入、Verifier 判定、Gatekeeper 结果、tool invocation。
建议 fixture 场景:
| 场景 | 预期 |
|---|---|
| 伪造不存在的 invocation id | Gatekeeper fail,最终不可 PASS |
| 引用真实 id 但 tool_name 不匹配 | Gatekeeper fail,倾向 REJECT |
| excerpt 与真实工具输出明显不符 | Gatekeeper fail 或 warn,最终不可 PASS |
| hypothesis 写成 root_cause claim | Verifier 判 `overstated` 或 `unsupported` |
| Composer 输入不含某错误码但输出包含该错误码 | 测试失败 |
阶段退出条件:
- 关键回归测试已进入 CI 可运行测试集合。
- eval fixture 覆盖证据归因幻觉的主路径。
- 文档、OpenSpec、devflow 与实际代码行为一致。
---
## 22. 兼容与回滚策略
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,102 @@
# Decisions: executor-gatekeeper-hook
## sm-flow Progress
### Clarify
Entry summary: implement stage two of Executor Structured Output V2 by adding deterministic Gatekeeper checks between Executor and Verifier.
Slug: `executor-gatekeeper-hook`
Scale: complex program, stage-specific standard slice. It changes internal verifier payload and audit persistence but does not change external API or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor now emits `executor_evidence_v2` without final-expression fields.
- `chat-verifier-agent`: Verifier input is assembled explicitly by `VerifierInputHook` and persisted through `ChatService`.
- `evidence-trace-hardening`: tool invocation rows are the source of truth for evidence trace references.
Current code shape:
- `VerifierInputHook` parses Executor raw output, builds `tool_trace_summary`, and assembles Verifier payload.
- `VerifierContextHolder` stores parse status, structured output, and trace summary for later persistence.
- `ChatService.persistVerifierEvaluation(...)` writes verifier snapshots into `diagnosis_session.self_evaluation.verifier_evaluation`.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` can provide the valid invocation pool.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Gatekeeper live in Executor hook? | evidence-driven | No. User explicitly chose Verifier hook for this version. |
| Should Gatekeeper retry Executor? | evidence-driven | No. Retry is deferred; this phase only validates and audits. |
| Which rules are mandatory now? | evidence-driven | Implement schema and invocation reference rules. Excerpt similarity and utilization can follow later. |
| Where should audit data persist? | evidence-driven | Existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
| Does this require database migration? | evidence-driven | No. Use existing JSON self_evaluation and existing tool_invocation rows. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Gatekeeper output, rules, integration, persistence, and risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for payload, persistence, and initial rules.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- Gatekeeper introduces an internal verifier payload extension but no external API or database schema change.
- `VerifierInputHook` remains the integration point, matching the user decision to keep Gatekeeper in the Verifier hook.
- Gatekeeper needs repository access to validate invocation ids; placing that access in a service keeps hook logic small.
- Failed Gatekeeper results are audit signals in this phase and do not trigger retry.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage two | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design rules and payload | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier payload: L2 internal extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation result:
- Added `ExecutorGatekeeperService` with initial `schema.executor_v2` and `evidence.invocation_ref` checks.
- Integrated Gatekeeper into `VerifierInputHook` after parse and trace summary construction.
- Added `gatekeeper_result` to Verifier payload and `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated `chat-verifier-prompt.md` so Gatekeeper failures must not produce PASS.
- Added and updated focused tests for service validation, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Validation:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed with 23 tests.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
Known limits:
- No Executor retry is triggered by Gatekeeper failure in this phase.
- `evidence.excerpt_similarity`, `claim.hallucination_phrases`, and `claim.evidence_utilization` remain deferred by design.
@@ -0,0 +1,119 @@
## Context
Stage one introduced `executor_evidence_v2`, but `VerifierInputHook` still only parses JSON and passes the structured object to Verifier. A valid JSON object can still contain:
- removed fields such as `diagnosis_summary` or `user_facing_answer`
- missing or wrong `answer_version`
- empty `claims[].evidence_bindings`
- fabricated `source_invocation_ids`
- `tool_name` values that do not match the real `tool_invocation`
Gatekeeper handles these deterministic failures before Verifier performs semantic reasoning.
## Goals / Non-Goals
Goals:
- Add Gatekeeper into `VerifierInputHook`.
- Produce a small `gatekeeper_result` object.
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` in verifier evaluation.
- Implement initial rules: schema and invocation reference.
Non-goals:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## Gatekeeper Output
Gatekeeper returns:
```json
{
"status": "fail",
"failed_rules": ["evidence.invocation_ref"],
"warnings": [],
"errors": [
{
"rule_id": "evidence.invocation_ref",
"target": "claims[0].evidence_bindings[0]",
"message": "source_invocation_ids not found in current session"
}
]
}
```
Status calculation:
```text
any fail -> fail
else any warning -> warn
else pass
```
## Initial Rules
### schema.executor_v2
Fail when:
- structured output is absent after parse status is valid
- `answer_version` is not `executor_evidence_v2`
- `claims` is not an array
- removed fields `diagnosis_summary` or `user_facing_answer` are present
- any claim misses required fields
- any claim has empty `evidence_bindings`
- `hypotheses`, `recommended_actions`, or `missing_info` are missing or non-array
For this phase, missing optional arrays may be normalized only if implementation remains simple. If not normalized, missing arrays fail schema to keep behavior deterministic.
### evidence.invocation_ref
Fail when:
- `claims[].evidence_bindings[].source_invocation_ids` is missing or empty
- any referenced invocation id is not in the current session's `tool_invocation` rows
- `tool_name` does not match the referenced invocation's real `tool_name`
Recommended action evidence bindings remain optional and are not hard-fail checked in this phase.
## Integration
`VerifierInputHook.beforeModel(...)` flow becomes:
```text
parse Executor output
build tool_trace_summary
Gatekeeper.validate(sessionId, structuredOutput, parseStatus)
put gatekeeper_result into verifier payload
store gatekeeper_result in VerifierContextHolder
```
`ChatService.persistVerifierEvaluation(...)` adds:
```text
gatekeeper_result: VerifierContextHolder.getGatekeeperResult()
```
If Gatekeeper throws unexpectedly, hook should fail closed with a minimal `fail` result in payload rather than dropping validation silently.
## Interface Impact
- Verifier payload: L2 internal contract extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension under existing `self_evaluation`.
- No external API or database schema change.
## Risks / Mitigations
- Risk: tests or code instantiate `VerifierInputHook` with the old constructor.
- Mitigation: keep a compatibility constructor that uses a no-op/pass Gatekeeper, or update tests explicitly.
- Risk: Gatekeeper fails because no session id exists.
- Mitigation: return fail with a clear `schema.executor_v2` or `evidence.invocation_ref` error only when validation cannot establish current-session references.
- Risk: Verifier prompt may ignore Gatekeeper.
- Mitigation: payload and audit are still authoritative for later phases; Verifier prompt update can be minimal in this phase.
@@ -0,0 +1,37 @@
## Why
Executor V2 makes the output structured, but structure alone does not prevent physical hallucinations such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names. These failures should be caught deterministically before the Verifier reasons over claims.
This phase adds Gatekeeper inside `VerifierInputHook` as a deterministic pre-verifier quality gate and persists its result for audit.
## What Changes
- Add a Gatekeeper validation service for Executor structured output.
- Add initial rules:
- `schema.executor_v2`
- `evidence.invocation_ref`
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Keep Gatekeeper in `VerifierInputHook`; do not move it to Executor hook.
- Keep retry behavior unchanged; failed Gatekeeper results do not trigger Executor retry in this phase.
- Keep Verifier prompt/output behavior unchanged except that it can see `gatekeeper_result`.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Verifier input now includes deterministic Gatekeeper results for Executor structured output.
## Impact
- Affected hook: `VerifierInputHook`.
- Affected service/runtime: new Gatekeeper service and `VerifierContextHolder`.
- Affected persistence: `ChatService.persistVerifierEvaluation(...)` writes `gatekeeper_result` into existing `self_evaluation`.
- Affected repository access: Gatekeeper reads current-session `tool_invocation` rows via `ToolInvocationRepository`.
- Affected tests: `VerifierInputHookTest` and `ChatServiceSequentialAgentTest`.
- Database schema: no change.
@@ -0,0 +1,51 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
## ADDED Requirements
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
@@ -0,0 +1,31 @@
## 1. Gatekeeper Core
- [x] 1.1 Add a Gatekeeper validation service with a small result shape: `status`, `failed_rules`, `warnings`, `errors`.
- [x] 1.2 Implement `schema.executor_v2` rule.
- [x] 1.3 Implement `evidence.invocation_ref` rule using current-session `tool_invocation` rows.
- [x] 1.4 Keep recommended action evidence bindings out of hard-fail validation for this phase.
## 2. Hook Integration
- [x] 2.1 Inject Gatekeeper into `VerifierInputHook`.
- [x] 2.2 Add `gatekeeper_result` to Verifier payload.
- [x] 2.3 Store `gatekeeper_result` in `VerifierContextHolder`.
- [x] 2.4 Preserve parse-only boundary in `VerifierInputHook`.
## 3. Persistence
- [x] 3.1 Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [x] 3.2 Preserve existing verifier evaluation fields.
## 4. Prompt Compatibility
- [x] 4.1 Update `chat-verifier-prompt.md` minimally so Verifier sees `gatekeeper_result` and must not output PASS when it fails.
## 5. Tests And Verification
- [x] 5.1 Add Gatekeeper unit tests for schema failures and pass cases.
- [x] 5.2 Add hook tests proving payload contains `gatekeeper_result`.
- [x] 5.3 Add tests for fabricated invocation id and tool name mismatch.
- [x] 5.4 Add ChatService persistence test for `gatekeeper_result`.
- [x] 5.5 Run targeted tests.
- [x] 5.6 Validate this OpenSpec change.
@@ -129,6 +129,11 @@ The Verifier's verdict SHALL be persisted for observability.
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
@@ -200,6 +205,12 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
@@ -315,3 +326,32 @@ The system SHALL tolerate malformed or absent structured Executor output without
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
@@ -8,6 +8,7 @@ import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.service.ExecutorGatekeeperService;
import com.superbiz.agent.service.ToolTraceSummaryService;
import com.superbiz.agent.util.SessionContextHolder;
import com.superbiz.agent.util.VerifierContextHolder;
@@ -28,12 +29,19 @@ import java.util.Map;
public class VerifierInputHook extends MessagesModelHook {
private final ToolTraceSummaryService toolTraceSummaryService;
private final ExecutorGatekeeperService executorGatekeeperService;
private final ObjectMapper objectMapper = new ObjectMapper();
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
};
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService) {
this(toolTraceSummaryService, null);
}
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService,
ExecutorGatekeeperService executorGatekeeperService) {
this.toolTraceSummaryService = toolTraceSummaryService;
this.executorGatekeeperService = executorGatekeeperService;
}
@Override
@@ -60,12 +68,16 @@ public class VerifierInputHook extends MessagesModelHook {
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
Map<String, Object> verifierInput = new LinkedHashMap<>();
verifierInput.put("original_query", VerifierContextHolder.getOriginalQuery());
verifierInput.put("executor_final_answer", executorFinalAnswer);
verifierInput.put("executor_structured_output", parseResult.structuredOutput());
verifierInput.put("executor_output_parse_status", parseResult.status());
verifierInput.put("tool_trace_summary", toolTraceSummary);
verifierInput.put("gatekeeper_result", gatekeeperResult);
verifierInput.put("retry_context", VerifierContextHolder.getRetryContext());
String payload = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(verifierInput);
@@ -76,6 +88,29 @@ public class VerifierInputHook extends MessagesModelHook {
}
}
private Map<String, Object> runGatekeeper(String sessionId, ExecutorOutputParseResult parseResult) {
if (executorGatekeeperService == null) {
return passGatekeeperResult();
}
try {
return executorGatekeeperService.validate(sessionId, parseResult.structuredOutput(), parseResult.status());
} catch (Exception e) {
log.error("Gatekeeper validation failed unexpectedly", e);
return executorGatekeeperService.fail("gatekeeper.internal_error",
"gatekeeper",
e.getMessage() == null ? "gatekeeper validation failed" : e.getMessage());
}
}
private Map<String, Object> passGatekeeperResult() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", "pass");
result.put("failed_rules", List.of());
result.put("warnings", List.of());
result.put("errors", List.of());
return result;
}
private ExecutorOutputParseResult parseExecutorOutput(String executorFinalAnswer) {
if (executorFinalAnswer == null || executorFinalAnswer.isBlank()) {
return new ExecutorOutputParseResult(null, status("missing", "executor_final_answer is blank"));
@@ -112,6 +112,9 @@ public class ChatService {
@Autowired
private SelfEvaluationMergeService selfEvaluationMergeService;
@Autowired
private ExecutorGatekeeperService executorGatekeeperService;
@Value("${verifier.low-confidence-threshold:0.5}")
private double verifierLowConfidenceThreshold;
@@ -541,7 +544,7 @@ public class ChatService {
.model(chatModel)
.systemPrompt(chatVerifierPrompt)
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"),
new VerifierInputHook(toolTraceSummaryService))
new VerifierInputHook(toolTraceSummaryService, executorGatekeeperService))
.outputKey("verifier_output")
.build();
}
@@ -755,6 +758,9 @@ public class ChatService {
verifierEvaluation.put("executor_structured_output", VerifierContextHolder.getExecutorStructuredOutput());
verifierEvaluation.put("tool_trace_summary",
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
verifierEvaluation.put("gatekeeper_result",
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
.orElse(Map.of("status", "pass", "failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(session.getSelfEvaluation(), verifierEvaluation);
session.setSelfEvaluation(merged);
@@ -0,0 +1,237 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.springframework.stereotype.Service;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* Deterministic checks for Executor structured output before verifier reasoning.
*/
@Service
public class ExecutorGatekeeperService {
public static final String STATUS_PASS = "pass";
public static final String STATUS_WARN = "warn";
public static final String STATUS_FAIL = "fail";
public static final String RULE_SCHEMA = "schema.executor_v2";
public static final String RULE_INVOCATION_REF = "evidence.invocation_ref";
private final ToolInvocationRepository toolInvocationRepository;
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
this.toolInvocationRepository = toolInvocationRepository;
}
public Map<String, Object> validate(String sessionId,
Map<String, Object> structuredOutput,
Map<String, Object> parseStatus) {
GatekeeperResult result = new GatekeeperResult();
validateSchema(structuredOutput, parseStatus, result);
if (structuredOutput != null) {
validateInvocationRefs(sessionId, structuredOutput, result);
}
return result.toMap();
}
public Map<String, Object> pass() {
return new GatekeeperResult().toMap();
}
public Map<String, Object> fail(String ruleId, String target, String message) {
GatekeeperResult result = new GatekeeperResult();
result.fail(ruleId, target, message);
return result.toMap();
}
private void validateSchema(Map<String, Object> structuredOutput,
Map<String, Object> parseStatus,
GatekeeperResult result) {
String status = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
if (structuredOutput == null) {
if ("valid".equals(status)) {
result.fail(RULE_SCHEMA, "executor_structured_output", "structured output is missing after valid parse");
}
return;
}
if (!"executor_evidence_v2".equals(String.valueOf(structuredOutput.get("answer_version")))) {
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2");
}
if (structuredOutput.containsKey("diagnosis_summary")) {
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2");
}
if (structuredOutput.containsKey("user_facing_answer")) {
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2");
}
Object claimsValue = structuredOutput.get("claims");
if (!(claimsValue instanceof List<?> claims)) {
result.fail(RULE_SCHEMA, "claims", "claims must be an array");
return;
}
for (int i = 0; i < claims.size(); i++) {
String target = "claims[" + i + "]";
Object claimValue = claims.get(i);
if (!(claimValue instanceof Map<?, ?> claim)) {
result.fail(RULE_SCHEMA, target, "claim must be an object");
continue;
}
requireString(claim, "claim_id", target, result);
requireString(claim, "claim_type", target, result);
requireString(claim, "claim_text", target, result);
String supportLevel = stringValue(claim.get("support_level"));
if (!"direct".equals(supportLevel) && !"indirect".equals(supportLevel)) {
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect");
}
Object bindings = claim.get("evidence_bindings");
if (!(bindings instanceof List<?> bindingList) || bindingList.isEmpty()) {
result.fail(RULE_SCHEMA, target + ".evidence_bindings", "claims must include non-empty evidence_bindings");
}
}
requireArray(structuredOutput, "hypotheses", result);
requireArray(structuredOutput, "recommended_actions", result);
requireArray(structuredOutput, "missing_info", result);
}
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
if (sessionId == null || sessionId.isBlank()) {
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_ids");
return;
}
Map<Long, ToolInvocation> validInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)
.stream()
.filter(invocation -> invocation.getId() != null)
.collect(Collectors.toMap(ToolInvocation::getId, Function.identity(), (left, right) -> left));
Object claimsValue = structuredOutput.get("claims");
if (!(claimsValue instanceof List<?> claims)) {
return;
}
for (int claimIndex = 0; claimIndex < claims.size(); claimIndex++) {
Object claimValue = claims.get(claimIndex);
if (!(claimValue instanceof Map<?, ?> claim)) {
continue;
}
Object bindingsValue = claim.get("evidence_bindings");
if (!(bindingsValue instanceof List<?> bindings)) {
continue;
}
for (int bindingIndex = 0; bindingIndex < bindings.size(); bindingIndex++) {
String target = "claims[" + claimIndex + "].evidence_bindings[" + bindingIndex + "]";
Object bindingValue = bindings.get(bindingIndex);
if (!(bindingValue instanceof Map<?, ?> binding)) {
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object");
continue;
}
validateBindingInvocationIds(binding, validInvocations, target, result);
}
}
}
private void validateBindingInvocationIds(Map<?, ?> binding,
Map<Long, ToolInvocation> validInvocations,
String target,
GatekeeperResult result) {
Object idsValue = binding.get("source_invocation_ids");
if (!(idsValue instanceof List<?> ids) || ids.isEmpty()) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must be a non-empty array");
return;
}
String claimedToolName = stringValue(binding.get("tool_name"));
if (claimedToolName.isBlank()) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required");
}
Set<Long> checkedIds = new HashSet<>();
for (Object idValue : ids) {
Long id = asLong(idValue);
if (id == null) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must contain numeric ids");
continue;
}
if (!checkedIds.add(id)) {
continue;
}
ToolInvocation invocation = validInvocations.get(id);
if (invocation == null) {
result.fail(RULE_INVOCATION_REF, target, "source_invocation_ids not found in current session: " + id);
continue;
}
if (!claimedToolName.isBlank() && !Objects.equals(claimedToolName, invocation.getToolName())) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
}
}
}
private void requireArray(Map<String, Object> output, String field, GatekeeperResult result) {
if (!(output.get(field) instanceof List<?>)) {
result.fail(RULE_SCHEMA, field, field + " must be an array");
}
}
private void requireString(Map<?, ?> object, String field, String target, GatekeeperResult result) {
if (stringValue(object.get(field)).isBlank()) {
result.fail(RULE_SCHEMA, target + "." + field, field + " is required");
}
}
private String stringValue(Object value) {
return value == null ? "" : String.valueOf(value);
}
private Long asLong(Object value) {
if (value instanceof Number number) {
return number.longValue();
}
if (value instanceof String text) {
try {
return Long.parseLong(text);
} catch (NumberFormatException ignored) {
return null;
}
}
return null;
}
private static final class GatekeeperResult {
private final List<String> failedRules = new ArrayList<>();
private final List<String> warnings = new ArrayList<>();
private final List<Map<String, Object>> errors = new ArrayList<>();
void fail(String ruleId, String target, String message) {
if (!failedRules.contains(ruleId)) {
failedRules.add(ruleId);
}
Map<String, Object> error = new LinkedHashMap<>();
error.put("rule_id", ruleId);
error.put("target", target);
error.put("message", message);
errors.add(error);
}
Map<String, Object> toMap() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", failedRules.isEmpty() ? (warnings.isEmpty() ? STATUS_PASS : STATUS_WARN) : STATUS_FAIL);
result.put("failed_rules", failedRules);
result.put("warnings", warnings);
result.put("errors", errors);
return result;
}
}
}
@@ -14,6 +14,7 @@ public final class VerifierContextHolder {
private static final ThreadLocal<Map<String, Object>> EXECUTOR_STRUCTURED_OUTPUT = new ThreadLocal<>();
private static final ThreadLocal<Map<String, Object>> EXECUTOR_OUTPUT_PARSE_STATUS = new ThreadLocal<>();
private static final ThreadLocal<List<Map<String, Object>>> TOOL_TRACE_SUMMARY = new ThreadLocal<>();
private static final ThreadLocal<Map<String, Object>> GATEKEEPER_RESULT = new ThreadLocal<>();
private VerifierContextHolder() {
}
@@ -66,6 +67,14 @@ public final class VerifierContextHolder {
return TOOL_TRACE_SUMMARY.get();
}
public static void setGatekeeperResult(Map<String, Object> gatekeeperResult) {
GATEKEEPER_RESULT.set(gatekeeperResult);
}
public static Map<String, Object> getGatekeeperResult() {
return GATEKEEPER_RESULT.get();
}
public static void clear() {
ORIGINAL_QUERY.remove();
RETRY_CONTEXT.remove();
@@ -73,5 +82,6 @@ public final class VerifierContextHolder {
EXECUTOR_STRUCTURED_OUTPUT.remove();
EXECUTOR_OUTPUT_PARSE_STATUS.remove();
TOOL_TRACE_SUMMARY.remove();
GATEKEEPER_RESULT.remove();
}
}
@@ -20,6 +20,7 @@
- `input_summary`
- `output_summary`
- `evidence_level`
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`failed_rules`、`warnings`、`errors`
- `retry_context`:第二轮可选输入;若为空,按首轮处理
## 任务步骤
@@ -88,6 +89,11 @@
### 步骤四:生成 verdict
严格使用以下判定矩阵:
0. 若 `gatekeeper_result.status="fail"`
- 不得输出 `PASS`
- 若 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
- 否则至少输出 `LOW_CONFID`
1. 若任一关键事实(`is_critical=true`)为 `contradicted`
- `verdict = "REJECT"`
- `groundedness_score = 0.0`
@@ -4,6 +4,9 @@ import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.service.ExecutorGatekeeperService;
import com.superbiz.agent.service.ToolTraceSummaryService;
import com.superbiz.agent.util.VerifierContextHolder;
import org.junit.jupiter.api.AfterEach;
@@ -90,7 +93,12 @@ class VerifierInputHookTest {
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
Map.of("trace_ref", "trace-1", "tool_name", "query_metrics")
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService);
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("structured-v2-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("structured-v2-session").toolName("query_metrics").build()
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
VerifierContextHolder.setOriginalQuery("分析 MySQL 连接池耗尽");
String executorOutput = """
@@ -131,6 +139,55 @@ class VerifierInputHookTest {
assertFalse(payload.path("executor_structured_output").has("user_facing_answer"));
assertEquals("连接池 active 达到上限",
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
assertEquals("pass", payload.path("gatekeeper_result").path("status").asText());
assertEquals("pass", VerifierContextHolder.getGatekeeperResult().get("status"));
}
@Test
void beforeModelAddsFailingGatekeeperResultForFabricatedInvocationId() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("fabricated-invocation-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("fabricated-invocation-session").toolName("query_metrics").build()
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_ids": [999],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
AgentCommand command = hook.beforeModel(
List.of(new AssistantMessage(executorOutput)),
RunnableConfig.builder().addMetadata("sessionId", "fabricated-invocation-session").build()
);
JsonNode payload = readPayload(command);
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
assertEquals("evidence.invocation_ref",
payload.path("gatekeeper_result").path("failed_rules").get(0).asText());
}
@Test
@@ -8,6 +8,7 @@ import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
@@ -21,6 +22,7 @@ import org.springframework.ai.chat.model.Generation;
import org.springframework.ai.chat.prompt.Prompt;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.test.util.ReflectionTestUtils;
import org.mockito.ArgumentCaptor;
import java.util.List;
import java.util.Map;
@@ -33,6 +35,8 @@ import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.ArgumentMatchers.any;
import static org.mockito.ArgumentMatchers.anyString;
import static org.mockito.ArgumentMatchers.isNull;
import static org.mockito.Mockito.verify;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
@@ -330,6 +334,63 @@ class ChatServiceSequentialAgentTest {
assertFalse(result.answer().contains("executor_evidence_v2"));
}
@Test
void executeChatComplexPersistsGatekeeperResultInVerifierEvaluation() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-gatekeeper-persist-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-gatekeeper-persist-session")
.toolName("query_metrics")
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-gatekeeper-persist-session"
);
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
Map<String, Object> verifierEvaluation = captor.getValue();
assertTrue(verifierEvaluation.containsKey("gatekeeper_result"));
@SuppressWarnings("unchecked")
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
assertEquals("pass", gatekeeperResult.get("status"));
}
@Test
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
ChatService chatService = new ChatService();
@@ -432,6 +493,7 @@ class ChatServiceSequentialAgentTest {
when(toolTraceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
SelfEvaluationMergeService selfEvaluationMergeService = mock(SelfEvaluationMergeService.class);
when(selfEvaluationMergeService.mergeVerifierEvaluation(any(), any())).thenReturn("{}");
ExecutorGatekeeperService executorGatekeeperService = new ExecutorGatekeeperService(toolInvocationRepository);
ReflectionTestUtils.setField(chatService, "dateTimeTools", new DateTimeTools());
ReflectionTestUtils.setField(chatService, "lookupKnowledgeTool", new LookupKnowledgeTool());
@@ -444,6 +506,7 @@ class ChatServiceSequentialAgentTest {
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);
ReflectionTestUtils.setField(chatService, "toolTraceSummaryService", toolTraceSummaryService);
ReflectionTestUtils.setField(chatService, "selfEvaluationMergeService", selfEvaluationMergeService);
ReflectionTestUtils.setField(chatService, "executorGatekeeperService", executorGatekeeperService);
ReflectionTestUtils.setField(chatService, "verifierLowConfidenceThreshold", 0.5d);
ReflectionTestUtils.setField(chatService, "chatPlannerPrompt", "PLANNER_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatExecutorPrompt", "EXECUTOR_TEST_PROMPT");
@@ -0,0 +1,98 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
class ExecutorGatekeeperServiceTest {
@Test
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
}
@Test
void validateFailsWhenRemovedFieldsArePresent() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of());
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> output = validOutput(101L, "query_metrics");
output.put("user_facing_answer", "旧版最终答案");
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("schema.executor_v2"));
}
@Test
void validateFailsForFabricatedInvocationId() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(999L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@Test
void validateFailsForToolNameMismatch() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_logs").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@SuppressWarnings("unchecked")
private Map<String, Object> validOutput(Long invocationId, String toolName) {
return new java.util.LinkedHashMap<>(Map.of(
"answer_version", "executor_evidence_v2",
"claims", List.of(Map.of(
"claim_id", "claim-1",
"claim_type", "symptom",
"claim_text", "连接池 active 达到上限",
"support_level", "direct",
"evidence_bindings", List.of(Map.of(
"source_type", "tool_trace",
"source_id", "trace-1",
"tool_name", toolName,
"source_invocation_ids", List.of(invocationId),
"evidence_excerpt", "active=50 max=50"
))
)),
"hypotheses", List.of(),
"recommended_actions", List.of(),
"missing_info", List.of()
));
}
}