feat: add chat verifier agent

This commit is contained in:
zhuyongxin
2026-07-03 10:54:33 +08:00
parent 4f5316d473
commit 9050487307
28 changed files with 3200 additions and 209 deletions
@@ -0,0 +1 @@
ready
@@ -0,0 +1,42 @@
{
"id": "chat-verifier-agent",
"metadata": {
"status": "archived",
"created_at": "2026-07-02",
"updated_at": "2026-07-03",
"archive_readiness": "archived",
"implementation_status": "archived"
},
"summary": "Add a verifier agent to the complex chat path and persist auditable verifier decisions with evidence traceability.",
"artifacts": {
"proposal": "proposal.md",
"design": "design.md",
"tasks": "tasks.md",
"specs": [
"specs/chat-verifier-agent/spec.md"
],
"devflow": "devflow/projects/2026-07-02-chat-verifier-agent"
},
"tasks": [
"Verifier prompt",
"VerifierInputHook explicit payload",
"ChatService planner-executor-verifier orchestration",
"Verdict routing and fixed user output templates",
"Tool trace summary and evidence_refs traceability",
"self_evaluation merge semantics",
"Compile and runtime verification"
],
"verification": [
{
"type": "script",
"command": "mvn -q -DskipTests compile",
"result": "passed"
},
{
"type": "runtime",
"command": "POST /api/chat",
"session_id": "9138f064",
"result": "planner, executor, and verifier executed; verifier_evaluation contains evidence_refs and tool_trace_summary source_invocation_ids"
}
]
}
@@ -0,0 +1,349 @@
## Context
Chat 多 Agent 链路当前由 ChatService 驱动 Planner → Executor,答案输出前无质量门禁。Verifier Agent 作为 Executor 后置质量门禁,在 Executor 输出后做事实核查。
前序 change `executor-action-memory-relevance` 已在 Executor 侧构建了行动记忆和检索质量归一化,Verifier 不需要重复验证检索质量。
## Goals / Non-Goals
**Goals:**
- Verifier 作为无工具 ReactAgent,由 ChatService 显式调用
- Verifier 输出 verdict (PASS/LOW_CONFID/REJECT) + groundedness_score + facts_checked
- ChatService 负责单轮显式编排:Planner → Executor → Verifier
- ChatService 外层根据 Verifier 判决做轮次路由:PASS→输出,LOW_CONFID≥0.5→带声明输出,LOW_CONFID<0.5→补充一轮,REJECT→降级
- Verifier 判决写入 diagnosis_session.self_evaluation JSON 容器中的 `verifier_evaluation` 槽位做可观测
**Non-Goals:**
- Verifier 不调用工具
- 不改动单 Agent 链路
- 不修改 Executor 的输出内容
- 不涉及数据库表结构变更
- Verifier 不继承 Executor 的中间推理过程(通过 MessagesModelHook 过滤)
## Decisions
| 决策 | 选择 | 放弃方案 | 原因 |
|------|------|---------|------|
| Verifier 是否有工具 | 无工具 ReactAgent | 有工具的 Agent | 职责单一,只核查不检索 |
| 判决分类 | PASS / LOW_CONFID / REJECT | PASS / FAIL 二分类 | LOW_CONFID 提供了弹性输出路径 |
| 回调机制 | ChatService 外层控制最多两轮 | 全交给 Supervisor / 不回调 | 轮次上限需要硬控制,不能只靠 prompt 记忆 |
| 可观测方案 | 写入 self_evaluation JSON 容器 | agent_step / tool_invocation / 新表 | 不改表结构,同时避免与 evidence_score 覆盖冲突 |
| 输入隔离 | 显式状态输入 + MessagesModelHook 裁剪噪音 | 仅靠原始消息过滤 / 数据库注入 | Verifier 需要稳定读取 query、工具摘要、最终答案,不能依赖消息格式猜测 |
| 阈值配置 | yml 配置化 | 硬编码 | 方便运维调整,不需改代码 |
## Verifier 输入契约
Verifier 的业务输入由 `ChatService` 显式组装,不依赖原始 conversation messages 的隐式结构。
### 必选输入
- `original_query`:用户原始问题
- `executor_final_answer`:本轮 Executor 最终答案
- `tool_trace_summary`:由工具调用事实整理出的半结构化摘要
### 条件输入
- `retry_context`:仅第二轮注入,描述上一轮 verifier 发现的证据缺口和补充约束
### tool_trace_summary 最小结构
```json
[
{
"tool_name": "lookup_knowledge",
"success": true,
"input_summary": "查询 ERR_TIMEOUT",
"output_summary": "命中 payment/errors.md,返回错误码定义",
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": false,
"input_summary": "按 traceId 查询日志",
"output_summary": "日志服务超时",
"evidence_level": "none"
}
]
```
约束:
- `tool_trace_summary` 只纳入证据型工具调用,不纳入纯辅助或无业务事实意义的工具
- `tool_trace_summary` 来源于工具调用事实,不直接透传原始日志全文
- Verifier 基于摘要做事实核查,不直接读取数据库
- 若某工具调用失败,仍需记录在摘要中,供 Verifier 判断证据缺口
### 证据型工具边界
默认纳入 `tool_trace_summary` 的工具:
- `lookup_knowledge`
- `query_logs`
- `query_metrics`
- `query_order` 或其他业务事实查询类工具
- 其他只读、能提供客观事实的工具
默认不纳入:
- `getCurrentDateTime`
- 纯格式化、转换、控制类工具
- 与事实核查无关的辅助工具
### 摘要压缩规则
- 每次调用只保留“最小证据摘要”,不透传原始返回全文
- `output_summary` 控制为 1-3 句,重点描述“这次调用证明了什么 / 没能证明什么”
- 失败调用必须保留,但统一标记:
- `success=false`
- `evidence_level=none`
- 同一工具、同一主题域、同一轮次的重复调用可以折叠为一条合并摘要
- 合并摘要至少保留:
- 首次有效命中结果
- 额外重复次数 / 未命中次数 / 失败次数
### 截断优先级
若 `tool_trace_summary` 过长,优先保留:
1. 被 `executor_final_answer` 直接引用的证据
2. 支撑根因结论的证据
3. 支撑修复结论的证据
4. 与上一轮 `retry_context` 缺口直接相关的证据
低优先级、与最终答案无关的辅助性工具摘要可被截断。
### retry_context 最小结构
```json
{
"round": 1,
"missing_evidence_facts": [
"“根因是连接池耗尽”缺少直接证据",
"“错误码 ERR_TIMEOUT 来自支付网关”只有间接支持"
],
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
}
```
### MessagesModelHook 职责边界
- 可以:移除 Planner/Executor 中间推理、无关闲聊和冗余 message
- 不可以:作为 Verifier 核心业务输入的唯一来源
- 目标:降噪,而非拼装业务事实
## Verifier 判决矩阵
Verifier 先提取并校验 `facts_checked`,再依据矩阵生成 verdict,避免只靠模型主观判断。
### facts_checked 分类
每条事实仅允许以下四类之一:
- `direct_evidence`:工具结果中有明确直接证据
- `indirect_support`:可由工具结果合理推导,但不是直接陈述
- `no_evidence`:工具结果中没有足够信息支撑
- `contradicted`:工具结果与该事实冲突,或该事实编造了不存在的关键实体/错误码/结论
### 关键事实范围
Verifier 优先校验关键事实,至少包括:
- 根因结论(root cause)
- 错误码 / 接口 / 组件归属
- 证据来源陈述(如“日志显示”“文档说明”)
- 明确修复结论
一般性建议、风险提示、非事实性表述默认不纳入关键事实,除非答案明确声称“已被证据证明”。
### verdict 规则
- `REJECT`
- 任意关键事实为 `contradicted`
- 或答案编造了工具/日志/文档中不存在的关键实体、错误码、结论
- `PASS`
- 所有关键事实均为 `direct_evidence` 或 `indirect_support`
- 且至少一条关键事实为 `direct_evidence`
- 且不存在 `contradicted`
- `LOW_CONFID`
- 不存在 `contradicted`
- 但存在关键事实为 `no_evidence`
- 或所有关键事实都只有 `indirect_support`,缺少直接锚点
一句话归纳:
- `REJECT` = 有冲突
- `LOW_CONFID` = 无冲突但缺关键证据
- `PASS` = 无冲突且关键事实均有支撑
### groundedness_score 计算
`groundedness_score` 不由模型自由打分,而由关键事实分类映射得到:
```text
direct_evidence = 1.0
indirect_support = 0.6
no_evidence = 0.0
contradicted = 0.0
```
规则:
- 仅对关键事实计分
- 取平均值后截断到 `[0.0, 1.0]`
- 若存在任意关键事实为 `contradicted`,直接 verdict=`REJECT`,且 `groundedness_score=0.0`
### 第二轮补证据范围
第二轮 `retry_context` 仅回灌以下关键缺口:
- 关键事实为 `no_evidence`
- 关键事实为 `indirect_support`,但仍缺直接证据锚点
`REJECT` 不进入第二轮补证据,直接降级输出。
## 用户侧输出协议
Verifier 的内部判决与用户侧最终输出类型分离:
- `PASS` → `NORMAL`
- `LOW_CONFID` → `LOW_CONFID_WITH_DISCLAIMER`
- `REJECT` → `DEGRADED`
### LOW_CONFID_WITH_DISCLAIMER
适用场景:
- 第一轮 `LOW_CONFID` 且 `groundedness_score >= threshold`
- 第二轮后仍为 `LOW_CONFID`
输出规则:
- 使用固定免责声明前缀
- 免责声明后拼接 `executor_final_answer`
- 可选附加“当前证据缺口”列表,但来源必须是 verifier 的关键缺口,不得自由扩写
建议模板:
```text
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
{executor_final_answer}
当前缺口:
- ...
- ...
```
### DEGRADED
适用场景:
- 任意一轮 `REJECT`
- 系统无法基于现有证据形成可靠结论
输出规则:
- 不透传原始 `executor_final_answer`
- 使用固定降级模板
- 仅允许包含:
- 已确认信息
- 证据缺口
- 下一步建议
建议模板:
```text
当前无法基于已获取证据生成可靠结论,建议人工介入。
已确认信息:
- ...
证据缺口:
- ...
建议下一步:
- ...
```
### 输出边界
- `LOW_CONFID_WITH_DISCLAIMER` 可以带出原始答案,但必须加固定免责声明
- `DEGRADED` 不得透传未经验证的原始答案
- 用户侧输出模板由代码层拼装,不依赖 Verifier 自由生成
## self_evaluation 存储约定
`diagnosis_session.self_evaluation` 统一定义为 JSON 容器对象,而不是单一评估结果:
```json
{
"rule_evaluation": {
"evidence_score": 65,
"source": "rule",
"factors": []
},
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.42,
"facts_checked": [],
"rationale": "...",
"round": 1
}
}
```
写入约束:
- `EvaluationService` 只负责写 `rule_evaluation`
- `ChatService` 只负责写 `verifier_evaluation`
- 两侧都必须使用 read-modify-write,保留另一侧已有内容
- 禁止整段覆盖 `self_evaluation`,除非初始化为空对象
## Risks / Trade-offs
- [Risk] Verifier 误判导致好答案被降级 → Mitigation: REJECT 仅用于明显编造场景,LOW_CONFID 为主要输出路径
- [Risk] callback Planner 后新答案质量不一定提升 → Mitigation: 仅回调一次,Token 成本可控
- [Risk] 第二轮仍可能产出 REJECT → Mitigation: 第二轮 REJECT 仍降级,不透传
- [Risk] `self_evaluation` 被异步 evidence_score 覆盖 → Mitigation: 定义 JSON 容器槽位,统一 read-modify-write
- [Risk] Verifier 增加 Token 消耗 → Mitigation: 单次轻量 LLM 调用,估算 <500 token
- [Risk] 消息过滤可能导致输入契约漂移 → Mitigation: 主输入由显式状态输入提供,Hook 仅用于剔除中间推理和无关噪音
- [Risk] 判决边界主观化,导致不同模型输出不稳定 → Mitigation: 用 facts_checked 分类 + verdict 矩阵 + 映射分数约束输出
- [Risk] 最终用户文案随模型漂移,导致产品行为不稳定 → Mitigation: LOW_CONFID/DEGRADED 使用固定输出协议和模板
## Implementation Plan
1. 创建 `chat-verifier-prompt.md`
2. 新建 `VerifierInputHook.java`(MessagesModelHook 实现,BEFORE_MODEL 时裁剪 messages,只保留必要上下文)
3. `ChatService.java` 新增 `buildChatVerifierAgent()` 方法(ReactAgent,无工具,带 hook)
4. 添加 `verifier.low-confidence-threshold: 0.5` 到 application.yml
5. 在 `ChatService.executeChatComplex()` 中显式调用 `Planner → Executor → Verifier`
6. 保留 `SupervisorAgent` 构造作为 legacy residue,不再依赖 prompt-only supervisor sequencing 保证 Verifier 执行
7. 在 `ChatService.executeChatComplex()` 外层实现最多两轮调用控制
8. 组装 Verifier 显式状态输入:`original_query` / `executor_final_answer` / `tool_trace_summary` / `retry_context`
9. 将 `self_evaluation` 升级为 JSON 容器读写:`rule_evaluation` / `verifier_evaluation`
10. 读取 Verifier 判决写入 `verifier_evaluation`
## Implementation Notes
### Explicit orchestration
The final implementation uses `ChatService` to call `planner -> executor -> verifier` directly in each outer round. This replaces the earlier prompt-only dependency on `SupervisorAgent` for verifier execution. The supervisor construction remains in the code as legacy residue, but runtime correctness is driven by explicit `callAgent(...)` ordering.
### Traceability model
The implemented verifier input and persisted evaluation include an evidence index:
- `tool_trace_summary[*].trace_ref`
- `tool_trace_summary[*].source_invocation_ids`
- `tool_trace_summary[*].query_samples`
- `tool_trace_summary[*].retrieval_layers`
- `tool_trace_summary[*].relevance_levels`
- `tool_trace_summary[*].source_documents`
Each verifier fact may carry `facts_checked[*].evidence_refs`, which points back to `trace_ref` and the underlying `tool_invocation` ids. This closes the audit gap where verifier could list many checked facts but the reviewer could not tell which facts related to which tool calls.
### Observability adjustment
`agent_step.thought` is now intentionally concise for verifier steps. Full verifier judgment belongs in `diagnosis_session.self_evaluation.verifier_evaluation`, with `model_output` retaining the model output snapshot.
@@ -0,0 +1,31 @@
# Proposal: chat-verifier-agent
## Why
Chat 多 Agent 链路缺少出口质量门禁。Executor 输出答案后会直接返回给用户,无法在返回前拦截缺证据、低置信或明显编造的结论。
## What Changes
- 新增无工具 Verifier Agent,在 Executor 输出后读取答案和工具调用证据摘要,产出 `PASS` / `LOW_CONFID` / `REJECT` 判决。
- `ChatService` 显式编排 `Planner -> Executor -> Verifier`,并根据 Verifier 判决控制最终输出或最多一次补充轮次。
- Verifier 输入使用显式状态块:`original_query`、`executor_final_answer`、`tool_trace_summary`、第二轮可选 `retry_context`。
- `diagnosis_session.self_evaluation` 作为 JSON 容器保存 `rule_evaluation` 与 `verifier_evaluation`,避免异步评分覆盖 Verifier 结果。
- Verifier 结果增加可追溯证据引用:`tool_trace_summary[*].trace_ref`、`source_invocation_ids` 与 `facts_checked[*].evidence_refs`。
- `LOW_CONFID` 和 `REJECT` 用户侧输出使用固定协议,`REJECT` 不透传未经验证的原始答案。
## Capabilities
### New Capabilities
- `chat-verifier-agent`: Chat 多 Agent 出口事实核查、判决路由、观测存储和证据可追溯能力。
### Modified Capabilities
- None.
## Impact
- Affected code: `ChatService`, chat verifier prompt, verifier input assembly, self-evaluation persistence, multi-agent runtime orchestration.
- Affected runtime behavior: complex chat path now runs a Verifier gate after Executor and may perform one bounded retry for low-confidence evidence gaps.
- No database schema change is required; `self_evaluation` remains the persistence container.
- Non-goals: Verifier 不调用工具、不改写 Executor 答案、不影响单 Agent 链路、不支持超过两轮的补充编排。
@@ -0,0 +1,189 @@
## ADDED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
- **AND** at least one critical fact has direct evidence
- **AND** no critical fact is contradicted
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical fact contradicts the tool results
- **AND** some critical facts have no supporting evidence
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only indirect support
- **WHEN** no critical fact contradicts the tool results
- **AND** all critical facts are only indirectly supported
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
#### Scenario: Output format validation
- **WHEN** the Verifier completes its analysis
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
#### Scenario: strict schema output
- **WHEN** the Verifier returns its result
- **THEN** it SHALL output exactly one JSON object
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
- **AND** the JSON object SHALL include `critical_fact_count`
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
Each checked fact SHALL be labeled using a fixed evidence classification.
#### Scenario: fact classification values
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
### Requirement: groundedness_score SHALL be derived from fact classifications
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
#### Scenario: contradicted fact forces reject
- **WHEN** any critical fact is labeled `contradicted`
- **THEN** the Verifier SHALL output verdict="REJECT"
- **AND** groundedness_score SHALL be `0.0`
#### Scenario: score derived from supported facts
- **WHEN** no critical fact is contradicted
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS → direct output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** the system SHALL output the Executor's answer directly
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence补充
- **AND** it SHALL produce a degraded output directly
#### Scenario: REJECT → degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
### Requirement: User-facing verifier outputs SHALL follow fixed templates
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
#### Scenario: LOW_CONFID uses disclaimer template
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
#### Scenario: rule evaluation stored separately
- **WHEN** the rule-based evidence scoring completes
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
- **AND** existing `verifier_evaluation` data SHALL be preserved
#### Scenario: verifier evaluation stored separately
- **WHEN** the Verifier completes
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: no whole-object overwrite after initialization
- **WHEN** either evaluation channel updates `self_evaluation`
- **THEN** the implementation SHALL use read-modify-write semantics
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
#### Scenario: tool trace summary preserves invocation references
- **WHEN** the system prepares verifier inputs
- **THEN** each summary item SHALL include a stable `trace_ref`
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
#### Scenario: only evidence-bearing tools included
- **WHEN** the system generates `tool_trace_summary`
- **THEN** it SHALL include only evidence-bearing tool invocations
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
#### Scenario: failed evidence calls preserved as evidence gaps
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
- **THEN** the summary SHALL still include that invocation
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
#### Scenario: repeated tool calls may be compacted
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
- **THEN** the system MAY compact them into a merged summary entry
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
#### Scenario: raw outputs not passed through in full
- **WHEN** a tool invocation returns large raw content
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
#### Scenario: MessagesModelHook used only for noise reduction
- **WHEN** a MessagesModelHook is used for the Verifier
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
#### Scenario: facts_checked contains evidence refs
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL include `evidence_refs`
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** the ChatService persists `verifier_evaluation`
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
@@ -0,0 +1,58 @@
# Tasks: chat-verifier-agent
## 1. Verifier Prompt
- [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`.
- [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`.
- [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping.
- [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`.
- [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
- [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence.
## 2. Verifier Input Hook
- [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`.
- [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`.
- [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage.
## 3. ChatService Integration
- [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`.
- [x] 3.2 Add configurable `verifier.low-confidence-threshold`.
- [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call.
- [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round.
- [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid.
- [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
## 4. Verdict Routing And User Output
- [x] 4.1 Route `PASS` to the executor answer.
- [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer.
- [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer.
- [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
- [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
## 5. Trace Summary And Observability
- [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`.
- [x] 5.2 Include only evidence-bearing tools by default.
- [x] 5.3 Compact repeated calls by tool and topic domain.
- [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels.
- [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`.
- [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`.
- [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`.
## 6. self_evaluation Merge Semantics
- [x] 6.1 Add `SelfEvaluationMergeService`.
- [x] 6.2 Write verifier results under `verifier_evaluation`.
- [x] 6.3 Write rule scoring under `rule_evaluation`.
- [x] 6.4 Preserve the other channel with read-modify-write semantics.
## 7. Verification
- [x] 7.1 Compile verification: `mvn -q -DskipTests compile`.
- [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`.
- [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
- [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`.
- [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.
+192
View File
@@ -0,0 +1,192 @@
# chat-verifier-agent Specification
## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
## Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
- **AND** at least one critical fact has direct evidence
- **AND** no critical fact is contradicted
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical fact contradicts the tool results
- **AND** some critical facts have no supporting evidence
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only indirect support
- **WHEN** no critical fact contradicts the tool results
- **AND** all critical facts are only indirectly supported
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
#### Scenario: Output format validation
- **WHEN** the Verifier completes its analysis
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
#### Scenario: strict schema output
- **WHEN** the Verifier returns its result
- **THEN** it SHALL output exactly one JSON object
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
- **AND** the JSON object SHALL include `critical_fact_count`
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
Each checked fact SHALL be labeled using a fixed evidence classification.
#### Scenario: fact classification values
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
### Requirement: groundedness_score SHALL be derived from fact classifications
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
#### Scenario: contradicted fact forces reject
- **WHEN** any critical fact is labeled `contradicted`
- **THEN** the Verifier SHALL output verdict="REJECT"
- **AND** groundedness_score SHALL be `0.0`
#### Scenario: score derived from supported facts
- **WHEN** no critical fact is contradicted
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS → direct output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** the system SHALL output the Executor's answer directly
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence补充
- **AND** it SHALL produce a degraded output directly
#### Scenario: REJECT → degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
### Requirement: User-facing verifier outputs SHALL follow fixed templates
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
#### Scenario: LOW_CONFID uses disclaimer template
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
#### Scenario: rule evaluation stored separately
- **WHEN** the rule-based evidence scoring completes
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
- **AND** existing `verifier_evaluation` data SHALL be preserved
#### Scenario: verifier evaluation stored separately
- **WHEN** the Verifier completes
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: no whole-object overwrite after initialization
- **WHEN** either evaluation channel updates `self_evaluation`
- **THEN** the implementation SHALL use read-modify-write semantics
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
#### Scenario: tool trace summary preserves invocation references
- **WHEN** the system prepares verifier inputs
- **THEN** each summary item SHALL include a stable `trace_ref`
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
#### Scenario: only evidence-bearing tools included
- **WHEN** the system generates `tool_trace_summary`
- **THEN** it SHALL include only evidence-bearing tool invocations
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
#### Scenario: failed evidence calls preserved as evidence gaps
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
- **THEN** the summary SHALL still include that invocation
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
#### Scenario: repeated tool calls may be compacted
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
- **THEN** the system MAY compact them into a merged summary entry
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
#### Scenario: raw outputs not passed through in full
- **WHEN** a tool invocation returns large raw content
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
#### Scenario: MessagesModelHook used only for noise reduction
- **WHEN** a MessagesModelHook is used for the Verifier
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
#### Scenario: facts_checked contains evidence refs
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL include `evidence_refs`
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** the ChatService persists `verifier_evaluation`
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier