feat(agent): add executor evidence v2 contract

This commit is contained in:
aruo
2026-07-08 01:37:15 +08:00
parent a6afbfaa9d
commit 050cbc8fee
21 changed files with 2707 additions and 20 deletions
+1
View File
@@ -6,6 +6,7 @@
|---|---|---|---|---|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
@@ -0,0 +1,41 @@
# Acceptance: executor-v2-output-contract
## Implementation Result
Completed stage one of Executor Structured Output V2.
- Executor prompt now emits `executor_evidence_v2`.
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
- ChatService PASS path renders V2 structured output into readable Chinese.
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
## Static Verification
- `cmd /c openspec validate executor-v2-output-contract`
- Result: passed.
- Coverage: OpenSpec syntax and change validity.
## Script Verification
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
## Browser / Manual Verification
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
## Remaining Work
- Phase two: Gatekeeper in `VerifierInputHook`.
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer.
- Phase five: eval fixtures and full audit closure.
@@ -0,0 +1,32 @@
# Brief: executor-v2-output-contract
## Background
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
## Goal
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
## Scope
- Update Chat Executor prompt to emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
- Preserve V1 `user_facing_answer` extraction for compatibility.
## Non-Goals
- No Gatekeeper implementation.
- No Verifier V2 `claim_checks`.
- No Composer.
- No Planner changes.
- No database schema changes.
- No evidence tool signature changes.
## Related OpenSpec
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
@@ -0,0 +1,39 @@
# Decisions: executor-v2-output-contract
## Key Decisions
### Executor V2 removes final-expression fields
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
### Temporary renderer bridges the gap before Composer
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
### V1 compatibility remains
Decision: Existing V1 `user_facing_answer` extraction remains.
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
### Gatekeeper and Verifier V2 are deferred
Decision: This phase does not add Gatekeeper or `claim_checks`.
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
## Interface Impact
- Internal Agent output contract: L4, because fields are removed.
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
- External Chat answer: compatible intent; users still get readable Chinese.
## Risks
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
@@ -0,0 +1,21 @@
# Evidence: executor-v2-output-contract
## Context Used
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
## Code Evidence
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
## Key Finding
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
@@ -0,0 +1,514 @@
# Current Chat Agent Data Contracts
**状态**:当前实现
**日期**:2026-07-07
**范围**:当前 Chat 复杂诊断链路的数据结构定义
当前代码实现是三 Agent 顺序链路:
```text
chat_planner -> chat_executor -> chat_verifier
```
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
---
## 1. Workflow Input
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
```text
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
--- 用户问题 ---
{question}
--- retry_context ---
{retry_context}
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `question` | 用户输入 | 用户本轮原始问题 |
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
---
## 2. chat_planner
### 2.1 Input
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
```json
{
"question": "用户原始问题",
"history": [],
"available_knowledge_domains": "...",
"skill_catalog": {},
"retry_context": null
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `question` | workflow input | 用户原始问题 |
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
### 2.2 Output:`planner_plan`
当前 prompt 要求输出 JSON:
```json
{
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
"reasoning": "规划思路说明"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
| `selection_reason` | string | skill 选择理由 |
| `plan` | array | 给 Executor 的执行步骤 |
| `reasoning` | string | 规划思路说明 |
运行态输出 key:
```text
planner_plan
```
---
## 3. chat_executor
### 3.1 Input
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
```json
{
"planner_plan": {},
"history": [],
"retry_context": null,
"tool_permissions": {
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
"tool_callbacks": []
}
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
| `retry_context` | `ChatService` | 本轮补证据约束 |
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
### 3.2 Output:`executor_feedback`
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "root_cause",
"claim_text": "事实断言或有限结论",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
"source_invocation_ids": [],
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "未被证实但值得排查的方向",
"basis": "它基于哪些已知证据或为什么只是推测",
"needed_evidence": ["需要补充的证据"]
}
],
"recommended_actions": [
{
"action_text": "建议动作",
"reason": "为什么建议做这个动作",
"evidence_bindings": []
}
],
"missing_info": [
"导致无法确认完整根因的证据缺口"
],
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
| `claims[].claim_id` | string | claim 标识 |
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
| `claims[].claim_text` | string | 事实断言文本 |
| `claims[].support_level` | string | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
| `evidence_bindings[].tool_name` | string | 来源工具名 |
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
| `hypotheses` | array | 未证实但值得排查的方向 |
| `hypotheses[].hypothesis_text` | string | 假设文本 |
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
| `recommended_actions` | array | 建议动作 |
| `recommended_actions[].action_text` | string | 建议动作文本 |
| `recommended_actions[].reason` | string | 建议原因 |
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
| `missing_info` | array | 证据缺口 |
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
运行态输出 key:
```text
executor_feedback
```
---
## 4. chat_verifier
### 4.1 Input
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
```json
{
"original_query": "用户原始问题",
"executor_final_answer": "{...executor_feedback raw text...}",
"executor_structured_output": {},
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"tool_trace_summary": [],
"retry_context": null
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
### 4.2 `tool_trace_summary`
`ToolTraceSummaryService` 聚合 evidence tools:
```text
lookup_knowledge, query_logs, query_metrics, query_order
```
输出项结构:
```json
{
"trace_ref": "trace-1",
"tool_name": "query_logs",
"success": true,
"input_summary": "query=payment-service timeout",
"output_summary": "log_evidence: ...",
"evidence_level": "direct",
"topic_domain": "general",
"source_invocation_ids": [394],
"invocation_count": 1,
"failed_invocation_count": 0,
"no_hit_invocation_count": 0,
"query_samples": ["payment-service timeout"],
"retrieval_layers": [],
"relevance_levels": [],
"source_documents": []
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
| `tool_name` | string | 聚合后的工具名 |
| `success` | boolean | 是否存在可用证据 |
| `input_summary` | string | 工具输入摘要 |
| `output_summary` | string | 工具输出摘要 |
| `evidence_level` | string | `direct` / `indirect` / `none` |
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
| `invocation_count` | number | 聚合调用次数 |
| `failed_invocation_count` | number | 失败调用次数 |
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
| `query_samples` | array | 查询样例 |
| `retrieval_layers` | array | 检索层级 |
| `relevance_levels` | array | 相关性等级 |
| `source_documents` | array | 来源文档标签 |
### 4.3 Output:`verifier_output`
当前 `chat-verifier-prompt.md` 要求输出:
```json
{
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [
{
"fact": "ERR_TIMEOUT 表示请求超时",
"is_critical": true,
"verification": "direct_evidence",
"detail": "知识库文档明确给出该错误码定义",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
}
]
}
],
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | number | 关键事实证据支撑评分 |
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
| `facts_checked` | array | Verifier 校验过的事实列表 |
| `facts_checked[].fact` | string | 被校验事实 |
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
| `facts_checked[].detail` | string | 校验说明 |
| `facts_checked[].evidence_refs` | array | 证据引用 |
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
| `evidence_refs[].tool_name` | string | 引用工具 |
| `evidence_refs[].topic_domain` | string | 引用主题域 |
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
| `evidence_refs[].note` | string | 引用说明 |
| `rationale` | string | verdict 判定理由 |
运行态输出 key:
```text
verifier_output
```
---
## 5. VerifierDecision
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
```json
{
"verdict": "LOW_CONFID",
"groundednessScore": 0.5,
"criticalFactCount": 2,
"factsChecked": [],
"rationale": "证据不足",
"round": 1
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | Verifier verdict |
| `groundednessScore` | number | groundedness score |
| `criticalFactCount` | number | 关键事实数量 |
| `factsChecked` | array | 解析后的 facts_checked |
| `rationale` | string | 判定理由 |
| `round` | number | 当前验证轮次 |
---
## 6. retry_context
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
```json
{
"round": 1,
"missing_evidence_facts": [
"某关键事实:缺少直接证据"
],
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `round` | number | 触发 retry 的轮次 |
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
| `instruction` | string | 补证据约束 |
---
## 7. diagnosis_session.self_evaluation.verifier_evaluation
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
```json
{
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.5,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "证据不足",
"round": 1,
"traceability_version": "v1",
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"executor_structured_output": {},
"tool_trace_summary": []
}
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verifier_evaluation.verdict` | string | Verifier verdict |
| `verifier_evaluation.groundedness_score` | number | groundedness score |
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
| `verifier_evaluation.rationale` | string | 判定理由 |
| `verifier_evaluation.round` | number | 验证轮次 |
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
---
## 8. Final Answer Rendering
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
| Verdict | 当前行为 |
|---|---|
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
低置信模板使用:
```text
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
已确认信息:
- ...
当前缺口:
- ...
建议下一步:
- ...
```
拒绝模板使用:
```text
当前无法基于已获取证据生成可靠结论。
已确认信息:
- ...
证据缺口:
- ...
建议下一步:
- ...
```
---
## 9. Trace Persistence Data
### 9.1 diagnosis_session
| 字段 | 类型 | 定义 |
|---|---|---|
| `session_id` | string | 会话 id |
| `query` | text | 用户问题 |
| `status` | string | 会话状态 |
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
| `total_duration_ms` | number | 总耗时 |
| `total_token_count` | number | 总 token |
| `step_count` | number | agent step 数 |
| `tool_call_count` | number | tool invocation 数 |
| `answer` | longtext | 最终用户答案 |
| `self_evaluation` | json | 包含 verifier_evaluation |
| `feedback` | string | 用户反馈 |
### 9.2 agent_step
| 字段 | 类型 | 定义 |
|---|---|---|
| `session_id` | string | 会话 id |
| `step_index` | number | 步骤序号 |
| `agent_name` | string | `planner` / `executor` / `verifier` |
| `model_input` | text | 模型输入摘要 |
| `model_output` | text | 模型输出摘要 |
| `thought` | text | hook 记录的摘要信息 |
| `has_tool_call` | boolean | 是否包含工具调用 |
| `duration_ms` | number | 模型调用耗时 |
| `token_count` | number | token 数 |
### 9.3 tool_invocation
| 字段 | 类型 | 定义 |
|---|---|---|
| `id` | number | 工具调用 id |
| `session_id` | string | 会话 id |
| `step_id` | number | 对应 agent_step id |
| `tool_name` | string | 工具名 |
| `input_params` | json | 工具输入参数 |
| `output_preview` | text | 工具输出预览 |
| `output_length` | number | 原始输出长度 |
| `retrieval_layer` | string | 检索层 |
| `l0_match_count` | number | L0 命中数 |
| `l1_match_count` | number | L1 命中数 |
| `is_truncated` | boolean | 输出是否截断 |
| `relevance_level` | string | 相关性等级 |
| `dedup_reason` | string | 去重原因 |
| `retrieval_details` | json | 检索细节 |
| `duration_ms` | number | 工具耗时 |
| `success` | boolean | 是否成功 |
| `error_message` | text | 错误信息 |
File diff suppressed because it is too large Load Diff
@@ -0,0 +1 @@
archive-ready
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,112 @@
# Decisions: executor-v2-output-contract
## sm-flow Progress
### Clarify
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
Slug: `executor-v2-output-contract`
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
### Context
Relevant devflow history:
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
Conflict with historical decision:
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
Current code shape:
- `chat-executor-prompt.md` defines the V1 Executor output contract.
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
### Specify
OpenSpec artifacts:
- `proposal.md`: scope and compatibility boundary for stage one.
- `design.md`: V2 Executor contract and temporary rendering strategy.
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
- No database schema or evidence tool signature changes are required.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue background / stage one | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design contract and rendering bridge | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
### Commit
Commit gate result: passed.
- `proposal.md` exists and explains why this phase is needed.
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
- `tasks.md` contains executable implementation and verification tasks.
- `cmd /c openspec validate executor-v2-output-contract` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation summary:
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
- Preserved V1 `user_facing_answer` extraction for compatibility.
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
Verification:
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-v2-output-contract` passed.
Known limitations:
- Gatekeeper is not implemented in this phase.
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
- The V2 renderer is temporary and should be replaced by Composer in a later phase.
@@ -0,0 +1,112 @@
## Context
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
- `diagnosis_summary`
- `user_facing_answer`
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
This stage implements only the first slice of the V2 migration:
```text
Executor V2 output contract
-> existing VerifierInputHook parsing
-> existing Verifier
-> temporary ChatService structured renderer
```
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
## Goals / Non-Goals
Goals:
- Make Chat Executor emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
- Preserve evidence-binding requirements for confirmed claims.
- Prevent PASS routing from returning raw JSON to normal Chat users.
Non-goals:
- No Gatekeeper implementation.
- No Verifier prompt rewrite to `claim_checks`.
- No Composer agent.
- No database schema changes.
- No Planner changes.
- No evidence tool signature changes.
- No retry behavior changes.
## Executor V2 Contract
Executor final output SHALL be one JSON object:
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "payment-service 出现请求超时日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_logs",
"source_invocation_ids": [394],
"evidence_excerpt": "request timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
Removed fields:
- `diagnosis_summary`
- `user_facing_answer`
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
## Temporary Rendering Strategy
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
When Verifier returns `PASS`:
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
- `claims[].claim_text`
- `hypotheses[].hypothesis_text`
- `missing_info[]`
- `recommended_actions[].action_text` and `reason`
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
## Parser Boundary
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
## Interface Impact
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
## Risks / Mitigations
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
- Mitigation: add temporary V2 renderer in `ChatService`.
- Risk: current Verifier prompt still mentions `user_facing_answer`.
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
- Risk: tests assume V1 fields.
- Mitigation: update/add focused tests for V2 output without final-expression fields.
@@ -0,0 +1,33 @@
## Why
The current Chat Executor evidence contract still mixes diagnostic material with final user-facing prose through `diagnosis_summary` and `user_facing_answer`. This keeps the Executor in a "diagnose and narrate" mode, so unsupported details can be smuggled into the final answer before later Gatekeeper, Verifier V2, and Composer phases exist.
This first phase narrows Executor to structured diagnostic material only and adds a temporary safe rendering path so normal Chat responses do not expose raw Executor JSON while later phases are implemented.
## What Changes
- **BREAKING internal Agent contract**: Chat Executor output changes from `executor_evidence_v1` to `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from the Chat Executor final JSON contract.
- Keep the existing structured arrays: `claims`, `hypotheses`, `recommended_actions`, and `missing_info`.
- Preserve `claims[].evidence_bindings` and current-session evidence attribution rules.
- Adjust runtime final-answer handling so a PASS result with V2 Executor output is rendered into readable Chinese from structured fields instead of returning raw JSON.
- Keep Planner, Verifier, Gatekeeper, retry behavior, database schema, and tool signatures unchanged in this phase.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: The Executor evidence-attribution contract is tightened so V2 structured output no longer contains final-expression fields. Verifier still receives `executor_final_answer` as raw text and `executor_structured_output` when parseable.
## Impact
- Affected prompt: `src/main/resources/prompts/chat-executor-prompt.md`.
- Affected runtime: `ChatService` PASS answer extraction/rendering for Executor V2.
- Affected parser boundary: `VerifierInputHook` should continue parsing JSON but must not treat schema validation as its own responsibility in this phase.
- Affected tests: ChatService sequential flow tests and VerifierInputHook parsing tests for V2 output without `user_facing_answer`.
- Interface impact: L4 for internal Agent output contract because fields are removed from Executor JSON; external HTTP/chat answer behavior must remain readable Chinese and must not expose raw JSON.
@@ -0,0 +1,50 @@
## MODIFIED Requirements
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
@@ -0,0 +1,24 @@
## 1. Executor V2 Prompt
- [x] 1.1 Update `src/main/resources/prompts/chat-executor-prompt.md` so the final contract uses `answer_version="executor_evidence_v2"`.
- [x] 1.2 Remove `diagnosis_summary` and `user_facing_answer` from the required Executor output schema and validation rules.
- [x] 1.3 Keep claims, hypotheses, recommended actions, missing information, and evidence-binding rules.
- [x] 1.4 Keep Planner and tool-use behavior unchanged.
## 2. Runtime Rendering Safety
- [x] 2.1 Update `ChatService` PASS handling so V2 structured output is rendered into readable Chinese instead of raw JSON.
- [x] 2.2 Preserve V1 `user_facing_answer` extraction for compatibility.
- [x] 2.3 Ensure fallback behavior does not expose raw Executor JSON when structured rendering fails.
## 3. Parser Compatibility
- [x] 3.1 Keep `VerifierInputHook` parse-only behavior compatible with V2 output.
- [x] 3.2 Add or update tests proving V2 output without `user_facing_answer` parses into `executor_structured_output`.
## 4. Tests And Verification
- [x] 4.1 Add or update ChatService tests for PASS with `executor_evidence_v2`.
- [x] 4.2 Add or update tests proving final user answer does not contain raw JSON contract text.
- [x] 4.3 Run targeted tests for ChatService and VerifierInputHook.
- [x] 4.4 Validate this OpenSpec change.
+16 -10
View File
@@ -256,20 +256,24 @@ When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate co
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor final output contains required top-level fields
- **WHEN** Executor completes a Chat diagnosis step
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
@@ -286,10 +290,11 @@ The Chat Executor SHALL produce a machine-checkable final output that separates
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: User-facing answer is available
- **WHEN** Executor emits structured output
- **THEN** `user_facing_answer` SHALL be written in Chinese
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
@@ -309,3 +314,4 @@ The system SHALL tolerate malformed or absent structured Executor output without
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
@@ -444,8 +444,12 @@ public class ChatService {
}
if ("PASS".equals(finalDecision.verdict())) {
answer = extractUserFacingAnswer(answer)
.orElse(answer == null || answer.isBlank() ? "抱歉,多 Agent 分析未能生成有效结论。" : answer);
String executorAnswer = answer;
answer = extractUserFacingAnswer(executorAnswer)
.or(() -> renderStructuredExecutorAnswer(executorAnswer))
.orElse(executorAnswer == null || executorAnswer.isBlank()
? "抱歉,多 Agent 分析未能生成有效结论。"
: executorAnswer);
persistVerifierEvaluation(session, finalDecision, round);
break;
}
@@ -816,6 +820,98 @@ public class ChatService {
return Optional.empty();
}
private Optional<String> renderStructuredExecutorAnswer(String executorAnswer) {
if (executorAnswer == null || executorAnswer.isBlank()) {
return Optional.empty();
}
try {
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(executorAnswer));
if (!"executor_evidence_v2".equals(root.path("answer_version").asText(""))) {
return Optional.empty();
}
StringBuilder output = new StringBuilder();
appendTextArraySection(output, "已确认信息", root.path("claims"), "claim_text", "暂无可稳定确认的信息");
appendHypothesesSection(output, root.path("hypotheses"));
appendStringArraySection(output, "当前缺口", root.path("missing_info"), "当前缺少足够的直接证据支撑完整结论");
appendRecommendedActionsSection(output, root.path("recommended_actions"));
String rendered = output.toString().trim();
return rendered.isBlank() ? Optional.empty() : Optional.of(rendered);
} catch (Exception e) {
logger.debug("Failed to render executor_evidence_v2 answer", e);
return Optional.empty();
}
}
private void appendTextArraySection(StringBuilder output, String title, JsonNode items,
String fieldName, String emptyText) {
output.append(title).append(":");
if (!items.isArray() || items.isEmpty()) {
output.append("\n- ").append(emptyText);
return;
}
for (JsonNode item : items) {
String text = item.path(fieldName).asText("");
if (!text.isBlank()) {
output.append("\n- ").append(text);
}
}
if (output.charAt(output.length() - 1) == ':') {
output.append("\n- ").append(emptyText);
}
}
private void appendHypothesesSection(StringBuilder output, JsonNode hypotheses) {
if (!hypotheses.isArray() || hypotheses.isEmpty()) {
return;
}
output.append("\n\n可能方向:");
for (JsonNode hypothesis : hypotheses) {
String text = hypothesis.path("hypothesis_text").asText("");
if (text.isBlank()) {
continue;
}
String basis = hypothesis.path("basis").asText("");
output.append("\n- ").append(text);
if (!basis.isBlank()) {
output.append("(").append(basis).append(")");
}
}
}
private void appendStringArraySection(StringBuilder output, String title, JsonNode items, String emptyText) {
output.append("\n\n").append(title).append(":");
if (!items.isArray() || items.isEmpty()) {
output.append("\n- ").append(emptyText);
return;
}
for (JsonNode item : items) {
String text = item.asText("");
if (!text.isBlank()) {
output.append("\n- ").append(text);
}
}
}
private void appendRecommendedActionsSection(StringBuilder output, JsonNode actions) {
output.append("\n\n建议下一步:");
if (!actions.isArray() || actions.isEmpty()) {
output.append("\n- 围绕上述证据缺口补充只读查询,再由人工复核最终结论");
return;
}
for (JsonNode action : actions) {
String text = action.path("action_text").asText("");
if (text.isBlank()) {
continue;
}
String reason = action.path("reason").asText("");
output.append("\n- ").append(text);
if (!reason.isBlank()) {
output.append(":").append(reason);
}
}
}
private String buildDegradedOutput(VerifierDecision decision) {
StringBuilder output = new StringBuilder(DEGRADED_PREFIX);
@@ -51,7 +51,6 @@
支持等级:
- `direct`:工具返回中有直接事实。
- `indirect`:工具返回可支撑方向,但没有直接陈述完整结论。
- `none`:不能放入 `claims`,应放入 `hypotheses`、`recommended_actions` 或 `missing_info`。
### hypotheses
`hypotheses` 用来放合理怀疑但未被工具证实的方向。
@@ -71,8 +70,7 @@
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
@@ -106,15 +104,16 @@
],
"missing_info": [
"导致无法确认完整根因的证据缺口"
],
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致,不得额外加入未绑定证据的确认式事实。"
]
}
```
## 输出校验
- `answer_version` 必须是 `executor_evidence_v2`。
- 不得输出 `diagnosis_summary`。
- 不得输出 `user_facing_answer`。
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
- `claims[*].evidence_bindings` 不能为空。
- `evidence_excerpt` 必须来自工具返回,不允许编造。
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
- `user_facing_answer` 不得出现 `claims` 中没有、且又被写成确认结论的事实。
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
@@ -10,7 +10,7 @@
- `original_query`:用户原始问题
- `executor_final_answer`:本轮 Executor 最终答案
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`、`user_facing_answer`
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
- `trace_ref`
@@ -31,7 +31,7 @@
- 必须检查 claim 的 `evidence_bindings` 是否能对应到 `tool_trace_summary` 中真实存在的 trace、tool 或 source_invocation_ids
- 如果 claim 声称 direct/indirect 支撑,但 evidence binding 不存在、无法定位、或 excerpt 与工具摘要不匹配,不得判为 `direct_evidence`
然后必须扫描 `executor_structured_output.user_facing_answer`:
如果 `executor_structured_output.user_facing_answer` 存在,则必须扫描它:
- 如果其中出现 confirmed-sounding facts(确认式事实、根因、指标值、错误码、服务名、修复结论)
- 且这些事实没有出现在 `executor_structured_output.claims`
- 必须额外加入 `facts_checked` 并按工具证据校验
@@ -84,6 +84,55 @@ class VerifierInputHookTest {
assertEquals("valid", VerifierContextHolder.getExecutorOutputParseStatus().get("status"));
}
@Test
void beforeModelAddsStructuredExecutorOutputWhenV2ContractHasNoUserFacingAnswer() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
Map.of("trace_ref", "trace-1", "tool_name", "query_metrics")
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService);
VerifierContextHolder.setOriginalQuery("分析 MySQL 连接池耗尽");
String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
AgentCommand command = hook.beforeModel(
List.of(new AssistantMessage(executorOutput)),
RunnableConfig.builder().addMetadata("sessionId", "structured-v2-session").build()
);
JsonNode payload = readPayload(command);
assertEquals("valid", payload.path("executor_output_parse_status").path("status").asText());
assertEquals("executor_evidence_v2",
payload.path("executor_structured_output").path("answer_version").asText());
assertFalse(payload.path("executor_structured_output").has("user_facing_answer"));
assertEquals("连接池 active 达到上限",
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
}
@Test
void beforeModelExtractsStructuredOutputFromPrefixedJsonFence() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
@@ -273,6 +273,63 @@ class ChatServiceSequentialAgentTest {
assertTrue(chatModel.verifierPromptText.contains("连接池 active 达到上限"));
}
@Test
void executeChatComplexRendersExecutorEvidenceV2InsteadOfRawJsonOnPass() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "连接泄漏可能参与了连接池耗尽",
"basis": "已有连接池满载证据,但缺少泄漏检测日志",
"needed_evidence": ["连接泄漏检测日志"]
}
],
"recommended_actions": [
{
"action_text": "补充查询连接池泄漏检测日志",
"reason": "用于确认是否存在连接未释放"
}
],
"missing_info": ["缺少连接泄漏检测日志"]
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-v2-render-session"
);
assertTrue(result.answer().contains("已确认信息"));
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertTrue(result.answer().contains("可能方向"));
assertTrue(result.answer().contains("建议下一步"));
assertFalse(result.answer().contains("\"answer_version\""));
assertFalse(result.answer().contains("executor_evidence_v2"));
}
@Test
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
ChatService chatService = new ChatService();