feat(agent): add executor evidence v2 contract
This commit is contained in:
@@ -6,6 +6,7 @@
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
# Acceptance: executor-v2-output-contract
|
||||
|
||||
## Implementation Result
|
||||
|
||||
Completed stage one of Executor Structured Output V2.
|
||||
|
||||
- Executor prompt now emits `executor_evidence_v2`.
|
||||
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
|
||||
- ChatService PASS path renders V2 structured output into readable Chinese.
|
||||
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `cmd /c openspec validate executor-v2-output-contract`
|
||||
- Result: passed.
|
||||
- Coverage: OpenSpec syntax and change validity.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
|
||||
|
||||
## Browser / Manual Verification
|
||||
|
||||
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
|
||||
|
||||
## Unverified
|
||||
|
||||
- Full live application run with a real LLM.
|
||||
- MySQL trace inspection after a real chat session.
|
||||
|
||||
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- Phase two: Gatekeeper in `VerifierInputHook`.
|
||||
- Phase three: Verifier V2 `claim_checks`.
|
||||
- Phase four: Composer.
|
||||
- Phase five: eval fixtures and full audit closure.
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: executor-v2-output-contract
|
||||
|
||||
## Background
|
||||
|
||||
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
|
||||
|
||||
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
|
||||
|
||||
## Goal
|
||||
|
||||
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
|
||||
|
||||
## Scope
|
||||
|
||||
- Update Chat Executor prompt to emit `executor_evidence_v2`.
|
||||
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
|
||||
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
|
||||
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
|
||||
- Preserve V1 `user_facing_answer` extraction for compatibility.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No Gatekeeper implementation.
|
||||
- No Verifier V2 `claim_checks`.
|
||||
- No Composer.
|
||||
- No Planner changes.
|
||||
- No database schema changes.
|
||||
- No evidence tool signature changes.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
|
||||
@@ -0,0 +1,39 @@
|
||||
# Decisions: executor-v2-output-contract
|
||||
|
||||
## Key Decisions
|
||||
|
||||
### Executor V2 removes final-expression fields
|
||||
|
||||
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
|
||||
|
||||
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
|
||||
|
||||
### Temporary renderer bridges the gap before Composer
|
||||
|
||||
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
|
||||
|
||||
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
|
||||
|
||||
### V1 compatibility remains
|
||||
|
||||
Decision: Existing V1 `user_facing_answer` extraction remains.
|
||||
|
||||
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
|
||||
|
||||
### Gatekeeper and Verifier V2 are deferred
|
||||
|
||||
Decision: This phase does not add Gatekeeper or `claim_checks`.
|
||||
|
||||
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Internal Agent output contract: L4, because fields are removed.
|
||||
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
|
||||
- External Chat answer: compatible intent; users still get readable Chinese.
|
||||
|
||||
## Risks
|
||||
|
||||
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
|
||||
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Evidence: executor-v2-output-contract
|
||||
|
||||
## Context Used
|
||||
|
||||
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
|
||||
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
|
||||
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
|
||||
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||
|
||||
## Code Evidence
|
||||
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
|
||||
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
|
||||
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
|
||||
|
||||
## Key Finding
|
||||
|
||||
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
|
||||
|
||||
@@ -0,0 +1,514 @@
|
||||
# Current Chat Agent Data Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**日期**:2026-07-07
|
||||
**范围**:当前 Chat 复杂诊断链路的数据结构定义
|
||||
|
||||
当前代码实现是三 Agent 顺序链路:
|
||||
|
||||
```text
|
||||
chat_planner -> chat_executor -> chat_verifier
|
||||
```
|
||||
|
||||
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
|
||||
|
||||
---
|
||||
|
||||
## 1. Workflow Input
|
||||
|
||||
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
|
||||
|
||||
```text
|
||||
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
|
||||
|
||||
--- 用户问题 ---
|
||||
{question}
|
||||
|
||||
--- retry_context ---
|
||||
{retry_context}
|
||||
|
||||
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | 用户输入 | 用户本轮原始问题 |
|
||||
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
|
||||
|
||||
---
|
||||
|
||||
## 2. chat_planner
|
||||
|
||||
### 2.1 Input
|
||||
|
||||
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
|
||||
|
||||
```json
|
||||
{
|
||||
"question": "用户原始问题",
|
||||
"history": [],
|
||||
"available_knowledge_domains": "...",
|
||||
"skill_catalog": {},
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | workflow input | 用户原始问题 |
|
||||
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
|
||||
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
|
||||
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
|
||||
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
|
||||
|
||||
### 2.2 Output:`planner_plan`
|
||||
|
||||
当前 prompt 要求输出 JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
|
||||
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
|
||||
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
|
||||
"reasoning": "规划思路说明"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
|
||||
| `selection_reason` | string | skill 选择理由 |
|
||||
| `plan` | array | 给 Executor 的执行步骤 |
|
||||
| `reasoning` | string | 规划思路说明 |
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
planner_plan
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. chat_executor
|
||||
|
||||
### 3.1 Input
|
||||
|
||||
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
|
||||
|
||||
```json
|
||||
{
|
||||
"planner_plan": {},
|
||||
"history": [],
|
||||
"retry_context": null,
|
||||
"tool_permissions": {
|
||||
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
|
||||
"tool_callbacks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
|
||||
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
|
||||
| `retry_context` | `ChatService` | 本轮补证据约束 |
|
||||
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
|
||||
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
|
||||
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
|
||||
|
||||
### 3.2 Output:`executor_feedback`
|
||||
|
||||
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "事实断言或有限结论",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
|
||||
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
|
||||
"source_invocation_ids": [],
|
||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "未被证实但值得排查的方向",
|
||||
"basis": "它基于哪些已知证据或为什么只是推测",
|
||||
"needed_evidence": ["需要补充的证据"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "建议动作",
|
||||
"reason": "为什么建议做这个动作",
|
||||
"evidence_bindings": []
|
||||
}
|
||||
],
|
||||
"missing_info": [
|
||||
"导致无法确认完整根因的证据缺口"
|
||||
],
|
||||
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
|
||||
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
|
||||
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
|
||||
| `claims[].claim_id` | string | claim 标识 |
|
||||
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
|
||||
| `claims[].claim_text` | string | 事实断言文本 |
|
||||
| `claims[].support_level` | string | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
|
||||
| `evidence_bindings[].tool_name` | string | 来源工具名 |
|
||||
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
|
||||
| `hypotheses` | array | 未证实但值得排查的方向 |
|
||||
| `hypotheses[].hypothesis_text` | string | 假设文本 |
|
||||
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
|
||||
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
|
||||
| `recommended_actions` | array | 建议动作 |
|
||||
| `recommended_actions[].action_text` | string | 建议动作文本 |
|
||||
| `recommended_actions[].reason` | string | 建议原因 |
|
||||
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
|
||||
| `missing_info` | array | 证据缺口 |
|
||||
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
executor_feedback
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. chat_verifier
|
||||
|
||||
### 4.1 Input
|
||||
|
||||
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor_feedback raw text...}",
|
||||
"executor_structured_output": {},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
|
||||
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
|
||||
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
|
||||
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
|
||||
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
|
||||
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
|
||||
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
|
||||
|
||||
### 4.2 `tool_trace_summary`
|
||||
|
||||
`ToolTraceSummaryService` 聚合 evidence tools:
|
||||
|
||||
```text
|
||||
lookup_knowledge, query_logs, query_metrics, query_order
|
||||
```
|
||||
|
||||
输出项结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"input_summary": "query=payment-service timeout",
|
||||
"output_summary": "log_evidence: ...",
|
||||
"evidence_level": "direct",
|
||||
"topic_domain": "general",
|
||||
"source_invocation_ids": [394],
|
||||
"invocation_count": 1,
|
||||
"failed_invocation_count": 0,
|
||||
"no_hit_invocation_count": 0,
|
||||
"query_samples": ["payment-service timeout"],
|
||||
"retrieval_layers": [],
|
||||
"relevance_levels": [],
|
||||
"source_documents": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
|
||||
| `tool_name` | string | 聚合后的工具名 |
|
||||
| `success` | boolean | 是否存在可用证据 |
|
||||
| `input_summary` | string | 工具输入摘要 |
|
||||
| `output_summary` | string | 工具输出摘要 |
|
||||
| `evidence_level` | string | `direct` / `indirect` / `none` |
|
||||
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
|
||||
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
|
||||
| `invocation_count` | number | 聚合调用次数 |
|
||||
| `failed_invocation_count` | number | 失败调用次数 |
|
||||
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
|
||||
| `query_samples` | array | 查询样例 |
|
||||
| `retrieval_layers` | array | 检索层级 |
|
||||
| `relevance_levels` | array | 相关性等级 |
|
||||
| `source_documents` | array | 来源文档标签 |
|
||||
|
||||
### 4.3 Output:`verifier_output`
|
||||
|
||||
当前 `chat-verifier-prompt.md` 要求输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "ERR_TIMEOUT 表示请求超时",
|
||||
"is_critical": true,
|
||||
"verification": "direct_evidence",
|
||||
"detail": "知识库文档明确给出该错误码定义",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "lookup_knowledge",
|
||||
"topic_domain": "api",
|
||||
"source_invocation_ids": [101, 104],
|
||||
"note": "trace-1 的文档摘要直接给出错误码定义"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | number | 关键事实证据支撑评分 |
|
||||
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
|
||||
| `facts_checked` | array | Verifier 校验过的事实列表 |
|
||||
| `facts_checked[].fact` | string | 被校验事实 |
|
||||
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
|
||||
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
|
||||
| `facts_checked[].detail` | string | 校验说明 |
|
||||
| `facts_checked[].evidence_refs` | array | 证据引用 |
|
||||
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
|
||||
| `evidence_refs[].tool_name` | string | 引用工具 |
|
||||
| `evidence_refs[].topic_domain` | string | 引用主题域 |
|
||||
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
|
||||
| `evidence_refs[].note` | string | 引用说明 |
|
||||
| `rationale` | string | verdict 判定理由 |
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
verifier_output
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. VerifierDecision
|
||||
|
||||
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundednessScore": 0.5,
|
||||
"criticalFactCount": 2,
|
||||
"factsChecked": [],
|
||||
"rationale": "证据不足",
|
||||
"round": 1
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | Verifier verdict |
|
||||
| `groundednessScore` | number | groundedness score |
|
||||
| `criticalFactCount` | number | 关键事实数量 |
|
||||
| `factsChecked` | array | 解析后的 facts_checked |
|
||||
| `rationale` | string | 判定理由 |
|
||||
| `round` | number | 当前验证轮次 |
|
||||
|
||||
---
|
||||
|
||||
## 6. retry_context
|
||||
|
||||
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"某关键事实:缺少直接证据"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `round` | number | 触发 retry 的轮次 |
|
||||
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
|
||||
| `instruction` | string | 补证据约束 |
|
||||
|
||||
---
|
||||
|
||||
## 7. diagnosis_session.self_evaluation.verifier_evaluation
|
||||
|
||||
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.5,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [],
|
||||
"rationale": "证据不足",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"executor_structured_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verifier_evaluation.verdict` | string | Verifier verdict |
|
||||
| `verifier_evaluation.groundedness_score` | number | groundedness score |
|
||||
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
|
||||
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
|
||||
| `verifier_evaluation.rationale` | string | 判定理由 |
|
||||
| `verifier_evaluation.round` | number | 验证轮次 |
|
||||
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
|
||||
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
|
||||
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
|
||||
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
|
||||
|
||||
---
|
||||
|
||||
## 8. Final Answer Rendering
|
||||
|
||||
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
|
||||
|
||||
| Verdict | 当前行为 |
|
||||
|---|---|
|
||||
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
|
||||
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
|
||||
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
|
||||
|
||||
低置信模板使用:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
拒绝模板使用:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Trace Persistence Data
|
||||
|
||||
### 9.1 diagnosis_session
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `query` | text | 用户问题 |
|
||||
| `status` | string | 会话状态 |
|
||||
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
|
||||
| `total_duration_ms` | number | 总耗时 |
|
||||
| `total_token_count` | number | 总 token |
|
||||
| `step_count` | number | agent step 数 |
|
||||
| `tool_call_count` | number | tool invocation 数 |
|
||||
| `answer` | longtext | 最终用户答案 |
|
||||
| `self_evaluation` | json | 包含 verifier_evaluation |
|
||||
| `feedback` | string | 用户反馈 |
|
||||
|
||||
### 9.2 agent_step
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_index` | number | 步骤序号 |
|
||||
| `agent_name` | string | `planner` / `executor` / `verifier` |
|
||||
| `model_input` | text | 模型输入摘要 |
|
||||
| `model_output` | text | 模型输出摘要 |
|
||||
| `thought` | text | hook 记录的摘要信息 |
|
||||
| `has_tool_call` | boolean | 是否包含工具调用 |
|
||||
| `duration_ms` | number | 模型调用耗时 |
|
||||
| `token_count` | number | token 数 |
|
||||
|
||||
### 9.3 tool_invocation
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `id` | number | 工具调用 id |
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_id` | number | 对应 agent_step id |
|
||||
| `tool_name` | string | 工具名 |
|
||||
| `input_params` | json | 工具输入参数 |
|
||||
| `output_preview` | text | 工具输出预览 |
|
||||
| `output_length` | number | 原始输出长度 |
|
||||
| `retrieval_layer` | string | 检索层 |
|
||||
| `l0_match_count` | number | L0 命中数 |
|
||||
| `l1_match_count` | number | L1 命中数 |
|
||||
| `is_truncated` | boolean | 输出是否截断 |
|
||||
| `relevance_level` | string | 相关性等级 |
|
||||
| `dedup_reason` | string | 去重原因 |
|
||||
| `retrieval_details` | json | 检索细节 |
|
||||
| `duration_ms` | number | 工具耗时 |
|
||||
| `success` | boolean | 是否成功 |
|
||||
| `error_message` | text | 错误信息 |
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
archive-ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-07
|
||||
@@ -0,0 +1,112 @@
|
||||
# Decisions: executor-v2-output-contract
|
||||
|
||||
## sm-flow Progress
|
||||
|
||||
### Clarify
|
||||
|
||||
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
|
||||
|
||||
Slug: `executor-v2-output-contract`
|
||||
|
||||
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
|
||||
|
||||
### Context
|
||||
|
||||
Relevant devflow history:
|
||||
|
||||
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
|
||||
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
|
||||
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
|
||||
|
||||
Conflict with historical decision:
|
||||
|
||||
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
|
||||
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
|
||||
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
|
||||
|
||||
Current code shape:
|
||||
|
||||
- `chat-executor-prompt.md` defines the V1 Executor output contract.
|
||||
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
|
||||
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
|
||||
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
|
||||
|
||||
### Grill
|
||||
|
||||
Question pool:
|
||||
|
||||
| Question | Mode | Resolution |
|
||||
|---|---|---|
|
||||
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
|
||||
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
|
||||
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
|
||||
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
|
||||
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
|
||||
|
||||
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
|
||||
|
||||
### Specify
|
||||
|
||||
OpenSpec artifacts:
|
||||
|
||||
- `proposal.md`: scope and compatibility boundary for stage one.
|
||||
- `design.md`: V2 Executor contract and temporary rendering strategy.
|
||||
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
|
||||
- `tasks.md`: executable implementation and verification checklist.
|
||||
|
||||
### Audit
|
||||
|
||||
Architecture risk summary:
|
||||
|
||||
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
|
||||
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
|
||||
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
|
||||
- No database schema or evidence tool signature changes are required.
|
||||
|
||||
Cross-artifact alignment:
|
||||
|
||||
| Source | Target | Status |
|
||||
|---|---|---|
|
||||
| issue background / stage one | proposal | aligned |
|
||||
| proposal scope / non-goals | design | aligned |
|
||||
| design contract and rendering bridge | specs | aligned |
|
||||
| specs observable behavior | tasks | aligned |
|
||||
|
||||
Interface impact:
|
||||
|
||||
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
|
||||
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
|
||||
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
|
||||
|
||||
### Commit
|
||||
|
||||
Commit gate result: passed.
|
||||
|
||||
- `proposal.md` exists and explains why this phase is needed.
|
||||
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
|
||||
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
|
||||
- `tasks.md` contains executable implementation and verification tasks.
|
||||
- `cmd /c openspec validate executor-v2-output-contract` passed.
|
||||
- No unresolved user-interview questions remain for this stage.
|
||||
|
||||
### Apply
|
||||
|
||||
Implementation summary:
|
||||
|
||||
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
|
||||
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
|
||||
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
|
||||
- Preserved V1 `user_facing_answer` extraction for compatibility.
|
||||
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
|
||||
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
|
||||
|
||||
Verification:
|
||||
|
||||
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
|
||||
- `cmd /c openspec validate executor-v2-output-contract` passed.
|
||||
|
||||
Known limitations:
|
||||
|
||||
- Gatekeeper is not implemented in this phase.
|
||||
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
|
||||
- The V2 renderer is temporary and should be replaced by Composer in a later phase.
|
||||
@@ -0,0 +1,112 @@
|
||||
## Context
|
||||
|
||||
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
|
||||
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
|
||||
|
||||
This stage implements only the first slice of the V2 migration:
|
||||
|
||||
```text
|
||||
Executor V2 output contract
|
||||
-> existing VerifierInputHook parsing
|
||||
-> existing Verifier
|
||||
-> temporary ChatService structured renderer
|
||||
```
|
||||
|
||||
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
Goals:
|
||||
|
||||
- Make Chat Executor emit `executor_evidence_v2`.
|
||||
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
|
||||
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
|
||||
- Preserve evidence-binding requirements for confirmed claims.
|
||||
- Prevent PASS routing from returning raw JSON to normal Chat users.
|
||||
|
||||
Non-goals:
|
||||
|
||||
- No Gatekeeper implementation.
|
||||
- No Verifier prompt rewrite to `claim_checks`.
|
||||
- No Composer agent.
|
||||
- No database schema changes.
|
||||
- No Planner changes.
|
||||
- No evidence tool signature changes.
|
||||
- No retry behavior changes.
|
||||
|
||||
## Executor V2 Contract
|
||||
|
||||
Executor final output SHALL be one JSON object:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "payment-service 出现请求超时日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_ids": [394],
|
||||
"evidence_excerpt": "request timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
Removed fields:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
|
||||
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
|
||||
|
||||
## Temporary Rendering Strategy
|
||||
|
||||
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
|
||||
|
||||
When Verifier returns `PASS`:
|
||||
|
||||
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
|
||||
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
|
||||
- `claims[].claim_text`
|
||||
- `hypotheses[].hypothesis_text`
|
||||
- `missing_info[]`
|
||||
- `recommended_actions[].action_text` and `reason`
|
||||
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
|
||||
|
||||
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
|
||||
|
||||
## Parser Boundary
|
||||
|
||||
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
|
||||
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
|
||||
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
|
||||
|
||||
## Risks / Mitigations
|
||||
|
||||
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
|
||||
- Mitigation: add temporary V2 renderer in `ChatService`.
|
||||
- Risk: current Verifier prompt still mentions `user_facing_answer`.
|
||||
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
|
||||
- Risk: tests assume V1 fields.
|
||||
- Mitigation: update/add focused tests for V2 output without final-expression fields.
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
## Why
|
||||
|
||||
The current Chat Executor evidence contract still mixes diagnostic material with final user-facing prose through `diagnosis_summary` and `user_facing_answer`. This keeps the Executor in a "diagnose and narrate" mode, so unsupported details can be smuggled into the final answer before later Gatekeeper, Verifier V2, and Composer phases exist.
|
||||
|
||||
This first phase narrows Executor to structured diagnostic material only and adds a temporary safe rendering path so normal Chat responses do not expose raw Executor JSON while later phases are implemented.
|
||||
|
||||
## What Changes
|
||||
|
||||
- **BREAKING internal Agent contract**: Chat Executor output changes from `executor_evidence_v1` to `executor_evidence_v2`.
|
||||
- Remove `diagnosis_summary` and `user_facing_answer` from the Chat Executor final JSON contract.
|
||||
- Keep the existing structured arrays: `claims`, `hypotheses`, `recommended_actions`, and `missing_info`.
|
||||
- Preserve `claims[].evidence_bindings` and current-session evidence attribution rules.
|
||||
- Adjust runtime final-answer handling so a PASS result with V2 Executor output is rendered into readable Chinese from structured fields instead of returning raw JSON.
|
||||
- Keep Planner, Verifier, Gatekeeper, retry behavior, database schema, and tool signatures unchanged in this phase.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `chat-verifier-agent`: The Executor evidence-attribution contract is tightened so V2 structured output no longer contains final-expression fields. Verifier still receives `executor_final_answer` as raw text and `executor_structured_output` when parseable.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected prompt: `src/main/resources/prompts/chat-executor-prompt.md`.
|
||||
- Affected runtime: `ChatService` PASS answer extraction/rendering for Executor V2.
|
||||
- Affected parser boundary: `VerifierInputHook` should continue parsing JSON but must not treat schema validation as its own responsibility in this phase.
|
||||
- Affected tests: ChatService sequential flow tests and VerifierInputHook parsing tests for V2 output without `user_facing_answer`.
|
||||
- Interface impact: L4 for internal Agent output contract because fields are removed from Executor JSON; external HTTP/chat answer behavior must remain readable Chinese and must not expose raw JSON.
|
||||
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Executor SHALL output an evidence-attribution contract
|
||||
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
|
||||
|
||||
#### Scenario: Executor V2 final output contains only structured diagnostic fields
|
||||
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
|
||||
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
|
||||
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
|
||||
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
|
||||
- **AND** the output SHALL NOT contain `diagnosis_summary`
|
||||
- **AND** the output SHALL NOT contain `user_facing_answer`
|
||||
|
||||
#### Scenario: Confirmed claims carry evidence bindings
|
||||
- **WHEN** Executor emits an item under `claims`
|
||||
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
|
||||
- **AND** `support_level` SHALL be one of `direct` or `indirect`
|
||||
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
|
||||
|
||||
#### Scenario: Evidence bindings support multiple tool types
|
||||
- **WHEN** Executor binds evidence to a claim
|
||||
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
|
||||
- **AND** the binding MAY include `source_id`
|
||||
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
|
||||
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
|
||||
|
||||
#### Scenario: Unsupported conclusions are not confirmed claims
|
||||
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
|
||||
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
|
||||
- **AND** Executor SHALL NOT present it as a confirmed claim
|
||||
|
||||
#### Scenario: Runbook and skill guidance do not become incident facts
|
||||
- **WHEN** Executor uses runbook, skill, or historical-case guidance
|
||||
- **THEN** the guidance MAY influence `recommended_actions`
|
||||
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
|
||||
|
||||
### Requirement: User-facing Chat answers SHALL remain readable Chinese
|
||||
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
|
||||
|
||||
#### Scenario: V2 machine contract is not exposed as normal user answer
|
||||
- **WHEN** Executor emits `executor_evidence_v2`
|
||||
- **AND** Verifier returns `PASS`
|
||||
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
|
||||
- **AND** normal user output SHALL NOT be the raw Executor JSON object
|
||||
|
||||
#### Scenario: Machine contract remains available for trace inspection
|
||||
- **WHEN** the Chat trace or verifier evaluation is inspected
|
||||
- **THEN** the structured Executor contract MAY be shown for debugging or audit
|
||||
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
## 1. Executor V2 Prompt
|
||||
|
||||
- [x] 1.1 Update `src/main/resources/prompts/chat-executor-prompt.md` so the final contract uses `answer_version="executor_evidence_v2"`.
|
||||
- [x] 1.2 Remove `diagnosis_summary` and `user_facing_answer` from the required Executor output schema and validation rules.
|
||||
- [x] 1.3 Keep claims, hypotheses, recommended actions, missing information, and evidence-binding rules.
|
||||
- [x] 1.4 Keep Planner and tool-use behavior unchanged.
|
||||
|
||||
## 2. Runtime Rendering Safety
|
||||
|
||||
- [x] 2.1 Update `ChatService` PASS handling so V2 structured output is rendered into readable Chinese instead of raw JSON.
|
||||
- [x] 2.2 Preserve V1 `user_facing_answer` extraction for compatibility.
|
||||
- [x] 2.3 Ensure fallback behavior does not expose raw Executor JSON when structured rendering fails.
|
||||
|
||||
## 3. Parser Compatibility
|
||||
|
||||
- [x] 3.1 Keep `VerifierInputHook` parse-only behavior compatible with V2 output.
|
||||
- [x] 3.2 Add or update tests proving V2 output without `user_facing_answer` parses into `executor_structured_output`.
|
||||
|
||||
## 4. Tests And Verification
|
||||
|
||||
- [x] 4.1 Add or update ChatService tests for PASS with `executor_evidence_v2`.
|
||||
- [x] 4.2 Add or update tests proving final user answer does not contain raw JSON contract text.
|
||||
- [x] 4.3 Run targeted tests for ChatService and VerifierInputHook.
|
||||
- [x] 4.4 Validate this OpenSpec change.
|
||||
@@ -256,20 +256,24 @@ When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate co
|
||||
### Requirement: Executor SHALL output an evidence-attribution contract
|
||||
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
|
||||
|
||||
#### Scenario: Executor final output contains required top-level fields
|
||||
- **WHEN** Executor completes a Chat diagnosis step
|
||||
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
|
||||
#### Scenario: Executor V2 final output contains only structured diagnostic fields
|
||||
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
|
||||
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
|
||||
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
|
||||
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
|
||||
- **AND** the output SHALL NOT contain `diagnosis_summary`
|
||||
- **AND** the output SHALL NOT contain `user_facing_answer`
|
||||
|
||||
#### Scenario: Confirmed claims carry evidence bindings
|
||||
- **WHEN** Executor emits an item under `claims`
|
||||
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
|
||||
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
|
||||
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
|
||||
- **AND** `support_level` SHALL be one of `direct` or `indirect`
|
||||
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
|
||||
|
||||
#### Scenario: Evidence bindings support multiple tool types
|
||||
- **WHEN** Executor binds evidence to a claim
|
||||
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
|
||||
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
|
||||
- **AND** the binding MAY include `source_id`
|
||||
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
|
||||
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
|
||||
|
||||
@@ -286,10 +290,11 @@ The Chat Executor SHALL produce a machine-checkable final output that separates
|
||||
### Requirement: User-facing Chat answers SHALL remain readable Chinese
|
||||
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
|
||||
|
||||
#### Scenario: User-facing answer is available
|
||||
- **WHEN** Executor emits structured output
|
||||
- **THEN** `user_facing_answer` SHALL be written in Chinese
|
||||
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
|
||||
#### Scenario: V2 machine contract is not exposed as normal user answer
|
||||
- **WHEN** Executor emits `executor_evidence_v2`
|
||||
- **AND** Verifier returns `PASS`
|
||||
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
|
||||
- **AND** normal user output SHALL NOT be the raw Executor JSON object
|
||||
|
||||
#### Scenario: Machine contract remains available for trace inspection
|
||||
- **WHEN** the Chat trace or verifier evaluation is inspected
|
||||
@@ -309,3 +314,4 @@ The system SHALL tolerate malformed or absent structured Executor output without
|
||||
- **WHEN** Executor output parsing fails
|
||||
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
|
||||
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
|
||||
|
||||
|
||||
@@ -444,8 +444,12 @@ public class ChatService {
|
||||
}
|
||||
|
||||
if ("PASS".equals(finalDecision.verdict())) {
|
||||
answer = extractUserFacingAnswer(answer)
|
||||
.orElse(answer == null || answer.isBlank() ? "抱歉,多 Agent 分析未能生成有效结论。" : answer);
|
||||
String executorAnswer = answer;
|
||||
answer = extractUserFacingAnswer(executorAnswer)
|
||||
.or(() -> renderStructuredExecutorAnswer(executorAnswer))
|
||||
.orElse(executorAnswer == null || executorAnswer.isBlank()
|
||||
? "抱歉,多 Agent 分析未能生成有效结论。"
|
||||
: executorAnswer);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
@@ -816,6 +820,98 @@ public class ChatService {
|
||||
return Optional.empty();
|
||||
}
|
||||
|
||||
private Optional<String> renderStructuredExecutorAnswer(String executorAnswer) {
|
||||
if (executorAnswer == null || executorAnswer.isBlank()) {
|
||||
return Optional.empty();
|
||||
}
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(executorAnswer));
|
||||
if (!"executor_evidence_v2".equals(root.path("answer_version").asText(""))) {
|
||||
return Optional.empty();
|
||||
}
|
||||
|
||||
StringBuilder output = new StringBuilder();
|
||||
appendTextArraySection(output, "已确认信息", root.path("claims"), "claim_text", "暂无可稳定确认的信息");
|
||||
appendHypothesesSection(output, root.path("hypotheses"));
|
||||
appendStringArraySection(output, "当前缺口", root.path("missing_info"), "当前缺少足够的直接证据支撑完整结论");
|
||||
appendRecommendedActionsSection(output, root.path("recommended_actions"));
|
||||
String rendered = output.toString().trim();
|
||||
return rendered.isBlank() ? Optional.empty() : Optional.of(rendered);
|
||||
} catch (Exception e) {
|
||||
logger.debug("Failed to render executor_evidence_v2 answer", e);
|
||||
return Optional.empty();
|
||||
}
|
||||
}
|
||||
|
||||
private void appendTextArraySection(StringBuilder output, String title, JsonNode items,
|
||||
String fieldName, String emptyText) {
|
||||
output.append(title).append(":");
|
||||
if (!items.isArray() || items.isEmpty()) {
|
||||
output.append("\n- ").append(emptyText);
|
||||
return;
|
||||
}
|
||||
for (JsonNode item : items) {
|
||||
String text = item.path(fieldName).asText("");
|
||||
if (!text.isBlank()) {
|
||||
output.append("\n- ").append(text);
|
||||
}
|
||||
}
|
||||
if (output.charAt(output.length() - 1) == ':') {
|
||||
output.append("\n- ").append(emptyText);
|
||||
}
|
||||
}
|
||||
|
||||
private void appendHypothesesSection(StringBuilder output, JsonNode hypotheses) {
|
||||
if (!hypotheses.isArray() || hypotheses.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
output.append("\n\n可能方向:");
|
||||
for (JsonNode hypothesis : hypotheses) {
|
||||
String text = hypothesis.path("hypothesis_text").asText("");
|
||||
if (text.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
String basis = hypothesis.path("basis").asText("");
|
||||
output.append("\n- ").append(text);
|
||||
if (!basis.isBlank()) {
|
||||
output.append("(").append(basis).append(")");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private void appendStringArraySection(StringBuilder output, String title, JsonNode items, String emptyText) {
|
||||
output.append("\n\n").append(title).append(":");
|
||||
if (!items.isArray() || items.isEmpty()) {
|
||||
output.append("\n- ").append(emptyText);
|
||||
return;
|
||||
}
|
||||
for (JsonNode item : items) {
|
||||
String text = item.asText("");
|
||||
if (!text.isBlank()) {
|
||||
output.append("\n- ").append(text);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private void appendRecommendedActionsSection(StringBuilder output, JsonNode actions) {
|
||||
output.append("\n\n建议下一步:");
|
||||
if (!actions.isArray() || actions.isEmpty()) {
|
||||
output.append("\n- 围绕上述证据缺口补充只读查询,再由人工复核最终结论");
|
||||
return;
|
||||
}
|
||||
for (JsonNode action : actions) {
|
||||
String text = action.path("action_text").asText("");
|
||||
if (text.isBlank()) {
|
||||
continue;
|
||||
}
|
||||
String reason = action.path("reason").asText("");
|
||||
output.append("\n- ").append(text);
|
||||
if (!reason.isBlank()) {
|
||||
output.append(":").append(reason);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private String buildDegradedOutput(VerifierDecision decision) {
|
||||
StringBuilder output = new StringBuilder(DEGRADED_PREFIX);
|
||||
|
||||
|
||||
@@ -51,7 +51,6 @@
|
||||
支持等级:
|
||||
- `direct`:工具返回中有直接事实。
|
||||
- `indirect`:工具返回可支撑方向,但没有直接陈述完整结论。
|
||||
- `none`:不能放入 `claims`,应放入 `hypotheses`、`recommended_actions` 或 `missing_info`。
|
||||
|
||||
### hypotheses
|
||||
`hypotheses` 用来放合理怀疑但未被工具证实的方向。
|
||||
@@ -71,8 +70,7 @@
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
@@ -106,15 +104,16 @@
|
||||
],
|
||||
"missing_info": [
|
||||
"导致无法确认完整根因的证据缺口"
|
||||
],
|
||||
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致,不得额外加入未绑定证据的确认式事实。"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## 输出校验
|
||||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||||
- 不得输出 `diagnosis_summary`。
|
||||
- 不得输出 `user_facing_answer`。
|
||||
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
|
||||
- `claims[*].evidence_bindings` 不能为空。
|
||||
- `evidence_excerpt` 必须来自工具返回,不允许编造。
|
||||
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
|
||||
- `user_facing_answer` 不得出现 `claims` 中没有、且又被写成确认结论的事实。
|
||||
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
|
||||
- `original_query`:用户原始问题
|
||||
- `executor_final_answer`:本轮 Executor 最终答案
|
||||
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`、`user_facing_answer`
|
||||
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
|
||||
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
|
||||
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
|
||||
- `trace_ref`
|
||||
@@ -31,7 +31,7 @@
|
||||
- 必须检查 claim 的 `evidence_bindings` 是否能对应到 `tool_trace_summary` 中真实存在的 trace、tool 或 source_invocation_ids
|
||||
- 如果 claim 声称 direct/indirect 支撑,但 evidence binding 不存在、无法定位、或 excerpt 与工具摘要不匹配,不得判为 `direct_evidence`
|
||||
|
||||
然后必须扫描 `executor_structured_output.user_facing_answer`:
|
||||
如果 `executor_structured_output.user_facing_answer` 存在,则必须扫描它:
|
||||
- 如果其中出现 confirmed-sounding facts(确认式事实、根因、指标值、错误码、服务名、修复结论)
|
||||
- 且这些事实没有出现在 `executor_structured_output.claims`
|
||||
- 必须额外加入 `facts_checked` 并按工具证据校验
|
||||
|
||||
@@ -84,6 +84,55 @@ class VerifierInputHookTest {
|
||||
assertEquals("valid", VerifierContextHolder.getExecutorOutputParseStatus().get("status"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void beforeModelAddsStructuredExecutorOutputWhenV2ContractHasNoUserFacingAnswer() throws Exception {
|
||||
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
|
||||
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
|
||||
Map.of("trace_ref", "trace-1", "tool_name", "query_metrics")
|
||||
));
|
||||
VerifierInputHook hook = new VerifierInputHook(traceSummaryService);
|
||||
VerifierContextHolder.setOriginalQuery("分析 MySQL 连接池耗尽");
|
||||
|
||||
String executorOutput = """
|
||||
{
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "连接池 active 达到上限",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "trace-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_ids": [101],
|
||||
"evidence_excerpt": "active=50 max=50"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
""";
|
||||
|
||||
AgentCommand command = hook.beforeModel(
|
||||
List.of(new AssistantMessage(executorOutput)),
|
||||
RunnableConfig.builder().addMetadata("sessionId", "structured-v2-session").build()
|
||||
);
|
||||
|
||||
JsonNode payload = readPayload(command);
|
||||
assertEquals("valid", payload.path("executor_output_parse_status").path("status").asText());
|
||||
assertEquals("executor_evidence_v2",
|
||||
payload.path("executor_structured_output").path("answer_version").asText());
|
||||
assertFalse(payload.path("executor_structured_output").has("user_facing_answer"));
|
||||
assertEquals("连接池 active 达到上限",
|
||||
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void beforeModelExtractsStructuredOutputFromPrefixedJsonFence() throws Exception {
|
||||
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
|
||||
|
||||
@@ -273,6 +273,63 @@ class ChatServiceSequentialAgentTest {
|
||||
assertTrue(chatModel.verifierPromptText.contains("连接池 active 达到上限"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void executeChatComplexRendersExecutorEvidenceV2InsteadOfRawJsonOnPass() throws Exception {
|
||||
ChatService chatService = createChatService();
|
||||
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||
chatModel.executorOutput = """
|
||||
{
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "连接池 active 达到上限",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "trace-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_ids": [101],
|
||||
"evidence_excerpt": "active=50 max=50"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "连接泄漏可能参与了连接池耗尽",
|
||||
"basis": "已有连接池满载证据,但缺少泄漏检测日志",
|
||||
"needed_evidence": ["连接泄漏检测日志"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询连接池泄漏检测日志",
|
||||
"reason": "用于确认是否存在连接未释放"
|
||||
}
|
||||
],
|
||||
"missing_info": ["缺少连接泄漏检测日志"]
|
||||
}
|
||||
""";
|
||||
|
||||
ChatService.ChatResult result = chatService.executeChatComplex(
|
||||
chatModel,
|
||||
new ToolCallback[0],
|
||||
"请分析 MySQL 连接池耗尽",
|
||||
List.of(),
|
||||
"sequential-v2-render-session"
|
||||
);
|
||||
|
||||
assertTrue(result.answer().contains("已确认信息"));
|
||||
assertTrue(result.answer().contains("连接池 active 达到上限"));
|
||||
assertTrue(result.answer().contains("可能方向"));
|
||||
assertTrue(result.answer().contains("建议下一步"));
|
||||
assertFalse(result.answer().contains("\"answer_version\""));
|
||||
assertFalse(result.answer().contains("executor_evidence_v2"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
|
||||
ChatService chatService = new ChatService();
|
||||
|
||||
Reference in New Issue
Block a user