Files
SuperBizAgent-java/openspec/changes/archive/2026-07-07-executor-v2-output-contract/design.md
T

113 lines
3.9 KiB
Markdown

## Context
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
- `diagnosis_summary`
- `user_facing_answer`
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
This stage implements only the first slice of the V2 migration:
```text
Executor V2 output contract
-> existing VerifierInputHook parsing
-> existing Verifier
-> temporary ChatService structured renderer
```
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
## Goals / Non-Goals
Goals:
- Make Chat Executor emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
- Preserve evidence-binding requirements for confirmed claims.
- Prevent PASS routing from returning raw JSON to normal Chat users.
Non-goals:
- No Gatekeeper implementation.
- No Verifier prompt rewrite to `claim_checks`.
- No Composer agent.
- No database schema changes.
- No Planner changes.
- No evidence tool signature changes.
- No retry behavior changes.
## Executor V2 Contract
Executor final output SHALL be one JSON object:
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "payment-service 出现请求超时日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_logs",
"source_invocation_ids": [394],
"evidence_excerpt": "request timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
Removed fields:
- `diagnosis_summary`
- `user_facing_answer`
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
## Temporary Rendering Strategy
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
When Verifier returns `PASS`:
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
- `claims[].claim_text`
- `hypotheses[].hypothesis_text`
- `missing_info[]`
- `recommended_actions[].action_text` and `reason`
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
## Parser Boundary
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
## Interface Impact
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
## Risks / Mitigations
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
- Mitigation: add temporary V2 renderer in `ChatService`.
- Risk: current Verifier prompt still mentions `user_facing_answer`.
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
- Risk: tests assume V1 fields.
- Mitigation: update/add focused tests for V2 output without final-expression fields.