Files
SuperBizAgent-java/openspec/changes/archive/2026-07-07-executor-v2-output-contract/decisions.md
T

113 lines
5.7 KiB
Markdown

# Decisions: executor-v2-output-contract
## sm-flow Progress
### Clarify
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
Slug: `executor-v2-output-contract`
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
### Context
Relevant devflow history:
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
Conflict with historical decision:
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
Current code shape:
- `chat-executor-prompt.md` defines the V1 Executor output contract.
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
### Specify
OpenSpec artifacts:
- `proposal.md`: scope and compatibility boundary for stage one.
- `design.md`: V2 Executor contract and temporary rendering strategy.
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
- No database schema or evidence tool signature changes are required.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue background / stage one | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design contract and rendering bridge | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
### Commit
Commit gate result: passed.
- `proposal.md` exists and explains why this phase is needed.
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
- `tasks.md` contains executable implementation and verification tasks.
- `cmd /c openspec validate executor-v2-output-contract` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation summary:
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
- Preserved V1 `user_facing_answer` extraction for compatibility.
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
Verification:
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-v2-output-contract` passed.
Known limitations:
- Gatekeeper is not implemented in this phase.
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
- The V2 renderer is temporary and should be replaced by Composer in a later phase.