5.7 KiB
Decisions: executor-v2-output-contract
sm-flow Progress
Clarify
Entry summary: implement stage one of Executor Structured Output V2: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
Slug: executor-v2-output-contract
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
Context
Relevant devflow history:
chat-verifier-agent: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.evidence-trace-hardening: evidence-bearing tool traces are persisted and summarized throughToolTraceSummaryService.executor-evidence-output-contract: V1 introducedexecutor_evidence_v1withdiagnosis_summary, structured claims, anduser_facing_answer.
Conflict with historical decision:
- Previous
executor-evidence-output-contractdeliberately keptuser_facing_answerin Executor output. - New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor
user_facing_answer.
Current code shape:
chat-executor-prompt.mddefines the V1 Executor output contract.VerifierInputHookparses Executor JSON and setsexecutor_structured_output.ChatService.extractUserFacingAnswer(...)currently readsuser_facing_answeron PASS.- If no replacement is added, PASS may expose raw Executor JSON after V2 removes
user_facing_answer.
Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
Can user_facing_answer remain temporarily in Executor? |
evidence-driven | No. The V2 design requires removing it in stage one. |
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
Specify
OpenSpec artifacts:
proposal.md: scope and compatibility boundary for stage one.design.md: V2 Executor contract and temporary rendering strategy.specs/chat-verifier-agent/spec.md: delta requirements for the Executor contract.tasks.md: executable implementation and verification checklist.
Audit
Architecture risk summary:
- The first-stage change deliberately breaks the internal Executor JSON contract by removing
diagnosis_summaryanduser_facing_answer. - External Chat answers must remain readable Chinese, so
ChatServiceneeds a temporary V2 renderer before Composer exists. VerifierInputHookshould remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.- No database schema or evidence tool signature changes are required.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue background / stage one | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design contract and rendering bridge | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Internal Agent output contract: L4, because
diagnosis_summaryanduser_facing_answerare removed. - Verifier payload: L2, because
executor_final_answerremains raw text andexecutor_structured_outputremains optional. - External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
Commit
Commit gate result: passed.
proposal.mdexists and explains why this phase is needed.design.mdrecords the V2 contract, temporary rendering strategy, non-goals, and interface impact.specs/chat-verifier-agent/spec.mdexpresses observable behavior for Executor V2 and user-facing rendering safety.tasks.mdcontains executable implementation and verification tasks.cmd /c openspec validate executor-v2-output-contractpassed.- No unresolved user-interview questions remain for this stage.
Apply
Implementation summary:
- Updated
chat-executor-prompt.mdto requireanswer_version="executor_evidence_v2". - Removed
diagnosis_summaryanduser_facing_answerfrom the Executor final output schema and output validation rules. - Added a temporary
ChatServicestructured renderer for PASS +executor_evidence_v2so normal users receive readable Chinese instead of raw JSON. - Preserved V1
user_facing_answerextraction for compatibility. - Kept
VerifierInputHookparse-only behavior compatible with V2 output. - Adjusted
chat-verifier-prompt.mdwording souser_facing_answeris treated as a compatibility field, not a V2 required field.
Verification:
mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" testpassed.cmd /c openspec validate executor-v2-output-contractpassed.
Known limitations:
- Gatekeeper is not implemented in this phase.
- Verifier still outputs
facts_checked;claim_checksbelongs to a later phase. - The V2 renderer is temporary and should be replaced by Composer in a later phase.