Files

5.7 KiB

Decisions: executor-v2-output-contract

sm-flow Progress

Clarify

Entry summary: implement stage one of Executor Structured Output V2: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.

Slug: executor-v2-output-contract

Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.

Context

Relevant devflow history:

  • chat-verifier-agent: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
  • evidence-trace-hardening: evidence-bearing tool traces are persisted and summarized through ToolTraceSummaryService.
  • executor-evidence-output-contract: V1 introduced executor_evidence_v1 with diagnosis_summary, structured claims, and user_facing_answer.

Conflict with historical decision:

  • Previous executor-evidence-output-contract deliberately kept user_facing_answer in Executor output.
  • New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
  • For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor user_facing_answer.

Current code shape:

  • chat-executor-prompt.md defines the V1 Executor output contract.
  • VerifierInputHook parses Executor JSON and sets executor_structured_output.
  • ChatService.extractUserFacingAnswer(...) currently reads user_facing_answer on PASS.
  • If no replacement is added, PASS may expose raw Executor JSON after V2 removes user_facing_answer.

Grill

Question pool:

Question Mode Resolution
Does stage one include Gatekeeper? evidence-driven No. The issue splits Gatekeeper into stage two.
Does stage one change Planner? evidence-driven No. Planner is explicitly out of scope.
Can user_facing_answer remain temporarily in Executor? evidence-driven No. The V2 design requires removing it in stage one.
How do users get readable output before Composer exists? evidence-driven ChatService must use a temporary structured renderer for V2 PASS output.
Is the internal Agent contract breaking? evidence-driven Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable.

No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.

Specify

OpenSpec artifacts:

  • proposal.md: scope and compatibility boundary for stage one.
  • design.md: V2 Executor contract and temporary rendering strategy.
  • specs/chat-verifier-agent/spec.md: delta requirements for the Executor contract.
  • tasks.md: executable implementation and verification checklist.

Audit

Architecture risk summary:

  • The first-stage change deliberately breaks the internal Executor JSON contract by removing diagnosis_summary and user_facing_answer.
  • External Chat answers must remain readable Chinese, so ChatService needs a temporary V2 renderer before Composer exists.
  • VerifierInputHook should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
  • No database schema or evidence tool signature changes are required.

Cross-artifact alignment:

Source Target Status
issue background / stage one proposal aligned
proposal scope / non-goals design aligned
design contract and rendering bridge specs aligned
specs observable behavior tasks aligned

Interface impact:

  • Internal Agent output contract: L4, because diagnosis_summary and user_facing_answer are removed.
  • Verifier payload: L2, because executor_final_answer remains raw text and executor_structured_output remains optional.
  • External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.

Commit

Commit gate result: passed.

  • proposal.md exists and explains why this phase is needed.
  • design.md records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
  • specs/chat-verifier-agent/spec.md expresses observable behavior for Executor V2 and user-facing rendering safety.
  • tasks.md contains executable implementation and verification tasks.
  • cmd /c openspec validate executor-v2-output-contract passed.
  • No unresolved user-interview questions remain for this stage.

Apply

Implementation summary:

  • Updated chat-executor-prompt.md to require answer_version="executor_evidence_v2".
  • Removed diagnosis_summary and user_facing_answer from the Executor final output schema and output validation rules.
  • Added a temporary ChatService structured renderer for PASS + executor_evidence_v2 so normal users receive readable Chinese instead of raw JSON.
  • Preserved V1 user_facing_answer extraction for compatibility.
  • Kept VerifierInputHook parse-only behavior compatible with V2 output.
  • Adjusted chat-verifier-prompt.md wording so user_facing_answer is treated as a compatibility field, not a V2 required field.

Verification:

  • mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test passed.
  • cmd /c openspec validate executor-v2-output-contract passed.

Known limitations:

  • Gatekeeper is not implemented in this phase.
  • Verifier still outputs facts_checked; claim_checks belongs to a later phase.
  • The V2 renderer is temporary and should be replaced by Composer in a later phase.