Files
SuperBizAgent-java/openspec/changes/archive/2026-07-07-executor-v2-output-contract/design.md
T

3.9 KiB

Context

The prior executor_evidence_v1 contract made Executor responsible for both evidence attribution and final answer wording:

  • diagnosis_summary
  • user_facing_answer

That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.

This stage implements only the first slice of the V2 migration:

Executor V2 output contract
  -> existing VerifierInputHook parsing
  -> existing Verifier
  -> temporary ChatService structured renderer

Gatekeeper, Verifier V2 claim_checks, and Composer are later phases.

Goals / Non-Goals

Goals:

  • Make Chat Executor emit executor_evidence_v2.
  • Remove diagnosis_summary and user_facing_answer from Executor output.
  • Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
  • Preserve evidence-binding requirements for confirmed claims.
  • Prevent PASS routing from returning raw JSON to normal Chat users.

Non-goals:

  • No Gatekeeper implementation.
  • No Verifier prompt rewrite to claim_checks.
  • No Composer agent.
  • No database schema changes.
  • No Planner changes.
  • No evidence tool signature changes.
  • No retry behavior changes.

Executor V2 Contract

Executor final output SHALL be one JSON object:

{
  "answer_version": "executor_evidence_v2",
  "claims": [
    {
      "claim_id": "claim-1",
      "claim_type": "symptom",
      "claim_text": "payment-service 出现请求超时日志。",
      "support_level": "direct",
      "evidence_bindings": [
        {
          "source_type": "tool_trace",
          "source_id": "trace-1",
          "tool_name": "query_logs",
          "source_invocation_ids": [394],
          "evidence_excerpt": "request timeout"
        }
      ]
    }
  ],
  "hypotheses": [],
  "recommended_actions": [],
  "missing_info": []
}

Removed fields:

  • diagnosis_summary
  • user_facing_answer

hypotheses, recommended_actions, and missing_info SHOULD be present as arrays. They may be empty.

Temporary Rendering Strategy

Before Composer exists, ChatService needs a safe PASS fallback for V2 output.

When Verifier returns PASS:

  1. If Executor output has user_facing_answer, keep the existing V1 behavior.
  2. Else, if Executor output is executor_evidence_v2, render a readable Chinese answer from:
    • claims[].claim_text
    • hypotheses[].hypothesis_text
    • missing_info[]
    • recommended_actions[].action_text and reason
  3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.

The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.

Parser Boundary

VerifierInputHook may continue parsing raw Executor output into executor_structured_output when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.

Interface Impact

  • Internal Agent output contract: L4, because two fields are removed from Executor JSON.
  • Verifier payload: L2, because existing executor_final_answer remains raw text and executor_structured_output remains optional.
  • External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.

Risks / Mitigations

  • Risk: existing PASS path exposes raw JSON because user_facing_answer is gone.
    • Mitigation: add temporary V2 renderer in ChatService.
  • Risk: current Verifier prompt still mentions user_facing_answer.
    • Mitigation: stage one keeps Verifier behavior compatible; it should verify claims when structured output is valid and simply find no extra user_facing_answer.
  • Risk: tests assume V1 fields.
    • Mitigation: update/add focused tests for V2 output without final-expression fields.