Files
SuperBizAgent-java/openspec/changes/executor-composer-final-answer/design.md
T

5.6 KiB

Context

After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:

chat_planner
  -> chat_executor
  -> VerifierInputHook + Gatekeeper
  -> chat_verifier
  -> ChatService final rendering

Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits claim_checks. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.

Stage four adds the final expression layer:

VerifierDecision + Executor structured output
  -> ChatService filters allowed material
  -> chat_composer
  -> Composer JSON
  -> final diagnosis_session.answer

Goals / Non-Goals

Goals:

  • Add a Composer prompt and model call after Verifier.
  • Ensure Composer receives only filtered material derived from Verifier decisions.
  • Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor user_facing_answer, raw Executor JSON, or raw tool output.
  • Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
  • Persist Composer output for audit under verifier_evaluation.composer_output.

Non-Goals:

  • No Planner changes.
  • No Executor retry changes.
  • No Gatekeeper rule expansion.
  • No Verifier verification-class expansion.
  • No database schema migration.
  • No stage-five fixture expansion.

Decisions

Composer is an expression layer, not a diagnosis layer

Composer SHALL receive only filtered material:

  • original_query
  • verdict
  • allowed_claims
  • allowed_hypotheses
  • missing_info
  • recommended_actions
  • rationale

It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.

Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.

ChatService owns filtering

ChatService filters Executor material using Verifier claim_checks before invoking Composer.

Suggested mapping:

Verifier classification Composer handling
direct_observation include in allowed_claims as confirmed material
reasonable_inference include in allowed_claims, but do not allow “唯一根因” wording unless claim type already supports root cause
overstated do not include as confirmed; may become allowed_hypotheses or missing_info
unsupported do not include as confirmed; may become missing_info
external_unknown do not include as confirmed; may become missing_info
contradicted do not include as confirmed; favor REJECT-safe output

For REJECT, allowed_hypotheses SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.

Composer output is strict JSON with safe fallback

Composer SHALL output:

{
  "answer_summary": "...",
  "recommended_actions": [
    {
      "action_text": "...",
      "reason": "..."
    }
  ],
  "user_facing_answer": "..."
}

If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor user_facing_answer.

Audit stays in self_evaluation

No new table is needed. ChatService persists a minimal Composer audit snapshot:

{
  "verifier_evaluation": {
    "composer_output": {
      "status": "valid",
      "answer_summary": "...",
      "recommended_actions": [],
      "user_facing_answer": "..."
    }
  }
}

When Composer fails, status should be malformed or fallback, with a short detail. The audit should remain compact and avoid storing full prompt copies.

Interface impact is L2 internal

The external Chat API still returns a final answer string. Internally, ChatService gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.

Risks / Trade-offs

  • Risk: Composer introduces another LLM call and can fail formatting.
    • Mitigation: strict JSON contract plus deterministic fallback templates.
  • Risk: filtering is too strict and PASS answers become terse.
    • Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
  • Risk: legacy tests expect PASS to use Executor output.
    • Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
  • Risk: malformed Composer output could leak raw JSON.
    • Mitigation: parse output before exposing it; fallback only from filtered material.

Migration Plan

  1. Add chat-composer-prompt.md.
  2. Add Composer input assembly and output parsing in ChatService.
  3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
  4. Persist composer_output under verifier_evaluation.
  5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.

Rollback:

  • Keep fallback templates available if Composer is disabled or malformed.
  • Do not roll back to Executor user_facing_answer; that would reintroduce the original evidence-attribution risk.

Open Questions

None blocking. Stage four defaults:

  • Composer is implemented as chat_composer.
  • Composer does not call tools.
  • Composer output failure falls back to fixed safe templates.
  • No database schema changes.