## Context After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is: ```text chat_planner -> chat_executor -> VerifierInputHook + Gatekeeper -> chat_verifier -> ChatService final rendering ``` Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers. Stage four adds the final expression layer: ```text VerifierDecision + Executor structured output -> ChatService filters allowed material -> chat_composer -> Composer JSON -> final diagnosis_session.answer ``` ## Goals / Non-Goals **Goals:** - Add a Composer prompt and model call after Verifier. - Ensure Composer receives only filtered material derived from Verifier decisions. - Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output. - Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed. - Persist Composer output for audit under `verifier_evaluation.composer_output`. **Non-Goals:** - No Planner changes. - No Executor retry changes. - No Gatekeeper rule expansion. - No Verifier verification-class expansion. - No database schema migration. - No stage-five fixture expansion. ## Decisions ### Composer is an expression layer, not a diagnosis layer Composer SHALL receive only filtered material: - `original_query` - `verdict` - `allowed_claims` - `allowed_hypotheses` - `missing_info` - `recommended_actions` - `rationale` It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts. Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier. ### ChatService owns filtering ChatService filters Executor material using Verifier `claim_checks` before invoking Composer. Suggested mapping: | Verifier classification | Composer handling | |---|---| | `direct_observation` | include in `allowed_claims` as confirmed material | | `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause | | `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` | | `unsupported` | do not include as confirmed; may become `missing_info` | | `external_unknown` | do not include as confirmed; may become `missing_info` | | `contradicted` | do not include as confirmed; favor REJECT-safe output | For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain. ### Composer output is strict JSON with safe fallback Composer SHALL output: ```json { "answer_summary": "...", "recommended_actions": [ { "action_text": "...", "reason": "..." } ], "user_facing_answer": "..." } ``` If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`. ### Audit stays in self_evaluation No new table is needed. ChatService persists a minimal Composer audit snapshot: ```json { "verifier_evaluation": { "composer_output": { "status": "valid", "answer_summary": "...", "recommended_actions": [], "user_facing_answer": "..." } } } ``` When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies. ### Interface impact is L2 internal The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material. ## Risks / Trade-offs - Risk: Composer introduces another LLM call and can fail formatting. - Mitigation: strict JSON contract plus deterministic fallback templates. - Risk: filtering is too strict and PASS answers become terse. - Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints. - Risk: legacy tests expect PASS to use Executor output. - Mitigation: update tests to assert Composer or safe fallback is the only final-answer source. - Risk: malformed Composer output could leak raw JSON. - Mitigation: parse output before exposing it; fallback only from filtered material. ## Migration Plan 1. Add `chat-composer-prompt.md`. 2. Add Composer input assembly and output parsing in `ChatService`. 3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering. 4. Persist `composer_output` under `verifier_evaluation`. 5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback. Rollback: - Keep fallback templates available if Composer is disabled or malformed. - Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk. ## Open Questions None blocking. Stage four defaults: - Composer is implemented as `chat_composer`. - Composer does not call tools. - Composer output failure falls back to fixed safe templates. - No database schema changes.