5.6 KiB
Context
After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> ChatService final rendering
Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits claim_checks. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.
Stage four adds the final expression layer:
VerifierDecision + Executor structured output
-> ChatService filters allowed material
-> chat_composer
-> Composer JSON
-> final diagnosis_session.answer
Goals / Non-Goals
Goals:
- Add a Composer prompt and model call after Verifier.
- Ensure Composer receives only filtered material derived from Verifier decisions.
- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor
user_facing_answer, raw Executor JSON, or raw tool output. - Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
- Persist Composer output for audit under
verifier_evaluation.composer_output.
Non-Goals:
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier verification-class expansion.
- No database schema migration.
- No stage-five fixture expansion.
Decisions
Composer is an expression layer, not a diagnosis layer
Composer SHALL receive only filtered material:
original_queryverdictallowed_claimsallowed_hypothesesmissing_inforecommended_actionsrationale
It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.
Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.
ChatService owns filtering
ChatService filters Executor material using Verifier claim_checks before invoking Composer.
Suggested mapping:
| Verifier classification | Composer handling |
|---|---|
direct_observation |
include in allowed_claims as confirmed material |
reasonable_inference |
include in allowed_claims, but do not allow “唯一根因” wording unless claim type already supports root cause |
overstated |
do not include as confirmed; may become allowed_hypotheses or missing_info |
unsupported |
do not include as confirmed; may become missing_info |
external_unknown |
do not include as confirmed; may become missing_info |
contradicted |
do not include as confirmed; favor REJECT-safe output |
For REJECT, allowed_hypotheses SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.
Composer output is strict JSON with safe fallback
Composer SHALL output:
{
"answer_summary": "...",
"recommended_actions": [
{
"action_text": "...",
"reason": "..."
}
],
"user_facing_answer": "..."
}
If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor user_facing_answer.
Audit stays in self_evaluation
No new table is needed. ChatService persists a minimal Composer audit snapshot:
{
"verifier_evaluation": {
"composer_output": {
"status": "valid",
"answer_summary": "...",
"recommended_actions": [],
"user_facing_answer": "..."
}
}
}
When Composer fails, status should be malformed or fallback, with a short detail. The audit should remain compact and avoid storing full prompt copies.
Interface impact is L2 internal
The external Chat API still returns a final answer string. Internally, ChatService gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.
Risks / Trade-offs
- Risk: Composer introduces another LLM call and can fail formatting.
- Mitigation: strict JSON contract plus deterministic fallback templates.
- Risk: filtering is too strict and PASS answers become terse.
- Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
- Risk: legacy tests expect PASS to use Executor output.
- Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
- Risk: malformed Composer output could leak raw JSON.
- Mitigation: parse output before exposing it; fallback only from filtered material.
Migration Plan
- Add
chat-composer-prompt.md. - Add Composer input assembly and output parsing in
ChatService. - Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
- Persist
composer_outputunderverifier_evaluation. - Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.
Rollback:
- Keep fallback templates available if Composer is disabled or malformed.
- Do not roll back to Executor
user_facing_answer; that would reintroduce the original evidence-attribution risk.
Open Questions
None blocking. Stage four defaults:
- Composer is implemented as
chat_composer. - Composer does not call tools.
- Composer output failure falls back to fixed safe templates.
- No database schema changes.