153 lines
5.6 KiB
Markdown
153 lines
5.6 KiB
Markdown
## Context
|
|
|
|
After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:
|
|
|
|
```text
|
|
chat_planner
|
|
-> chat_executor
|
|
-> VerifierInputHook + Gatekeeper
|
|
-> chat_verifier
|
|
-> ChatService final rendering
|
|
```
|
|
|
|
Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.
|
|
|
|
Stage four adds the final expression layer:
|
|
|
|
```text
|
|
VerifierDecision + Executor structured output
|
|
-> ChatService filters allowed material
|
|
-> chat_composer
|
|
-> Composer JSON
|
|
-> final diagnosis_session.answer
|
|
```
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Add a Composer prompt and model call after Verifier.
|
|
- Ensure Composer receives only filtered material derived from Verifier decisions.
|
|
- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output.
|
|
- Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
|
|
- Persist Composer output for audit under `verifier_evaluation.composer_output`.
|
|
|
|
**Non-Goals:**
|
|
|
|
- No Planner changes.
|
|
- No Executor retry changes.
|
|
- No Gatekeeper rule expansion.
|
|
- No Verifier verification-class expansion.
|
|
- No database schema migration.
|
|
- No stage-five fixture expansion.
|
|
|
|
## Decisions
|
|
|
|
### Composer is an expression layer, not a diagnosis layer
|
|
|
|
Composer SHALL receive only filtered material:
|
|
|
|
- `original_query`
|
|
- `verdict`
|
|
- `allowed_claims`
|
|
- `allowed_hypotheses`
|
|
- `missing_info`
|
|
- `recommended_actions`
|
|
- `rationale`
|
|
|
|
It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.
|
|
|
|
Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.
|
|
|
|
### ChatService owns filtering
|
|
|
|
ChatService filters Executor material using Verifier `claim_checks` before invoking Composer.
|
|
|
|
Suggested mapping:
|
|
|
|
| Verifier classification | Composer handling |
|
|
|---|---|
|
|
| `direct_observation` | include in `allowed_claims` as confirmed material |
|
|
| `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause |
|
|
| `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` |
|
|
| `unsupported` | do not include as confirmed; may become `missing_info` |
|
|
| `external_unknown` | do not include as confirmed; may become `missing_info` |
|
|
| `contradicted` | do not include as confirmed; favor REJECT-safe output |
|
|
|
|
For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.
|
|
|
|
### Composer output is strict JSON with safe fallback
|
|
|
|
Composer SHALL output:
|
|
|
|
```json
|
|
{
|
|
"answer_summary": "...",
|
|
"recommended_actions": [
|
|
{
|
|
"action_text": "...",
|
|
"reason": "..."
|
|
}
|
|
],
|
|
"user_facing_answer": "..."
|
|
}
|
|
```
|
|
|
|
If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
|
|
|
### Audit stays in self_evaluation
|
|
|
|
No new table is needed. ChatService persists a minimal Composer audit snapshot:
|
|
|
|
```json
|
|
{
|
|
"verifier_evaluation": {
|
|
"composer_output": {
|
|
"status": "valid",
|
|
"answer_summary": "...",
|
|
"recommended_actions": [],
|
|
"user_facing_answer": "..."
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies.
|
|
|
|
### Interface impact is L2 internal
|
|
|
|
The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- Risk: Composer introduces another LLM call and can fail formatting.
|
|
- Mitigation: strict JSON contract plus deterministic fallback templates.
|
|
- Risk: filtering is too strict and PASS answers become terse.
|
|
- Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
|
|
- Risk: legacy tests expect PASS to use Executor output.
|
|
- Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
|
|
- Risk: malformed Composer output could leak raw JSON.
|
|
- Mitigation: parse output before exposing it; fallback only from filtered material.
|
|
|
|
## Migration Plan
|
|
|
|
1. Add `chat-composer-prompt.md`.
|
|
2. Add Composer input assembly and output parsing in `ChatService`.
|
|
3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
|
|
4. Persist `composer_output` under `verifier_evaluation`.
|
|
5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.
|
|
|
|
Rollback:
|
|
|
|
- Keep fallback templates available if Composer is disabled or malformed.
|
|
- Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk.
|
|
|
|
## Open Questions
|
|
|
|
None blocking. Stage four defaults:
|
|
|
|
- Composer is implemented as `chat_composer`.
|
|
- Composer does not call tools.
|
|
- Composer output failure falls back to fixed safe templates.
|
|
- No database schema changes.
|