docs(openspec): propose executor composer final answer
This commit is contained in:
@@ -0,0 +1,152 @@
|
||||
## Context
|
||||
|
||||
After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook + Gatekeeper
|
||||
-> chat_verifier
|
||||
-> ChatService final rendering
|
||||
```
|
||||
|
||||
Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.
|
||||
|
||||
Stage four adds the final expression layer:
|
||||
|
||||
```text
|
||||
VerifierDecision + Executor structured output
|
||||
-> ChatService filters allowed material
|
||||
-> chat_composer
|
||||
-> Composer JSON
|
||||
-> final diagnosis_session.answer
|
||||
```
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add a Composer prompt and model call after Verifier.
|
||||
- Ensure Composer receives only filtered material derived from Verifier decisions.
|
||||
- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output.
|
||||
- Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
|
||||
- Persist Composer output for audit under `verifier_evaluation.composer_output`.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- No Planner changes.
|
||||
- No Executor retry changes.
|
||||
- No Gatekeeper rule expansion.
|
||||
- No Verifier verification-class expansion.
|
||||
- No database schema migration.
|
||||
- No stage-five fixture expansion.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Composer is an expression layer, not a diagnosis layer
|
||||
|
||||
Composer SHALL receive only filtered material:
|
||||
|
||||
- `original_query`
|
||||
- `verdict`
|
||||
- `allowed_claims`
|
||||
- `allowed_hypotheses`
|
||||
- `missing_info`
|
||||
- `recommended_actions`
|
||||
- `rationale`
|
||||
|
||||
It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.
|
||||
|
||||
Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.
|
||||
|
||||
### ChatService owns filtering
|
||||
|
||||
ChatService filters Executor material using Verifier `claim_checks` before invoking Composer.
|
||||
|
||||
Suggested mapping:
|
||||
|
||||
| Verifier classification | Composer handling |
|
||||
|---|---|
|
||||
| `direct_observation` | include in `allowed_claims` as confirmed material |
|
||||
| `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause |
|
||||
| `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` |
|
||||
| `unsupported` | do not include as confirmed; may become `missing_info` |
|
||||
| `external_unknown` | do not include as confirmed; may become `missing_info` |
|
||||
| `contradicted` | do not include as confirmed; favor REJECT-safe output |
|
||||
|
||||
For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.
|
||||
|
||||
### Composer output is strict JSON with safe fallback
|
||||
|
||||
Composer SHALL output:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_summary": "...",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "...",
|
||||
"reason": "..."
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "..."
|
||||
}
|
||||
```
|
||||
|
||||
If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||
|
||||
### Audit stays in self_evaluation
|
||||
|
||||
No new table is needed. ChatService persists a minimal Composer audit snapshot:
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "...",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "..."
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies.
|
||||
|
||||
### Interface impact is L2 internal
|
||||
|
||||
The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Risk: Composer introduces another LLM call and can fail formatting.
|
||||
- Mitigation: strict JSON contract plus deterministic fallback templates.
|
||||
- Risk: filtering is too strict and PASS answers become terse.
|
||||
- Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
|
||||
- Risk: legacy tests expect PASS to use Executor output.
|
||||
- Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
|
||||
- Risk: malformed Composer output could leak raw JSON.
|
||||
- Mitigation: parse output before exposing it; fallback only from filtered material.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add `chat-composer-prompt.md`.
|
||||
2. Add Composer input assembly and output parsing in `ChatService`.
|
||||
3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
|
||||
4. Persist `composer_output` under `verifier_evaluation`.
|
||||
5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.
|
||||
|
||||
Rollback:
|
||||
|
||||
- Keep fallback templates available if Composer is disabled or malformed.
|
||||
- Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None blocking. Stage four defaults:
|
||||
|
||||
- Composer is implemented as `chat_composer`.
|
||||
- Composer does not call tools.
|
||||
- Composer output failure falls back to fixed safe templates.
|
||||
- No database schema changes.
|
||||
Reference in New Issue
Block a user