diff --git a/openspec/changes/executor-composer-final-answer/.committed b/openspec/changes/executor-composer-final-answer/.committed new file mode 100644 index 0000000..be211ab --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/.committed @@ -0,0 +1,6 @@ +Committed OpenSpec for executor-composer-final-answer. + +Commit gate passed on 2026-07-08. + +Validated with: +- cmd /c openspec validate executor-composer-final-answer diff --git a/openspec/changes/executor-composer-final-answer/.openspec.yaml b/openspec/changes/executor-composer-final-answer/.openspec.yaml new file mode 100644 index 0000000..aee4ef1 --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-07-07 diff --git a/openspec/changes/executor-composer-final-answer/decisions.md b/openspec/changes/executor-composer-final-answer/decisions.md new file mode 100644 index 0000000..a01da78 --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/decisions.md @@ -0,0 +1,43 @@ +# Decisions + +## Context Collection + +- `devflow/index.md` confirms `executor-v2-output-contract`, `executor-gatekeeper-hook`, and `executor-verifier-claim-checks` are archived. +- `openspec/changes` has no active changes before this phase. +- `openspec/specs/chat-verifier-agent/spec.md` is the existing capability that owns verifier routing and audit behavior. +- No existing `chat-composer-agent` capability exists, so this change introduces it. + +## Question Pool + +| Question | Type | Resolution | +|---|---|---| +| Should Composer be a new capability or folded into `chat-verifier-agent`? | evidence-driven | New `chat-composer-agent` capability plus modified `chat-verifier-agent` routing. Composer is a distinct expression layer, while ChatService routing remains part of the verifier chain. | +| Can Composer call tools or inspect raw tool output? | evidence-driven | No. The issue and prior design require Composer to receive only Verifier-allowed material. | +| Should Gatekeeper, Verifier, or Planner change in this phase? | evidence-driven | No. Stage four is limited to Composer final-answer generation and routing. | +| What is the interface impact level? | evidence-driven | L2 internal contract change: ChatService internal final-answer semantics change, but no external API or database schema changes. | + +## Key Decisions + +- Composer is implemented as `chat_composer` or an equivalent model call after Verifier. +- Composer input is assembled by ChatService, not by the model. +- `claim_checks` are authoritative for filtering allowed material. +- REJECT input always has `allowed_hypotheses=[]`. +- Composer malformed output falls back to fixed safe templates. +- Fallback must never use Executor `user_facing_answer` or raw Executor JSON. +- Composer audit is persisted under `verifier_evaluation.composer_output`. + +## Cross-Artifact Alignment + +| Source | Alignment | +|---|---| +| Issue objective | Stage four in `mvp/issues/executor-structured-output-v2.md` requires Composer output final answer from Verifier-allowed material. Covered by proposal, design, specs, and tasks. | +| Proposal -> design | Proposal says Composer owns final expression; design defines input filtering, output parsing, fallback, and audit. | +| Design -> specs | Design decisions are reflected in `chat-composer-agent` requirements and modified `chat-verifier-agent` routing requirements. | +| Specs -> tasks | Each required behavior has implementation and test tasks, including malformed fallback and no raw output leakage. | + +## Commit Gate Notes + +- Interface impact: L2 internal. +- No unresolved user-interview question identified. +- No database migration required. +- No OpenSpec/devflow conflict found. diff --git a/openspec/changes/executor-composer-final-answer/design.md b/openspec/changes/executor-composer-final-answer/design.md new file mode 100644 index 0000000..0af5017 --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/design.md @@ -0,0 +1,152 @@ +## Context + +After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is: + +```text +chat_planner + -> chat_executor + -> VerifierInputHook + Gatekeeper + -> chat_verifier + -> ChatService final rendering +``` + +Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers. + +Stage four adds the final expression layer: + +```text +VerifierDecision + Executor structured output + -> ChatService filters allowed material + -> chat_composer + -> Composer JSON + -> final diagnosis_session.answer +``` + +## Goals / Non-Goals + +**Goals:** + +- Add a Composer prompt and model call after Verifier. +- Ensure Composer receives only filtered material derived from Verifier decisions. +- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output. +- Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed. +- Persist Composer output for audit under `verifier_evaluation.composer_output`. + +**Non-Goals:** + +- No Planner changes. +- No Executor retry changes. +- No Gatekeeper rule expansion. +- No Verifier verification-class expansion. +- No database schema migration. +- No stage-five fixture expansion. + +## Decisions + +### Composer is an expression layer, not a diagnosis layer + +Composer SHALL receive only filtered material: + +- `original_query` +- `verdict` +- `allowed_claims` +- `allowed_hypotheses` +- `missing_info` +- `recommended_actions` +- `rationale` + +It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts. + +Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier. + +### ChatService owns filtering + +ChatService filters Executor material using Verifier `claim_checks` before invoking Composer. + +Suggested mapping: + +| Verifier classification | Composer handling | +|---|---| +| `direct_observation` | include in `allowed_claims` as confirmed material | +| `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause | +| `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` | +| `unsupported` | do not include as confirmed; may become `missing_info` | +| `external_unknown` | do not include as confirmed; may become `missing_info` | +| `contradicted` | do not include as confirmed; favor REJECT-safe output | + +For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain. + +### Composer output is strict JSON with safe fallback + +Composer SHALL output: + +```json +{ + "answer_summary": "...", + "recommended_actions": [ + { + "action_text": "...", + "reason": "..." + } + ], + "user_facing_answer": "..." +} +``` + +If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`. + +### Audit stays in self_evaluation + +No new table is needed. ChatService persists a minimal Composer audit snapshot: + +```json +{ + "verifier_evaluation": { + "composer_output": { + "status": "valid", + "answer_summary": "...", + "recommended_actions": [], + "user_facing_answer": "..." + } + } +} +``` + +When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies. + +### Interface impact is L2 internal + +The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material. + +## Risks / Trade-offs + +- Risk: Composer introduces another LLM call and can fail formatting. + - Mitigation: strict JSON contract plus deterministic fallback templates. +- Risk: filtering is too strict and PASS answers become terse. + - Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints. +- Risk: legacy tests expect PASS to use Executor output. + - Mitigation: update tests to assert Composer or safe fallback is the only final-answer source. +- Risk: malformed Composer output could leak raw JSON. + - Mitigation: parse output before exposing it; fallback only from filtered material. + +## Migration Plan + +1. Add `chat-composer-prompt.md`. +2. Add Composer input assembly and output parsing in `ChatService`. +3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering. +4. Persist `composer_output` under `verifier_evaluation`. +5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback. + +Rollback: + +- Keep fallback templates available if Composer is disabled or malformed. +- Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk. + +## Open Questions + +None blocking. Stage four defaults: + +- Composer is implemented as `chat_composer`. +- Composer does not call tools. +- Composer output failure falls back to fixed safe templates. +- No database schema changes. diff --git a/openspec/changes/executor-composer-final-answer/proposal.md b/openspec/changes/executor-composer-final-answer/proposal.md new file mode 100644 index 0000000..6dc9711 --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/proposal.md @@ -0,0 +1,47 @@ +## Why + +Stage one removed final-answer fields from Executor, stage two added deterministic Gatekeeper checks, and stage three made Verifier claim-oriented. The remaining gap is final answer rendering: ChatService can still rely on temporary V2 rendering paths instead of a dedicated expression layer, which risks letting unverified Executor material shape user-facing answers. + +This phase introduces a Composer expression layer so final Chat answers are generated only from Verifier-allowed material. + +## What Changes + +- Add a `chat_composer` final-answer generation step after Verifier. +- Add a strict Composer prompt and JSON output contract with `answer_summary`, `recommended_actions`, and `user_facing_answer`. +- Build Composer input in `ChatService` from filtered Verifier results: + - `allowed_claims` + - `allowed_hypotheses` + - `missing_info` + - `recommended_actions` + - `rationale` +- Ensure PASS, LOW_CONFID, and REJECT final answers no longer read Executor `user_facing_answer` or raw Executor JSON. +- Persist Composer output in `diagnosis_session.self_evaluation.verifier_evaluation.composer_output`. +- Add safe fallback templates for malformed Composer output that never expose raw Executor JSON or raw Composer JSON. + +## Capabilities + +### New Capabilities + +- `chat-composer-agent`: Final expression layer that turns Verifier-allowed structured material into a readable Chinese user answer without introducing new facts. + +### Modified Capabilities + +- `chat-verifier-agent`: ChatService final routing changes so Verifier verdicts feed Composer or safe fixed templates instead of direct Executor answer paths. + +## Impact + +- Affected prompt: add `src/main/resources/prompts/chat-composer-prompt.md`. +- Affected service: `ChatService` final rendering after Verifier, Composer input assembly, Composer output parsing, malformed-output fallback. +- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `composer_output`. +- Affected tests: `ChatServiceSequentialAgentTest` and focused tests around Composer input filtering, final answer routing, and malformed Composer fallback. +- Database schema: no table or column change. +- External API: no endpoint contract change; final answer text semantics become stricter because unverified Executor material is no longer a source for user-visible answers. + +## Non-Goals + +- No Planner changes. +- No Gatekeeper rule expansion. +- No Verifier classification changes. +- No Executor retry behavior change. +- No database migration. +- No stage-five eval fixture expansion in this phase. diff --git a/openspec/changes/executor-composer-final-answer/specs/chat-composer-agent/spec.md b/openspec/changes/executor-composer-final-answer/specs/chat-composer-agent/spec.md new file mode 100644 index 0000000..e75da3d --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/specs/chat-composer-agent/spec.md @@ -0,0 +1,80 @@ +## ADDED Requirements + +### Requirement: Composer SHALL generate final user-facing Chat answers +The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material. + +#### Scenario: Composer receives only filtered material +- **WHEN** ChatService invokes Composer +- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale` +- **AND** the Composer input SHALL NOT contain raw tool output +- **AND** the Composer input SHALL NOT contain the full unscreened Executor output +- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer` + +#### Scenario: Composer outputs strict JSON +- **WHEN** Composer completes +- **THEN** it SHALL output exactly one JSON object +- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer` +- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object + +#### Scenario: Composer does not introduce new facts +- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer` +- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input +- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text + +### Requirement: Composer input SHALL honor Verifier claim checks +ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`. + +#### Scenario: Passing claims become allowed claims +- **WHEN** a claim check verification is `direct_observation` +- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims` + +#### Scenario: Reasonable inferences remain bounded +- **WHEN** a claim check verification is `reasonable_inference` +- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims` +- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS` + +#### Scenario: Overstated claims are not confirmed findings +- **WHEN** a claim check verification is `overstated` +- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims` +- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info` + +#### Scenario: Unsupported or external claims are withheld +- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted` +- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims` +- **AND** the final user-facing answer SHALL NOT present that claim as confirmed + +### Requirement: Composer SHALL respect verdict-specific wording +Composer SHALL phrase final answers according to the effective Verifier verdict. + +#### Scenario: PASS answer uses confirmed material +- **WHEN** the effective verdict is `PASS` +- **THEN** the final answer MAY state confirmed findings from `allowed_claims` +- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present + +#### Scenario: LOW_CONFID answer separates findings and gaps +- **WHEN** the effective verdict is `LOW_CONFID` +- **THEN** the final answer SHALL distinguish confirmed information from possible directions +- **AND** it SHALL mention evidence gaps from `missing_info` +- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings + +#### Scenario: REJECT answer avoids root-cause conclusions +- **WHEN** the effective verdict is `REJECT` +- **THEN** Composer input SHALL have `allowed_hypotheses=[]` +- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion +- **AND** the final answer SHALL NOT include a root-cause conclusion + +### Requirement: Composer failures SHALL degrade safely +The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material. + +#### Scenario: malformed Composer output falls back safely +- **WHEN** Composer returns malformed JSON or omits required fields +- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material +- **AND** the final answer SHALL NOT expose raw Composer output +- **AND** the final answer SHALL NOT expose raw Executor output +- **AND** the final answer SHALL NOT use Executor `user_facing_answer` + +#### Scenario: Composer audit is persisted +- **WHEN** ChatService persists verifier evaluation +- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used +- **AND** the audit SHALL include the parsed Composer fields when valid +- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output diff --git a/openspec/changes/executor-composer-final-answer/specs/chat-verifier-agent/spec.md b/openspec/changes/executor-composer-final-answer/specs/chat-verifier-agent/spec.md new file mode 100644 index 0000000..224499a --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/specs/chat-verifier-agent/spec.md @@ -0,0 +1,80 @@ +## MODIFIED Requirements + +### Requirement: ChatService SHALL route based on Verifier verdict +The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed. + +#### Scenario: PASS -> Composer output +- **WHEN** Verifier outputs verdict="PASS" +- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template +- **AND** the final user-facing answer SHALL NOT pass through raw Executor output +- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer` + +#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty +- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5 +- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template +- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps +- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions + +#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round +- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback +- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence +- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template +- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output + +#### Scenario: REJECT does not enter retry round +- **WHEN** Verifier outputs verdict="REJECT" +- **THEN** the system SHALL NOT start a retry round for evidence supplementation +- **AND** it SHALL produce a degraded output directly through Composer-safe rendering + +#### Scenario: REJECT -> degraded output +- **WHEN** Verifier outputs verdict="REJECT" +- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated +- **AND** it SHALL NOT pass through the raw Executor answer +- **AND** it SHALL NOT include a root-cause conclusion + +### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols +The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses. + +#### Scenario: PASS uses Composer-safe output +- **WHEN** the final verdict is `PASS` +- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template +- **AND** it SHALL NOT use Executor `user_facing_answer` +- **AND** it SHALL NOT expose raw Executor JSON + +#### Scenario: LOW_CONFID uses Composer-safe uncertainty output +- **WHEN** the final verdict is `LOW_CONFID` +- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material +- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps +- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions + +#### Scenario: REJECT uses degraded template +- **WHEN** the final verdict is `REJECT` +- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output +- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions +- **AND** it SHALL NOT include unverified raw answer content +- **AND** it SHALL NOT include a root-cause conclusion + +### Requirement: Verifier SHALL be observable +The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability. + +#### Scenario: claim checks written to self_evaluation +- **WHEN** the Verifier evaluation is persisted +- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks` +- **AND** it SHALL continue to include compatibility `facts_checked` +- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved + +#### Scenario: composer output written to self_evaluation +- **WHEN** final answer composition completes +- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output` +- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used +- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved + +#### Scenario: verdict written to self_evaluation +- **WHEN** the Verifier produces a verdict +- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation` +- **AND** existing `rule_evaluation` data SHALL be preserved + +#### Scenario: gatekeeper result written to self_evaluation +- **WHEN** the Verifier evaluation is persisted +- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result` +- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved diff --git a/openspec/changes/executor-composer-final-answer/tasks.md b/openspec/changes/executor-composer-final-answer/tasks.md new file mode 100644 index 0000000..798cd6d --- /dev/null +++ b/openspec/changes/executor-composer-final-answer/tasks.md @@ -0,0 +1,44 @@ +## 1. Composer Prompt Contract + +- [ ] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`. +- [ ] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`. +- [ ] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`. +- [ ] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT. +- [ ] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output. + +## 2. Composer Invocation And Input Filtering + +- [ ] 2.1 Load the Composer prompt in `ChatService`. +- [ ] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision. +- [ ] 2.3 Build Composer input from Verifier decision and Executor structured output. +- [ ] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`. +- [ ] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output. +- [ ] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`. +- [ ] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`. +- [ ] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`. + +## 3. Composer Output Parsing, Fallback, And Audit + +- [ ] 3.1 Parse Composer strict JSON output. +- [ ] 3.2 Add safe fallback rendering for malformed Composer output. +- [ ] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`. +- [ ] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`. +- [ ] 3.5 Preserve existing verifier audit fields when writing Composer output. + +## 4. Final Answer Routing + +- [ ] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering. +- [ ] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material. +- [ ] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer. +- [ ] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path. + +## 5. Tests And Verification + +- [ ] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output. +- [ ] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content. +- [ ] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists. +- [ ] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps. +- [ ] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion. +- [ ] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage. +- [ ] 5.7 Run targeted Maven tests. +- [ ] 5.8 Validate this OpenSpec change and all specs.