docs(openspec): propose executor composer final answer

This commit is contained in:
aruo
2026-07-08 02:49:46 +08:00
parent 39c0c5f8be
commit a5b4502c72
8 changed files with 454 additions and 0 deletions
@@ -0,0 +1,80 @@
## ADDED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** ChatService invokes Composer
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer SHALL respect verdict-specific wording
Composer SHALL phrase final answers according to the effective Verifier verdict.
#### Scenario: PASS answer uses confirmed material
- **WHEN** the effective verdict is `PASS`
- **THEN** the final answer MAY state confirmed findings from `allowed_claims`
- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present
#### Scenario: LOW_CONFID answer separates findings and gaps
- **WHEN** the effective verdict is `LOW_CONFID`
- **THEN** the final answer SHALL distinguish confirmed information from possible directions
- **AND** it SHALL mention evidence gaps from `missing_info`
- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings
#### Scenario: REJECT answer avoids root-cause conclusions
- **WHEN** the effective verdict is `REJECT`
- **THEN** Composer input SHALL have `allowed_hypotheses=[]`
- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion
- **AND** the final answer SHALL NOT include a root-cause conclusion
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer returns malformed JSON or omits required fields
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** ChatService persists verifier evaluation
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
- **AND** the audit SHALL include the parsed Composer fields when valid
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,80 @@
## MODIFIED Requirements
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS -> Composer output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
#### Scenario: REJECT -> degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
#### Scenario: PASS uses Composer-safe output
- **WHEN** the final verdict is `PASS`
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- **AND** it SHALL NOT use Executor `user_facing_answer`
- **AND** it SHALL NOT expose raw Executor JSON
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved