docs(openspec): propose executor composer final answer
This commit is contained in:
@@ -0,0 +1,6 @@
|
|||||||
|
Committed OpenSpec for executor-composer-final-answer.
|
||||||
|
|
||||||
|
Commit gate passed on 2026-07-08.
|
||||||
|
|
||||||
|
Validated with:
|
||||||
|
- cmd /c openspec validate executor-composer-final-answer
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-07
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
# Decisions
|
||||||
|
|
||||||
|
## Context Collection
|
||||||
|
|
||||||
|
- `devflow/index.md` confirms `executor-v2-output-contract`, `executor-gatekeeper-hook`, and `executor-verifier-claim-checks` are archived.
|
||||||
|
- `openspec/changes` has no active changes before this phase.
|
||||||
|
- `openspec/specs/chat-verifier-agent/spec.md` is the existing capability that owns verifier routing and audit behavior.
|
||||||
|
- No existing `chat-composer-agent` capability exists, so this change introduces it.
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| Question | Type | Resolution |
|
||||||
|
|---|---|---|
|
||||||
|
| Should Composer be a new capability or folded into `chat-verifier-agent`? | evidence-driven | New `chat-composer-agent` capability plus modified `chat-verifier-agent` routing. Composer is a distinct expression layer, while ChatService routing remains part of the verifier chain. |
|
||||||
|
| Can Composer call tools or inspect raw tool output? | evidence-driven | No. The issue and prior design require Composer to receive only Verifier-allowed material. |
|
||||||
|
| Should Gatekeeper, Verifier, or Planner change in this phase? | evidence-driven | No. Stage four is limited to Composer final-answer generation and routing. |
|
||||||
|
| What is the interface impact level? | evidence-driven | L2 internal contract change: ChatService internal final-answer semantics change, but no external API or database schema changes. |
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Composer is implemented as `chat_composer` or an equivalent model call after Verifier.
|
||||||
|
- Composer input is assembled by ChatService, not by the model.
|
||||||
|
- `claim_checks` are authoritative for filtering allowed material.
|
||||||
|
- REJECT input always has `allowed_hypotheses=[]`.
|
||||||
|
- Composer malformed output falls back to fixed safe templates.
|
||||||
|
- Fallback must never use Executor `user_facing_answer` or raw Executor JSON.
|
||||||
|
- Composer audit is persisted under `verifier_evaluation.composer_output`.
|
||||||
|
|
||||||
|
## Cross-Artifact Alignment
|
||||||
|
|
||||||
|
| Source | Alignment |
|
||||||
|
|---|---|
|
||||||
|
| Issue objective | Stage four in `mvp/issues/executor-structured-output-v2.md` requires Composer output final answer from Verifier-allowed material. Covered by proposal, design, specs, and tasks. |
|
||||||
|
| Proposal -> design | Proposal says Composer owns final expression; design defines input filtering, output parsing, fallback, and audit. |
|
||||||
|
| Design -> specs | Design decisions are reflected in `chat-composer-agent` requirements and modified `chat-verifier-agent` routing requirements. |
|
||||||
|
| Specs -> tasks | Each required behavior has implementation and test tasks, including malformed fallback and no raw output leakage. |
|
||||||
|
|
||||||
|
## Commit Gate Notes
|
||||||
|
|
||||||
|
- Interface impact: L2 internal.
|
||||||
|
- No unresolved user-interview question identified.
|
||||||
|
- No database migration required.
|
||||||
|
- No OpenSpec/devflow conflict found.
|
||||||
@@ -0,0 +1,152 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_planner
|
||||||
|
-> chat_executor
|
||||||
|
-> VerifierInputHook + Gatekeeper
|
||||||
|
-> chat_verifier
|
||||||
|
-> ChatService final rendering
|
||||||
|
```
|
||||||
|
|
||||||
|
Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.
|
||||||
|
|
||||||
|
Stage four adds the final expression layer:
|
||||||
|
|
||||||
|
```text
|
||||||
|
VerifierDecision + Executor structured output
|
||||||
|
-> ChatService filters allowed material
|
||||||
|
-> chat_composer
|
||||||
|
-> Composer JSON
|
||||||
|
-> final diagnosis_session.answer
|
||||||
|
```
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Add a Composer prompt and model call after Verifier.
|
||||||
|
- Ensure Composer receives only filtered material derived from Verifier decisions.
|
||||||
|
- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output.
|
||||||
|
- Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
|
||||||
|
- Persist Composer output for audit under `verifier_evaluation.composer_output`.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- No Planner changes.
|
||||||
|
- No Executor retry changes.
|
||||||
|
- No Gatekeeper rule expansion.
|
||||||
|
- No Verifier verification-class expansion.
|
||||||
|
- No database schema migration.
|
||||||
|
- No stage-five fixture expansion.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### Composer is an expression layer, not a diagnosis layer
|
||||||
|
|
||||||
|
Composer SHALL receive only filtered material:
|
||||||
|
|
||||||
|
- `original_query`
|
||||||
|
- `verdict`
|
||||||
|
- `allowed_claims`
|
||||||
|
- `allowed_hypotheses`
|
||||||
|
- `missing_info`
|
||||||
|
- `recommended_actions`
|
||||||
|
- `rationale`
|
||||||
|
|
||||||
|
It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.
|
||||||
|
|
||||||
|
Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.
|
||||||
|
|
||||||
|
### ChatService owns filtering
|
||||||
|
|
||||||
|
ChatService filters Executor material using Verifier `claim_checks` before invoking Composer.
|
||||||
|
|
||||||
|
Suggested mapping:
|
||||||
|
|
||||||
|
| Verifier classification | Composer handling |
|
||||||
|
|---|---|
|
||||||
|
| `direct_observation` | include in `allowed_claims` as confirmed material |
|
||||||
|
| `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause |
|
||||||
|
| `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` |
|
||||||
|
| `unsupported` | do not include as confirmed; may become `missing_info` |
|
||||||
|
| `external_unknown` | do not include as confirmed; may become `missing_info` |
|
||||||
|
| `contradicted` | do not include as confirmed; favor REJECT-safe output |
|
||||||
|
|
||||||
|
For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.
|
||||||
|
|
||||||
|
### Composer output is strict JSON with safe fallback
|
||||||
|
|
||||||
|
Composer SHALL output:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"answer_summary": "...",
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "...",
|
||||||
|
"reason": "..."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"user_facing_answer": "..."
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||||
|
|
||||||
|
### Audit stays in self_evaluation
|
||||||
|
|
||||||
|
No new table is needed. ChatService persists a minimal Composer audit snapshot:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"composer_output": {
|
||||||
|
"status": "valid",
|
||||||
|
"answer_summary": "...",
|
||||||
|
"recommended_actions": [],
|
||||||
|
"user_facing_answer": "..."
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies.
|
||||||
|
|
||||||
|
### Interface impact is L2 internal
|
||||||
|
|
||||||
|
The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- Risk: Composer introduces another LLM call and can fail formatting.
|
||||||
|
- Mitigation: strict JSON contract plus deterministic fallback templates.
|
||||||
|
- Risk: filtering is too strict and PASS answers become terse.
|
||||||
|
- Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
|
||||||
|
- Risk: legacy tests expect PASS to use Executor output.
|
||||||
|
- Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
|
||||||
|
- Risk: malformed Composer output could leak raw JSON.
|
||||||
|
- Mitigation: parse output before exposing it; fallback only from filtered material.
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. Add `chat-composer-prompt.md`.
|
||||||
|
2. Add Composer input assembly and output parsing in `ChatService`.
|
||||||
|
3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
|
||||||
|
4. Persist `composer_output` under `verifier_evaluation`.
|
||||||
|
5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.
|
||||||
|
|
||||||
|
Rollback:
|
||||||
|
|
||||||
|
- Keep fallback templates available if Composer is disabled or malformed.
|
||||||
|
- Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk.
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
None blocking. Stage four defaults:
|
||||||
|
|
||||||
|
- Composer is implemented as `chat_composer`.
|
||||||
|
- Composer does not call tools.
|
||||||
|
- Composer output failure falls back to fixed safe templates.
|
||||||
|
- No database schema changes.
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
Stage one removed final-answer fields from Executor, stage two added deterministic Gatekeeper checks, and stage three made Verifier claim-oriented. The remaining gap is final answer rendering: ChatService can still rely on temporary V2 rendering paths instead of a dedicated expression layer, which risks letting unverified Executor material shape user-facing answers.
|
||||||
|
|
||||||
|
This phase introduces a Composer expression layer so final Chat answers are generated only from Verifier-allowed material.
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Add a `chat_composer` final-answer generation step after Verifier.
|
||||||
|
- Add a strict Composer prompt and JSON output contract with `answer_summary`, `recommended_actions`, and `user_facing_answer`.
|
||||||
|
- Build Composer input in `ChatService` from filtered Verifier results:
|
||||||
|
- `allowed_claims`
|
||||||
|
- `allowed_hypotheses`
|
||||||
|
- `missing_info`
|
||||||
|
- `recommended_actions`
|
||||||
|
- `rationale`
|
||||||
|
- Ensure PASS, LOW_CONFID, and REJECT final answers no longer read Executor `user_facing_answer` or raw Executor JSON.
|
||||||
|
- Persist Composer output in `diagnosis_session.self_evaluation.verifier_evaluation.composer_output`.
|
||||||
|
- Add safe fallback templates for malformed Composer output that never expose raw Executor JSON or raw Composer JSON.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `chat-composer-agent`: Final expression layer that turns Verifier-allowed structured material into a readable Chinese user answer without introducing new facts.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `chat-verifier-agent`: ChatService final routing changes so Verifier verdicts feed Composer or safe fixed templates instead of direct Executor answer paths.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affected prompt: add `src/main/resources/prompts/chat-composer-prompt.md`.
|
||||||
|
- Affected service: `ChatService` final rendering after Verifier, Composer input assembly, Composer output parsing, malformed-output fallback.
|
||||||
|
- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `composer_output`.
|
||||||
|
- Affected tests: `ChatServiceSequentialAgentTest` and focused tests around Composer input filtering, final answer routing, and malformed Composer fallback.
|
||||||
|
- Database schema: no table or column change.
|
||||||
|
- External API: no endpoint contract change; final answer text semantics become stricter because unverified Executor material is no longer a source for user-visible answers.
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- No Planner changes.
|
||||||
|
- No Gatekeeper rule expansion.
|
||||||
|
- No Verifier classification changes.
|
||||||
|
- No Executor retry behavior change.
|
||||||
|
- No database migration.
|
||||||
|
- No stage-five eval fixture expansion in this phase.
|
||||||
@@ -0,0 +1,80 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Composer SHALL generate final user-facing Chat answers
|
||||||
|
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
|
||||||
|
|
||||||
|
#### Scenario: Composer receives only filtered material
|
||||||
|
- **WHEN** ChatService invokes Composer
|
||||||
|
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
|
||||||
|
- **AND** the Composer input SHALL NOT contain raw tool output
|
||||||
|
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
|
||||||
|
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: Composer outputs strict JSON
|
||||||
|
- **WHEN** Composer completes
|
||||||
|
- **THEN** it SHALL output exactly one JSON object
|
||||||
|
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
|
||||||
|
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||||
|
|
||||||
|
#### Scenario: Composer does not introduce new facts
|
||||||
|
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
|
||||||
|
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
|
||||||
|
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
|
||||||
|
|
||||||
|
### Requirement: Composer input SHALL honor Verifier claim checks
|
||||||
|
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
|
||||||
|
|
||||||
|
#### Scenario: Passing claims become allowed claims
|
||||||
|
- **WHEN** a claim check verification is `direct_observation`
|
||||||
|
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
|
||||||
|
|
||||||
|
#### Scenario: Reasonable inferences remain bounded
|
||||||
|
- **WHEN** a claim check verification is `reasonable_inference`
|
||||||
|
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
|
||||||
|
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Overstated claims are not confirmed findings
|
||||||
|
- **WHEN** a claim check verification is `overstated`
|
||||||
|
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
|
||||||
|
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
|
||||||
|
|
||||||
|
#### Scenario: Unsupported or external claims are withheld
|
||||||
|
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
|
||||||
|
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
|
||||||
|
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
|
||||||
|
|
||||||
|
### Requirement: Composer SHALL respect verdict-specific wording
|
||||||
|
Composer SHALL phrase final answers according to the effective Verifier verdict.
|
||||||
|
|
||||||
|
#### Scenario: PASS answer uses confirmed material
|
||||||
|
- **WHEN** the effective verdict is `PASS`
|
||||||
|
- **THEN** the final answer MAY state confirmed findings from `allowed_claims`
|
||||||
|
- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID answer separates findings and gaps
|
||||||
|
- **WHEN** the effective verdict is `LOW_CONFID`
|
||||||
|
- **THEN** the final answer SHALL distinguish confirmed information from possible directions
|
||||||
|
- **AND** it SHALL mention evidence gaps from `missing_info`
|
||||||
|
- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings
|
||||||
|
|
||||||
|
#### Scenario: REJECT answer avoids root-cause conclusions
|
||||||
|
- **WHEN** the effective verdict is `REJECT`
|
||||||
|
- **THEN** Composer input SHALL have `allowed_hypotheses=[]`
|
||||||
|
- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion
|
||||||
|
- **AND** the final answer SHALL NOT include a root-cause conclusion
|
||||||
|
|
||||||
|
### Requirement: Composer failures SHALL degrade safely
|
||||||
|
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
|
||||||
|
|
||||||
|
#### Scenario: malformed Composer output falls back safely
|
||||||
|
- **WHEN** Composer returns malformed JSON or omits required fields
|
||||||
|
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
|
||||||
|
- **AND** the final answer SHALL NOT expose raw Composer output
|
||||||
|
- **AND** the final answer SHALL NOT expose raw Executor output
|
||||||
|
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: Composer audit is persisted
|
||||||
|
- **WHEN** ChatService persists verifier evaluation
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
|
||||||
|
- **AND** the audit SHALL include the parsed Composer fields when valid
|
||||||
|
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
|
||||||
@@ -0,0 +1,80 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||||
|
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||||
|
|
||||||
|
#### Scenario: PASS -> Composer output
|
||||||
|
- **WHEN** Verifier outputs verdict="PASS"
|
||||||
|
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||||
|
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||||
|
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
|
||||||
|
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
|
||||||
|
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||||
|
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||||
|
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
|
||||||
|
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||||
|
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
|
||||||
|
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
|
||||||
|
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
|
||||||
|
|
||||||
|
#### Scenario: REJECT does not enter retry round
|
||||||
|
- **WHEN** Verifier outputs verdict="REJECT"
|
||||||
|
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
|
||||||
|
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
|
||||||
|
|
||||||
|
#### Scenario: REJECT -> degraded output
|
||||||
|
- **WHEN** Verifier outputs verdict="REJECT"
|
||||||
|
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||||
|
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||||
|
- **AND** it SHALL NOT include a root-cause conclusion
|
||||||
|
|
||||||
|
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
|
||||||
|
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
|
||||||
|
|
||||||
|
#### Scenario: PASS uses Composer-safe output
|
||||||
|
- **WHEN** the final verdict is `PASS`
|
||||||
|
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
|
||||||
|
- **AND** it SHALL NOT use Executor `user_facing_answer`
|
||||||
|
- **AND** it SHALL NOT expose raw Executor JSON
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
|
||||||
|
- **WHEN** the final verdict is `LOW_CONFID`
|
||||||
|
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
|
||||||
|
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
|
||||||
|
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||||
|
|
||||||
|
#### Scenario: REJECT uses degraded template
|
||||||
|
- **WHEN** the final verdict is `REJECT`
|
||||||
|
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
|
||||||
|
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||||
|
- **AND** it SHALL NOT include unverified raw answer content
|
||||||
|
- **AND** it SHALL NOT include a root-cause conclusion
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL be observable
|
||||||
|
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||||
|
|
||||||
|
#### Scenario: claim checks written to self_evaluation
|
||||||
|
- **WHEN** the Verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||||
|
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||||
|
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||||
|
|
||||||
|
#### Scenario: composer output written to self_evaluation
|
||||||
|
- **WHEN** final answer composition completes
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
|
||||||
|
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
|
||||||
|
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
|
||||||
|
|
||||||
|
#### Scenario: verdict written to self_evaluation
|
||||||
|
- **WHEN** the Verifier produces a verdict
|
||||||
|
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||||
|
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||||
|
|
||||||
|
#### Scenario: gatekeeper result written to self_evaluation
|
||||||
|
- **WHEN** the Verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||||
|
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
## 1. Composer Prompt Contract
|
||||||
|
|
||||||
|
- [ ] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
|
||||||
|
- [ ] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
|
||||||
|
- [ ] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
|
||||||
|
- [ ] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
|
||||||
|
- [ ] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
|
||||||
|
|
||||||
|
## 2. Composer Invocation And Input Filtering
|
||||||
|
|
||||||
|
- [ ] 2.1 Load the Composer prompt in `ChatService`.
|
||||||
|
- [ ] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
|
||||||
|
- [ ] 2.3 Build Composer input from Verifier decision and Executor structured output.
|
||||||
|
- [ ] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
|
||||||
|
- [ ] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
|
||||||
|
- [ ] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
|
||||||
|
- [ ] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
|
||||||
|
- [ ] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
|
||||||
|
|
||||||
|
## 3. Composer Output Parsing, Fallback, And Audit
|
||||||
|
|
||||||
|
- [ ] 3.1 Parse Composer strict JSON output.
|
||||||
|
- [ ] 3.2 Add safe fallback rendering for malformed Composer output.
|
||||||
|
- [ ] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||||
|
- [ ] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||||
|
- [ ] 3.5 Preserve existing verifier audit fields when writing Composer output.
|
||||||
|
|
||||||
|
## 4. Final Answer Routing
|
||||||
|
|
||||||
|
- [ ] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
|
||||||
|
- [ ] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
|
||||||
|
- [ ] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
|
||||||
|
- [ ] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
|
||||||
|
|
||||||
|
## 5. Tests And Verification
|
||||||
|
|
||||||
|
- [ ] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
|
||||||
|
- [ ] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
|
||||||
|
- [ ] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
|
||||||
|
- [ ] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
|
||||||
|
- [ ] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
|
||||||
|
- [ ] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
|
||||||
|
- [ ] 5.7 Run targeted Maven tests.
|
||||||
|
- [ ] 5.8 Validate this OpenSpec change and all specs.
|
||||||
Reference in New Issue
Block a user