feat(agent): add composer final answer

This commit is contained in:
aruo
2026-07-08 09:51:07 +08:00
parent a5b4502c72
commit 6015bcbf6f
20 changed files with 909 additions and 316 deletions
@@ -41,3 +41,34 @@
- No unresolved user-interview question identified.
- No database migration required.
- No OpenSpec/devflow conflict found.
## Apply Notes
- Implemented `chat_composer` as a post-Verifier expression Agent in `ChatService`.
- `ChatService` now builds Composer input from `VerifierDecision` plus parsed Executor structured output, not from raw Executor answer text.
- PASS, LOW_CONFID, and REJECT final-answer paths now use Composer output or a deterministic safe fallback.
- Verifier-missing or Verifier-malformed paths use fixed fallback directly because there is no trustworthy Verifier decision for Composer.
- Composer audit is persisted under `verifier_evaluation.composer_output` without adding a database table.
## Test Drift And Fixes
- Initial test compilation failed because `ChatServiceSequentialAgentTest.java` had a UTF-8 BOM at the file start. Removed the BOM.
- Existing sequential-flow tests still expected the old three-Agent call sequence and temporary V2 renderer behavior. Updated them to include `chat_composer` when a valid Verifier decision exists.
- Tests that exercise fixed fallback now force malformed Composer output so they verify no raw JSON or Executor final answer leakage.
- LOW_CONFID tests were updated to allow indirect support to appear as a possible direction rather than requiring it to disappear from all final-answer text.
## Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -32,28 +32,6 @@ The system SHALL use ChatService for explicit single-round `Planner -> Executor
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
#### Scenario: PASS uses Composer-safe output
- **WHEN** the final verdict is `PASS`
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- **AND** it SHALL NOT use Executor `user_facing_answer`
- **AND** it SHALL NOT expose raw Executor JSON
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
@@ -78,3 +56,34 @@ The Verifier's verdict and downstream final-answer composition SHALL be persiste
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
## ADDED Requirements
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
#### Scenario: PASS uses Composer-safe output
- **WHEN** the final verdict is `PASS`
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- **AND** it SHALL NOT use Executor `user_facing_answer`
- **AND** it SHALL NOT expose raw Executor JSON
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
## REMOVED Requirements
### Requirement: User-facing verifier outputs SHALL follow fixed templates
**Reason**: final user-facing answers are no longer owned by legacy Verifier templates. They must be generated from Verifier-allowed material through Composer or deterministic safe fallback.
**Migration**: use "User-facing verifier outputs SHALL follow Composer-safe protocols" and the new `chat-composer-agent` capability.
@@ -0,0 +1,44 @@
## 1. Composer Prompt Contract
- [x] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
- [x] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
- [x] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
- [x] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
- [x] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
## 2. Composer Invocation And Input Filtering
- [x] 2.1 Load the Composer prompt in `ChatService`.
- [x] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
- [x] 2.3 Build Composer input from Verifier decision and Executor structured output.
- [x] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
- [x] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
- [x] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
- [x] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
- [x] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
## 3. Composer Output Parsing, Fallback, And Audit
- [x] 3.1 Parse Composer strict JSON output.
- [x] 3.2 Add safe fallback rendering for malformed Composer output.
- [x] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- [x] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [x] 3.5 Preserve existing verifier audit fields when writing Composer output.
## 4. Final Answer Routing
- [x] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
- [x] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
- [x] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
- [x] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
## 5. Tests And Verification
- [x] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
- [x] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
- [x] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
- [x] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
- [x] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
- [x] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
- [x] 5.7 Run targeted Maven tests.
- [x] 5.8 Validate this OpenSpec change and all specs.
@@ -1,44 +0,0 @@
## 1. Composer Prompt Contract
- [ ] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
- [ ] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
- [ ] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
- [ ] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
- [ ] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
## 2. Composer Invocation And Input Filtering
- [ ] 2.1 Load the Composer prompt in `ChatService`.
- [ ] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
- [ ] 2.3 Build Composer input from Verifier decision and Executor structured output.
- [ ] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
- [ ] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
- [ ] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
- [ ] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
- [ ] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
## 3. Composer Output Parsing, Fallback, And Audit
- [ ] 3.1 Parse Composer strict JSON output.
- [ ] 3.2 Add safe fallback rendering for malformed Composer output.
- [ ] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- [ ] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [ ] 3.5 Preserve existing verifier audit fields when writing Composer output.
## 4. Final Answer Routing
- [ ] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
- [ ] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
- [ ] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
- [ ] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
## 5. Tests And Verification
- [ ] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
- [ ] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
- [ ] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
- [ ] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
- [ ] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
- [ ] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
- [ ] 5.7 Run targeted Maven tests.
- [ ] 5.8 Validate this OpenSpec change and all specs.