feat(agent): add composer final answer
This commit is contained in:
@@ -9,6 +9,7 @@
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
# Acceptance
|
||||
|
||||
## Implementation Result
|
||||
|
||||
Implemented stage four of Executor Structured Output V2:
|
||||
|
||||
- Added `chat_composer` prompt and Agent.
|
||||
- Final answers for PASS, LOW_CONFID, and REJECT now use Composer when Verifier decision is valid.
|
||||
- Composer input is filtered from Verifier decision and Executor structured output.
|
||||
- Unsupported, external-unknown, and contradicted claims are excluded from confirmed final-answer material.
|
||||
- REJECT Composer input has `allowed_hypotheses=[]`.
|
||||
- Malformed Composer output uses deterministic safe fallback.
|
||||
- Fallback does not expose raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||
- `composer_output` is persisted in verifier audit.
|
||||
|
||||
## Static Verification
|
||||
|
||||
- Reviewed implementation diff for stage-four scope.
|
||||
- `cmd /c openspec validate executor-composer-final-answer` passed.
|
||||
- `cmd /c openspec validate --specs` passed before archive.
|
||||
|
||||
## Script Verification
|
||||
|
||||
Passed:
|
||||
|
||||
```powershell
|
||||
mvn "-Dtest=ChatServiceSequentialAgentTest" test
|
||||
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
Coverage:
|
||||
|
||||
- Composer prompt loading and invocation.
|
||||
- valid Composer output as final answer source.
|
||||
- malformed Composer output fallback.
|
||||
- PASS no raw Executor JSON leakage.
|
||||
- LOW_CONFID separation of confirmed material, possible directions, and gaps.
|
||||
- REJECT safe output without raw Executor answer.
|
||||
- Gatekeeper and Verifier stage compatibility.
|
||||
|
||||
## Browser / Manual Verification
|
||||
|
||||
Not run. This stage changes backend prompt, routing, parser, audit, and tests only.
|
||||
|
||||
## OpenSpec Archive Status
|
||||
|
||||
Archived:
|
||||
|
||||
```text
|
||||
openspec/changes/archive/2026-07-08-executor-composer-final-answer
|
||||
```
|
||||
|
||||
## Remaining Risks
|
||||
|
||||
- Stage five still needs broader eval fixture coverage for full evidence-attribution regressions.
|
||||
- Composer prompt quality can be improved after real run traces are collected.
|
||||
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
|
||||
@@ -0,0 +1,54 @@
|
||||
# Executor Composer Final Answer
|
||||
|
||||
## Background
|
||||
|
||||
Stages one to three moved the Chat diagnosis chain to structured Executor output, deterministic Gatekeeper validation, and Verifier `claim_checks`.
|
||||
|
||||
Before this stage, `ChatService` still owned final answer rendering. PASS paths could use a temporary V2 renderer, while LOW_CONFID and REJECT paths used templates. That left final user-facing expression too close to Executor material and made it harder to prove that only Verifier-allowed claims reached the user.
|
||||
|
||||
## Goal
|
||||
|
||||
Add a Composer expression layer after Verifier:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook + Gatekeeper
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
```
|
||||
|
||||
Composer produces user-facing answers from filtered material only:
|
||||
|
||||
- `allowed_claims`
|
||||
- `allowed_hypotheses`
|
||||
- `missing_info`
|
||||
- `recommended_actions`
|
||||
- `rationale`
|
||||
|
||||
## Scope
|
||||
|
||||
- Added `chat-composer-prompt.md`.
|
||||
- Added `chat_composer` Agent construction in `ChatService`.
|
||||
- Added Composer input filtering from Verifier decision and Executor structured output.
|
||||
- Replaced PASS temporary V2 renderer usage with Composer-or-safe-fallback rendering.
|
||||
- Routed LOW_CONFID and REJECT final answers through Composer when Verifier output is valid.
|
||||
- Added deterministic fallback for malformed Composer output.
|
||||
- Persisted `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
- Updated sequential workflow tests.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No Planner changes.
|
||||
- No Executor retry changes.
|
||||
- No Gatekeeper rule expansion.
|
||||
- No Verifier classification expansion.
|
||||
- No database schema migration.
|
||||
- No stage-five eval fixture expansion.
|
||||
|
||||
## OpenSpec
|
||||
|
||||
- Active change before archive: `openspec/changes/executor-composer-final-answer`
|
||||
- Capabilities: `chat-composer-agent`, `chat-verifier-agent`
|
||||
- Scale: standard
|
||||
@@ -0,0 +1,70 @@
|
||||
# Decisions
|
||||
|
||||
## Scope Decision
|
||||
|
||||
Stage four is limited to Composer final-answer generation and routing.
|
||||
|
||||
Reason: Gatekeeper and Verifier contracts were stabilized in earlier stages; this phase should only close the final-expression path.
|
||||
|
||||
## Composer Responsibility
|
||||
|
||||
Composer is an expression layer, not a diagnosis layer.
|
||||
|
||||
It may rephrase and organize only filtered material. It must not call tools, introduce new facts, rejudge root cause, or read raw Executor/tool output.
|
||||
|
||||
## Filtering Decision
|
||||
|
||||
`ChatService` owns Composer input filtering:
|
||||
|
||||
| Verifier classification | Composer handling |
|
||||
|---|---|
|
||||
| `direct_observation` | `allowed_claims` |
|
||||
| `reasonable_inference` | `allowed_claims`, with bounded wording |
|
||||
| `overstated` | `allowed_hypotheses` or `missing_info` |
|
||||
| `unsupported` | `missing_info` |
|
||||
| `external_unknown` | `missing_info` |
|
||||
| `contradicted` | `missing_info` / REJECT-safe output |
|
||||
|
||||
For REJECT, `allowed_hypotheses` is always empty.
|
||||
|
||||
## Fallback Decision
|
||||
|
||||
Malformed Composer output falls back to deterministic rendering from filtered Composer input.
|
||||
|
||||
Fallback must never expose:
|
||||
|
||||
- raw Composer JSON;
|
||||
- raw Executor JSON;
|
||||
- Executor `user_facing_answer`;
|
||||
- full unscreened tool output.
|
||||
|
||||
## Audit Decision
|
||||
|
||||
No new table is added. Composer output is persisted under:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.verifier_evaluation.composer_output
|
||||
```
|
||||
|
||||
The audit snapshot is intentionally compact and stores status plus parsed user-facing fields.
|
||||
|
||||
## Apply Fix Record
|
||||
|
||||
Initial targeted verification exposed test drift:
|
||||
|
||||
- test file had a UTF-8 BOM and failed Java compilation;
|
||||
- scripted chat model did not recognize `COMPOSER_TEST_PROMPT`;
|
||||
- older tests expected three-Agent execution and temporary V2 renderer behavior;
|
||||
- LOW_CONFID assertions required indirect support to disappear instead of appearing as a possible direction.
|
||||
|
||||
Resolution: remove BOM, add Composer script branch, and update assertions to match the committed Composer contract.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
L2 internal behavior change:
|
||||
|
||||
- external Chat API still returns a final answer string;
|
||||
- internal final-answer source changes from Executor/temporary renderer to Composer or safe fallback;
|
||||
- audit JSON gains `composer_output` under existing `self_evaluation`.
|
||||
|
||||
No database schema change.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Evidence
|
||||
|
||||
## Relevant History
|
||||
|
||||
- `executor-v2-output-contract`: Executor emits structured diagnostic material and no final-expression fields.
|
||||
- `executor-gatekeeper-hook`: Gatekeeper validates deterministic evidence failures before Verifier.
|
||||
- `executor-verifier-claim-checks`: Verifier emits `claim_checks` and effective verdict guardrails.
|
||||
|
||||
## Code Evidence
|
||||
|
||||
- `src/main/resources/prompts/chat-composer-prompt.md`: defines Composer as an expression layer with strict JSON output.
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`: loads Composer prompt, invokes `chat_composer`, filters Composer input, parses Composer output, falls back safely, and persists Composer audit.
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers Composer invocation, fallback, REJECT/LOW_CONFID behavior, and no raw JSON leakage.
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
- Composer must be after Verifier because Verifier `claim_checks` are the authority for allowed final-answer material.
|
||||
- Composer must not receive raw tool output or full unscreened Executor output because that would re-open the evidence attribution problem.
|
||||
- Verifier malformed/missing output should not invoke Composer because there is no trustworthy decision to filter with.
|
||||
- Fixed fallback remains necessary because Composer is an LLM call with a strict JSON contract and can produce malformed output.
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
Passed:
|
||||
|
||||
```powershell
|
||||
mvn "-Dtest=ChatServiceSequentialAgentTest" test
|
||||
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
cmd /c openspec validate executor-composer-final-answer
|
||||
cmd /c openspec validate --specs
|
||||
```
|
||||
|
||||
Known existing warnings:
|
||||
|
||||
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
|
||||
- Existing Lombok `@Builder` default warnings remain.
|
||||
Reference in New Issue
Block a user