Compare commits

..
6 Commits
65 changed files with 6563 additions and 273 deletions
+4
View File
@@ -6,6 +6,10 @@
|---|---|---|---|---|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
@@ -0,0 +1,48 @@
# Acceptance: executor-gatekeeper-hook
## Implementation Result
Completed stage two of Executor Structured Output V2.
- Added `ExecutorGatekeeperService`.
- Added initial `schema.executor_v2` and `evidence.invocation_ref` rules.
- Added `gatekeeper_result` to Verifier payload.
- Stored `gatekeeper_result` in `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated Verifier prompt so Gatekeeper fail must not produce PASS.
- Added focused tests for schema failure, valid pass, fabricated invocation ids, tool name mismatch, hook payload, and persistence.
## Static Verification
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
- Coverage: OpenSpec change validity.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests for Gatekeeper service, hook integration, and ChatService persistence.
## Browser / Manual Verification
Not run. This stage changes backend validation and audit behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
- Excerpt similarity or phrase/utilization rules.
Reason: This phase intentionally covers deterministic schema and invocation-reference validation. Full live verification is better after Verifier V2 and Composer are implemented.
## Remaining Work
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer final-answer generation.
- Phase five: eval fixtures and full audit closure.
## Archive Status
Devflow archive files created for stage two. OpenSpec archive is expected before moving to stage three.
@@ -0,0 +1,44 @@
# Brief: executor-gatekeeper-hook
## Background
Stage one of Executor Structured Output V2 changed Chat Executor output to `executor_evidence_v2`, removing final-expression fields from Executor. That made the output structured, but it did not yet prevent deterministic evidence attribution failures such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names.
## Goal
Add a deterministic Gatekeeper between Executor output parsing and Verifier model execution.
The Gatekeeper should:
- Validate the initial Executor V2 schema.
- Validate `claims[].evidence_bindings[].source_invocation_ids` against current-session `tool_invocation` rows.
- Validate evidence binding `tool_name` against the persisted invocation tool name.
- Expose a small `gatekeeper_result` to Verifier and audit persistence.
## Scope
Included:
- New Gatekeeper validation service.
- `schema.executor_v2` initial rule.
- `evidence.invocation_ref` initial rule.
- `VerifierInputHook` payload integration.
- `VerifierContextHolder` storage.
- `ChatService` verifier evaluation persistence.
- Minimal verifier prompt update.
- Focused tests for Gatekeeper, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Excluded:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## OpenSpec
- Change: `openspec/changes/executor-gatekeeper-hook`
- Parent stage: `openspec/changes/archive/2026-07-07-executor-v2-output-contract`
@@ -0,0 +1,51 @@
# Decisions: executor-gatekeeper-hook
## Key Decisions
### Gatekeeper stays in VerifierInputHook
Decision: Gatekeeper is integrated inside `VerifierInputHook`, after Executor output parsing and before Verifier model execution.
Reason: The user explicitly chose to keep this version in the Verifier hook and not move validation into Executor hook. This preserves the current workflow orchestration.
### No retry in this phase
Decision: Gatekeeper failure does not trigger automatic Executor retry.
Reason: Retry behavior is intentionally deferred. This phase only validates, exposes, and audits deterministic failures.
### Initial rule set is intentionally small
Decision: Stage two implements only `schema.executor_v2` and `evidence.invocation_ref` as hard checks.
Reason: These rules catch the highest-confidence physical failures with low implementation risk. Excerpt similarity, hallucination phrases, and evidence utilization remain later enhancements.
### No new database schema
Decision: Persist `gatekeeper_result` in existing `diagnosis_session.self_evaluation.verifier_evaluation`.
Reason: The user asked to keep database fields minimal. Existing JSON audit storage is enough for this phase.
### Internal interface impact
Decision: This is an L2 internal interface extension.
Impact:
- Verifier payload gains `gatekeeper_result`.
- `VerifierContextHolder` gains Gatekeeper result storage.
- `self_evaluation.verifier_evaluation` gains `gatekeeper_result`.
- No external API, DTO, database table, or schema migration changes.
## Deferred Decisions
- Whether Gatekeeper should later trigger Executor retry.
- Whether `evidence.excerpt_similarity` should be hard fail or warn-only.
- Whether hallucination phrase and evidence utilization rules should be config-driven from metadata files.
- How Verifier V2 `claim_checks` should enforce Gatekeeper failures in code, beyond prompt instruction.
## Remaining Risks
- Verifier prompt compliance is not a deterministic guarantee; stage three should make Gatekeeper fail incompatible with PASS in Verifier V2 behavior.
- Excerpt authenticity is not checked in this phase, so real invocation ids can still be paired with misleading excerpt text until a later rule is implemented.
@@ -0,0 +1,54 @@
# Evidence: executor-gatekeeper-hook
## Context Evidence
- `executor-v2-output-contract` established `executor_evidence_v2` and removed Executor final-expression fields.
- `VerifierInputHook` is the existing integration point for explicit Verifier payload construction.
- `ChatService.persistVerifierEvaluation(...)` is the existing persistence path for verifier audit snapshots.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` provides the current-session invocation pool used by Gatekeeper.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- Implements `schema.executor_v2`.
- Implements `evidence.invocation_ref`.
- Returns `status`, `failed_rules`, `warnings`, and `errors`.
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Runs Gatekeeper after parsing Executor output and building trace summary.
- Adds `gatekeeper_result` to Verifier payload.
- Stores `gatekeeper_result` in `VerifierContextHolder`.
- `src/main/java/com/superbiz/agent/util/VerifierContextHolder.java`
- Stores per-request Gatekeeper result for later persistence.
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Wires `ExecutorGatekeeperService` into verifier hook construction.
- Persists `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- `src/main/resources/prompts/chat-verifier-prompt.md`
- Documents `gatekeeper_result` as an input.
- States Gatekeeper fail must not produce PASS.
## Test Evidence
- `src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java`
- Covers schema failure and valid pass behavior.
- Covers fabricated invocation ids and tool name mismatch.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`
- Covers Verifier payload containing `gatekeeper_result`.
- Covers hook behavior for fabricated invocation ids.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- Covers persistence of `gatekeeper_result` into verifier evaluation.
## Validation Evidence
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests.
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
@@ -0,0 +1,41 @@
# Acceptance: executor-v2-output-contract
## Implementation Result
Completed stage one of Executor Structured Output V2.
- Executor prompt now emits `executor_evidence_v2`.
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
- ChatService PASS path renders V2 structured output into readable Chinese.
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
## Static Verification
- `cmd /c openspec validate executor-v2-output-contract`
- Result: passed.
- Coverage: OpenSpec syntax and change validity.
## Script Verification
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
## Browser / Manual Verification
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
## Remaining Work
- Phase two: Gatekeeper in `VerifierInputHook`.
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer.
- Phase five: eval fixtures and full audit closure.
@@ -0,0 +1,32 @@
# Brief: executor-v2-output-contract
## Background
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
## Goal
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
## Scope
- Update Chat Executor prompt to emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
- Preserve V1 `user_facing_answer` extraction for compatibility.
## Non-Goals
- No Gatekeeper implementation.
- No Verifier V2 `claim_checks`.
- No Composer.
- No Planner changes.
- No database schema changes.
- No evidence tool signature changes.
## Related OpenSpec
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
@@ -0,0 +1,39 @@
# Decisions: executor-v2-output-contract
## Key Decisions
### Executor V2 removes final-expression fields
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
### Temporary renderer bridges the gap before Composer
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
### V1 compatibility remains
Decision: Existing V1 `user_facing_answer` extraction remains.
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
### Gatekeeper and Verifier V2 are deferred
Decision: This phase does not add Gatekeeper or `claim_checks`.
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
## Interface Impact
- Internal Agent output contract: L4, because fields are removed.
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
- External Chat answer: compatible intent; users still get readable Chinese.
## Risks
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
@@ -0,0 +1,21 @@
# Evidence: executor-v2-output-contract
## Context Used
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
## Code Evidence
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
## Key Finding
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
@@ -0,0 +1,58 @@
# Acceptance
## Implementation Result
Implemented stage three of Executor Structured Output V2:
- Verifier prompt now validates claim derivability rather than scanning final natural-language output.
- Verifier output supports `claim_checks`.
- `ChatService` derives compatibility `facts_checked` from `claim_checks`.
- `ChatService` persists both `claim_checks` and `facts_checked`.
- Effective verdict guardrails prevent Gatekeeper failures and malformed structured output from remaining `PASS`.
- Verifier logging summarizes `claim_checks`.
## Static Verification
- Reviewed `git diff --stat` and changed files are scoped to stage three implementation, tests, OpenSpec/devflow, and the issue handoff document.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- After OpenSpec archive, `cmd /c openspec validate --specs` passed.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Verifier payload and Gatekeeper hook behavior.
- Gatekeeper schema/invocation validation.
- `claim_checks` parsing and persistence.
- `claim_checks` to `facts_checked` compatibility mapping.
- all claim verification mapping classes.
- Gatekeeper failure downgrade from model `PASS`.
- malformed Executor output downgrade from model `PASS`.
The targeted Maven test command was re-run after OpenSpec archive and passed.
## Browser / Manual Verification
Not run. This stage changes backend prompt, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-07-executor-verifier-claim-checks
```
Archive follow-up: the generated canonical `chat-verifier-agent` spec was reviewed and amended to preserve pre-existing verifier input and Gatekeeper schema scenarios while adding the new claim-check scenarios.
## Remaining Risks
- Composer is not implemented in this stage; final PASS rendering still uses the temporary V2 renderer until stage four.
- Full eval fixture expansion is deferred to stage five.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,41 @@
# Executor Verifier Claim Checks
## Background
Stage one moved Chat Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. After those stages, Verifier still primarily used the legacy `facts_checked` contract and could still treat `executor_final_answer` as a fact source.
That left two risks:
- Verifier could still extract extra confirmed facts from natural-language Executor output.
- Downstream audit and retry consumers could not distinguish V2 claim-level verification from legacy fact checks.
## Goal
Make Verifier V2 claim-oriented:
- verify `executor_structured_output.claims` as the primary target;
- emit `claim_checks` as the authoritative V2 result;
- keep `facts_checked` only as a compatibility projection;
- enforce code-side guardrails so Gatekeeper failures or malformed structured output cannot remain effective `PASS`.
## Scope
- Updated `chat-verifier-prompt.md` to frame verification as claim derivability.
- Extended `ChatService` to parse, normalize, map, and persist `claim_checks`.
- Added effective verdict guardrails for Gatekeeper failure and malformed/missing Executor structured output.
- Updated verifier logging summaries to count `claim_checks`.
- Updated sequential workflow tests to cover claim mapping and downgrade behavior.
## Non-Goals
- No Composer integration in this phase.
- No final-answer material filtering beyond existing templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
## OpenSpec
- Active change before archive: `openspec/changes/executor-verifier-claim-checks`
- Capability: `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,57 @@
# Decisions
## Scope Decision
Stage three is limited to Verifier V2 claim checks. Composer is explicitly deferred to stage four.
Reason: Composer requires stable verifier output and allowed-material filtering; mixing it into this stage would make rollback and acceptance unclear.
## Contract Decision
`claim_checks` is the authoritative V2 verifier output.
`facts_checked` remains as a compatibility projection generated from `claim_checks` when present.
Reason: existing low-confidence rendering, retry context, trace output, and evaluation code still depend on `facts_checked`.
## Mapping Decision
Claim verification maps to legacy facts as follows:
| claim verification | legacy facts_checked verification |
|---|---|
| `direct_observation` | `direct_evidence` |
| `reasonable_inference` | `indirect_support` |
| `overstated` | `indirect_support` |
| `unsupported` | `no_evidence` |
| `external_unknown` | `no_evidence` |
| `contradicted` | `contradicted` |
## Guardrail Decision
Effective verdict is enforced in code:
- `gatekeeper_result.status=fail` cannot remain `PASS`.
- `evidence.invocation_ref` failure downgrades to `REJECT`.
- Other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- missing/malformed Executor structured output cannot remain `PASS`.
Reason: prompt compliance is not deterministic enough for safety-critical evidence attribution.
## Apply Fix Record
Initial targeted Maven verification failed because older tests expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
Classification: test drift from the committed OpenSpec, not a design blocker.
Resolution: update tests to use valid Executor V2 output for PASS paths and assert downgrade behavior for malformed or Gatekeeper-failed outputs.
## Interface Impact
L2 internal contract extension:
- Verifier output gains `claim_checks`.
- Existing `facts_checked` remains available.
- Persistence JSON gains `claim_checks` under existing `self_evaluation`.
No external API or database schema changes.
@@ -0,0 +1,35 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Existing Verifier used `facts_checked`, low-confidence rendering, retry context, and verifier audit.
## Code Evidence
- `src/main/resources/prompts/chat-verifier-prompt.md`: Verifier prompt now makes `executor_structured_output.claims` primary and treats `executor_final_answer` as debug/fallback only.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: parses `claim_checks`, maps them to compatibility `facts_checked`, persists both, and applies effective verdict guardrails.
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`: verifier thought summaries now include `claim_checks`.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers claim-check mapping, Gatekeeper downgrade, malformed output downgrade, and V2 PASS paths.
## Evidence-Driven Conclusions
- `facts_checked` cannot be removed yet because existing low-confidence templates, retry context, trace tooling, and eval paths still consume it.
- Prompt-only prevention is insufficient for Gatekeeper failures; `ChatService` must enforce effective verdict downgrades in code.
- No database schema migration is needed because `claim_checks` is persisted inside existing `diagnosis_session.self_evaluation`.
- Composer remains stage four and must not be mixed into this stage.
## Verification Evidence
Script verification passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-verifier-claim-checks
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,57 @@
# Acceptance
## Implementation Result
Implemented stage four of Executor Structured Output V2:
- Added `chat_composer` prompt and Agent.
- Final answers for PASS, LOW_CONFID, and REJECT now use Composer when Verifier decision is valid.
- Composer input is filtered from Verifier decision and Executor structured output.
- Unsupported, external-unknown, and contradicted claims are excluded from confirmed final-answer material.
- REJECT Composer input has `allowed_hypotheses=[]`.
- Malformed Composer output uses deterministic safe fallback.
- Fallback does not expose raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- `composer_output` is persisted in verifier audit.
## Static Verification
- Reviewed implementation diff for stage-four scope.
- `cmd /c openspec validate executor-composer-final-answer` passed.
- `cmd /c openspec validate --specs` passed before archive.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Composer prompt loading and invocation.
- valid Composer output as final answer source.
- malformed Composer output fallback.
- PASS no raw Executor JSON leakage.
- LOW_CONFID separation of confirmed material, possible directions, and gaps.
- REJECT safe output without raw Executor answer.
- Gatekeeper and Verifier stage compatibility.
## Browser / Manual Verification
Not run. This stage changes backend prompt, routing, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-08-executor-composer-final-answer
```
## Remaining Risks
- Stage five still needs broader eval fixture coverage for full evidence-attribution regressions.
- Composer prompt quality can be improved after real run traces are collected.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,54 @@
# Executor Composer Final Answer
## Background
Stages one to three moved the Chat diagnosis chain to structured Executor output, deterministic Gatekeeper validation, and Verifier `claim_checks`.
Before this stage, `ChatService` still owned final answer rendering. PASS paths could use a temporary V2 renderer, while LOW_CONFID and REJECT paths used templates. That left final user-facing expression too close to Executor material and made it harder to prove that only Verifier-allowed claims reached the user.
## Goal
Add a Composer expression layer after Verifier:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> chat_composer
-> final answer
```
Composer produces user-facing answers from filtered material only:
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
## Scope
- Added `chat-composer-prompt.md`.
- Added `chat_composer` Agent construction in `ChatService`.
- Added Composer input filtering from Verifier decision and Executor structured output.
- Replaced PASS temporary V2 renderer usage with Composer-or-safe-fallback rendering.
- Routed LOW_CONFID and REJECT final answers through Composer when Verifier output is valid.
- Added deterministic fallback for malformed Composer output.
- Persisted `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated sequential workflow tests.
## Non-Goals
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier classification expansion.
- No database schema migration.
- No stage-five eval fixture expansion.
## OpenSpec
- Active change before archive: `openspec/changes/executor-composer-final-answer`
- Capabilities: `chat-composer-agent`, `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,70 @@
# Decisions
## Scope Decision
Stage four is limited to Composer final-answer generation and routing.
Reason: Gatekeeper and Verifier contracts were stabilized in earlier stages; this phase should only close the final-expression path.
## Composer Responsibility
Composer is an expression layer, not a diagnosis layer.
It may rephrase and organize only filtered material. It must not call tools, introduce new facts, rejudge root cause, or read raw Executor/tool output.
## Filtering Decision
`ChatService` owns Composer input filtering:
| Verifier classification | Composer handling |
|---|---|
| `direct_observation` | `allowed_claims` |
| `reasonable_inference` | `allowed_claims`, with bounded wording |
| `overstated` | `allowed_hypotheses` or `missing_info` |
| `unsupported` | `missing_info` |
| `external_unknown` | `missing_info` |
| `contradicted` | `missing_info` / REJECT-safe output |
For REJECT, `allowed_hypotheses` is always empty.
## Fallback Decision
Malformed Composer output falls back to deterministic rendering from filtered Composer input.
Fallback must never expose:
- raw Composer JSON;
- raw Executor JSON;
- Executor `user_facing_answer`;
- full unscreened tool output.
## Audit Decision
No new table is added. Composer output is persisted under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.composer_output
```
The audit snapshot is intentionally compact and stores status plus parsed user-facing fields.
## Apply Fix Record
Initial targeted verification exposed test drift:
- test file had a UTF-8 BOM and failed Java compilation;
- scripted chat model did not recognize `COMPOSER_TEST_PROMPT`;
- older tests expected three-Agent execution and temporary V2 renderer behavior;
- LOW_CONFID assertions required indirect support to disappear instead of appearing as a possible direction.
Resolution: remove BOM, add Composer script branch, and update assertions to match the committed Composer contract.
## Interface Impact
L2 internal behavior change:
- external Chat API still returns a final answer string;
- internal final-answer source changes from Executor/temporary renderer to Composer or safe fallback;
- audit JSON gains `composer_output` under existing `self_evaluation`.
No database schema change.
@@ -0,0 +1,36 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits structured diagnostic material and no final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates deterministic evidence failures before Verifier.
- `executor-verifier-claim-checks`: Verifier emits `claim_checks` and effective verdict guardrails.
## Code Evidence
- `src/main/resources/prompts/chat-composer-prompt.md`: defines Composer as an expression layer with strict JSON output.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: loads Composer prompt, invokes `chat_composer`, filters Composer input, parses Composer output, falls back safely, and persists Composer audit.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers Composer invocation, fallback, REJECT/LOW_CONFID behavior, and no raw JSON leakage.
## Evidence-Driven Conclusions
- Composer must be after Verifier because Verifier `claim_checks` are the authority for allowed final-answer material.
- Composer must not receive raw tool output or full unscreened Executor output because that would re-open the evidence attribution problem.
- Verifier malformed/missing output should not invoke Composer because there is no trustworthy decision to filter with.
- Fixed fallback remains necessary because Composer is an LLM call with a strict JSON contract and can produce malformed output.
## Verification Evidence
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,514 @@
# Current Chat Agent Data Contracts
**状态**:当前实现
**日期**:2026-07-07
**范围**:当前 Chat 复杂诊断链路的数据结构定义
当前代码实现是三 Agent 顺序链路:
```text
chat_planner -> chat_executor -> chat_verifier
```
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
---
## 1. Workflow Input
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
```text
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
--- 用户问题 ---
{question}
--- retry_context ---
{retry_context}
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `question` | 用户输入 | 用户本轮原始问题 |
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
---
## 2. chat_planner
### 2.1 Input
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
```json
{
"question": "用户原始问题",
"history": [],
"available_knowledge_domains": "...",
"skill_catalog": {},
"retry_context": null
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `question` | workflow input | 用户原始问题 |
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
### 2.2 Output:`planner_plan`
当前 prompt 要求输出 JSON:
```json
{
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
"reasoning": "规划思路说明"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `selected_skill` | string/null | Planner 选择的诊断 skill 名称 |
| `selection_reason` | string | skill 选择理由 |
| `plan` | array | 给 Executor 的执行步骤 |
| `reasoning` | string | 规划思路说明 |
运行态输出 key:
```text
planner_plan
```
---
## 3. chat_executor
### 3.1 Input
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
```json
{
"planner_plan": {},
"history": [],
"retry_context": null,
"tool_permissions": {
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
"tool_callbacks": []
}
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
| `retry_context` | `ChatService` | 本轮补证据约束 |
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
### 3.2 Output:`executor_feedback`
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "root_cause",
"claim_text": "事实断言或有限结论",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
"source_invocation_ids": [],
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "未被证实但值得排查的方向",
"basis": "它基于哪些已知证据或为什么只是推测",
"needed_evidence": ["需要补充的证据"]
}
],
"recommended_actions": [
{
"action_text": "建议动作",
"reason": "为什么建议做这个动作",
"evidence_bindings": []
}
],
"missing_info": [
"导致无法确认完整根因的证据缺口"
],
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
| `claims[].claim_id` | string | claim 标识 |
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
| `claims[].claim_text` | string | 事实断言文本 |
| `claims[].support_level` | string | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
| `evidence_bindings[].tool_name` | string | 来源工具名 |
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
| `hypotheses` | array | 未证实但值得排查的方向 |
| `hypotheses[].hypothesis_text` | string | 假设文本 |
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
| `recommended_actions` | array | 建议动作 |
| `recommended_actions[].action_text` | string | 建议动作文本 |
| `recommended_actions[].reason` | string | 建议原因 |
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
| `missing_info` | array | 证据缺口 |
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
运行态输出 key:
```text
executor_feedback
```
---
## 4. chat_verifier
### 4.1 Input
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
```json
{
"original_query": "用户原始问题",
"executor_final_answer": "{...executor_feedback raw text...}",
"executor_structured_output": {},
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"tool_trace_summary": [],
"retry_context": null
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
### 4.2 `tool_trace_summary`
`ToolTraceSummaryService` 聚合 evidence tools:
```text
lookup_knowledge, query_logs, query_metrics, query_order
```
输出项结构:
```json
{
"trace_ref": "trace-1",
"tool_name": "query_logs",
"success": true,
"input_summary": "query=payment-service timeout",
"output_summary": "log_evidence: ...",
"evidence_level": "direct",
"topic_domain": "general",
"source_invocation_ids": [394],
"invocation_count": 1,
"failed_invocation_count": 0,
"no_hit_invocation_count": 0,
"query_samples": ["payment-service timeout"],
"retrieval_layers": [],
"relevance_levels": [],
"source_documents": []
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
| `tool_name` | string | 聚合后的工具名 |
| `success` | boolean | 是否存在可用证据 |
| `input_summary` | string | 工具输入摘要 |
| `output_summary` | string | 工具输出摘要 |
| `evidence_level` | string | `direct` / `indirect` / `none` |
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
| `invocation_count` | number | 聚合调用次数 |
| `failed_invocation_count` | number | 失败调用次数 |
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
| `query_samples` | array | 查询样例 |
| `retrieval_layers` | array | 检索层级 |
| `relevance_levels` | array | 相关性等级 |
| `source_documents` | array | 来源文档标签 |
### 4.3 Output:`verifier_output`
当前 `chat-verifier-prompt.md` 要求输出:
```json
{
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [
{
"fact": "ERR_TIMEOUT 表示请求超时",
"is_critical": true,
"verification": "direct_evidence",
"detail": "知识库文档明确给出该错误码定义",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
}
]
}
],
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | number | 关键事实证据支撑评分 |
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
| `facts_checked` | array | Verifier 校验过的事实列表 |
| `facts_checked[].fact` | string | 被校验事实 |
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
| `facts_checked[].detail` | string | 校验说明 |
| `facts_checked[].evidence_refs` | array | 证据引用 |
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
| `evidence_refs[].tool_name` | string | 引用工具 |
| `evidence_refs[].topic_domain` | string | 引用主题域 |
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
| `evidence_refs[].note` | string | 引用说明 |
| `rationale` | string | verdict 判定理由 |
运行态输出 key:
```text
verifier_output
```
---
## 5. VerifierDecision
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
```json
{
"verdict": "LOW_CONFID",
"groundednessScore": 0.5,
"criticalFactCount": 2,
"factsChecked": [],
"rationale": "证据不足",
"round": 1
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | Verifier verdict |
| `groundednessScore` | number | groundedness score |
| `criticalFactCount` | number | 关键事实数量 |
| `factsChecked` | array | 解析后的 facts_checked |
| `rationale` | string | 判定理由 |
| `round` | number | 当前验证轮次 |
---
## 6. retry_context
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
```json
{
"round": 1,
"missing_evidence_facts": [
"某关键事实:缺少直接证据"
],
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `round` | number | 触发 retry 的轮次 |
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
| `instruction` | string | 补证据约束 |
---
## 7. diagnosis_session.self_evaluation.verifier_evaluation
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
```json
{
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.5,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "证据不足",
"round": 1,
"traceability_version": "v1",
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"executor_structured_output": {},
"tool_trace_summary": []
}
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verifier_evaluation.verdict` | string | Verifier verdict |
| `verifier_evaluation.groundedness_score` | number | groundedness score |
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
| `verifier_evaluation.rationale` | string | 判定理由 |
| `verifier_evaluation.round` | number | 验证轮次 |
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
---
## 8. Final Answer Rendering
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
| Verdict | 当前行为 |
|---|---|
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
低置信模板使用:
```text
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
已确认信息:
- ...
当前缺口:
- ...
建议下一步:
- ...
```
拒绝模板使用:
```text
当前无法基于已获取证据生成可靠结论。
已确认信息:
- ...
证据缺口:
- ...
建议下一步:
- ...
```
---
## 9. Trace Persistence Data
### 9.1 diagnosis_session
| 字段 | 类型 | 定义 |
|---|---|---|
| `session_id` | string | 会话 id |
| `query` | text | 用户问题 |
| `status` | string | 会话状态 |
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
| `total_duration_ms` | number | 总耗时 |
| `total_token_count` | number | 总 token |
| `step_count` | number | agent step 数 |
| `tool_call_count` | number | tool invocation 数 |
| `answer` | longtext | 最终用户答案 |
| `self_evaluation` | json | 包含 verifier_evaluation |
| `feedback` | string | 用户反馈 |
### 9.2 agent_step
| 字段 | 类型 | 定义 |
|---|---|---|
| `session_id` | string | 会话 id |
| `step_index` | number | 步骤序号 |
| `agent_name` | string | `planner` / `executor` / `verifier` |
| `model_input` | text | 模型输入摘要 |
| `model_output` | text | 模型输出摘要 |
| `thought` | text | hook 记录的摘要信息 |
| `has_tool_call` | boolean | 是否包含工具调用 |
| `duration_ms` | number | 模型调用耗时 |
| `token_count` | number | token 数 |
### 9.3 tool_invocation
| 字段 | 类型 | 定义 |
|---|---|---|
| `id` | number | 工具调用 id |
| `session_id` | string | 会话 id |
| `step_id` | number | 对应 agent_step id |
| `tool_name` | string | 工具名 |
| `input_params` | json | 工具输入参数 |
| `output_preview` | text | 工具输出预览 |
| `output_length` | number | 原始输出长度 |
| `retrieval_layer` | string | 检索层 |
| `l0_match_count` | number | L0 命中数 |
| `l1_match_count` | number | L1 命中数 |
| `is_truncated` | boolean | 输出是否截断 |
| `relevance_level` | string | 相关性等级 |
| `dedup_reason` | string | 去重原因 |
| `retrieval_details` | json | 检索细节 |
| `duration_ms` | number | 工具耗时 |
| `success` | boolean | 是否成功 |
| `error_message` | text | 错误信息 |
File diff suppressed because it is too large Load Diff
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,102 @@
# Decisions: executor-gatekeeper-hook
## sm-flow Progress
### Clarify
Entry summary: implement stage two of Executor Structured Output V2 by adding deterministic Gatekeeper checks between Executor and Verifier.
Slug: `executor-gatekeeper-hook`
Scale: complex program, stage-specific standard slice. It changes internal verifier payload and audit persistence but does not change external API or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor now emits `executor_evidence_v2` without final-expression fields.
- `chat-verifier-agent`: Verifier input is assembled explicitly by `VerifierInputHook` and persisted through `ChatService`.
- `evidence-trace-hardening`: tool invocation rows are the source of truth for evidence trace references.
Current code shape:
- `VerifierInputHook` parses Executor raw output, builds `tool_trace_summary`, and assembles Verifier payload.
- `VerifierContextHolder` stores parse status, structured output, and trace summary for later persistence.
- `ChatService.persistVerifierEvaluation(...)` writes verifier snapshots into `diagnosis_session.self_evaluation.verifier_evaluation`.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` can provide the valid invocation pool.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Gatekeeper live in Executor hook? | evidence-driven | No. User explicitly chose Verifier hook for this version. |
| Should Gatekeeper retry Executor? | evidence-driven | No. Retry is deferred; this phase only validates and audits. |
| Which rules are mandatory now? | evidence-driven | Implement schema and invocation reference rules. Excerpt similarity and utilization can follow later. |
| Where should audit data persist? | evidence-driven | Existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
| Does this require database migration? | evidence-driven | No. Use existing JSON self_evaluation and existing tool_invocation rows. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Gatekeeper output, rules, integration, persistence, and risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for payload, persistence, and initial rules.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- Gatekeeper introduces an internal verifier payload extension but no external API or database schema change.
- `VerifierInputHook` remains the integration point, matching the user decision to keep Gatekeeper in the Verifier hook.
- Gatekeeper needs repository access to validate invocation ids; placing that access in a service keeps hook logic small.
- Failed Gatekeeper results are audit signals in this phase and do not trigger retry.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage two | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design rules and payload | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier payload: L2 internal extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation result:
- Added `ExecutorGatekeeperService` with initial `schema.executor_v2` and `evidence.invocation_ref` checks.
- Integrated Gatekeeper into `VerifierInputHook` after parse and trace summary construction.
- Added `gatekeeper_result` to Verifier payload and `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated `chat-verifier-prompt.md` so Gatekeeper failures must not produce PASS.
- Added and updated focused tests for service validation, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Validation:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed with 23 tests.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
Known limits:
- No Executor retry is triggered by Gatekeeper failure in this phase.
- `evidence.excerpt_similarity`, `claim.hallucination_phrases`, and `claim.evidence_utilization` remain deferred by design.
@@ -0,0 +1,119 @@
## Context
Stage one introduced `executor_evidence_v2`, but `VerifierInputHook` still only parses JSON and passes the structured object to Verifier. A valid JSON object can still contain:
- removed fields such as `diagnosis_summary` or `user_facing_answer`
- missing or wrong `answer_version`
- empty `claims[].evidence_bindings`
- fabricated `source_invocation_ids`
- `tool_name` values that do not match the real `tool_invocation`
Gatekeeper handles these deterministic failures before Verifier performs semantic reasoning.
## Goals / Non-Goals
Goals:
- Add Gatekeeper into `VerifierInputHook`.
- Produce a small `gatekeeper_result` object.
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` in verifier evaluation.
- Implement initial rules: schema and invocation reference.
Non-goals:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## Gatekeeper Output
Gatekeeper returns:
```json
{
"status": "fail",
"failed_rules": ["evidence.invocation_ref"],
"warnings": [],
"errors": [
{
"rule_id": "evidence.invocation_ref",
"target": "claims[0].evidence_bindings[0]",
"message": "source_invocation_ids not found in current session"
}
]
}
```
Status calculation:
```text
any fail -> fail
else any warning -> warn
else pass
```
## Initial Rules
### schema.executor_v2
Fail when:
- structured output is absent after parse status is valid
- `answer_version` is not `executor_evidence_v2`
- `claims` is not an array
- removed fields `diagnosis_summary` or `user_facing_answer` are present
- any claim misses required fields
- any claim has empty `evidence_bindings`
- `hypotheses`, `recommended_actions`, or `missing_info` are missing or non-array
For this phase, missing optional arrays may be normalized only if implementation remains simple. If not normalized, missing arrays fail schema to keep behavior deterministic.
### evidence.invocation_ref
Fail when:
- `claims[].evidence_bindings[].source_invocation_ids` is missing or empty
- any referenced invocation id is not in the current session's `tool_invocation` rows
- `tool_name` does not match the referenced invocation's real `tool_name`
Recommended action evidence bindings remain optional and are not hard-fail checked in this phase.
## Integration
`VerifierInputHook.beforeModel(...)` flow becomes:
```text
parse Executor output
build tool_trace_summary
Gatekeeper.validate(sessionId, structuredOutput, parseStatus)
put gatekeeper_result into verifier payload
store gatekeeper_result in VerifierContextHolder
```
`ChatService.persistVerifierEvaluation(...)` adds:
```text
gatekeeper_result: VerifierContextHolder.getGatekeeperResult()
```
If Gatekeeper throws unexpectedly, hook should fail closed with a minimal `fail` result in payload rather than dropping validation silently.
## Interface Impact
- Verifier payload: L2 internal contract extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension under existing `self_evaluation`.
- No external API or database schema change.
## Risks / Mitigations
- Risk: tests or code instantiate `VerifierInputHook` with the old constructor.
- Mitigation: keep a compatibility constructor that uses a no-op/pass Gatekeeper, or update tests explicitly.
- Risk: Gatekeeper fails because no session id exists.
- Mitigation: return fail with a clear `schema.executor_v2` or `evidence.invocation_ref` error only when validation cannot establish current-session references.
- Risk: Verifier prompt may ignore Gatekeeper.
- Mitigation: payload and audit are still authoritative for later phases; Verifier prompt update can be minimal in this phase.
@@ -0,0 +1,37 @@
## Why
Executor V2 makes the output structured, but structure alone does not prevent physical hallucinations such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names. These failures should be caught deterministically before the Verifier reasons over claims.
This phase adds Gatekeeper inside `VerifierInputHook` as a deterministic pre-verifier quality gate and persists its result for audit.
## What Changes
- Add a Gatekeeper validation service for Executor structured output.
- Add initial rules:
- `schema.executor_v2`
- `evidence.invocation_ref`
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Keep Gatekeeper in `VerifierInputHook`; do not move it to Executor hook.
- Keep retry behavior unchanged; failed Gatekeeper results do not trigger Executor retry in this phase.
- Keep Verifier prompt/output behavior unchanged except that it can see `gatekeeper_result`.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Verifier input now includes deterministic Gatekeeper results for Executor structured output.
## Impact
- Affected hook: `VerifierInputHook`.
- Affected service/runtime: new Gatekeeper service and `VerifierContextHolder`.
- Affected persistence: `ChatService.persistVerifierEvaluation(...)` writes `gatekeeper_result` into existing `self_evaluation`.
- Affected repository access: Gatekeeper reads current-session `tool_invocation` rows via `ToolInvocationRepository`.
- Affected tests: `VerifierInputHookTest` and `ChatServiceSequentialAgentTest`.
- Database schema: no change.
@@ -0,0 +1,51 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
## ADDED Requirements
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
@@ -0,0 +1,31 @@
## 1. Gatekeeper Core
- [x] 1.1 Add a Gatekeeper validation service with a small result shape: `status`, `failed_rules`, `warnings`, `errors`.
- [x] 1.2 Implement `schema.executor_v2` rule.
- [x] 1.3 Implement `evidence.invocation_ref` rule using current-session `tool_invocation` rows.
- [x] 1.4 Keep recommended action evidence bindings out of hard-fail validation for this phase.
## 2. Hook Integration
- [x] 2.1 Inject Gatekeeper into `VerifierInputHook`.
- [x] 2.2 Add `gatekeeper_result` to Verifier payload.
- [x] 2.3 Store `gatekeeper_result` in `VerifierContextHolder`.
- [x] 2.4 Preserve parse-only boundary in `VerifierInputHook`.
## 3. Persistence
- [x] 3.1 Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [x] 3.2 Preserve existing verifier evaluation fields.
## 4. Prompt Compatibility
- [x] 4.1 Update `chat-verifier-prompt.md` minimally so Verifier sees `gatekeeper_result` and must not output PASS when it fails.
## 5. Tests And Verification
- [x] 5.1 Add Gatekeeper unit tests for schema failures and pass cases.
- [x] 5.2 Add hook tests proving payload contains `gatekeeper_result`.
- [x] 5.3 Add tests for fabricated invocation id and tool name mismatch.
- [x] 5.4 Add ChatService persistence test for `gatekeeper_result`.
- [x] 5.5 Run targeted tests.
- [x] 5.6 Validate this OpenSpec change.
@@ -0,0 +1 @@
archive-ready
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,112 @@
# Decisions: executor-v2-output-contract
## sm-flow Progress
### Clarify
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
Slug: `executor-v2-output-contract`
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
### Context
Relevant devflow history:
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
Conflict with historical decision:
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
Current code shape:
- `chat-executor-prompt.md` defines the V1 Executor output contract.
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
### Specify
OpenSpec artifacts:
- `proposal.md`: scope and compatibility boundary for stage one.
- `design.md`: V2 Executor contract and temporary rendering strategy.
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
- No database schema or evidence tool signature changes are required.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue background / stage one | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design contract and rendering bridge | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
### Commit
Commit gate result: passed.
- `proposal.md` exists and explains why this phase is needed.
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
- `tasks.md` contains executable implementation and verification tasks.
- `cmd /c openspec validate executor-v2-output-contract` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation summary:
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
- Preserved V1 `user_facing_answer` extraction for compatibility.
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
Verification:
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-v2-output-contract` passed.
Known limitations:
- Gatekeeper is not implemented in this phase.
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
- The V2 renderer is temporary and should be replaced by Composer in a later phase.
@@ -0,0 +1,112 @@
## Context
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
- `diagnosis_summary`
- `user_facing_answer`
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
This stage implements only the first slice of the V2 migration:
```text
Executor V2 output contract
-> existing VerifierInputHook parsing
-> existing Verifier
-> temporary ChatService structured renderer
```
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
## Goals / Non-Goals
Goals:
- Make Chat Executor emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
- Preserve evidence-binding requirements for confirmed claims.
- Prevent PASS routing from returning raw JSON to normal Chat users.
Non-goals:
- No Gatekeeper implementation.
- No Verifier prompt rewrite to `claim_checks`.
- No Composer agent.
- No database schema changes.
- No Planner changes.
- No evidence tool signature changes.
- No retry behavior changes.
## Executor V2 Contract
Executor final output SHALL be one JSON object:
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "payment-service 出现请求超时日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_logs",
"source_invocation_ids": [394],
"evidence_excerpt": "request timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
Removed fields:
- `diagnosis_summary`
- `user_facing_answer`
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
## Temporary Rendering Strategy
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
When Verifier returns `PASS`:
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
- `claims[].claim_text`
- `hypotheses[].hypothesis_text`
- `missing_info[]`
- `recommended_actions[].action_text` and `reason`
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
## Parser Boundary
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
## Interface Impact
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
## Risks / Mitigations
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
- Mitigation: add temporary V2 renderer in `ChatService`.
- Risk: current Verifier prompt still mentions `user_facing_answer`.
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
- Risk: tests assume V1 fields.
- Mitigation: update/add focused tests for V2 output without final-expression fields.
@@ -0,0 +1,33 @@
## Why
The current Chat Executor evidence contract still mixes diagnostic material with final user-facing prose through `diagnosis_summary` and `user_facing_answer`. This keeps the Executor in a "diagnose and narrate" mode, so unsupported details can be smuggled into the final answer before later Gatekeeper, Verifier V2, and Composer phases exist.
This first phase narrows Executor to structured diagnostic material only and adds a temporary safe rendering path so normal Chat responses do not expose raw Executor JSON while later phases are implemented.
## What Changes
- **BREAKING internal Agent contract**: Chat Executor output changes from `executor_evidence_v1` to `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from the Chat Executor final JSON contract.
- Keep the existing structured arrays: `claims`, `hypotheses`, `recommended_actions`, and `missing_info`.
- Preserve `claims[].evidence_bindings` and current-session evidence attribution rules.
- Adjust runtime final-answer handling so a PASS result with V2 Executor output is rendered into readable Chinese from structured fields instead of returning raw JSON.
- Keep Planner, Verifier, Gatekeeper, retry behavior, database schema, and tool signatures unchanged in this phase.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: The Executor evidence-attribution contract is tightened so V2 structured output no longer contains final-expression fields. Verifier still receives `executor_final_answer` as raw text and `executor_structured_output` when parseable.
## Impact
- Affected prompt: `src/main/resources/prompts/chat-executor-prompt.md`.
- Affected runtime: `ChatService` PASS answer extraction/rendering for Executor V2.
- Affected parser boundary: `VerifierInputHook` should continue parsing JSON but must not treat schema validation as its own responsibility in this phase.
- Affected tests: ChatService sequential flow tests and VerifierInputHook parsing tests for V2 output without `user_facing_answer`.
- Interface impact: L4 for internal Agent output contract because fields are removed from Executor JSON; external HTTP/chat answer behavior must remain readable Chinese and must not expose raw JSON.
@@ -0,0 +1,50 @@
## MODIFIED Requirements
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
@@ -0,0 +1,24 @@
## 1. Executor V2 Prompt
- [x] 1.1 Update `src/main/resources/prompts/chat-executor-prompt.md` so the final contract uses `answer_version="executor_evidence_v2"`.
- [x] 1.2 Remove `diagnosis_summary` and `user_facing_answer` from the required Executor output schema and validation rules.
- [x] 1.3 Keep claims, hypotheses, recommended actions, missing information, and evidence-binding rules.
- [x] 1.4 Keep Planner and tool-use behavior unchanged.
## 2. Runtime Rendering Safety
- [x] 2.1 Update `ChatService` PASS handling so V2 structured output is rendered into readable Chinese instead of raw JSON.
- [x] 2.2 Preserve V1 `user_facing_answer` extraction for compatibility.
- [x] 2.3 Ensure fallback behavior does not expose raw Executor JSON when structured rendering fails.
## 3. Parser Compatibility
- [x] 3.1 Keep `VerifierInputHook` parse-only behavior compatible with V2 output.
- [x] 3.2 Add or update tests proving V2 output without `user_facing_answer` parses into `executor_structured_output`.
## 4. Tests And Verification
- [x] 4.1 Add or update ChatService tests for PASS with `executor_evidence_v2`.
- [x] 4.2 Add or update tests proving final user answer does not contain raw JSON contract text.
- [x] 4.3 Run targeted tests for ChatService and VerifierInputHook.
- [x] 4.4 Validate this OpenSpec change.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,110 @@
# Decisions: executor-verifier-claim-checks
## sm-flow Progress
### Clarify
Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.
Slug: `executor-verifier-claim-checks`
Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Current Verifier still primarily uses `facts_checked`.
Current code shape:
- `chat-verifier-prompt.md` still frames the task around `executor_final_answer` and `facts_checked`.
- `ChatService.parseVerifierDecision(...)` only parses `facts_checked`.
- `buildLowConfidenceOutput(...)` and `buildRetryContext(...)` consume `VerifierDecision.factsChecked()`.
- `persistVerifierEvaluation(...)` already persists parse status, structured output, trace summary, and gatekeeper result.
- `AgentLoggingHook` summarizes verifier output using `facts_checked`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Verifier still scan `executor_final_answer` when structured output is valid? | evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. |
| Should `facts_checked` be removed now? | evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. |
| Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. |
| Does this require database migration? | evidence-driven | No. `claim_checks` is stored under existing JSON self_evaluation. |
| Does this introduce Composer? | evidence-driven | No. Composer is stage four. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Verifier V2 output, compatibility mapping, verdict guardrails, risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for claim checks, compatibility facts, persistence, and guardrails.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- This is an L2 internal verifier contract extension.
- Existing consumers continue using `facts_checked`, which is now generated from `claim_checks` when present.
- Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
- No external API or database schema changes are introduced.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage three | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design output contract and guardrails | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier output: L2 internal extension with `claim_checks`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- `cmd /c openspec status --change executor-verifier-claim-checks` reports 4/4 artifacts complete.
- No unresolved user-interview questions remain for this stage.
### Apply
Capability source: OpenSpec CLI + `openspec-apply-change` protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, `rg`/diff inspection, and targeted tests.
Implemented changes:
- Updated `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
- Added `claim_checks` parsing, normalization, persistence, and compatibility mapping to legacy `facts_checked` in `ChatService`.
- Added code-side effective verdict guardrails:
- missing/malformed structured output cannot remain `PASS`;
- `gatekeeper_result.status=fail` cannot remain `PASS`;
- `evidence.invocation_ref` failures downgrade to `REJECT`;
- other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- Updated verifier logging summaries to account for `claim_checks`.
- Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.
Conflict / fix record:
- Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
- Classification: code/test drift from the committed OpenSpec, not a design blocker.
- Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.
Verification:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
@@ -0,0 +1,140 @@
## Context
Current state after stage two:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> ChatService final rendering
```
Verifier receives explicit inputs:
- `original_query`
- `executor_final_answer`
- `executor_structured_output`
- `executor_output_parse_status`
- `tool_trace_summary`
- `gatekeeper_result`
- `retry_context`
However, Verifier output is still primarily:
```json
{
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 1,
"facts_checked": [],
"rationale": "..."
}
```
Stage three introduces V2 output while preserving the old compatibility field.
## Verifier V2 Output
Verifier should output:
```json
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.62,
"critical_fact_count": 1,
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service CPU usage is high",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "query_metrics shows CPU=92%",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [394],
"note": "metrics summary contains CPU=92%"
}
]
}
],
"hypothesis_checks": [],
"facts_checked": [],
"rationale": "..."
}
```
`facts_checked` remains for compatibility. If Verifier does not emit it, `ChatService` must derive it from `claim_checks`.
## Claim Verification Set
`claim_checks[].verification` is limited to:
| Value | Meaning | Legacy mapping |
|---|---|---|
| `direct_observation` | Evidence directly observes the claim | `direct_evidence` |
| `reasonable_inference` | Evidence can reasonably support the claim, but not as direct observation | `indirect_support` |
| `overstated` | Evidence partially supports the claim, but the claim says too much | `indirect_support` |
| `unsupported` | Evidence is insufficient | `no_evidence` |
| `external_unknown` | Claim introduces evidence-external entity/value/root cause | `no_evidence` |
| `contradicted` | Claim conflicts with evidence | `contradicted` |
## Compatibility Mapping
`ChatService` must keep old downstream behavior alive by producing `facts_checked`.
Suggested mapping:
```text
facts_checked[].fact = "{claim_id}: {claim_text}"
facts_checked[].is_critical = claim_type in ["root_cause", "symptom", "impact", "risk"]
facts_checked[].verification = mapped legacy verification
facts_checked[].detail = claim_checks[].detail
facts_checked[].evidence_refs = claim_checks[].evidence_refs
```
If Verifier emits both `claim_checks` and `facts_checked`, `claim_checks` is authoritative. `facts_checked` may be replaced by the deterministic compatibility projection to avoid inconsistent audit data.
If Verifier emits only old `facts_checked`, ChatService keeps the old path.
## Verdict Guardrails
Gatekeeper fail:
- If `gatekeeper_result.status = fail`, effective verdict must not be `PASS`.
- If the model returns `PASS`, ChatService should downgrade the effective verdict to `LOW_CONFID` or `REJECT`.
- For this phase, `evidence.invocation_ref` failure should downgrade to `REJECT`; other Gatekeeper failures should downgrade to `LOW_CONFID`.
Malformed or missing structured output:
- If `executor_output_parse_status.status` is `missing` or `malformed`, Verifier should not use natural-language extraction to produce PASS.
- Effective verdict should be `LOW_CONFID`.
## Prompt Boundary
The prompt should say:
- Primary target is `executor_structured_output.claims`.
- Do not extract additional confirmed facts from `executor_final_answer` when structured output is valid.
- `executor_final_answer` is debug/fallback only.
- `claim_checks` is the primary output.
- `facts_checked` is compatibility output.
## Interface Impact
- L2 internal contract extension.
- No external API change.
- No database schema change.
- Audit JSON gains `claim_checks`.
## Risks / Mitigations
- Risk: old low-confidence templates rely on `facts_checked`.
- Mitigation: derive `facts_checked` from `claim_checks`.
- Risk: prompt-only Gatekeeper PASS prevention is insufficient.
- Mitigation: add code-side effective verdict guard.
- Risk: Agent logging only summarizes `facts_checked`.
- Mitigation: update logging to understand `claim_checks` while keeping old summary compatibility.
@@ -0,0 +1,47 @@
## Why
Stage one moved Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. The Verifier still mainly operates through the legacy `facts_checked` contract and the prompt still allows fallback extraction from `executor_final_answer`.
That keeps two problems alive:
- Verifier can still treat natural-language Executor output as a fact source.
- Downstream code cannot distinguish claim-level verification results from legacy natural-language fact checks.
This phase makes Verifier V2 claim-oriented: Verifier evaluates `executor_structured_output.claims` for whether each claim can be reasonably derived from evidence, emits `claim_checks`, and keeps `facts_checked` only as a compatibility projection.
## What Changes
- Update `chat-verifier-prompt.md` so the primary verification target is `executor_structured_output.claims`.
- Add Verifier V2 output field `claim_checks`.
- Preserve compatibility by mapping `claim_checks` into legacy `facts_checked`.
- Parse and persist `claim_checks` in `ChatService`.
- Ensure `gatekeeper_result.status=fail` cannot result in an effective `PASS`.
- Ensure missing or malformed Executor structured output does not fall back to natural-language fact extraction for PASS.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Verifier output now includes claim-level checks and uses claim derivability as the primary groundedness contract.
## Impact
- Affected prompt: `src/main/resources/prompts/chat-verifier-prompt.md`.
- Affected service: `ChatService.parseVerifierDecision(...)`, retry context generation, verifier persistence.
- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `claim_checks` and keeps `facts_checked`.
- Affected logging: verifier thought summaries may count `claim_checks`.
- Affected tests: `ChatServiceSequentialAgentTest` and focused verifier parsing tests.
- Database schema: no table or column change.
## Non-Goals
- No Composer in this phase.
- No final-answer material filtering in this phase beyond existing LOW_CONFID/REJECT templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
@@ -0,0 +1,109 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** `gatekeeper_result.status` is not `fail`
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts the tool results
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts the tool results
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: Malformed structured output cannot pass through natural language fallback
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
#### Scenario: claim-level verifier output is accepted
- **WHEN** the Verifier checks Executor V2 structured output
- **THEN** the output SHALL contain `verdict`, `groundedness_score`, `critical_fact_count`, `claim_checks`, `facts_checked`, and `rationale`
- **AND** `claim_checks` SHALL be the primary V2 verification result
- **AND** `facts_checked` SHALL remain available for compatibility
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** ChatService SHALL derive compatibility `facts_checked`
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: structured claims are the primary verification target
- **WHEN** `executor_output_parse_status.status` is `valid`
- **AND** `executor_structured_output.claims` is available
- **THEN** Verifier SHALL verify each claim through `claim_checks`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: malformed structured output cannot pass through natural language fallback
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: gatekeeper fail prevents PASS
- **WHEN** `gatekeeper_result.status` is `fail`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL NOT be `PASS`
#### Scenario: invocation reference failure downgrades to reject
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
## ADDED Requirements
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
#### Scenario: claim check verification values are constrained
- **WHEN** Verifier emits `claim_checks`
- **THEN** each item SHALL use one of `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, or `contradicted`
#### Scenario: claim check evidence references remain auditable
- **WHEN** Verifier emits `claim_checks`
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
@@ -0,0 +1,39 @@
## 1. Prompt Contract
- [x] 1.1 Update `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
- [x] 1.2 Remove the valid-structured-output path that scans `executor_final_answer` for extra confirmed facts.
- [x] 1.3 Add `claim_checks` and optional `hypothesis_checks` to the output contract.
- [x] 1.4 Keep `facts_checked` as compatibility output.
- [x] 1.5 State that missing/malformed structured output cannot produce PASS through natural-language fallback.
## 2. Parser And Compatibility Mapping
- [x] 2.1 Extend `VerifierDecision` to store `claim_checks`.
- [x] 2.2 Parse `claim_checks` from verifier output.
- [x] 2.3 Derive compatibility `facts_checked` from `claim_checks` when present.
- [x] 2.4 Preserve old `facts_checked` parsing when `claim_checks` is absent.
- [x] 2.5 Persist `claim_checks` in verifier evaluation.
## 3. Effective Verdict Guardrails
- [x] 3.1 Add code-side guard so `gatekeeper_result.status=fail` cannot result in effective PASS.
- [x] 3.2 Downgrade `evidence.invocation_ref` failures to REJECT.
- [x] 3.3 Downgrade other Gatekeeper failures to at least LOW_CONFID.
- [x] 3.4 Ensure missing/malformed Executor structured output cannot produce effective PASS.
## 4. Compatibility Consumers
- [x] 4.1 Ensure low-confidence rendering still uses compatibility `facts_checked`.
- [x] 4.2 Ensure `buildRetryContext(...)` still receives evidence gaps from compatibility `facts_checked`.
- [x] 4.3 Update verifier logging summary to account for `claim_checks`.
- [x] 4.4 Preserve existing traceability and Gatekeeper audit fields.
## 5. Tests And Verification
- [x] 5.1 Add or update tests for parsing `claim_checks`.
- [x] 5.2 Add tests for `claim_checks` to `facts_checked` compatibility mapping.
- [x] 5.3 Add tests for each verification mapping class: `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, `contradicted`.
- [x] 5.4 Add tests proving Gatekeeper fail cannot remain PASS.
- [x] 5.5 Add tests proving malformed/missing structured output cannot remain PASS.
- [x] 5.6 Run targeted tests.
- [x] 5.7 Validate this OpenSpec change.
@@ -0,0 +1,6 @@
Committed OpenSpec for executor-composer-final-answer.
Commit gate passed on 2026-07-08.
Validated with:
- cmd /c openspec validate executor-composer-final-answer
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,74 @@
# Decisions
## Context Collection
- `devflow/index.md` confirms `executor-v2-output-contract`, `executor-gatekeeper-hook`, and `executor-verifier-claim-checks` are archived.
- `openspec/changes` has no active changes before this phase.
- `openspec/specs/chat-verifier-agent/spec.md` is the existing capability that owns verifier routing and audit behavior.
- No existing `chat-composer-agent` capability exists, so this change introduces it.
## Question Pool
| Question | Type | Resolution |
|---|---|---|
| Should Composer be a new capability or folded into `chat-verifier-agent`? | evidence-driven | New `chat-composer-agent` capability plus modified `chat-verifier-agent` routing. Composer is a distinct expression layer, while ChatService routing remains part of the verifier chain. |
| Can Composer call tools or inspect raw tool output? | evidence-driven | No. The issue and prior design require Composer to receive only Verifier-allowed material. |
| Should Gatekeeper, Verifier, or Planner change in this phase? | evidence-driven | No. Stage four is limited to Composer final-answer generation and routing. |
| What is the interface impact level? | evidence-driven | L2 internal contract change: ChatService internal final-answer semantics change, but no external API or database schema changes. |
## Key Decisions
- Composer is implemented as `chat_composer` or an equivalent model call after Verifier.
- Composer input is assembled by ChatService, not by the model.
- `claim_checks` are authoritative for filtering allowed material.
- REJECT input always has `allowed_hypotheses=[]`.
- Composer malformed output falls back to fixed safe templates.
- Fallback must never use Executor `user_facing_answer` or raw Executor JSON.
- Composer audit is persisted under `verifier_evaluation.composer_output`.
## Cross-Artifact Alignment
| Source | Alignment |
|---|---|
| Issue objective | Stage four in `mvp/issues/executor-structured-output-v2.md` requires Composer output final answer from Verifier-allowed material. Covered by proposal, design, specs, and tasks. |
| Proposal -> design | Proposal says Composer owns final expression; design defines input filtering, output parsing, fallback, and audit. |
| Design -> specs | Design decisions are reflected in `chat-composer-agent` requirements and modified `chat-verifier-agent` routing requirements. |
| Specs -> tasks | Each required behavior has implementation and test tasks, including malformed fallback and no raw output leakage. |
## Commit Gate Notes
- Interface impact: L2 internal.
- No unresolved user-interview question identified.
- No database migration required.
- No OpenSpec/devflow conflict found.
## Apply Notes
- Implemented `chat_composer` as a post-Verifier expression Agent in `ChatService`.
- `ChatService` now builds Composer input from `VerifierDecision` plus parsed Executor structured output, not from raw Executor answer text.
- PASS, LOW_CONFID, and REJECT final-answer paths now use Composer output or a deterministic safe fallback.
- Verifier-missing or Verifier-malformed paths use fixed fallback directly because there is no trustworthy Verifier decision for Composer.
- Composer audit is persisted under `verifier_evaluation.composer_output` without adding a database table.
## Test Drift And Fixes
- Initial test compilation failed because `ChatServiceSequentialAgentTest.java` had a UTF-8 BOM at the file start. Removed the BOM.
- Existing sequential-flow tests still expected the old three-Agent call sequence and temporary V2 renderer behavior. Updated them to include `chat_composer` when a valid Verifier decision exists.
- Tests that exercise fixed fallback now force malformed Composer output so they verify no raw JSON or Executor final answer leakage.
- LOW_CONFID tests were updated to allow indirect support to appear as a possible direction rather than requiring it to disappear from all final-answer text.
## Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,152 @@
## Context
After the first three Executor Structured Output V2 stages, the Chat diagnosis chain is:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> ChatService final rendering
```
Executor now emits structured diagnostic material, Gatekeeper validates deterministic evidence failures, and Verifier emits `claim_checks`. The remaining risk is final answer rendering: the user-facing answer must be produced from Verifier-allowed material, not from raw Executor output or temporary V2 renderers.
Stage four adds the final expression layer:
```text
VerifierDecision + Executor structured output
-> ChatService filters allowed material
-> chat_composer
-> Composer JSON
-> final diagnosis_session.answer
```
## Goals / Non-Goals
**Goals:**
- Add a Composer prompt and model call after Verifier.
- Ensure Composer receives only filtered material derived from Verifier decisions.
- Ensure final user answers for PASS, LOW_CONFID, and REJECT do not read Executor `user_facing_answer`, raw Executor JSON, or raw tool output.
- Preserve safe LOW_CONFID/REJECT degradation when Composer output is malformed.
- Persist Composer output for audit under `verifier_evaluation.composer_output`.
**Non-Goals:**
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier verification-class expansion.
- No database schema migration.
- No stage-five fixture expansion.
## Decisions
### Composer is an expression layer, not a diagnosis layer
Composer SHALL receive only filtered material:
- `original_query`
- `verdict`
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
It SHALL NOT receive raw tool output or the full unscreened Executor output. This keeps diagnosis ownership with Executor + Verifier and prevents Composer from inventing new facts.
Alternative considered: render final answers with deterministic Java templates only. Rejected because PASS and LOW_CONFID answers still need natural, user-readable synthesis; fixed templates become rigid and would push semantic composition back into Executor or Verifier.
### ChatService owns filtering
ChatService filters Executor material using Verifier `claim_checks` before invoking Composer.
Suggested mapping:
| Verifier classification | Composer handling |
|---|---|
| `direct_observation` | include in `allowed_claims` as confirmed material |
| `reasonable_inference` | include in `allowed_claims`, but do not allow “唯一根因” wording unless claim type already supports root cause |
| `overstated` | do not include as confirmed; may become `allowed_hypotheses` or `missing_info` |
| `unsupported` | do not include as confirmed; may become `missing_info` |
| `external_unknown` | do not include as confirmed; may become `missing_info` |
| `contradicted` | do not include as confirmed; favor REJECT-safe output |
For `REJECT`, `allowed_hypotheses` SHALL be empty so the final answer does not keep speculating after a rejected evidence chain.
### Composer output is strict JSON with safe fallback
Composer SHALL output:
```json
{
"answer_summary": "...",
"recommended_actions": [
{
"action_text": "...",
"reason": "..."
}
],
"user_facing_answer": "..."
}
```
If output parsing fails or required fields are missing, ChatService SHALL use fixed safe templates from filtered material. The fallback SHALL NOT display raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
### Audit stays in self_evaluation
No new table is needed. ChatService persists a minimal Composer audit snapshot:
```json
{
"verifier_evaluation": {
"composer_output": {
"status": "valid",
"answer_summary": "...",
"recommended_actions": [],
"user_facing_answer": "..."
}
}
}
```
When Composer fails, `status` should be `malformed` or `fallback`, with a short `detail`. The audit should remain compact and avoid storing full prompt copies.
### Interface impact is L2 internal
The external Chat API still returns a final answer string. Internally, `ChatService` gains Composer input/output handling and final rendering semantics change. This is an internal behavioral contract change because existing PASS behavior can no longer direct-output Executor material.
## Risks / Trade-offs
- Risk: Composer introduces another LLM call and can fail formatting.
- Mitigation: strict JSON contract plus deterministic fallback templates.
- Risk: filtering is too strict and PASS answers become terse.
- Mitigation: include both direct observations and reasonable inferences, but preserve verdict-specific wording constraints.
- Risk: legacy tests expect PASS to use Executor output.
- Mitigation: update tests to assert Composer or safe fallback is the only final-answer source.
- Risk: malformed Composer output could leak raw JSON.
- Mitigation: parse output before exposing it; fallback only from filtered material.
## Migration Plan
1. Add `chat-composer-prompt.md`.
2. Add Composer input assembly and output parsing in `ChatService`.
3. Replace PASS temporary V2 rendering with Composer-or-safe-template rendering.
4. Persist `composer_output` under `verifier_evaluation`.
5. Update tests to cover PASS, LOW_CONFID, REJECT, filtered claims, and malformed output fallback.
Rollback:
- Keep fallback templates available if Composer is disabled or malformed.
- Do not roll back to Executor `user_facing_answer`; that would reintroduce the original evidence-attribution risk.
## Open Questions
None blocking. Stage four defaults:
- Composer is implemented as `chat_composer`.
- Composer does not call tools.
- Composer output failure falls back to fixed safe templates.
- No database schema changes.
@@ -0,0 +1,47 @@
## Why
Stage one removed final-answer fields from Executor, stage two added deterministic Gatekeeper checks, and stage three made Verifier claim-oriented. The remaining gap is final answer rendering: ChatService can still rely on temporary V2 rendering paths instead of a dedicated expression layer, which risks letting unverified Executor material shape user-facing answers.
This phase introduces a Composer expression layer so final Chat answers are generated only from Verifier-allowed material.
## What Changes
- Add a `chat_composer` final-answer generation step after Verifier.
- Add a strict Composer prompt and JSON output contract with `answer_summary`, `recommended_actions`, and `user_facing_answer`.
- Build Composer input in `ChatService` from filtered Verifier results:
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
- Ensure PASS, LOW_CONFID, and REJECT final answers no longer read Executor `user_facing_answer` or raw Executor JSON.
- Persist Composer output in `diagnosis_session.self_evaluation.verifier_evaluation.composer_output`.
- Add safe fallback templates for malformed Composer output that never expose raw Executor JSON or raw Composer JSON.
## Capabilities
### New Capabilities
- `chat-composer-agent`: Final expression layer that turns Verifier-allowed structured material into a readable Chinese user answer without introducing new facts.
### Modified Capabilities
- `chat-verifier-agent`: ChatService final routing changes so Verifier verdicts feed Composer or safe fixed templates instead of direct Executor answer paths.
## Impact
- Affected prompt: add `src/main/resources/prompts/chat-composer-prompt.md`.
- Affected service: `ChatService` final rendering after Verifier, Composer input assembly, Composer output parsing, malformed-output fallback.
- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `composer_output`.
- Affected tests: `ChatServiceSequentialAgentTest` and focused tests around Composer input filtering, final answer routing, and malformed Composer fallback.
- Database schema: no table or column change.
- External API: no endpoint contract change; final answer text semantics become stricter because unverified Executor material is no longer a source for user-visible answers.
## Non-Goals
- No Planner changes.
- No Gatekeeper rule expansion.
- No Verifier classification changes.
- No Executor retry behavior change.
- No database migration.
- No stage-five eval fixture expansion in this phase.
@@ -0,0 +1,80 @@
## ADDED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** ChatService invokes Composer
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer SHALL respect verdict-specific wording
Composer SHALL phrase final answers according to the effective Verifier verdict.
#### Scenario: PASS answer uses confirmed material
- **WHEN** the effective verdict is `PASS`
- **THEN** the final answer MAY state confirmed findings from `allowed_claims`
- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present
#### Scenario: LOW_CONFID answer separates findings and gaps
- **WHEN** the effective verdict is `LOW_CONFID`
- **THEN** the final answer SHALL distinguish confirmed information from possible directions
- **AND** it SHALL mention evidence gaps from `missing_info`
- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings
#### Scenario: REJECT answer avoids root-cause conclusions
- **WHEN** the effective verdict is `REJECT`
- **THEN** Composer input SHALL have `allowed_hypotheses=[]`
- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion
- **AND** the final answer SHALL NOT include a root-cause conclusion
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer returns malformed JSON or omits required fields
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** ChatService persists verifier evaluation
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
- **AND** the audit SHALL include the parsed Composer fields when valid
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,89 @@
## MODIFIED Requirements
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS -> Composer output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
#### Scenario: REJECT -> degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
## ADDED Requirements
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
#### Scenario: PASS uses Composer-safe output
- **WHEN** the final verdict is `PASS`
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- **AND** it SHALL NOT use Executor `user_facing_answer`
- **AND** it SHALL NOT expose raw Executor JSON
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
## REMOVED Requirements
### Requirement: User-facing verifier outputs SHALL follow fixed templates
**Reason**: final user-facing answers are no longer owned by legacy Verifier templates. They must be generated from Verifier-allowed material through Composer or deterministic safe fallback.
**Migration**: use "User-facing verifier outputs SHALL follow Composer-safe protocols" and the new `chat-composer-agent` capability.
@@ -0,0 +1,44 @@
## 1. Composer Prompt Contract
- [x] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
- [x] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
- [x] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
- [x] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
- [x] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
## 2. Composer Invocation And Input Filtering
- [x] 2.1 Load the Composer prompt in `ChatService`.
- [x] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
- [x] 2.3 Build Composer input from Verifier decision and Executor structured output.
- [x] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
- [x] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
- [x] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
- [x] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
- [x] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
## 3. Composer Output Parsing, Fallback, And Audit
- [x] 3.1 Parse Composer strict JSON output.
- [x] 3.2 Add safe fallback rendering for malformed Composer output.
- [x] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- [x] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [x] 3.5 Preserve existing verifier audit fields when writing Composer output.
## 4. Final Answer Routing
- [x] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
- [x] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
- [x] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
- [x] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
## 5. Tests And Verification
- [x] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
- [x] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
- [x] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
- [x] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
- [x] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
- [x] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
- [x] 5.7 Run targeted Maven tests.
- [x] 5.8 Validate this OpenSpec change and all specs.
@@ -0,0 +1,83 @@
# chat-composer-agent Specification
## Purpose
TBD - created by archiving change executor-composer-final-answer. Update Purpose after archive.
## Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** ChatService invokes Composer
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer SHALL respect verdict-specific wording
Composer SHALL phrase final answers according to the effective Verifier verdict.
#### Scenario: PASS answer uses confirmed material
- **WHEN** the effective verdict is `PASS`
- **THEN** the final answer MAY state confirmed findings from `allowed_claims`
- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present
#### Scenario: LOW_CONFID answer separates findings and gaps
- **WHEN** the effective verdict is `LOW_CONFID`
- **THEN** the final answer SHALL distinguish confirmed information from possible directions
- **AND** it SHALL mention evidence gaps from `missing_info`
- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings
#### Scenario: REJECT answer avoids root-cause conclusions
- **WHEN** the effective verdict is `REJECT`
- **THEN** Composer input SHALL have `allowed_hypotheses=[]`
- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion
- **AND** the final answer SHALL NOT include a root-cause conclusion
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer returns malformed JSON or omits required fields
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** ChatService persists verifier evaluation
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
- **AND** the audit SHALL include the parsed Composer fields when valid
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
+180 -68
View File
@@ -4,51 +4,54 @@
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
## Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
- **AND** at least one critical fact has direct evidence
- **AND** no critical fact is contradicted
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** `gatekeeper_result.status` is not `fail`
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical fact contradicts the tool results
- **AND** some critical facts have no supporting evidence
- **WHEN** no critical claim contradicts the tool results
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only indirect support
- **WHEN** no critical fact contradicts the tool results
- **AND** all critical facts are only indirectly supported
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts the tool results
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary`
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence`
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: Extra confirmed-sounding answer facts are still checked
- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims`
- **THEN** Verifier SHALL add those facts to `facts_checked`
- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
#### Scenario: Natural-language fallback remains available
#### Scenario: Malformed structured output cannot pass through natural language fallback
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer`
- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
#### Scenario: claim-level verifier output is accepted
- **WHEN** the Verifier checks Executor V2 structured output
- **THEN** the output SHALL contain `verdict`, `groundedness_score`, `critical_fact_count`, `claim_checks`, `facts_checked`, and `rationale`
- **AND** `claim_checks` SHALL be the primary V2 verification result
- **AND** `facts_checked` SHALL remain available for compatibility
#### Scenario: Output format validation
- **WHEN** the Verifier completes its analysis
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
- **THEN** the output SHALL contain "verdict", "groundedness_score", "claim_checks", "facts_checked", and "rationale" fields
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
@@ -57,14 +60,19 @@ The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_
- **THEN** it SHALL output exactly one JSON object
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
- **AND** the JSON object SHALL include `critical_fact_count`
- **AND** each `claim_checks` item SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
Each checked fact SHALL be labeled using a fixed evidence classification.
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
#### Scenario: fact classification values
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** ChatService SHALL derive compatibility `facts_checked`
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: groundedness_score SHALL be derived from fact classifications
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
@@ -81,54 +89,61 @@ The groundedness score SHALL be computed from critical fact classifications inst
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS → direct output
#### Scenario: PASS -> Composer output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** the system SHALL output the Executor's answer directly
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence补充
- **AND** it SHALL produce a degraded output directly
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
#### Scenario: REJECT → degraded output
#### Scenario: REJECT -> degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
### Requirement: User-facing verifier outputs SHALL follow fixed templates
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
#### Scenario: LOW_CONFID uses disclaimer template
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
@@ -161,7 +176,7 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th
#### Scenario: Verifier remains isolated from intermediate reasoning
- **WHEN** `executor_structured_output` is added to the verifier input
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, gatekeeper result, and retry context
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
@@ -200,6 +215,23 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
#### Scenario: structured claims are the primary verification target
- **WHEN** `executor_output_parse_status.status` is `valid`
- **AND** `executor_structured_output.claims` is available
- **THEN** Verifier SHALL verify each claim through `claim_checks`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: malformed structured output cannot pass through natural language fallback
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
@@ -256,20 +288,24 @@ When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate co
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor final output contains required top-level fields
- **WHEN** Executor completes a Chat diagnosis step
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
@@ -286,10 +322,11 @@ The Chat Executor SHALL produce a machine-checkable final output that separates
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: User-facing answer is available
- **WHEN** Executor emits structured output
- **THEN** `user_facing_answer` SHALL be written in Chinese
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
@@ -309,3 +346,78 @@ The system SHALL tolerate malformed or absent structured Executor output without
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper fail prevents PASS
- **WHEN** `gatekeeper_result.status` is `fail`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL NOT be `PASS`
#### Scenario: invocation reference failure downgrades to reject
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
#### Scenario: claim check verification values are constrained
- **WHEN** Verifier emits `claim_checks`
- **THEN** each item SHALL use one of `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, or `contradicted`
#### Scenario: claim check evidence references remain auditable
- **WHEN** Verifier emits `claim_checks`
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
#### Scenario: PASS uses Composer-safe output
- **WHEN** the final verdict is `PASS`
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- **AND** it SHALL NOT use Executor `user_facing_answer`
- **AND** it SHALL NOT expose raw Executor JSON
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
@@ -204,19 +204,22 @@ public class AgentLoggingHook extends MessagesModelHook {
private String summarizeVerifierThought(String verifierOutput) {
try {
JsonNode root = objectMapper.readTree(verifierOutput);
int claimCount = root.path("claim_checks").isArray() ? root.path("claim_checks").size() : 0;
int factCount = root.path("facts_checked").isArray() ? root.path("facts_checked").size() : 0;
int tracedFactCount = 0;
if (root.path("facts_checked").isArray()) {
for (JsonNode factNode : root.path("facts_checked")) {
if (factNode.path("evidence_refs").isArray() && factNode.path("evidence_refs").size() > 0) {
JsonNode tracedNodes = root.path("claim_checks").isArray() ? root.path("claim_checks") : root.path("facts_checked");
if (tracedNodes.isArray()) {
for (JsonNode node : tracedNodes) {
if (node.path("evidence_refs").isArray() && node.path("evidence_refs").size() > 0) {
tracedFactCount++;
}
}
}
return "verdict=%s, score=%s, critical_fact_count=%s, facts_checked=%d, traced_facts=%d".formatted(
return "verdict=%s, score=%s, critical_fact_count=%s, claim_checks=%d, facts_checked=%d, traced_facts=%d".formatted(
root.path("verdict").asText("UNKNOWN"),
root.path("groundedness_score").asText("0.0"),
root.path("critical_fact_count").asText("0"),
claimCount,
factCount,
tracedFactCount
);
@@ -8,6 +8,7 @@ import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.service.ExecutorGatekeeperService;
import com.superbiz.agent.service.ToolTraceSummaryService;
import com.superbiz.agent.util.SessionContextHolder;
import com.superbiz.agent.util.VerifierContextHolder;
@@ -28,12 +29,19 @@ import java.util.Map;
public class VerifierInputHook extends MessagesModelHook {
private final ToolTraceSummaryService toolTraceSummaryService;
private final ExecutorGatekeeperService executorGatekeeperService;
private final ObjectMapper objectMapper = new ObjectMapper();
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
};
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService) {
this(toolTraceSummaryService, null);
}
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService,
ExecutorGatekeeperService executorGatekeeperService) {
this.toolTraceSummaryService = toolTraceSummaryService;
this.executorGatekeeperService = executorGatekeeperService;
}
@Override
@@ -60,12 +68,16 @@ public class VerifierInputHook extends MessagesModelHook {
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
Map<String, Object> verifierInput = new LinkedHashMap<>();
verifierInput.put("original_query", VerifierContextHolder.getOriginalQuery());
verifierInput.put("executor_final_answer", executorFinalAnswer);
verifierInput.put("executor_structured_output", parseResult.structuredOutput());
verifierInput.put("executor_output_parse_status", parseResult.status());
verifierInput.put("tool_trace_summary", toolTraceSummary);
verifierInput.put("gatekeeper_result", gatekeeperResult);
verifierInput.put("retry_context", VerifierContextHolder.getRetryContext());
String payload = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(verifierInput);
@@ -76,6 +88,29 @@ public class VerifierInputHook extends MessagesModelHook {
}
}
private Map<String, Object> runGatekeeper(String sessionId, ExecutorOutputParseResult parseResult) {
if (executorGatekeeperService == null) {
return passGatekeeperResult();
}
try {
return executorGatekeeperService.validate(sessionId, parseResult.structuredOutput(), parseResult.status());
} catch (Exception e) {
log.error("Gatekeeper validation failed unexpectedly", e);
return executorGatekeeperService.fail("gatekeeper.internal_error",
"gatekeeper",
e.getMessage() == null ? "gatekeeper validation failed" : e.getMessage());
}
}
private Map<String, Object> passGatekeeperResult() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", "pass");
result.put("failed_rules", List.of());
result.put("warnings", List.of());
result.put("errors", List.of());
return result;
}
private ExecutorOutputParseResult parseExecutorOutput(String executorFinalAnswer) {
if (executorFinalAnswer == null || executorFinalAnswer.isBlank()) {
return new ExecutorOutputParseResult(null, status("missing", "executor_final_answer is blank"));
@@ -112,6 +112,9 @@ public class ChatService {
@Autowired
private SelfEvaluationMergeService selfEvaluationMergeService;
@Autowired
private ExecutorGatekeeperService executorGatekeeperService;
@Value("${verifier.low-confidence-threshold:0.5}")
private double verifierLowConfidenceThreshold;
@@ -122,6 +125,7 @@ public class ChatService {
private String chatPlannerPrompt;
private String chatExecutorPrompt;
private String chatVerifierPrompt;
private String chatComposerPrompt;
private final ObjectMapper objectMapper = new ObjectMapper();
@PostConstruct
@@ -137,6 +141,9 @@ public class ChatService {
chatVerifierPrompt = new String(
new ClassPathResource("prompts/chat-verifier-prompt.md").getInputStream().readAllBytes(),
StandardCharsets.UTF_8);
chatComposerPrompt = new String(
new ClassPathResource("prompts/chat-composer-prompt.md").getInputStream().readAllBytes(),
StandardCharsets.UTF_8);
logger.info("Chat 多 Agent Prompts 加载成功");
} catch (IOException e) {
logger.error("加载 Chat Prompt 文件失败", e);
@@ -417,8 +424,9 @@ public class ChatService {
Optional<OverAllState> stateOptional = workflow.invoke(workflowInput, config);
if (stateOptional.isEmpty()) {
finalDecision = buildVerifierFallbackDecision(round, "workflow 未返回有效状态");
answer = buildLowConfidenceOutput(answer, finalDecision);
persistVerifierEvaluation(session, finalDecision, round);
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
answer = renderResult.answer();
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
break;
}
@@ -438,21 +446,23 @@ public class ChatService {
if (finalDecision == null) {
finalDecision = buildVerifierFallbackDecision(round, "verifier_output 缺失或无法解析");
answer = buildLowConfidenceOutput(answer, finalDecision);
persistVerifierEvaluation(session, finalDecision, round);
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
answer = renderResult.answer();
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
break;
}
if ("PASS".equals(finalDecision.verdict())) {
answer = extractUserFacingAnswer(answer)
.orElse(answer == null || answer.isBlank() ? "抱歉,多 Agent 分析未能生成有效结论。" : answer);
persistVerifierEvaluation(session, finalDecision, round);
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
break;
}
if ("REJECT".equals(finalDecision.verdict())) {
answer = buildDegradedOutput(finalDecision);
persistVerifierEvaluation(session, finalDecision, round);
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
break;
}
@@ -460,8 +470,9 @@ public class ChatService {
&& finalDecision.groundednessScore() < verifierLowConfidenceThreshold
&& round < 2;
if (!shouldRetry) {
answer = buildLowConfidenceOutput(answer, finalDecision);
persistVerifierEvaluation(session, finalDecision, round);
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
answer = renderResult.answer();
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
break;
}
@@ -537,11 +548,22 @@ public class ChatService {
.model(chatModel)
.systemPrompt(chatVerifierPrompt)
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"),
new VerifierInputHook(toolTraceSummaryService))
new VerifierInputHook(toolTraceSummaryService, executorGatekeeperService))
.outputKey("verifier_output")
.build();
}
private ReactAgent buildChatComposerAgent(ChatModel chatModel) {
return ReactAgent.builder()
.name("chat_composer")
.description("负责将 Verifier 允许的材料组织为最终用户答案")
.model(chatModel)
.systemPrompt(chatComposerPrompt)
.hooks(new AgentLoggingHook(agentStepRepository, "composer"))
.outputKey("composer_output")
.build();
}
private ReactAgent buildChatExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
List<Map<String, String>> history, String retryContext) {
StringBuilder prompt = new StringBuilder(chatExecutorPrompt);
@@ -643,12 +665,18 @@ public class ChatService {
try {
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(verifierOutput));
List<Map<String, Object>> factsChecked = parseFactsChecked(root.path("facts_checked"));
List<Map<String, Object>> claimChecks = parseClaimChecks(root.path("claim_checks"));
List<Map<String, Object>> factsChecked = claimChecks.isEmpty()
? parseFactsChecked(root.path("facts_checked"))
: mapClaimChecksToFactsChecked(claimChecks);
String verdict = effectiveVerifierVerdict(root.path("verdict").asText("LOW_CONFID"));
int criticalFactCount = root.path("critical_fact_count").asInt(countCriticalFacts(factsChecked));
return new VerifierDecision(
root.path("verdict").asText("LOW_CONFID"),
verdict,
root.path("groundedness_score").asDouble(0.0),
root.path("critical_fact_count").asInt(0),
criticalFactCount,
claimChecks,
factsChecked,
root.path("rationale").asText(""),
round
@@ -671,6 +699,104 @@ public class ChatService {
return trimmed;
}
private String effectiveVerifierVerdict(String modelVerdict) {
String verdict = normalizeVerdict(modelVerdict);
Map<String, Object> parseStatus = VerifierContextHolder.getExecutorOutputParseStatus();
String parseState = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
if (("missing".equals(parseState) || "malformed".equals(parseState)) && "PASS".equals(verdict)) {
return "LOW_CONFID";
}
Map<String, Object> gatekeeperResult = VerifierContextHolder.getGatekeeperResult();
if (gatekeeperResult == null || !"fail".equals(String.valueOf(gatekeeperResult.get("status")))) {
return verdict;
}
if (containsRule(gatekeeperResult.get("failed_rules"), ExecutorGatekeeperService.RULE_INVOCATION_REF)) {
return "REJECT";
}
return "PASS".equals(verdict) ? "LOW_CONFID" : verdict;
}
private String normalizeVerdict(String verdict) {
if ("PASS".equals(verdict) || "LOW_CONFID".equals(verdict) || "REJECT".equals(verdict)) {
return verdict;
}
return "LOW_CONFID";
}
private boolean containsRule(Object rulesValue, String ruleId) {
if (!(rulesValue instanceof List<?> rules)) {
return false;
}
return rules.stream().anyMatch(rule -> ruleId.equals(String.valueOf(rule)));
}
private int countCriticalFacts(List<Map<String, Object>> factsChecked) {
return (int) factsChecked.stream()
.filter(fact -> Boolean.TRUE.equals(fact.get("is_critical")))
.count();
}
private List<Map<String, Object>> parseClaimChecks(JsonNode claimChecksNode) {
List<Map<String, Object>> claimChecks = new ArrayList<>();
if (!claimChecksNode.isArray()) {
return claimChecks;
}
for (JsonNode claimNode : claimChecksNode) {
Map<String, Object> claimCheck = new LinkedHashMap<>();
claimCheck.put("claim_id", claimNode.path("claim_id").asText(""));
claimCheck.put("claim_text", claimNode.path("claim_text").asText(""));
claimCheck.put("claim_type", claimNode.path("claim_type").asText(""));
claimCheck.put("verification", normalizeClaimVerification(claimNode.path("verification").asText("unsupported")));
claimCheck.put("detail", claimNode.path("detail").asText(""));
claimCheck.put("evidence_refs", parseEvidenceRefs(claimNode.path("evidence_refs")));
claimChecks.add(claimCheck);
}
return claimChecks;
}
private String normalizeClaimVerification(String verification) {
return switch (verification) {
case "direct_observation", "reasonable_inference", "overstated", "unsupported",
"external_unknown", "contradicted" -> verification;
default -> "unsupported";
};
}
private List<Map<String, Object>> mapClaimChecksToFactsChecked(List<Map<String, Object>> claimChecks) {
List<Map<String, Object>> factsChecked = new ArrayList<>();
for (Map<String, Object> claimCheck : claimChecks) {
String claimId = String.valueOf(claimCheck.getOrDefault("claim_id", ""));
String claimText = String.valueOf(claimCheck.getOrDefault("claim_text", ""));
String claimType = String.valueOf(claimCheck.getOrDefault("claim_type", ""));
Map<String, Object> fact = new LinkedHashMap<>();
fact.put("fact", claimId.isBlank() ? claimText : claimId + ": " + claimText);
fact.put("is_critical", isCriticalClaimType(claimType));
fact.put("verification", mapClaimVerificationToFactVerification(
String.valueOf(claimCheck.getOrDefault("verification", "unsupported"))));
fact.put("detail", claimCheck.getOrDefault("detail", ""));
fact.put("evidence_refs", claimCheck.getOrDefault("evidence_refs", List.of()));
factsChecked.add(fact);
}
return factsChecked;
}
private boolean isCriticalClaimType(String claimType) {
return "root_cause".equals(claimType)
|| "symptom".equals(claimType)
|| "impact".equals(claimType)
|| "risk".equals(claimType);
}
private String mapClaimVerificationToFactVerification(String verification) {
return switch (verification) {
case "direct_observation" -> "direct_evidence";
case "reasonable_inference", "overstated" -> "indirect_support";
case "contradicted" -> "contradicted";
default -> "no_evidence";
};
}
private List<Map<String, Object>> parseFactsChecked(JsonNode factsNode) {
List<Map<String, Object>> factsChecked = new ArrayList<>();
if (!factsNode.isArray()) {
@@ -716,7 +842,7 @@ public class ChatService {
}
private VerifierDecision buildVerifierFallbackDecision(int round, String rationale) {
return new VerifierDecision("LOW_CONFID", 0.0, 0, List.of(), rationale, round);
return new VerifierDecision("LOW_CONFID", 0.0, 0, List.of(), List.of(), rationale, round);
}
private String extractStateText(Optional<OverAllState> stateOptional, String key) {
@@ -734,6 +860,11 @@ public class ChatService {
}
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round) {
persistVerifierEvaluation(session, decision, round, null);
}
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round,
Map<String, Object> composerOutput) {
if (decision == null) {
return;
}
@@ -741,6 +872,7 @@ public class ChatService {
verifierEvaluation.put("verdict", decision.verdict());
verifierEvaluation.put("groundedness_score", decision.groundednessScore());
verifierEvaluation.put("critical_fact_count", decision.criticalFactCount());
verifierEvaluation.put("claim_checks", decision.claimChecks());
verifierEvaluation.put("facts_checked", decision.factsChecked());
verifierEvaluation.put("rationale", decision.rationale());
verifierEvaluation.put("round", round);
@@ -751,12 +883,340 @@ public class ChatService {
verifierEvaluation.put("executor_structured_output", VerifierContextHolder.getExecutorStructuredOutput());
verifierEvaluation.put("tool_trace_summary",
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
verifierEvaluation.put("gatekeeper_result",
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
.orElse(Map.of("status", "pass", "failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
if (composerOutput != null) {
verifierEvaluation.put("composer_output", composerOutput);
}
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(session.getSelfEvaluation(), verifierEvaluation);
session.setSelfEvaluation(merged);
diagnosisSessionRepository.save(session);
}
private ComposerRenderResult composeFinalAnswer(ChatModel chatModel, String originalQuery,
VerifierDecision decision, RunnableConfig config) {
Map<String, Object> composerInput = buildComposerInput(originalQuery, decision);
try {
ReactAgent composer = buildChatComposerAgent(chatModel);
String composerOutput = composer.call(objectMapper.writeValueAsString(composerInput), config).getText();
return parseComposerOutput(composerOutput, composerInput);
} catch (Exception e) {
logger.error("chat_composer 执行失败,使用安全降级模板", e);
return buildFixedFallbackAnswer(composerInput, "composer_exception");
}
}
private ComposerRenderResult buildFixedFallbackAnswer(String originalQuery, VerifierDecision decision) {
return buildFixedFallbackAnswer(buildComposerInput(originalQuery, decision), "fixed_fallback");
}
private Map<String, Object> buildComposerInput(String originalQuery, VerifierDecision decision) {
Map<String, Object> input = new LinkedHashMap<>();
input.put("original_query", originalQuery);
input.put("verdict", decision.verdict());
Map<String, Map<String, Object>> claimsById = indexExecutorClaims();
List<Map<String, Object>> allowedClaims = new ArrayList<>();
List<Map<String, Object>> allowedHypotheses = new ArrayList<>();
List<String> missingInfo = extractStructuredMissingInfo();
List<Map<String, Object>> recommendedActions = extractStructuredRecommendedActions();
if (decision.claimChecks().isEmpty()) {
addLegacyFactsToComposerInput(decision, allowedClaims, allowedHypotheses, missingInfo);
} else {
for (Map<String, Object> check : decision.claimChecks()) {
String verification = String.valueOf(check.getOrDefault("verification", "unsupported"));
String claimId = String.valueOf(check.getOrDefault("claim_id", ""));
String claimText = String.valueOf(check.getOrDefault("claim_text", ""));
String detail = String.valueOf(check.getOrDefault("detail", ""));
Map<String, Object> claim = buildComposerClaim(claimsById.get(claimId), check);
switch (verification) {
case "direct_observation", "reasonable_inference" -> allowedClaims.add(claim);
case "overstated" -> allowedHypotheses.add(Map.of(
"hypothesis_text", claimText,
"basis", detail.isBlank() ? "当前证据只能支持部分判断,不能作为确认结论" : detail
));
case "unsupported", "external_unknown", "contradicted" ->
addMissingInfo(missingInfo, claimText, detail);
default -> addMissingInfo(missingInfo, claimText, detail);
}
}
}
if ("REJECT".equals(decision.verdict())) {
allowedHypotheses = List.of();
}
input.put("allowed_claims", allowedClaims);
input.put("allowed_hypotheses", allowedHypotheses);
input.put("missing_info", missingInfo);
input.put("recommended_actions", recommendedActions);
input.put("rationale", decision.rationale());
return input;
}
private Map<String, Map<String, Object>> indexExecutorClaims() {
Map<String, Map<String, Object>> claimsById = new LinkedHashMap<>();
Map<String, Object> structuredOutput = VerifierContextHolder.getExecutorStructuredOutput();
Object claims = structuredOutput == null ? null : structuredOutput.get("claims");
if (claims instanceof List<?> claimList) {
for (Object item : claimList) {
if (item instanceof Map<?, ?> rawClaim) {
Map<String, Object> claim = new LinkedHashMap<>();
rawClaim.forEach((key, value) -> claim.put(String.valueOf(key), value));
String claimId = String.valueOf(claim.getOrDefault("claim_id", ""));
if (!claimId.isBlank()) {
claimsById.put(claimId, claim);
}
}
}
}
return claimsById;
}
private Map<String, Object> buildComposerClaim(Map<String, Object> executorClaim, Map<String, Object> claimCheck) {
Map<String, Object> claim = new LinkedHashMap<>();
claim.put("claim_id", valueFrom(executorClaim, claimCheck, "claim_id"));
claim.put("claim_type", valueFrom(executorClaim, claimCheck, "claim_type"));
claim.put("claim_text", valueFrom(executorClaim, claimCheck, "claim_text"));
claim.put("support_level", executorClaim == null ? "" : String.valueOf(executorClaim.getOrDefault("support_level", "")));
claim.put("verification", String.valueOf(claimCheck.getOrDefault("verification", "")));
claim.put("detail", String.valueOf(claimCheck.getOrDefault("detail", "")));
return claim;
}
private String valueFrom(Map<String, Object> primary, Map<String, Object> fallback, String key) {
Object value = primary == null ? null : primary.get(key);
if (value == null || String.valueOf(value).isBlank()) {
value = fallback.get(key);
}
return value == null ? "" : String.valueOf(value);
}
private List<String> extractStructuredMissingInfo() {
List<String> missingInfo = new ArrayList<>();
Map<String, Object> structuredOutput = VerifierContextHolder.getExecutorStructuredOutput();
Object missing = structuredOutput == null ? null : structuredOutput.get("missing_info");
if (missing instanceof List<?> missingList) {
for (Object item : missingList) {
String text = String.valueOf(item);
if (!text.isBlank() && !missingInfo.contains(text)) {
missingInfo.add(text);
}
}
}
return missingInfo;
}
private List<Map<String, Object>> extractStructuredRecommendedActions() {
List<Map<String, Object>> actions = new ArrayList<>();
Map<String, Object> structuredOutput = VerifierContextHolder.getExecutorStructuredOutput();
Object recommendedActions = structuredOutput == null ? null : structuredOutput.get("recommended_actions");
if (recommendedActions instanceof List<?> actionList) {
for (Object item : actionList) {
if (item instanceof Map<?, ?> rawAction) {
Map<String, Object> action = new LinkedHashMap<>();
action.put("action_text", textValue(rawAction.get("action_text")));
action.put("reason", textValue(rawAction.get("reason")));
if (!String.valueOf(action.get("action_text")).isBlank()) {
actions.add(action);
}
}
}
}
return actions;
}
private void addLegacyFactsToComposerInput(VerifierDecision decision, List<Map<String, Object>> allowedClaims,
List<Map<String, Object>> allowedHypotheses,
List<String> missingInfo) {
for (Map<String, Object> fact : decision.factsChecked()) {
String verification = String.valueOf(fact.getOrDefault("verification", ""));
String factText = String.valueOf(fact.getOrDefault("fact", ""));
String detail = String.valueOf(fact.getOrDefault("detail", ""));
if ("direct_evidence".equals(verification)) {
allowedClaims.add(Map.of(
"claim_id", "",
"claim_type", "",
"claim_text", factText,
"support_level", "direct",
"verification", verification,
"detail", detail
));
} else if ("indirect_support".equals(verification)) {
allowedHypotheses.add(Map.of(
"hypothesis_text", factText,
"basis", detail.isBlank() ? "当前仅有间接支持,不能作为确认结论" : detail
));
} else {
addMissingInfo(missingInfo, factText, detail);
}
}
}
private void addMissingInfo(List<String> missingInfo, String text, String detail) {
if (text == null || text.isBlank()) {
return;
}
String value = detail == null || detail.isBlank() ? text : text + ":" + detail;
if (!missingInfo.contains(value)) {
missingInfo.add(value);
}
}
private ComposerRenderResult parseComposerOutput(String composerOutput, Map<String, Object> composerInput) {
try {
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(composerOutput));
String answerSummary = root.path("answer_summary").asText("");
String userFacingAnswer = root.path("user_facing_answer").asText("");
if (answerSummary.isBlank() || userFacingAnswer.isBlank() || !root.path("recommended_actions").isArray()) {
return buildFixedFallbackAnswer(composerInput, "composer_schema_invalid");
}
Map<String, Object> audit = new LinkedHashMap<>();
audit.put("status", "valid");
audit.put("answer_summary", answerSummary);
audit.put("recommended_actions", parseComposerActions(root.path("recommended_actions")));
audit.put("user_facing_answer", userFacingAnswer);
return new ComposerRenderResult(userFacingAnswer, audit);
} catch (Exception e) {
logger.warn("解析 composer_output 失败,使用安全降级模板");
return buildFixedFallbackAnswer(composerInput, "composer_malformed");
}
}
private List<Map<String, Object>> parseComposerActions(JsonNode actionsNode) {
List<Map<String, Object>> actions = new ArrayList<>();
if (!actionsNode.isArray()) {
return actions;
}
for (JsonNode actionNode : actionsNode) {
Map<String, Object> action = new LinkedHashMap<>();
action.put("action_text", actionNode.path("action_text").asText(""));
action.put("reason", actionNode.path("reason").asText(""));
actions.add(action);
}
return actions;
}
private ComposerRenderResult buildFixedFallbackAnswer(Map<String, Object> composerInput, String status) {
String answer = renderSafeFallback(composerInput);
Map<String, Object> audit = new LinkedHashMap<>();
audit.put("status", status);
audit.put("detail", "used safe fallback rendering");
audit.put("answer_summary", firstSentence(answer));
audit.put("recommended_actions", composerInput.getOrDefault("recommended_actions", List.of()));
audit.put("user_facing_answer", answer);
return new ComposerRenderResult(answer, audit);
}
@SuppressWarnings("unchecked")
private String renderSafeFallback(Map<String, Object> composerInput) {
String verdict = String.valueOf(composerInput.getOrDefault("verdict", "LOW_CONFID"));
List<Map<String, Object>> allowedClaims =
(List<Map<String, Object>>) composerInput.getOrDefault("allowed_claims", List.of());
List<Map<String, Object>> allowedHypotheses =
(List<Map<String, Object>>) composerInput.getOrDefault("allowed_hypotheses", List.of());
List<String> missingInfo = (List<String>) composerInput.getOrDefault("missing_info", List.of());
List<Map<String, Object>> recommendedActions =
(List<Map<String, Object>>) composerInput.getOrDefault("recommended_actions", List.of());
StringBuilder output = new StringBuilder();
if ("REJECT".equals(verdict)) {
output.append(DEGRADED_PREFIX);
} else if ("LOW_CONFID".equals(verdict)) {
output.append(LOW_CONFID_DISCLAIMER);
}
output.append("\n\n已确认信息:");
if (allowedClaims.isEmpty()) {
output.append("\n- 暂无可稳定确认的信息");
} else {
for (Map<String, Object> claim : allowedClaims) {
String text = String.valueOf(claim.getOrDefault("claim_text", ""));
if (!text.isBlank()) {
output.append("\n- ").append(text);
}
}
}
if (!"REJECT".equals(verdict) && !allowedHypotheses.isEmpty()) {
output.append("\n\n可能方向:");
for (Map<String, Object> hypothesis : allowedHypotheses) {
output.append("\n- ").append(hypothesis.getOrDefault("hypothesis_text", ""));
String basis = String.valueOf(hypothesis.getOrDefault("basis", ""));
if (!basis.isBlank()) {
output.append("(").append(basis).append(")");
}
}
}
output.append("\n\n").append("REJECT".equals(verdict) ? "证据缺口:" : "当前缺口:");
if (missingInfo.isEmpty()) {
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
} else {
for (String gap : missingInfo) {
output.append("\n- ").append(gap);
}
}
output.append("\n\n建议下一步:");
if (recommendedActions.isEmpty()) {
for (String suggestion : buildNextStepSuggestionsFromTrace()) {
output.append("\n- ").append(suggestion);
}
} else {
for (Map<String, Object> action : recommendedActions) {
String text = String.valueOf(action.getOrDefault("action_text", ""));
if (!text.isBlank()) {
output.append("\n- ").append(text);
String reason = String.valueOf(action.getOrDefault("reason", ""));
if (!reason.isBlank()) {
output.append(":").append(reason);
}
}
}
}
return output.toString().trim();
}
private String firstSentence(String text) {
if (text == null || text.isBlank()) {
return "";
}
int end = text.indexOf('\n');
return end < 0 ? text : text.substring(0, end);
}
private String textValue(Object value) {
if (value == null) {
return "";
}
String text = String.valueOf(value);
return "null".equals(text) ? "" : text;
}
private List<String> buildNextStepSuggestionsFromTrace() {
List<String> suggestions = new ArrayList<>();
List<Map<String, Object>> toolSummary = toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null);
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
if (!hasKnowledgeTool) {
suggestions.add("补充知识库或业务文档检索结果,建立可引用的证据锚点");
}
if (hasFailedEvidence) {
suggestions.add("优先重试失败的证据型查询,补齐日志、指标或知识库侧证据");
}
if (suggestions.isEmpty()) {
suggestions.add("围绕上述证据缺口补充只读查询,再由人工复核最终结论");
}
return suggestions;
}
private String buildRetryContext(VerifierDecision decision) {
try {
List<String> missingFacts = extractEvidenceGaps(decision);
@@ -771,95 +1231,6 @@ public class ChatService {
}
}
private String buildLowConfidenceOutput(String executorAnswer, VerifierDecision decision) {
StringBuilder output = new StringBuilder(LOW_CONFID_DISCLAIMER);
List<String> confirmedFacts = extractConfirmedFacts(decision);
output.append("\n\n已确认信息:");
if (confirmedFacts.isEmpty()) {
output.append("\n- 暂无可稳定确认的信息");
} else {
for (String fact : confirmedFacts) {
output.append("\n- ").append(fact);
}
}
List<String> gaps = extractEvidenceGaps(decision);
output.append("\n\n当前缺口:");
if (!gaps.isEmpty()) {
for (String gap : gaps) {
output.append("\n- ").append(gap);
}
} else {
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
}
output.append("\n\n建议下一步:");
for (String suggestion : buildNextStepSuggestions(decision)) {
output.append("\n- ").append(suggestion);
}
return output.toString();
}
private Optional<String> extractUserFacingAnswer(String executorAnswer) {
if (executorAnswer == null || executorAnswer.isBlank()) {
return Optional.empty();
}
try {
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(executorAnswer));
String userFacingAnswer = root.path("user_facing_answer").asText("");
if (!userFacingAnswer.isBlank()) {
return Optional.of(userFacingAnswer);
}
} catch (Exception e) {
logger.debug("Executor answer is not structured JSON, keep raw answer");
}
return Optional.empty();
}
private String buildDegradedOutput(VerifierDecision decision) {
StringBuilder output = new StringBuilder(DEGRADED_PREFIX);
List<String> confirmedFacts = extractConfirmedFacts(decision);
List<String> gaps = extractEvidenceGaps(decision);
List<String> suggestions = buildNextStepSuggestions(decision);
output.append("\n\n已确认信息:");
if (confirmedFacts.isEmpty()) {
output.append("\n- 暂无可稳定确认的信息");
} else {
for (String fact : confirmedFacts) {
output.append("\n- ").append(fact);
}
}
output.append("\n\n证据缺口:");
if (gaps.isEmpty()) {
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
} else {
for (String gap : gaps) {
output.append("\n- ").append(gap);
}
}
output.append("\n\n建议下一步:");
for (String suggestion : suggestions) {
output.append("\n- ").append(suggestion);
}
return output.toString();
}
private List<String> extractConfirmedFacts(VerifierDecision decision) {
List<String> confirmedFacts = new ArrayList<>();
for (Map<String, Object> fact : decision.factsChecked()) {
String verification = String.valueOf(fact.get("verification"));
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
if (critical && "direct_evidence".equals(verification)) {
confirmedFacts.add(String.valueOf(fact.get("fact")));
}
}
return confirmedFacts;
}
private List<String> extractEvidenceGaps(VerifierDecision decision) {
List<String> gaps = new ArrayList<>();
for (Map<String, Object> fact : decision.factsChecked()) {
@@ -881,34 +1252,20 @@ public class ChatService {
return gaps;
}
private List<String> buildNextStepSuggestions(VerifierDecision decision) {
List<String> suggestions = new ArrayList<>();
List<Map<String, Object>> toolSummary = toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null);
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
if (!hasKnowledgeTool) {
suggestions.add("补充知识库或业务文档检索结果,建立可引用的证据锚点");
}
if (hasFailedEvidence) {
suggestions.add("优先重试失败的证据型查询,补齐日志、指标或知识库侧证据");
}
if (suggestions.isEmpty()) {
suggestions.add("围绕上述证据缺口补充只读查询,再由人工复核最终结论");
}
return suggestions;
}
private record VerifierDecision(
String verdict,
double groundednessScore,
int criticalFactCount,
List<Map<String, Object>> claimChecks,
List<Map<String, Object>> factsChecked,
String rationale,
int round
) {
}
private record ComposerRenderResult(String answer, Map<String, Object> audit) {
}
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
private void backfillSessionMetrics(DiagnosisSession session) {
try {
@@ -0,0 +1,237 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.springframework.stereotype.Service;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Objects;
import java.util.Set;
import java.util.function.Function;
import java.util.stream.Collectors;
/**
* Deterministic checks for Executor structured output before verifier reasoning.
*/
@Service
public class ExecutorGatekeeperService {
public static final String STATUS_PASS = "pass";
public static final String STATUS_WARN = "warn";
public static final String STATUS_FAIL = "fail";
public static final String RULE_SCHEMA = "schema.executor_v2";
public static final String RULE_INVOCATION_REF = "evidence.invocation_ref";
private final ToolInvocationRepository toolInvocationRepository;
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
this.toolInvocationRepository = toolInvocationRepository;
}
public Map<String, Object> validate(String sessionId,
Map<String, Object> structuredOutput,
Map<String, Object> parseStatus) {
GatekeeperResult result = new GatekeeperResult();
validateSchema(structuredOutput, parseStatus, result);
if (structuredOutput != null) {
validateInvocationRefs(sessionId, structuredOutput, result);
}
return result.toMap();
}
public Map<String, Object> pass() {
return new GatekeeperResult().toMap();
}
public Map<String, Object> fail(String ruleId, String target, String message) {
GatekeeperResult result = new GatekeeperResult();
result.fail(ruleId, target, message);
return result.toMap();
}
private void validateSchema(Map<String, Object> structuredOutput,
Map<String, Object> parseStatus,
GatekeeperResult result) {
String status = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
if (structuredOutput == null) {
if ("valid".equals(status)) {
result.fail(RULE_SCHEMA, "executor_structured_output", "structured output is missing after valid parse");
}
return;
}
if (!"executor_evidence_v2".equals(String.valueOf(structuredOutput.get("answer_version")))) {
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2");
}
if (structuredOutput.containsKey("diagnosis_summary")) {
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2");
}
if (structuredOutput.containsKey("user_facing_answer")) {
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2");
}
Object claimsValue = structuredOutput.get("claims");
if (!(claimsValue instanceof List<?> claims)) {
result.fail(RULE_SCHEMA, "claims", "claims must be an array");
return;
}
for (int i = 0; i < claims.size(); i++) {
String target = "claims[" + i + "]";
Object claimValue = claims.get(i);
if (!(claimValue instanceof Map<?, ?> claim)) {
result.fail(RULE_SCHEMA, target, "claim must be an object");
continue;
}
requireString(claim, "claim_id", target, result);
requireString(claim, "claim_type", target, result);
requireString(claim, "claim_text", target, result);
String supportLevel = stringValue(claim.get("support_level"));
if (!"direct".equals(supportLevel) && !"indirect".equals(supportLevel)) {
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect");
}
Object bindings = claim.get("evidence_bindings");
if (!(bindings instanceof List<?> bindingList) || bindingList.isEmpty()) {
result.fail(RULE_SCHEMA, target + ".evidence_bindings", "claims must include non-empty evidence_bindings");
}
}
requireArray(structuredOutput, "hypotheses", result);
requireArray(structuredOutput, "recommended_actions", result);
requireArray(structuredOutput, "missing_info", result);
}
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
if (sessionId == null || sessionId.isBlank()) {
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_ids");
return;
}
Map<Long, ToolInvocation> validInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)
.stream()
.filter(invocation -> invocation.getId() != null)
.collect(Collectors.toMap(ToolInvocation::getId, Function.identity(), (left, right) -> left));
Object claimsValue = structuredOutput.get("claims");
if (!(claimsValue instanceof List<?> claims)) {
return;
}
for (int claimIndex = 0; claimIndex < claims.size(); claimIndex++) {
Object claimValue = claims.get(claimIndex);
if (!(claimValue instanceof Map<?, ?> claim)) {
continue;
}
Object bindingsValue = claim.get("evidence_bindings");
if (!(bindingsValue instanceof List<?> bindings)) {
continue;
}
for (int bindingIndex = 0; bindingIndex < bindings.size(); bindingIndex++) {
String target = "claims[" + claimIndex + "].evidence_bindings[" + bindingIndex + "]";
Object bindingValue = bindings.get(bindingIndex);
if (!(bindingValue instanceof Map<?, ?> binding)) {
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object");
continue;
}
validateBindingInvocationIds(binding, validInvocations, target, result);
}
}
}
private void validateBindingInvocationIds(Map<?, ?> binding,
Map<Long, ToolInvocation> validInvocations,
String target,
GatekeeperResult result) {
Object idsValue = binding.get("source_invocation_ids");
if (!(idsValue instanceof List<?> ids) || ids.isEmpty()) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must be a non-empty array");
return;
}
String claimedToolName = stringValue(binding.get("tool_name"));
if (claimedToolName.isBlank()) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required");
}
Set<Long> checkedIds = new HashSet<>();
for (Object idValue : ids) {
Long id = asLong(idValue);
if (id == null) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must contain numeric ids");
continue;
}
if (!checkedIds.add(id)) {
continue;
}
ToolInvocation invocation = validInvocations.get(id);
if (invocation == null) {
result.fail(RULE_INVOCATION_REF, target, "source_invocation_ids not found in current session: " + id);
continue;
}
if (!claimedToolName.isBlank() && !Objects.equals(claimedToolName, invocation.getToolName())) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
}
}
}
private void requireArray(Map<String, Object> output, String field, GatekeeperResult result) {
if (!(output.get(field) instanceof List<?>)) {
result.fail(RULE_SCHEMA, field, field + " must be an array");
}
}
private void requireString(Map<?, ?> object, String field, String target, GatekeeperResult result) {
if (stringValue(object.get(field)).isBlank()) {
result.fail(RULE_SCHEMA, target + "." + field, field + " is required");
}
}
private String stringValue(Object value) {
return value == null ? "" : String.valueOf(value);
}
private Long asLong(Object value) {
if (value instanceof Number number) {
return number.longValue();
}
if (value instanceof String text) {
try {
return Long.parseLong(text);
} catch (NumberFormatException ignored) {
return null;
}
}
return null;
}
private static final class GatekeeperResult {
private final List<String> failedRules = new ArrayList<>();
private final List<String> warnings = new ArrayList<>();
private final List<Map<String, Object>> errors = new ArrayList<>();
void fail(String ruleId, String target, String message) {
if (!failedRules.contains(ruleId)) {
failedRules.add(ruleId);
}
Map<String, Object> error = new LinkedHashMap<>();
error.put("rule_id", ruleId);
error.put("target", target);
error.put("message", message);
errors.add(error);
}
Map<String, Object> toMap() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", failedRules.isEmpty() ? (warnings.isEmpty() ? STATUS_PASS : STATUS_WARN) : STATUS_FAIL);
result.put("failed_rules", failedRules);
result.put("warnings", warnings);
result.put("errors", errors);
return result;
}
}
}
@@ -14,6 +14,7 @@ public final class VerifierContextHolder {
private static final ThreadLocal<Map<String, Object>> EXECUTOR_STRUCTURED_OUTPUT = new ThreadLocal<>();
private static final ThreadLocal<Map<String, Object>> EXECUTOR_OUTPUT_PARSE_STATUS = new ThreadLocal<>();
private static final ThreadLocal<List<Map<String, Object>>> TOOL_TRACE_SUMMARY = new ThreadLocal<>();
private static final ThreadLocal<Map<String, Object>> GATEKEEPER_RESULT = new ThreadLocal<>();
private VerifierContextHolder() {
}
@@ -66,6 +67,14 @@ public final class VerifierContextHolder {
return TOOL_TRACE_SUMMARY.get();
}
public static void setGatekeeperResult(Map<String, Object> gatekeeperResult) {
GATEKEEPER_RESULT.set(gatekeeperResult);
}
public static Map<String, Object> getGatekeeperResult() {
return GATEKEEPER_RESULT.get();
}
public static void clear() {
ORIGINAL_QUERY.remove();
RETRY_CONTEXT.remove();
@@ -73,5 +82,6 @@ public final class VerifierContextHolder {
EXECUTOR_STRUCTURED_OUTPUT.remove();
EXECUTOR_OUTPUT_PARSE_STATUS.remove();
TOOL_TRACE_SUMMARY.remove();
GATEKEEPER_RESULT.remove();
}
}
@@ -0,0 +1,61 @@
你是 Answer Composer。你的职责是把 Verifier 允许输出的结构化材料组织成用户可读的中文答案。
边界约束:
- 你不是诊断 Agent。
- 你不调用工具。
- 你不重新判断根因。
- 你不补充输入中不存在的新事实。
- 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。
- 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。
## 输入字段
- `original_query`:用户原始问题
- `verdict`:PASS / LOW_CONFID / REJECT
- `allowed_claims`:允许作为已确认事实表达的结论
- `allowed_hypotheses`:允许作为可能方向表达的内容
- `missing_info`:证据缺口
- `recommended_actions`:建议动作
- `rationale`:Verifier 判定理由
## 表达规则
### PASS
- 可以表达确认结论。
- 只能使用 `allowed_claims` 和 `recommended_actions`。
- 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。
### LOW_CONFID
- 必须说明当前证据仍有缺口。
- 必须区分“已确认信息”和“可能方向”。
- 不得把 `allowed_hypotheses` 写成确认结论。
### REJECT
- 必须说明当前无法基于已获取证据生成可靠结论。
- 不得输出根因结论。
- 只能输出已确认信息、证据缺口和下一步建议。
## 输出协议
必须输出且只能输出以下 JSON 结构:
{
"answer_summary": "...",
"recommended_actions": [
{
"action_text": "...",
"reason": "..."
}
],
"user_facing_answer": "..."
}
输出要求:
- `answer_summary` 用 1-2 句话概括当前可表达结论。
- `recommended_actions` 可以为空数组,但字段不能缺失。
- `user_facing_answer` 是最终给用户看的中文答案。
- 不得输出 schema 之外的字段。
@@ -51,7 +51,6 @@
支持等级:
- `direct`:工具返回中有直接事实。
- `indirect`:工具返回可支撑方向,但没有直接陈述完整结论。
- `none`:不能放入 `claims`,应放入 `hypotheses`、`recommended_actions` 或 `missing_info`。
### hypotheses
`hypotheses` 用来放合理怀疑但未被工具证实的方向。
@@ -71,8 +70,7 @@
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
@@ -106,15 +104,16 @@
],
"missing_info": [
"导致无法确认完整根因的证据缺口"
],
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致,不得额外加入未绑定证据的确认式事实。"
]
}
```
## 输出校验
- `answer_version` 必须是 `executor_evidence_v2`。
- 不得输出 `diagnosis_summary`。
- 不得输出 `user_facing_answer`。
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
- `claims[*].evidence_bindings` 不能为空。
- `evidence_excerpt` 必须来自工具返回,不允许编造。
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
- `user_facing_answer` 不得出现 `claims` 中没有、且又被写成确认结论的事实。
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
@@ -1,4 +1,4 @@
你是质量闸 verifier。你的任务是对 `executor_final_answer` 做一次基于现有证据的事实校验。
你是质量闸 verifier。你的任务是对 Executor 的结构化 claims 做一次基于现有证据的可推导性校验。
边界约束:
- 不做新的检索
@@ -9,8 +9,8 @@
## 输入字段
- `original_query`:用户原始问题
- `executor_final_answer`:本轮 Executor 最终答案
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`、`user_facing_answer`
- `executor_final_answer`:Executor 原始输出,仅用于 debug/fallback;当结构化输出有效时,不得从这里抽取额外确认事实
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
- `trace_ref`
@@ -20,56 +20,52 @@
- `input_summary`
- `output_summary`
- `evidence_level`
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`failed_rules`、`warnings`、`errors`
- `retry_context`:第二轮可选输入;若为空,按首轮处理
## 任务步骤
### 步骤一:提取关键事实
### 步骤一:确定校验对象
如果 `executor_output_parse_status.status="valid"` 且 `executor_structured_output.claims` 存在:
- 优先逐条校验 `executor_structured_output.claims`
- 每个 claim 至少形成一条 `facts_checked`
- 每个 claim 至少形成一条 `claim_checks`
- 必须检查 claim 的 `evidence_bindings` 是否能对应到 `tool_trace_summary` 中真实存在的 trace、tool 或 source_invocation_ids
- 如果 claim 声称 direct/indirect 支撑,但 evidence binding 不存在、无法定位、或 excerpt 与工具摘要不匹配,不得判为 `direct_evidence`
- 不得从 `executor_final_answer` 中抽取不在 claims 里的额外确认事实
然后必须扫描 `executor_structured_output.user_facing_answer`:
- 如果其中出现 confirmed-sounding facts(确认式事实、根因、指标值、错误码、服务名、修复结论)
- 且这些事实没有出现在 `executor_structured_output.claims`
- 必须额外加入 `facts_checked` 并按工具证据校验
如果 structured output 缺失或 malformed:
- 不得通过扫描 `executor_final_answer` 生成 `PASS`
- 输出 `LOW_CONFID`
- `groundedness_score = 0.0`
- `claim_checks = []`
- `facts_checked = []`
- `rationale` 说明结构化输出不可用
如果 structured output 缺失或 malformed,则回退到旧逻辑:提取并校验 `executor_final_answer` 里的全部实质性结论。关键事实至少包括:
- 每一个根因结论
- 每一个错误码、接口、组件归属或语义判断
- 每一个明确的修复建议、参数建议、排查步骤
- 每一个“证据来源陈述”
覆盖要求:
- 不允许只抽取一个总括性事实替代整段答案
- 如果答案给出多个根因,必须逐条拆成多个 `fact`
- 如果答案给出多条修复建议,必须逐条拆成多个 `fact`
- 只有寒暄、流程衔接语、与结论无关的话,才可以不纳入 `facts_checked`
### 步骤二:逐条校验事实
每条事实必须输出:
- `fact`
- `is_critical`
### 步骤二:逐条校验 claim
每条 claim check 必须输出:
- `claim_id`
- `claim_text`
- `claim_type`
- `verification`
- `detail`
- `evidence_refs`
`verification` 只允许以下四个值:
- `direct_evidence`
- `indirect_support`
- `no_evidence`
`claim_checks[*].verification` 只允许以下六个值:
- `direct_observation`
- `reasonable_inference`
- `overstated`
- `unsupported`
- `external_unknown`
- `contradicted`
结构化 claim 的校验规则:
- claim 有真实 evidence binding,且工具摘要直接包含该事实 → `direct_evidence`
- claim 有真实 evidence binding,但工具摘要只能支持方向或背景 → `indirect_support`
- claim 无法绑定真实 trace、invocation 或 excerpt → `no_evidence`
- claim 与工具摘要冲突,或编造了不存在的关键实体、服务、错误码、指标值 → `contradicted`
- claim 有真实 evidence binding,且工具摘要直接包含该事实 → `direct_observation`
- claim 有真实 evidence binding,工具摘要没有逐字说明但可以合理推出 → `reasonable_inference`
- claim 有部分依据,但写成唯一根因、确认根因或说得过满 → `overstated`
- claim 无法绑定真实 trace、invocation 或 excerpt → `unsupported`
- claim 引入证据外的新服务名、订单号、错误码、指标值、根因 → `external_unknown`
- claim 与工具摘要冲突 → `contradicted`
`hypotheses` 和 `missing_info` 默认不是 confirmed facts,不应因为它们承认缺证据而惩罚。
但如果 `user_facing_answer` 把 hypothesis 写成确认结论,必须按 confirmed fact 校验。
### 步骤三:补齐 evidence_refs
`evidence_refs` 必须是数组,数组元素必须引用 `tool_trace_summary` 中真实存在的证据项。每个元素包含:
@@ -88,24 +84,31 @@
### 步骤四:生成 verdict
严格使用以下判定矩阵:
1. 若任一关键事实(`is_critical=true`)为 `contradicted`
0. 若 `gatekeeper_result.status="fail"`
- 不得输出 `PASS`
- 若 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
- 否则至少输出 `LOW_CONFID`
1. 若任一关键 claim 为 `contradicted`
- `verdict = "REJECT"`
- `groundedness_score = 0.0`
2. 否则,若所有关键事实均为 `direct_evidence` 或 `indirect_support`
且至少一条关键事实为 `direct_evidence`
2. 否则,若所有关键 claims 均为 `direct_observation` 或 `reasonable_inference`
且至少一条关键 claim 为 `direct_observation`
- `verdict = "PASS"`
3. 否则,若不存在 `contradicted`
且存在关键事实为 `no_evidence`
或所有关键事实都只有 `indirect_support`
且存在关键 claim 为 `unsupported` / `external_unknown` / `overstated`
或所有关键 claim 都只有 `reasonable_inference`
- `verdict = "LOW_CONFID"`
### 步骤五:计算 groundedness_score
只统计 `is_critical=true` 的事实,映射如下:
- `direct_evidence = 1.0`
- `indirect_support = 0.6`
- `no_evidence = 0.0`
只统计关键 claim,映射如下:
- `direct_observation = 1.0`
- `reasonable_inference = 0.6`
- `overstated = 0.3`
- `unsupported = 0.0`
- `external_unknown = 0.0`
- `contradicted = 0.0`
规则:
@@ -114,14 +117,28 @@
- 保留 2 位小数
- 分数范围必须在 `[0.0, 1.0]`
### 步骤六:PASS 前覆盖性自检
### 步骤六:facts_checked 兼容输出
你必须同时输出 `facts_checked`,用于旧链路兼容。
映射规则:
- `direct_observation` → `direct_evidence`
- `reasonable_inference` → `indirect_support`
- `overstated` → `indirect_support`
- `unsupported` → `no_evidence`
- `external_unknown` → `no_evidence`
- `contradicted` → `contradicted`
`facts_checked[*].fact` 使用 `{claim_id}: {claim_text}`。
### 步骤七:PASS 前覆盖性自检
在输出 `PASS` 前,必须再次检查:
- `facts_checked` 是否覆盖了 `executor_final_answer` 的全部实质性结论
- 是否遗漏了单独出现的根因、修复建议、参数建议、排查步骤
- `claim_checks` 是否覆盖了 `executor_structured_output.claims` 中的全部 claims
- 是否存在 `gatekeeper_result.status="fail"`
- 是否存在 malformed/missing structured output
如有明显遗漏,即使已校验事实都有证据,也不得输出 `PASS`。
### 步骤七:处理 retry_context
### 步骤八:处理 retry_context
若 `retry_context` 不为空:
- 优先检查上一轮缺失证据点是否已补足
- 不要扩展与缺口无关的新事实
@@ -135,9 +152,28 @@
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "ERR_TIMEOUT 表示请求超时",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "知识库文档明确给出该错误码定义",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
}
]
}
],
"hypothesis_checks": [],
"facts_checked": [
{
"fact": "ERR_TIMEOUT 表示请求超时",
"fact": "claim-1: ERR_TIMEOUT 表示请求超时",
"is_critical": true,
"verification": "direct_evidence",
"detail": "知识库文档明确给出该错误码定义",
@@ -158,7 +194,9 @@
输出要求:
- `verdict` 只能是 `PASS` / `LOW_CONFID` / `REJECT`
- `groundedness_score` 必须是 JSON number
- `critical_fact_count` 必须等于 `facts_checked` 中 `is_critical=true` 的数量
- `critical_fact_count` 必须等于关键 claim 的数量;兼容期也应等于 `facts_checked` 中 `is_critical=true` 的数量
- `claim_checks` 可以为空数组,但字段不能缺失
- `facts_checked` 可以为空数组,但字段不能缺失
- 每条 `claim_checks[*]` 都必须包含 `evidence_refs`
- 每条 `facts_checked[*]` 都必须包含 `evidence_refs`
- 不得输出 schema 之外的字段
@@ -4,6 +4,9 @@ import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.service.ExecutorGatekeeperService;
import com.superbiz.agent.service.ToolTraceSummaryService;
import com.superbiz.agent.util.VerifierContextHolder;
import org.junit.jupiter.api.AfterEach;
@@ -84,6 +87,109 @@ class VerifierInputHookTest {
assertEquals("valid", VerifierContextHolder.getExecutorOutputParseStatus().get("status"));
}
@Test
void beforeModelAddsStructuredExecutorOutputWhenV2ContractHasNoUserFacingAnswer() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
Map.of("trace_ref", "trace-1", "tool_name", "query_metrics")
));
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("structured-v2-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("structured-v2-session").toolName("query_metrics").build()
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
VerifierContextHolder.setOriginalQuery("分析 MySQL 连接池耗尽");
String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
AgentCommand command = hook.beforeModel(
List.of(new AssistantMessage(executorOutput)),
RunnableConfig.builder().addMetadata("sessionId", "structured-v2-session").build()
);
JsonNode payload = readPayload(command);
assertEquals("valid", payload.path("executor_output_parse_status").path("status").asText());
assertEquals("executor_evidence_v2",
payload.path("executor_structured_output").path("answer_version").asText());
assertFalse(payload.path("executor_structured_output").has("user_facing_answer"));
assertEquals("连接池 active 达到上限",
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
assertEquals("pass", payload.path("gatekeeper_result").path("status").asText());
assertEquals("pass", VerifierContextHolder.getGatekeeperResult().get("status"));
}
@Test
void beforeModelAddsFailingGatekeeperResultForFabricatedInvocationId() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("fabricated-invocation-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("fabricated-invocation-session").toolName("query_metrics").build()
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_ids": [999],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
AgentCommand command = hook.beforeModel(
List.of(new AssistantMessage(executorOutput)),
RunnableConfig.builder().addMetadata("sessionId", "fabricated-invocation-session").build()
);
JsonNode payload = readPayload(command);
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
assertEquals("evidence.invocation_ref",
payload.path("gatekeeper_result").path("failed_rules").get(0).asText());
}
@Test
void beforeModelExtractsStructuredOutputFromPrefixedJsonFence() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
@@ -8,6 +8,7 @@ import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
@@ -21,6 +22,7 @@ import org.springframework.ai.chat.model.Generation;
import org.springframework.ai.chat.prompt.Prompt;
import org.springframework.ai.tool.ToolCallback;
import org.springframework.test.util.ReflectionTestUtils;
import org.mockito.ArgumentCaptor;
import java.util.List;
import java.util.Map;
@@ -33,6 +35,8 @@ import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.ArgumentMatchers.any;
import static org.mockito.ArgumentMatchers.anyString;
import static org.mockito.ArgumentMatchers.isNull;
import static org.mockito.Mockito.verify;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
@@ -51,9 +55,10 @@ class ChatServiceSequentialAgentTest {
"sequential-test-session"
);
assertEquals("EXECUTOR_FINAL_ANSWER", result.answer());
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertEquals("sequential-test-session", result.sessionId());
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
assertTrue(chatModel.sawVerifierPrompt);
}
@@ -77,6 +82,7 @@ class ChatServiceSequentialAgentTest {
"rationale": "scripted low confidence"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
@@ -89,7 +95,7 @@ class ChatServiceSequentialAgentTest {
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
assertTrue(result.answer().contains("当前缺口"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
}
@Test
@@ -126,6 +132,7 @@ class ChatServiceSequentialAgentTest {
"rationale": "scripted low confidence"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
@@ -136,8 +143,9 @@ class ChatServiceSequentialAgentTest {
);
assertTrue(result.answer().contains("已确认信息:\n- 连接池耗尽 active=50/50"));
assertFalse(result.answer().contains("临时扩容连接池到 80"));
assertTrue(result.answer().contains("当前缺口:\n- OOM 导致连接泄漏:missing OOM log"));
assertTrue(result.answer().contains("80"));
assertTrue(result.answer().contains("suggestion inferred from evidence"));
assertTrue(result.answer().contains("missing OOM log"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
}
@@ -171,7 +179,6 @@ class ChatServiceSequentialAgentTest {
"sequential-invalid-verifier-session"
);
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
}
@@ -195,6 +202,7 @@ class ChatServiceSequentialAgentTest {
"rationale": "scripted reject"
}
""");
chatModel.composerOutput = "not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
@@ -221,8 +229,9 @@ class ChatServiceSequentialAgentTest {
"sequential-workflow-session"
);
assertEquals("EXECUTOR_FINAL_ANSWER", result.answer());
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
assertTrue(chatModel.sawVerifierPrompt);
}
@@ -232,8 +241,7 @@ class ChatServiceSequentialAgentTest {
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "已确认连接池 active 达到上限。",
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
@@ -253,8 +261,7 @@ class ChatServiceSequentialAgentTest {
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": [],
"user_facing_answer": "已确认连接池 active 达到上限。"
"missing_info": []
}
""";
@@ -266,13 +273,284 @@ class ChatServiceSequentialAgentTest {
"sequential-structured-executor-session"
);
assertEquals("已确认连接池 active 达到上限。", result.answer());
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertFalse(result.answer().contains("\"answer_version\""));
assertTrue(chatModel.verifierPromptText.contains("\"executor_structured_output\""));
assertTrue(chatModel.verifierPromptText.contains("\"executor_output_parse_status\""));
assertTrue(chatModel.verifierPromptText.contains("\"status\" : \"valid\""));
assertTrue(chatModel.verifierPromptText.contains("连接池 active 达到上限"));
}
@Test
void executeChatComplexRendersExecutorEvidenceV2InsteadOfRawJsonOnPass() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.composerOutput = "not-json";
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "连接泄漏可能参与了连接池耗尽",
"basis": "已有连接池满载证据,但缺少泄漏检测日志",
"needed_evidence": ["连接泄漏检测日志"]
}
],
"recommended_actions": [
{
"action_text": "补充查询连接池泄漏检测日志",
"reason": "用于确认是否存在连接未释放"
}
],
"missing_info": ["缺少连接泄漏检测日志"]
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-v2-render-session"
);
assertTrue(result.answer().contains("已确认信息"));
assertTrue(result.answer().contains("连接池 active 达到上限"));
assertTrue(result.answer().contains("建议下一步"));
assertFalse(result.answer().contains("\"answer_version\""));
assertFalse(result.answer().contains("executor_evidence_v2"));
}
@Test
void executeChatComplexPersistsGatekeeperResultInVerifierEvaluation() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-gatekeeper-persist-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-gatekeeper-persist-session")
.toolName("query_metrics")
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-gatekeeper-persist-session"
);
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
Map<String, Object> verifierEvaluation = captor.getValue();
assertTrue(verifierEvaluation.containsKey("gatekeeper_result"));
@SuppressWarnings("unchecked")
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
assertEquals("pass", gatekeeperResult.get("status"));
}
@Test
void executeChatComplexMapsClaimChecksToFactsCheckedAndPersistsBoth() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-claim-check-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-claim-check-session")
.toolName("query_metrics")
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.32,
"critical_fact_count": 6,
"claim_checks": [
{"claim_id":"claim-1","claim_text":"CPU 使用率 92%","claim_type":"symptom","verification":"direct_observation","detail":"direct","evidence_refs":[{"trace_ref":"trace-1","tool_name":"query_metrics","source_invocation_ids":[101],"note":"cpu"}]},
{"claim_id":"claim-2","claim_text":"CPU 过高可能导致超时","claim_type":"risk","verification":"reasonable_inference","detail":"inference","evidence_refs":[]},
{"claim_id":"claim-3","claim_text":"CPU 是唯一根因","claim_type":"root_cause","verification":"overstated","detail":"too strong","evidence_refs":[]},
{"claim_id":"claim-4","claim_text":"缺少线程池证据","claim_type":"symptom","verification":"unsupported","detail":"missing","evidence_refs":[]},
{"claim_id":"claim-5","claim_text":"出现证据外错误码 ERR_FAKE","claim_type":"symptom","verification":"external_unknown","detail":"external","evidence_refs":[]},
{"claim_id":"claim-6","claim_text":"证据显示 CPU 很低","claim_type":"symptom","verification":"contradicted","detail":"conflict","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "claim checks drive compatibility"
}
""");
chatModel.executorOutput = validExecutorV2Output();
chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-claim-check-session"
);
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
Map<String, Object> verifierEvaluation = captor.getValue();
@SuppressWarnings("unchecked")
List<Map<String, Object>> claimChecks = (List<Map<String, Object>>) verifierEvaluation.get("claim_checks");
@SuppressWarnings("unchecked")
List<Map<String, Object>> factsChecked = (List<Map<String, Object>>) verifierEvaluation.get("facts_checked");
assertEquals(6, claimChecks.size());
assertEquals(6, factsChecked.size());
assertEquals("direct_evidence", factsChecked.get(0).get("verification"));
assertEquals("indirect_support", factsChecked.get(1).get("verification"));
assertEquals("indirect_support", factsChecked.get(2).get("verification"));
assertEquals("no_evidence", factsChecked.get(3).get("verification"));
assertEquals("no_evidence", factsChecked.get(4).get("verification"));
assertEquals("contradicted", factsChecked.get(5).get("verification"));
assertTrue(String.valueOf(factsChecked.get(0).get("fact")).startsWith("claim-1:"));
}
@Test
void executeChatComplexDowngradesPassToRejectWhenGatekeeperInvocationRefFails() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ToolInvocationRepository invocationRepository =
(ToolInvocationRepository) ReflectionTestUtils.getField(chatService, "toolInvocationRepository");
when(invocationRepository.findBySessionIdOrderByIdAsc("sequential-gatekeeper-fail-session"))
.thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.sessionId("sequential-gatekeeper-fail-session")
.toolName("query_metrics")
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [
{"claim_id":"claim-1","claim_text":"连接池 active 达到上限","claim_type":"symptom","verification":"direct_observation","detail":"direct","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "model tried pass"
}
""");
chatModel.composerOutput = "not-json";
chatModel.executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_ids": [999],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-gatekeeper-fail-session"
);
assertTrue(result.answer().startsWith("当前无法基于已获取证据生成可靠结论"));
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
assertEquals("REJECT", captor.getValue().get("verdict"));
}
@Test
void executeChatComplexDowngradesPassToLowConfidenceWhenExecutorOutputMalformed() throws Exception {
ChatService chatService = createChatService();
SelfEvaluationMergeService mergeService =
(SelfEvaluationMergeService) ReflectionTestUtils.getField(chatService, "selfEvaluationMergeService");
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 0,
"claim_checks": [],
"facts_checked": [],
"rationale": "model tried pass"
}
""");
chatModel.composerOutput = "not-json";
chatModel.executorOutput = "{ not-json";
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-malformed-pass-session"
);
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
ArgumentCaptor<Map<String, Object>> captor = ArgumentCaptor.forClass(Map.class);
verify(mergeService).mergeVerifierEvaluation(isNull(), captor.capture());
assertEquals("LOW_CONFID", captor.getValue().get("verdict"));
}
@Test
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
ChatService chatService = new ChatService();
@@ -366,6 +644,10 @@ class ChatServiceSequentialAgentTest {
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L);
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.toolName("query_metrics")
.build()));
EvaluationService evaluationService = mock(EvaluationService.class);
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
@@ -375,6 +657,7 @@ class ChatServiceSequentialAgentTest {
when(toolTraceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
SelfEvaluationMergeService selfEvaluationMergeService = mock(SelfEvaluationMergeService.class);
when(selfEvaluationMergeService.mergeVerifierEvaluation(any(), any())).thenReturn("{}");
ExecutorGatekeeperService executorGatekeeperService = new ExecutorGatekeeperService(toolInvocationRepository);
ReflectionTestUtils.setField(chatService, "dateTimeTools", new DateTimeTools());
ReflectionTestUtils.setField(chatService, "lookupKnowledgeTool", new LookupKnowledgeTool());
@@ -387,20 +670,87 @@ class ChatServiceSequentialAgentTest {
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);
ReflectionTestUtils.setField(chatService, "toolTraceSummaryService", toolTraceSummaryService);
ReflectionTestUtils.setField(chatService, "selfEvaluationMergeService", selfEvaluationMergeService);
ReflectionTestUtils.setField(chatService, "executorGatekeeperService", executorGatekeeperService);
ReflectionTestUtils.setField(chatService, "verifierLowConfidenceThreshold", 0.5d);
ReflectionTestUtils.setField(chatService, "chatPlannerPrompt", "PLANNER_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatExecutorPrompt", "EXECUTOR_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatVerifierPrompt", "VERIFIER_TEST_PROMPT");
ReflectionTestUtils.setField(chatService, "chatComposerPrompt", "COMPOSER_TEST_PROMPT");
return chatService;
}
private String validExecutorV2Output() {
return """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
}
private static final class ScriptedChatModel implements ChatModel {
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
private String promptText = "";
private String plannerPromptText = "";
private String executorPromptText = "";
private String verifierPromptText = "";
private String executorOutput = "EXECUTOR_FINAL_ANSWER";
private String composerPromptText = "";
private String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "连接池 active 达到上限",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active=50 max=50"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
""";
private String composerOutput = """
{
"answer_summary": "已确认连接池 active 达到上限。",
"recommended_actions": [
{
"action_text": "补充查询连接池泄漏检测日志",
"reason": "用于确认是否存在连接未释放"
}
],
"user_facing_answer": "已确认连接池 active 达到上限。建议补充查询连接池泄漏检测日志。"
}
""";
private boolean sawVerifierPrompt;
private final java.util.List<String> verifierOutputs;
private int verifierOutputIndex;
@@ -411,15 +761,10 @@ class ChatServiceSequentialAgentTest {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"facts_checked": [
{
"fact": "executor answer generated",
"is_critical": true,
"verification": "direct_evidence",
"detail": "covered by scripted verifier",
"evidence_refs": []
}
"claim_checks": [
{"claim_id":"claim-1","claim_text":"连接池 active 达到上限","claim_type":"symptom","verification":"direct_observation","detail":"covered by scripted verifier","evidence_refs":[]}
],
"facts_checked": [],
"rationale": "scripted pass"
}
""");
@@ -452,6 +797,10 @@ class ChatServiceSequentialAgentTest {
int index = Math.min(verifierOutputIndex, verifierOutputs.size() - 1);
text = verifierOutputs.get(index);
verifierOutputIndex++;
} else if (promptText.contains("COMPOSER_TEST_PROMPT")) {
agentCalls.add("chat_composer");
composerPromptText = promptText;
text = composerOutput;
} else {
text = "UNEXPECTED_PROMPT";
}
@@ -0,0 +1,98 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when;
class ExecutorGatekeeperServiceTest {
@Test
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
}
@Test
void validateFailsWhenRemovedFieldsArePresent() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of());
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> output = validOutput(101L, "query_metrics");
output.put("user_facing_answer", "旧版最终答案");
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("schema.executor_v2"));
}
@Test
void validateFailsForFabricatedInvocationId() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(999L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@Test
void validateFailsForToolNameMismatch() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_logs").build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@SuppressWarnings("unchecked")
private Map<String, Object> validOutput(Long invocationId, String toolName) {
return new java.util.LinkedHashMap<>(Map.of(
"answer_version", "executor_evidence_v2",
"claims", List.of(Map.of(
"claim_id", "claim-1",
"claim_type", "symptom",
"claim_text", "连接池 active 达到上限",
"support_level", "direct",
"evidence_bindings", List.of(Map.of(
"source_type", "tool_trace",
"source_id", "trace-1",
"tool_name", toolName,
"source_invocation_ids", List.of(invocationId),
"evidence_excerpt", "active=50 max=50"
))
)),
"hypotheses", List.of(),
"recommended_actions", List.of(),
"missing_info", List.of()
));
}
}