Files
SuperBizAgent-java/openspec/changes/archive/2026-07-07-executor-verifier-claim-checks/decisions.md
T

111 lines
5.3 KiB
Markdown

# Decisions: executor-verifier-claim-checks
## sm-flow Progress
### Clarify
Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.
Slug: `executor-verifier-claim-checks`
Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Current Verifier still primarily uses `facts_checked`.
Current code shape:
- `chat-verifier-prompt.md` still frames the task around `executor_final_answer` and `facts_checked`.
- `ChatService.parseVerifierDecision(...)` only parses `facts_checked`.
- `buildLowConfidenceOutput(...)` and `buildRetryContext(...)` consume `VerifierDecision.factsChecked()`.
- `persistVerifierEvaluation(...)` already persists parse status, structured output, trace summary, and gatekeeper result.
- `AgentLoggingHook` summarizes verifier output using `facts_checked`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Verifier still scan `executor_final_answer` when structured output is valid? | evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. |
| Should `facts_checked` be removed now? | evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. |
| Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. |
| Does this require database migration? | evidence-driven | No. `claim_checks` is stored under existing JSON self_evaluation. |
| Does this introduce Composer? | evidence-driven | No. Composer is stage four. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Verifier V2 output, compatibility mapping, verdict guardrails, risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for claim checks, compatibility facts, persistence, and guardrails.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- This is an L2 internal verifier contract extension.
- Existing consumers continue using `facts_checked`, which is now generated from `claim_checks` when present.
- Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
- No external API or database schema changes are introduced.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage three | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design output contract and guardrails | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier output: L2 internal extension with `claim_checks`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- `cmd /c openspec status --change executor-verifier-claim-checks` reports 4/4 artifacts complete.
- No unresolved user-interview questions remain for this stage.
### Apply
Capability source: OpenSpec CLI + `openspec-apply-change` protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, `rg`/diff inspection, and targeted tests.
Implemented changes:
- Updated `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
- Added `claim_checks` parsing, normalization, persistence, and compatibility mapping to legacy `facts_checked` in `ChatService`.
- Added code-side effective verdict guardrails:
- missing/malformed structured output cannot remain `PASS`;
- `gatekeeper_result.status=fail` cannot remain `PASS`;
- `evidence.invocation_ref` failures downgrade to `REJECT`;
- other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- Updated verifier logging summaries to account for `claim_checks`.
- Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.
Conflict / fix record:
- Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
- Classification: code/test drift from the committed OpenSpec, not a design blocker.
- Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.
Verification:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.