# Decisions: executor-verifier-claim-checks ## sm-flow Progress ### Clarify Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented. Slug: `executor-verifier-claim-checks` Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema. ### Context Relevant history: - `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields. - `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`. - `chat-verifier-agent`: Current Verifier still primarily uses `facts_checked`. Current code shape: - `chat-verifier-prompt.md` still frames the task around `executor_final_answer` and `facts_checked`. - `ChatService.parseVerifierDecision(...)` only parses `facts_checked`. - `buildLowConfidenceOutput(...)` and `buildRetryContext(...)` consume `VerifierDecision.factsChecked()`. - `persistVerifierEvaluation(...)` already persists parse status, structured output, trace summary, and gatekeeper result. - `AgentLoggingHook` summarizes verifier output using `facts_checked`. ### Grill Question pool: | Question | Mode | Resolution | |---|---|---| | Should Verifier still scan `executor_final_answer` when structured output is valid? | evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. | | Should `facts_checked` be removed now? | evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. | | Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. | | Does this require database migration? | evidence-driven | No. `claim_checks` is stored under existing JSON self_evaluation. | | Does this introduce Composer? | evidence-driven | No. Composer is stage four. | No user-interview questions are open for this stage. ### Specify OpenSpec artifacts: - `proposal.md`: why and scope. - `design.md`: Verifier V2 output, compatibility mapping, verdict guardrails, risks. - `specs/chat-verifier-agent/spec.md`: observable requirements for claim checks, compatibility facts, persistence, and guardrails. - `tasks.md`: executable implementation and verification checklist. ### Audit Architecture risk summary: - This is an L2 internal verifier contract extension. - Existing consumers continue using `facts_checked`, which is now generated from `claim_checks` when present. - Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance. - No external API or database schema changes are introduced. Cross-artifact alignment: | Source | Target | Status | |---|---|---| | issue stage three | proposal | aligned | | proposal scope / non-goals | design | aligned | | design output contract and guardrails | specs | aligned | | specs observable behavior | tasks | aligned | Interface impact: - Verifier output: L2 internal extension with `claim_checks`. - Persistence JSON: L2 internal audit extension in existing `self_evaluation`. - External HTTP/API behavior: unchanged. ### Commit Commit gate result: passed. - `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist. - `cmd /c openspec validate executor-verifier-claim-checks` passed. - `cmd /c openspec status --change executor-verifier-claim-checks` reports 4/4 artifacts complete. - No unresolved user-interview questions remain for this stage. ### Apply Capability source: OpenSpec CLI + `openspec-apply-change` protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, `rg`/diff inspection, and targeted tests. Implemented changes: - Updated `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target. - Added `claim_checks` parsing, normalization, persistence, and compatibility mapping to legacy `facts_checked` in `ChatService`. - Added code-side effective verdict guardrails: - missing/malformed structured output cannot remain `PASS`; - `gatekeeper_result.status=fail` cannot remain `PASS`; - `evidence.invocation_ref` failures downgrade to `REJECT`; - other Gatekeeper failures downgrade at least to `LOW_CONFID`. - Updated verifier logging summaries to account for `claim_checks`. - Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output. Conflict / fix record: - Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 `user_facing_answer`. - Classification: code/test drift from the committed OpenSpec, not a design blocker. - Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback. Verification: - `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed. - `cmd /c openspec validate executor-verifier-claim-checks` passed.