5.3 KiB
Decisions: executor-verifier-claim-checks
sm-flow Progress
Clarify
Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.
Slug: executor-verifier-claim-checks
Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.
Context
Relevant history:
executor-v2-output-contract: Executor emitsexecutor_evidence_v2and no longer emits final-expression fields.executor-gatekeeper-hook: Gatekeeper validates schema and invocation references before Verifier and persistsgatekeeper_result.chat-verifier-agent: Current Verifier still primarily usesfacts_checked.
Current code shape:
chat-verifier-prompt.mdstill frames the task aroundexecutor_final_answerandfacts_checked.ChatService.parseVerifierDecision(...)only parsesfacts_checked.buildLowConfidenceOutput(...)andbuildRetryContext(...)consumeVerifierDecision.factsChecked().persistVerifierEvaluation(...)already persists parse status, structured output, trace summary, and gatekeeper result.AgentLoggingHooksummarizes verifier output usingfacts_checked.
Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
Should Verifier still scan executor_final_answer when structured output is valid? |
evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. |
Should facts_checked be removed now? |
evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. |
| Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. |
| Does this require database migration? | evidence-driven | No. claim_checks is stored under existing JSON self_evaluation. |
| Does this introduce Composer? | evidence-driven | No. Composer is stage four. |
No user-interview questions are open for this stage.
Specify
OpenSpec artifacts:
proposal.md: why and scope.design.md: Verifier V2 output, compatibility mapping, verdict guardrails, risks.specs/chat-verifier-agent/spec.md: observable requirements for claim checks, compatibility facts, persistence, and guardrails.tasks.md: executable implementation and verification checklist.
Audit
Architecture risk summary:
- This is an L2 internal verifier contract extension.
- Existing consumers continue using
facts_checked, which is now generated fromclaim_checkswhen present. - Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
- No external API or database schema changes are introduced.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage three | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design output contract and guardrails | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier output: L2 internal extension with
claim_checks. - Persistence JSON: L2 internal audit extension in existing
self_evaluation. - External HTTP/API behavior: unchanged.
Commit
Commit gate result: passed.
proposal.md,design.md,specs/chat-verifier-agent/spec.md, andtasks.mdexist.cmd /c openspec validate executor-verifier-claim-checkspassed.cmd /c openspec status --change executor-verifier-claim-checksreports 4/4 artifacts complete.- No unresolved user-interview questions remain for this stage.
Apply
Capability source: OpenSpec CLI + openspec-apply-change protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, rg/diff inspection, and targeted tests.
Implemented changes:
- Updated
chat-verifier-prompt.mdsoexecutor_structured_output.claimsis the primary verification target. - Added
claim_checksparsing, normalization, persistence, and compatibility mapping to legacyfacts_checkedinChatService. - Added code-side effective verdict guardrails:
- missing/malformed structured output cannot remain
PASS; gatekeeper_result.status=failcannot remainPASS;evidence.invocation_reffailures downgrade toREJECT;- other Gatekeeper failures downgrade at least to
LOW_CONFID.
- missing/malformed structured output cannot remain
- Updated verifier logging summaries to account for
claim_checks. - Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.
Conflict / fix record:
- Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1
user_facing_answer. - Classification: code/test drift from the committed OpenSpec, not a design blocker.
- Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.
Verification:
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" testpassed.cmd /c openspec validate executor-verifier-claim-checkspassed.