Files

5.3 KiB

Decisions: executor-verifier-claim-checks

sm-flow Progress

Clarify

Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.

Slug: executor-verifier-claim-checks

Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.

Context

Relevant history:

  • executor-v2-output-contract: Executor emits executor_evidence_v2 and no longer emits final-expression fields.
  • executor-gatekeeper-hook: Gatekeeper validates schema and invocation references before Verifier and persists gatekeeper_result.
  • chat-verifier-agent: Current Verifier still primarily uses facts_checked.

Current code shape:

  • chat-verifier-prompt.md still frames the task around executor_final_answer and facts_checked.
  • ChatService.parseVerifierDecision(...) only parses facts_checked.
  • buildLowConfidenceOutput(...) and buildRetryContext(...) consume VerifierDecision.factsChecked().
  • persistVerifierEvaluation(...) already persists parse status, structured output, trace summary, and gatekeeper result.
  • AgentLoggingHook summarizes verifier output using facts_checked.

Grill

Question pool:

Question Mode Resolution
Should Verifier still scan executor_final_answer when structured output is valid? evidence-driven No. Stage three explicitly makes claims the primary target and raw output debug/fallback only.
Should facts_checked be removed now? evidence-driven No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling.
Should Gatekeeper fail prevention be prompt-only? evidence-driven No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard.
Does this require database migration? evidence-driven No. claim_checks is stored under existing JSON self_evaluation.
Does this introduce Composer? evidence-driven No. Composer is stage four.

No user-interview questions are open for this stage.

Specify

OpenSpec artifacts:

  • proposal.md: why and scope.
  • design.md: Verifier V2 output, compatibility mapping, verdict guardrails, risks.
  • specs/chat-verifier-agent/spec.md: observable requirements for claim checks, compatibility facts, persistence, and guardrails.
  • tasks.md: executable implementation and verification checklist.

Audit

Architecture risk summary:

  • This is an L2 internal verifier contract extension.
  • Existing consumers continue using facts_checked, which is now generated from claim_checks when present.
  • Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
  • No external API or database schema changes are introduced.

Cross-artifact alignment:

Source Target Status
issue stage three proposal aligned
proposal scope / non-goals design aligned
design output contract and guardrails specs aligned
specs observable behavior tasks aligned

Interface impact:

  • Verifier output: L2 internal extension with claim_checks.
  • Persistence JSON: L2 internal audit extension in existing self_evaluation.
  • External HTTP/API behavior: unchanged.

Commit

Commit gate result: passed.

  • proposal.md, design.md, specs/chat-verifier-agent/spec.md, and tasks.md exist.
  • cmd /c openspec validate executor-verifier-claim-checks passed.
  • cmd /c openspec status --change executor-verifier-claim-checks reports 4/4 artifacts complete.
  • No unresolved user-interview questions remain for this stage.

Apply

Capability source: OpenSpec CLI + openspec-apply-change protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, rg/diff inspection, and targeted tests.

Implemented changes:

  • Updated chat-verifier-prompt.md so executor_structured_output.claims is the primary verification target.
  • Added claim_checks parsing, normalization, persistence, and compatibility mapping to legacy facts_checked in ChatService.
  • Added code-side effective verdict guardrails:
    • missing/malformed structured output cannot remain PASS;
    • gatekeeper_result.status=fail cannot remain PASS;
    • evidence.invocation_ref failures downgrade to REJECT;
    • other Gatekeeper failures downgrade at least to LOW_CONFID.
  • Updated verifier logging summaries to account for claim_checks.
  • Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.

Conflict / fix record:

  • Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 user_facing_answer.
  • Classification: code/test drift from the committed OpenSpec, not a design blocker.
  • Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.

Verification:

  • mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test passed.
  • cmd /c openspec validate executor-verifier-claim-checks passed.