# Decisions: executor-gatekeeper-hook ## sm-flow Progress ### Clarify Entry summary: implement stage two of Executor Structured Output V2 by adding deterministic Gatekeeper checks between Executor and Verifier. Slug: `executor-gatekeeper-hook` Scale: complex program, stage-specific standard slice. It changes internal verifier payload and audit persistence but does not change external API or database schema. ### Context Relevant history: - `executor-v2-output-contract`: Executor now emits `executor_evidence_v2` without final-expression fields. - `chat-verifier-agent`: Verifier input is assembled explicitly by `VerifierInputHook` and persisted through `ChatService`. - `evidence-trace-hardening`: tool invocation rows are the source of truth for evidence trace references. Current code shape: - `VerifierInputHook` parses Executor raw output, builds `tool_trace_summary`, and assembles Verifier payload. - `VerifierContextHolder` stores parse status, structured output, and trace summary for later persistence. - `ChatService.persistVerifierEvaluation(...)` writes verifier snapshots into `diagnosis_session.self_evaluation.verifier_evaluation`. - `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` can provide the valid invocation pool. ### Grill Question pool: | Question | Mode | Resolution | |---|---|---| | Should Gatekeeper live in Executor hook? | evidence-driven | No. User explicitly chose Verifier hook for this version. | | Should Gatekeeper retry Executor? | evidence-driven | No. Retry is deferred; this phase only validates and audits. | | Which rules are mandatory now? | evidence-driven | Implement schema and invocation reference rules. Excerpt similarity and utilization can follow later. | | Where should audit data persist? | evidence-driven | Existing `self_evaluation.verifier_evaluation.gatekeeper_result`. | | Does this require database migration? | evidence-driven | No. Use existing JSON self_evaluation and existing tool_invocation rows. | No user-interview questions are open for this stage. ### Specify OpenSpec artifacts: - `proposal.md`: why and scope. - `design.md`: Gatekeeper output, rules, integration, persistence, and risks. - `specs/chat-verifier-agent/spec.md`: observable requirements for payload, persistence, and initial rules. - `tasks.md`: executable implementation and verification checklist. ### Audit Architecture risk summary: - Gatekeeper introduces an internal verifier payload extension but no external API or database schema change. - `VerifierInputHook` remains the integration point, matching the user decision to keep Gatekeeper in the Verifier hook. - Gatekeeper needs repository access to validate invocation ids; placing that access in a service keeps hook logic small. - Failed Gatekeeper results are audit signals in this phase and do not trigger retry. Cross-artifact alignment: | Source | Target | Status | |---|---|---| | issue stage two | proposal | aligned | | proposal scope / non-goals | design | aligned | | design rules and payload | specs | aligned | | specs observable behavior | tasks | aligned | Interface impact: - Verifier payload: L2 internal extension with `gatekeeper_result`. - Persistence JSON: L2 internal audit extension in existing `self_evaluation`. - External HTTP/API behavior: unchanged. ### Commit Commit gate result: passed. - `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist. - `cmd /c openspec validate executor-gatekeeper-hook` passed. - No unresolved user-interview questions remain for this stage. ### Apply Implementation result: - Added `ExecutorGatekeeperService` with initial `schema.executor_v2` and `evidence.invocation_ref` checks. - Integrated Gatekeeper into `VerifierInputHook` after parse and trace summary construction. - Added `gatekeeper_result` to Verifier payload and `VerifierContextHolder`. - Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`. - Updated `chat-verifier-prompt.md` so Gatekeeper failures must not produce PASS. - Added and updated focused tests for service validation, hook payload, fabricated invocation ids, tool name mismatch, and persistence. Validation: - `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed with 23 tests. - `cmd /c openspec validate executor-gatekeeper-hook` passed. Known limits: - No Executor retry is triggered by Gatekeeper failure in this phase. - `evidence.excerpt_similarity`, `claim.hallucination_phrases`, and `claim.evidence_utilization` remain deferred by design.