feat(agent): add executor gatekeeper hook

This commit is contained in:
aruo
2026-07-08 02:01:49 +08:00
parent 050cbc8fee
commit c5e496e715
23 changed files with 1343 additions and 30 deletions
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,102 @@
# Decisions: executor-gatekeeper-hook
## sm-flow Progress
### Clarify
Entry summary: implement stage two of Executor Structured Output V2 by adding deterministic Gatekeeper checks between Executor and Verifier.
Slug: `executor-gatekeeper-hook`
Scale: complex program, stage-specific standard slice. It changes internal verifier payload and audit persistence but does not change external API or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor now emits `executor_evidence_v2` without final-expression fields.
- `chat-verifier-agent`: Verifier input is assembled explicitly by `VerifierInputHook` and persisted through `ChatService`.
- `evidence-trace-hardening`: tool invocation rows are the source of truth for evidence trace references.
Current code shape:
- `VerifierInputHook` parses Executor raw output, builds `tool_trace_summary`, and assembles Verifier payload.
- `VerifierContextHolder` stores parse status, structured output, and trace summary for later persistence.
- `ChatService.persistVerifierEvaluation(...)` writes verifier snapshots into `diagnosis_session.self_evaluation.verifier_evaluation`.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` can provide the valid invocation pool.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Gatekeeper live in Executor hook? | evidence-driven | No. User explicitly chose Verifier hook for this version. |
| Should Gatekeeper retry Executor? | evidence-driven | No. Retry is deferred; this phase only validates and audits. |
| Which rules are mandatory now? | evidence-driven | Implement schema and invocation reference rules. Excerpt similarity and utilization can follow later. |
| Where should audit data persist? | evidence-driven | Existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
| Does this require database migration? | evidence-driven | No. Use existing JSON self_evaluation and existing tool_invocation rows. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Gatekeeper output, rules, integration, persistence, and risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for payload, persistence, and initial rules.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- Gatekeeper introduces an internal verifier payload extension but no external API or database schema change.
- `VerifierInputHook` remains the integration point, matching the user decision to keep Gatekeeper in the Verifier hook.
- Gatekeeper needs repository access to validate invocation ids; placing that access in a service keeps hook logic small.
- Failed Gatekeeper results are audit signals in this phase and do not trigger retry.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage two | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design rules and payload | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier payload: L2 internal extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation result:
- Added `ExecutorGatekeeperService` with initial `schema.executor_v2` and `evidence.invocation_ref` checks.
- Integrated Gatekeeper into `VerifierInputHook` after parse and trace summary construction.
- Added `gatekeeper_result` to Verifier payload and `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated `chat-verifier-prompt.md` so Gatekeeper failures must not produce PASS.
- Added and updated focused tests for service validation, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Validation:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed with 23 tests.
- `cmd /c openspec validate executor-gatekeeper-hook` passed.
Known limits:
- No Executor retry is triggered by Gatekeeper failure in this phase.
- `evidence.excerpt_similarity`, `claim.hallucination_phrases`, and `claim.evidence_utilization` remain deferred by design.
@@ -0,0 +1,119 @@
## Context
Stage one introduced `executor_evidence_v2`, but `VerifierInputHook` still only parses JSON and passes the structured object to Verifier. A valid JSON object can still contain:
- removed fields such as `diagnosis_summary` or `user_facing_answer`
- missing or wrong `answer_version`
- empty `claims[].evidence_bindings`
- fabricated `source_invocation_ids`
- `tool_name` values that do not match the real `tool_invocation`
Gatekeeper handles these deterministic failures before Verifier performs semantic reasoning.
## Goals / Non-Goals
Goals:
- Add Gatekeeper into `VerifierInputHook`.
- Produce a small `gatekeeper_result` object.
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` in verifier evaluation.
- Implement initial rules: schema and invocation reference.
Non-goals:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## Gatekeeper Output
Gatekeeper returns:
```json
{
"status": "fail",
"failed_rules": ["evidence.invocation_ref"],
"warnings": [],
"errors": [
{
"rule_id": "evidence.invocation_ref",
"target": "claims[0].evidence_bindings[0]",
"message": "source_invocation_ids not found in current session"
}
]
}
```
Status calculation:
```text
any fail -> fail
else any warning -> warn
else pass
```
## Initial Rules
### schema.executor_v2
Fail when:
- structured output is absent after parse status is valid
- `answer_version` is not `executor_evidence_v2`
- `claims` is not an array
- removed fields `diagnosis_summary` or `user_facing_answer` are present
- any claim misses required fields
- any claim has empty `evidence_bindings`
- `hypotheses`, `recommended_actions`, or `missing_info` are missing or non-array
For this phase, missing optional arrays may be normalized only if implementation remains simple. If not normalized, missing arrays fail schema to keep behavior deterministic.
### evidence.invocation_ref
Fail when:
- `claims[].evidence_bindings[].source_invocation_ids` is missing or empty
- any referenced invocation id is not in the current session's `tool_invocation` rows
- `tool_name` does not match the referenced invocation's real `tool_name`
Recommended action evidence bindings remain optional and are not hard-fail checked in this phase.
## Integration
`VerifierInputHook.beforeModel(...)` flow becomes:
```text
parse Executor output
build tool_trace_summary
Gatekeeper.validate(sessionId, structuredOutput, parseStatus)
put gatekeeper_result into verifier payload
store gatekeeper_result in VerifierContextHolder
```
`ChatService.persistVerifierEvaluation(...)` adds:
```text
gatekeeper_result: VerifierContextHolder.getGatekeeperResult()
```
If Gatekeeper throws unexpectedly, hook should fail closed with a minimal `fail` result in payload rather than dropping validation silently.
## Interface Impact
- Verifier payload: L2 internal contract extension with `gatekeeper_result`.
- Persistence JSON: L2 internal audit extension under existing `self_evaluation`.
- No external API or database schema change.
## Risks / Mitigations
- Risk: tests or code instantiate `VerifierInputHook` with the old constructor.
- Mitigation: keep a compatibility constructor that uses a no-op/pass Gatekeeper, or update tests explicitly.
- Risk: Gatekeeper fails because no session id exists.
- Mitigation: return fail with a clear `schema.executor_v2` or `evidence.invocation_ref` error only when validation cannot establish current-session references.
- Risk: Verifier prompt may ignore Gatekeeper.
- Mitigation: payload and audit are still authoritative for later phases; Verifier prompt update can be minimal in this phase.
@@ -0,0 +1,37 @@
## Why
Executor V2 makes the output structured, but structure alone does not prevent physical hallucinations such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names. These failures should be caught deterministically before the Verifier reasons over claims.
This phase adds Gatekeeper inside `VerifierInputHook` as a deterministic pre-verifier quality gate and persists its result for audit.
## What Changes
- Add a Gatekeeper validation service for Executor structured output.
- Add initial rules:
- `schema.executor_v2`
- `evidence.invocation_ref`
- Add `gatekeeper_result` to Verifier payload.
- Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Keep Gatekeeper in `VerifierInputHook`; do not move it to Executor hook.
- Keep retry behavior unchanged; failed Gatekeeper results do not trigger Executor retry in this phase.
- Keep Verifier prompt/output behavior unchanged except that it can see `gatekeeper_result`.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Verifier input now includes deterministic Gatekeeper results for Executor structured output.
## Impact
- Affected hook: `VerifierInputHook`.
- Affected service/runtime: new Gatekeeper service and `VerifierContextHolder`.
- Affected persistence: `ChatService.persistVerifierEvaluation(...)` writes `gatekeeper_result` into existing `self_evaluation`.
- Affected repository access: Gatekeeper reads current-session `tool_invocation` rows via `ToolInvocationRepository`.
- Affected tests: `VerifierInputHookTest` and `ChatServiceSequentialAgentTest`.
- Database schema: no change.
@@ -0,0 +1,51 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: gatekeeper result available to Verifier
- **WHEN** the system prepares verifier inputs from Executor output
- **THEN** the payload SHALL include `gatekeeper_result`
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
## ADDED Requirements
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing invocation id
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
@@ -0,0 +1,31 @@
## 1. Gatekeeper Core
- [x] 1.1 Add a Gatekeeper validation service with a small result shape: `status`, `failed_rules`, `warnings`, `errors`.
- [x] 1.2 Implement `schema.executor_v2` rule.
- [x] 1.3 Implement `evidence.invocation_ref` rule using current-session `tool_invocation` rows.
- [x] 1.4 Keep recommended action evidence bindings out of hard-fail validation for this phase.
## 2. Hook Integration
- [x] 2.1 Inject Gatekeeper into `VerifierInputHook`.
- [x] 2.2 Add `gatekeeper_result` to Verifier payload.
- [x] 2.3 Store `gatekeeper_result` in `VerifierContextHolder`.
- [x] 2.4 Preserve parse-only boundary in `VerifierInputHook`.
## 3. Persistence
- [x] 3.1 Persist `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- [x] 3.2 Preserve existing verifier evaluation fields.
## 4. Prompt Compatibility
- [x] 4.1 Update `chat-verifier-prompt.md` minimally so Verifier sees `gatekeeper_result` and must not output PASS when it fails.
## 5. Tests And Verification
- [x] 5.1 Add Gatekeeper unit tests for schema failures and pass cases.
- [x] 5.2 Add hook tests proving payload contains `gatekeeper_result`.
- [x] 5.3 Add tests for fabricated invocation id and tool name mismatch.
- [x] 5.4 Add ChatService persistence test for `gatekeeper_result`.
- [x] 5.5 Run targeted tests.
- [x] 5.6 Validate this OpenSpec change.