120 lines
3.8 KiB
Markdown
120 lines
3.8 KiB
Markdown
## Context
|
|
|
|
Stage one introduced `executor_evidence_v2`, but `VerifierInputHook` still only parses JSON and passes the structured object to Verifier. A valid JSON object can still contain:
|
|
|
|
- removed fields such as `diagnosis_summary` or `user_facing_answer`
|
|
- missing or wrong `answer_version`
|
|
- empty `claims[].evidence_bindings`
|
|
- fabricated `source_invocation_ids`
|
|
- `tool_name` values that do not match the real `tool_invocation`
|
|
|
|
Gatekeeper handles these deterministic failures before Verifier performs semantic reasoning.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
Goals:
|
|
|
|
- Add Gatekeeper into `VerifierInputHook`.
|
|
- Produce a small `gatekeeper_result` object.
|
|
- Add `gatekeeper_result` to Verifier payload.
|
|
- Persist `gatekeeper_result` in verifier evaluation.
|
|
- Implement initial rules: schema and invocation reference.
|
|
|
|
Non-goals:
|
|
|
|
- No Executor retry on Gatekeeper failure.
|
|
- No excerpt similarity rule in this phase.
|
|
- No hallucination phrase or evidence utilization rule in this phase.
|
|
- No Verifier V2 `claim_checks`.
|
|
- No Composer.
|
|
- No database schema changes.
|
|
|
|
## Gatekeeper Output
|
|
|
|
Gatekeeper returns:
|
|
|
|
```json
|
|
{
|
|
"status": "fail",
|
|
"failed_rules": ["evidence.invocation_ref"],
|
|
"warnings": [],
|
|
"errors": [
|
|
{
|
|
"rule_id": "evidence.invocation_ref",
|
|
"target": "claims[0].evidence_bindings[0]",
|
|
"message": "source_invocation_ids not found in current session"
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
Status calculation:
|
|
|
|
```text
|
|
any fail -> fail
|
|
else any warning -> warn
|
|
else pass
|
|
```
|
|
|
|
## Initial Rules
|
|
|
|
### schema.executor_v2
|
|
|
|
Fail when:
|
|
|
|
- structured output is absent after parse status is valid
|
|
- `answer_version` is not `executor_evidence_v2`
|
|
- `claims` is not an array
|
|
- removed fields `diagnosis_summary` or `user_facing_answer` are present
|
|
- any claim misses required fields
|
|
- any claim has empty `evidence_bindings`
|
|
- `hypotheses`, `recommended_actions`, or `missing_info` are missing or non-array
|
|
|
|
For this phase, missing optional arrays may be normalized only if implementation remains simple. If not normalized, missing arrays fail schema to keep behavior deterministic.
|
|
|
|
### evidence.invocation_ref
|
|
|
|
Fail when:
|
|
|
|
- `claims[].evidence_bindings[].source_invocation_ids` is missing or empty
|
|
- any referenced invocation id is not in the current session's `tool_invocation` rows
|
|
- `tool_name` does not match the referenced invocation's real `tool_name`
|
|
|
|
Recommended action evidence bindings remain optional and are not hard-fail checked in this phase.
|
|
|
|
## Integration
|
|
|
|
`VerifierInputHook.beforeModel(...)` flow becomes:
|
|
|
|
```text
|
|
parse Executor output
|
|
build tool_trace_summary
|
|
Gatekeeper.validate(sessionId, structuredOutput, parseStatus)
|
|
put gatekeeper_result into verifier payload
|
|
store gatekeeper_result in VerifierContextHolder
|
|
```
|
|
|
|
`ChatService.persistVerifierEvaluation(...)` adds:
|
|
|
|
```text
|
|
gatekeeper_result: VerifierContextHolder.getGatekeeperResult()
|
|
```
|
|
|
|
If Gatekeeper throws unexpectedly, hook should fail closed with a minimal `fail` result in payload rather than dropping validation silently.
|
|
|
|
## Interface Impact
|
|
|
|
- Verifier payload: L2 internal contract extension with `gatekeeper_result`.
|
|
- Persistence JSON: L2 internal audit extension under existing `self_evaluation`.
|
|
- No external API or database schema change.
|
|
|
|
## Risks / Mitigations
|
|
|
|
- Risk: tests or code instantiate `VerifierInputHook` with the old constructor.
|
|
- Mitigation: keep a compatibility constructor that uses a no-op/pass Gatekeeper, or update tests explicitly.
|
|
- Risk: Gatekeeper fails because no session id exists.
|
|
- Mitigation: return fail with a clear `schema.executor_v2` or `evidence.invocation_ref` error only when validation cannot establish current-session references.
|
|
- Risk: Verifier prompt may ignore Gatekeeper.
|
|
- Mitigation: payload and audit are still authoritative for later phases; Verifier prompt update can be minimal in this phase.
|
|
|