feat(agent): add verifier claim checks

This commit is contained in:
aruo
2026-07-08 02:33:02 +08:00
parent c5e496e715
commit 1b31e78be5
19 changed files with 1247 additions and 115 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,110 @@
# Decisions: executor-verifier-claim-checks
## sm-flow Progress
### Clarify
Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.
Slug: `executor-verifier-claim-checks`
Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.
### Context
Relevant history:
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Current Verifier still primarily uses `facts_checked`.
Current code shape:
- `chat-verifier-prompt.md` still frames the task around `executor_final_answer` and `facts_checked`.
- `ChatService.parseVerifierDecision(...)` only parses `facts_checked`.
- `buildLowConfidenceOutput(...)` and `buildRetryContext(...)` consume `VerifierDecision.factsChecked()`.
- `persistVerifierEvaluation(...)` already persists parse status, structured output, trace summary, and gatekeeper result.
- `AgentLoggingHook` summarizes verifier output using `facts_checked`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Should Verifier still scan `executor_final_answer` when structured output is valid? | evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. |
| Should `facts_checked` be removed now? | evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. |
| Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. |
| Does this require database migration? | evidence-driven | No. `claim_checks` is stored under existing JSON self_evaluation. |
| Does this introduce Composer? | evidence-driven | No. Composer is stage four. |
No user-interview questions are open for this stage.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope.
- `design.md`: Verifier V2 output, compatibility mapping, verdict guardrails, risks.
- `specs/chat-verifier-agent/spec.md`: observable requirements for claim checks, compatibility facts, persistence, and guardrails.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- This is an L2 internal verifier contract extension.
- Existing consumers continue using `facts_checked`, which is now generated from `claim_checks` when present.
- Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
- No external API or database schema changes are introduced.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue stage three | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design output contract and guardrails | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Verifier output: L2 internal extension with `claim_checks`.
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
- External HTTP/API behavior: unchanged.
### Commit
Commit gate result: passed.
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- `cmd /c openspec status --change executor-verifier-claim-checks` reports 4/4 artifacts complete.
- No unresolved user-interview questions remain for this stage.
### Apply
Capability source: OpenSpec CLI + `openspec-apply-change` protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, `rg`/diff inspection, and targeted tests.
Implemented changes:
- Updated `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
- Added `claim_checks` parsing, normalization, persistence, and compatibility mapping to legacy `facts_checked` in `ChatService`.
- Added code-side effective verdict guardrails:
- missing/malformed structured output cannot remain `PASS`;
- `gatekeeper_result.status=fail` cannot remain `PASS`;
- `evidence.invocation_ref` failures downgrade to `REJECT`;
- other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- Updated verifier logging summaries to account for `claim_checks`.
- Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.
Conflict / fix record:
- Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
- Classification: code/test drift from the committed OpenSpec, not a design blocker.
- Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.
Verification:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
@@ -0,0 +1,140 @@
## Context
Current state after stage two:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> ChatService final rendering
```
Verifier receives explicit inputs:
- `original_query`
- `executor_final_answer`
- `executor_structured_output`
- `executor_output_parse_status`
- `tool_trace_summary`
- `gatekeeper_result`
- `retry_context`
However, Verifier output is still primarily:
```json
{
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 1,
"facts_checked": [],
"rationale": "..."
}
```
Stage three introduces V2 output while preserving the old compatibility field.
## Verifier V2 Output
Verifier should output:
```json
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.62,
"critical_fact_count": 1,
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service CPU usage is high",
"claim_type": "symptom",
"verification": "direct_observation",
"detail": "query_metrics shows CPU=92%",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [394],
"note": "metrics summary contains CPU=92%"
}
]
}
],
"hypothesis_checks": [],
"facts_checked": [],
"rationale": "..."
}
```
`facts_checked` remains for compatibility. If Verifier does not emit it, `ChatService` must derive it from `claim_checks`.
## Claim Verification Set
`claim_checks[].verification` is limited to:
| Value | Meaning | Legacy mapping |
|---|---|---|
| `direct_observation` | Evidence directly observes the claim | `direct_evidence` |
| `reasonable_inference` | Evidence can reasonably support the claim, but not as direct observation | `indirect_support` |
| `overstated` | Evidence partially supports the claim, but the claim says too much | `indirect_support` |
| `unsupported` | Evidence is insufficient | `no_evidence` |
| `external_unknown` | Claim introduces evidence-external entity/value/root cause | `no_evidence` |
| `contradicted` | Claim conflicts with evidence | `contradicted` |
## Compatibility Mapping
`ChatService` must keep old downstream behavior alive by producing `facts_checked`.
Suggested mapping:
```text
facts_checked[].fact = "{claim_id}: {claim_text}"
facts_checked[].is_critical = claim_type in ["root_cause", "symptom", "impact", "risk"]
facts_checked[].verification = mapped legacy verification
facts_checked[].detail = claim_checks[].detail
facts_checked[].evidence_refs = claim_checks[].evidence_refs
```
If Verifier emits both `claim_checks` and `facts_checked`, `claim_checks` is authoritative. `facts_checked` may be replaced by the deterministic compatibility projection to avoid inconsistent audit data.
If Verifier emits only old `facts_checked`, ChatService keeps the old path.
## Verdict Guardrails
Gatekeeper fail:
- If `gatekeeper_result.status = fail`, effective verdict must not be `PASS`.
- If the model returns `PASS`, ChatService should downgrade the effective verdict to `LOW_CONFID` or `REJECT`.
- For this phase, `evidence.invocation_ref` failure should downgrade to `REJECT`; other Gatekeeper failures should downgrade to `LOW_CONFID`.
Malformed or missing structured output:
- If `executor_output_parse_status.status` is `missing` or `malformed`, Verifier should not use natural-language extraction to produce PASS.
- Effective verdict should be `LOW_CONFID`.
## Prompt Boundary
The prompt should say:
- Primary target is `executor_structured_output.claims`.
- Do not extract additional confirmed facts from `executor_final_answer` when structured output is valid.
- `executor_final_answer` is debug/fallback only.
- `claim_checks` is the primary output.
- `facts_checked` is compatibility output.
## Interface Impact
- L2 internal contract extension.
- No external API change.
- No database schema change.
- Audit JSON gains `claim_checks`.
## Risks / Mitigations
- Risk: old low-confidence templates rely on `facts_checked`.
- Mitigation: derive `facts_checked` from `claim_checks`.
- Risk: prompt-only Gatekeeper PASS prevention is insufficient.
- Mitigation: add code-side effective verdict guard.
- Risk: Agent logging only summarizes `facts_checked`.
- Mitigation: update logging to understand `claim_checks` while keeping old summary compatibility.
@@ -0,0 +1,47 @@
## Why
Stage one moved Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. The Verifier still mainly operates through the legacy `facts_checked` contract and the prompt still allows fallback extraction from `executor_final_answer`.
That keeps two problems alive:
- Verifier can still treat natural-language Executor output as a fact source.
- Downstream code cannot distinguish claim-level verification results from legacy natural-language fact checks.
This phase makes Verifier V2 claim-oriented: Verifier evaluates `executor_structured_output.claims` for whether each claim can be reasonably derived from evidence, emits `claim_checks`, and keeps `facts_checked` only as a compatibility projection.
## What Changes
- Update `chat-verifier-prompt.md` so the primary verification target is `executor_structured_output.claims`.
- Add Verifier V2 output field `claim_checks`.
- Preserve compatibility by mapping `claim_checks` into legacy `facts_checked`.
- Parse and persist `claim_checks` in `ChatService`.
- Ensure `gatekeeper_result.status=fail` cannot result in an effective `PASS`.
- Ensure missing or malformed Executor structured output does not fall back to natural-language fact extraction for PASS.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Verifier output now includes claim-level checks and uses claim derivability as the primary groundedness contract.
## Impact
- Affected prompt: `src/main/resources/prompts/chat-verifier-prompt.md`.
- Affected service: `ChatService.parseVerifierDecision(...)`, retry context generation, verifier persistence.
- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `claim_checks` and keeps `facts_checked`.
- Affected logging: verifier thought summaries may count `claim_checks`.
- Affected tests: `ChatServiceSequentialAgentTest` and focused verifier parsing tests.
- Database schema: no table or column change.
## Non-Goals
- No Composer in this phase.
- No final-answer material filtering in this phase beyond existing LOW_CONFID/REJECT templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
@@ -0,0 +1,109 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** `gatekeeper_result.status` is not `fail`
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts the tool results
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts the tool results
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: Malformed structured output cannot pass through natural language fallback
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
#### Scenario: claim-level verifier output is accepted
- **WHEN** the Verifier checks Executor V2 structured output
- **THEN** the output SHALL contain `verdict`, `groundedness_score`, `critical_fact_count`, `claim_checks`, `facts_checked`, and `rationale`
- **AND** `claim_checks` SHALL be the primary V2 verification result
- **AND** `facts_checked` SHALL remain available for compatibility
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** ChatService SHALL derive compatibility `facts_checked`
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: claim checks written to self_evaluation
- **WHEN** the Verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: structured claims are the primary verification target
- **WHEN** `executor_output_parse_status.status` is `valid`
- **AND** `executor_structured_output.claims` is available
- **THEN** Verifier SHALL verify each claim through `claim_checks`
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
#### Scenario: malformed structured output cannot pass through natural language fallback
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
- **AND** the effective verdict SHALL be `LOW_CONFID`
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
#### Scenario: gatekeeper fail prevents PASS
- **WHEN** `gatekeeper_result.status` is `fail`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL NOT be `PASS`
#### Scenario: invocation reference failure downgrades to reject
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
- **AND** the Verifier model returns `verdict = "PASS"`
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
## ADDED Requirements
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
#### Scenario: claim check verification values are constrained
- **WHEN** Verifier emits `claim_checks`
- **THEN** each item SHALL use one of `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, or `contradicted`
#### Scenario: claim check evidence references remain auditable
- **WHEN** Verifier emits `claim_checks`
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
@@ -0,0 +1,39 @@
## 1. Prompt Contract
- [x] 1.1 Update `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
- [x] 1.2 Remove the valid-structured-output path that scans `executor_final_answer` for extra confirmed facts.
- [x] 1.3 Add `claim_checks` and optional `hypothesis_checks` to the output contract.
- [x] 1.4 Keep `facts_checked` as compatibility output.
- [x] 1.5 State that missing/malformed structured output cannot produce PASS through natural-language fallback.
## 2. Parser And Compatibility Mapping
- [x] 2.1 Extend `VerifierDecision` to store `claim_checks`.
- [x] 2.2 Parse `claim_checks` from verifier output.
- [x] 2.3 Derive compatibility `facts_checked` from `claim_checks` when present.
- [x] 2.4 Preserve old `facts_checked` parsing when `claim_checks` is absent.
- [x] 2.5 Persist `claim_checks` in verifier evaluation.
## 3. Effective Verdict Guardrails
- [x] 3.1 Add code-side guard so `gatekeeper_result.status=fail` cannot result in effective PASS.
- [x] 3.2 Downgrade `evidence.invocation_ref` failures to REJECT.
- [x] 3.3 Downgrade other Gatekeeper failures to at least LOW_CONFID.
- [x] 3.4 Ensure missing/malformed Executor structured output cannot produce effective PASS.
## 4. Compatibility Consumers
- [x] 4.1 Ensure low-confidence rendering still uses compatibility `facts_checked`.
- [x] 4.2 Ensure `buildRetryContext(...)` still receives evidence gaps from compatibility `facts_checked`.
- [x] 4.3 Update verifier logging summary to account for `claim_checks`.
- [x] 4.4 Preserve existing traceability and Gatekeeper audit fields.
## 5. Tests And Verification
- [x] 5.1 Add or update tests for parsing `claim_checks`.
- [x] 5.2 Add tests for `claim_checks` to `facts_checked` compatibility mapping.
- [x] 5.3 Add tests for each verification mapping class: `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, `contradicted`.
- [x] 5.4 Add tests proving Gatekeeper fail cannot remain PASS.
- [x] 5.5 Add tests proving malformed/missing structured output cannot remain PASS.
- [x] 5.6 Run targeted tests.
- [x] 5.7 Validate this OpenSpec change.