feat(agent): add verifier claim checks
This commit is contained in:
@@ -4,51 +4,54 @@
|
||||
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
|
||||
- **AND** at least one critical fact has direct evidence
|
||||
- **AND** no critical fact is contradicted
|
||||
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
|
||||
- **AND** at least one critical claim has direct observation
|
||||
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
|
||||
- **AND** `gatekeeper_result.status` is not `fail`
|
||||
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** some critical facts have no supporting evidence
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only indirect support
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** all critical facts are only indirectly supported
|
||||
#### Scenario: LOW_CONFID verdict with only inference support
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **AND** all critical claims are only `reasonable_inference`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
|
||||
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
|
||||
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
#### Scenario: Structured Executor claims are verified first
|
||||
- **WHEN** `executor_structured_output.claims` is present and valid
|
||||
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary`
|
||||
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
|
||||
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
|
||||
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence`
|
||||
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
|
||||
#### Scenario: Extra confirmed-sounding answer facts are still checked
|
||||
- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims`
|
||||
- **THEN** Verifier SHALL add those facts to `facts_checked`
|
||||
- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
|
||||
|
||||
#### Scenario: Natural-language fallback remains available
|
||||
#### Scenario: Malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** Executor does not return parseable structured output
|
||||
- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer`
|
||||
- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: claim-level verifier output is accepted
|
||||
- **WHEN** the Verifier checks Executor V2 structured output
|
||||
- **THEN** the output SHALL contain `verdict`, `groundedness_score`, `critical_fact_count`, `claim_checks`, `facts_checked`, and `rationale`
|
||||
- **AND** `claim_checks` SHALL be the primary V2 verification result
|
||||
- **AND** `facts_checked` SHALL remain available for compatibility
|
||||
|
||||
#### Scenario: Output format validation
|
||||
- **WHEN** the Verifier completes its analysis
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "claim_checks", "facts_checked", and "rationale" fields
|
||||
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
|
||||
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
|
||||
|
||||
@@ -57,14 +60,19 @@ The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
- **AND** the JSON object SHALL include `critical_fact_count`
|
||||
- **AND** each `claim_checks` item SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
Each checked fact SHALL be labeled using a fixed evidence classification.
|
||||
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
|
||||
|
||||
#### Scenario: fact classification values
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
|
||||
#### Scenario: claim checks are mapped to legacy facts
|
||||
- **WHEN** Verifier output contains `claim_checks`
|
||||
- **THEN** ChatService SHALL derive compatibility `facts_checked`
|
||||
- **AND** `direct_observation` SHALL map to `direct_evidence`
|
||||
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
|
||||
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
|
||||
- **AND** `contradicted` SHALL map to `contradicted`
|
||||
|
||||
### Requirement: groundedness_score SHALL be derived from fact classifications
|
||||
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
|
||||
@@ -124,6 +132,12 @@ The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facin
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
@@ -166,7 +180,7 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th
|
||||
#### Scenario: Verifier remains isolated from intermediate reasoning
|
||||
- **WHEN** `executor_structured_output` is added to the verifier input
|
||||
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
|
||||
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
|
||||
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, gatekeeper result, and retry context
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
@@ -211,6 +225,17 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th
|
||||
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
|
||||
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
|
||||
|
||||
#### Scenario: structured claims are the primary verification target
|
||||
- **WHEN** `executor_output_parse_status.status` is `valid`
|
||||
- **AND** `executor_structured_output.claims` is available
|
||||
- **THEN** Verifier SHALL verify each claim through `claim_checks`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
|
||||
#### Scenario: malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
|
||||
@@ -355,3 +380,26 @@ The system SHALL run deterministic Gatekeeper checks after Executor output parsi
|
||||
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `pass`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
|
||||
|
||||
#### Scenario: gatekeeper fail prevents PASS
|
||||
- **WHEN** `gatekeeper_result.status` is `fail`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||
- **AND** the effective verdict SHALL NOT be `PASS`
|
||||
|
||||
#### Scenario: invocation reference failure downgrades to reject
|
||||
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
|
||||
|
||||
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
|
||||
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
|
||||
|
||||
#### Scenario: claim check verification values are constrained
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each item SHALL use one of `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, or `contradicted`
|
||||
|
||||
#### Scenario: claim check evidence references remain auditable
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
|
||||
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
|
||||
|
||||
Reference in New Issue
Block a user