Files

16 KiB

MODIFIED Requirements

Requirement: Verifier SHALL fact-check Executor answers

The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.

Scenario: PASS verdict when all claims have evidence

  • WHEN all critical claims in verified_executor_output.claims have direct observation or reasonable inference support in verified_evidence
  • AND at least one critical claim has direct observation
  • AND no critical claim is contradicted, unsupported, external unknown, or overstated
  • AND verdict ceiling is PASS
  • THEN the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5

Scenario: LOW_CONFID verdict with partial evidence

  • WHEN no critical claim contradicts verified evidence
  • AND some critical claims are unsupported, external_unknown, or overstated
  • THEN the Verifier SHALL output model verdict="LOW_CONFID"

Scenario: LOW_CONFID verdict with only inference support

  • WHEN no critical claim contradicts verified evidence
  • AND all critical claims are only reasonable_inference
  • THEN the Verifier SHALL output model verdict="LOW_CONFID"

Scenario: REJECT verdict when claims contradict evidence

  • WHEN any critical claim in verified_executor_output.claims contradicts verified evidence
  • OR the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
  • THEN the Verifier SHALL output model verdict="REJECT"

Scenario: Verified structured claims are the only verification target

  • WHEN verified_executor_output.claims is present
  • THEN Verifier SHALL verify each structured claim against matching verified_evidence through claim_checks
  • AND each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
  • AND a claim without matching verified evidence SHALL NOT be classified as direct_observation
  • AND Verifier SHALL NOT receive or add confirmed facts from raw Executor text

Scenario: Executor output is invalid

  • WHEN Executor does not return a legal structured contract
  • THEN the Graph SHALL route directly to pre-verification Fallback
  • AND Verifier SHALL NOT execute or fabricate a diagnostic verdict

Requirement: facts_checked SHALL use a fixed classification set

The system SHALL continue to expose compatibility facts_checked using its fixed verification classification set.

Scenario: claim checks are mapped to legacy facts

  • WHEN Verifier output contains claim_checks
  • THEN the shared Verifier protocol parser SHALL derive compatibility facts_checked when the model did not provide them
  • AND direct_observation SHALL map to direct_evidence
  • AND reasonable_inference and overstated SHALL map to indirect_support
  • AND unsupported and external_unknown SHALL map to no_evidence
  • AND contradicted SHALL map to contradicted

Requirement: ChatService SHALL route based on Verifier verdict

The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.

Scenario: PASS routes to Composer

  • WHEN Verifier completes with effective verdict="PASS"
  • THEN the Graph SHALL invoke Composer with filtered Verifier-allowed material
  • AND the final user-facing answer SHALL NOT pass through raw Executor output
  • AND the final user-facing answer SHALL NOT read Executor user_facing_answer

Scenario: LOW_CONFID does not qualify for evidence retry

  • WHEN Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
  • THEN the Graph SHALL route to Composer without another Planner cycle
  • AND the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps

Scenario: LOW_CONFID qualifies for evidence retry

  • WHEN Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
  • THEN the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
  • AND it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry

Scenario: REJECT does not enter retry round

  • WHEN Verifier completes with effective verdict="REJECT"
  • THEN the Graph SHALL NOT start an evidence supplementation round
  • AND it SHALL route to Composer-safe output

Scenario: REJECT produces bounded output

  • WHEN effective verdict is REJECT
  • THEN the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
  • AND it SHALL NOT pass through raw Executor answer
  • AND it SHALL NOT include an unsupported root-cause conclusion

Scenario: Verifier execution fails

  • WHEN Verifier exhausts technical retry or returns a non-retryable failure
  • THEN the Graph SHALL route to pre-verification Fallback
  • AND no execution status string SHALL be used as model or effective verdict

Requirement: Verifier SHALL be observable

The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.

Scenario: claim checks written to self_evaluation

  • WHEN a completed Verifier evaluation is persisted
  • THEN diagnosis_run.self_evaluation.verifier_evaluation SHALL include claim_checks
  • AND it SHALL continue to include compatibility facts_checked
  • AND it SHALL include verifier_status, model_verdict, effective_verdict, verdict, groundedness_score, rationale, verified output/evidence, and Gatekeeper audit

Scenario: composer output written to self_evaluation

  • WHEN final answer composition completes
  • THEN diagnosis_run.self_evaluation.verifier_evaluation SHALL include compact composer_output when available
  • AND handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
  • AND existing claim/fact and Gatekeeper fields SHALL be preserved

Scenario: verdict written to self_evaluation

  • WHEN Verifier completes
  • THEN Graph result mapping SHALL write effective verdict under diagnosis_run.self_evaluation.verifier_evaluation.verdict
  • AND existing rule_evaluation and aiops_rule_evaluation channels SHALL be preserved

Scenario: pre-verification fallback is persisted

  • WHEN Graph reaches Fallback before Verifier completes
  • THEN verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
  • AND it SHALL NOT fabricate model_verdict or effective_verdict

Scenario: gatekeeper result written to self_evaluation

  • WHEN Graph result mapping persists available Gatekeeper state
  • THEN diagnosis_run.self_evaluation.verifier_evaluation SHALL include gatekeeper_result
  • AND the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper

Scenario: prompt audit written to verifier evaluation

  • WHEN a complex Chat Graph result is persisted
  • THEN the system SHALL include a prompt_audit object under diagnosis_run.self_evaluation.verifier_evaluation
  • AND prompt_audit.version SHALL identify the Chat Prompt audit catalog version
  • AND prompt_audit.prompts SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
  • AND full Prompt text SHALL NOT be persisted

Scenario: prompt audit available on fallback paths

  • WHEN Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
  • THEN the persisted verifier evaluation SHALL still include prompt_audit

Scenario: evaluation payload is inspected

  • WHEN Graph verifier evaluation is persisted
  • THEN it SHALL NOT contain raw Executor text or complete tool_trace_summary
  • AND compatibility executor_structured_output SHALL contain at most the verified projection

Requirement: Verifier SHALL consume explicit verification inputs

The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.

Scenario: explicit input blocks available to Verifier

  • WHEN the Verifier Graph Node starts
  • THEN the payload SHALL provide diagnosis_context, verified_executor_output, verified_evidence, gatekeeper_audit, and verdict_ceiling
  • AND permitted structured retry_context SHALL be provided only after evidence retry preparation

Scenario: Verifier remains isolated from intermediate and raw material

  • WHEN the Verifier input is serialized
  • THEN it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State

Scenario: only passed bindings are available

  • WHEN Gatekeeper returns mixed passed and failed checked bindings
  • THEN verified_executor_output and verified_evidence SHALL contain only claims/material matching passed bindings
  • AND the Verifier SHALL NOT receive failed or unreferenced tool material

Scenario: verified evidence preserves precise references

  • WHEN the system prepares Verifier input
  • THEN each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
  • AND the item SHALL be traceable to current-run Gatekeeper validation

Scenario: gatekeeper audit and ceiling are available

  • WHEN the system prepares Verifier input
  • THEN the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
  • AND a LOW_CONFID ceiling SHALL prevent effective PASS

Scenario: technical retry occurs

  • WHEN the first Verifier attempt returns invalid output or a retryable invocation failure
  • THEN the second attempt SHALL receive byte-identical serialized input
  • AND Executor, Gatekeeper, and tools SHALL NOT rerun

Requirement: Verifier facts SHALL be auditable

Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.

Scenario: claim checks contain evidence refs

  • WHEN the Verifier emits claim_checks
  • THEN each check SHALL include an evidence_refs array
  • AND any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
  • AND it SHALL NOT reference a failed or unverified binding

Scenario: verifier evaluation persists traceability snapshot

  • WHEN Graph result mapping persists verifier evaluation
  • THEN it SHALL include traceability_version
  • AND it SHALL include the bounded verified_evidence snapshot used by the Verifier
  • AND it SHALL NOT persist a complete tool trace summary as Verifier input

Requirement: Structured Executor output SHALL degrade safely

The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.

Scenario: Malformed Executor JSON is classified

  • WHEN Executor returns malformed JSON or text outside the expected contract
  • THEN Executor Node SHALL set INVALID_OUTPUT
  • AND the Graph SHALL route directly to deterministic pre-verification Fallback
  • AND Gatekeeper, Verifier, and model Composer SHALL NOT execute

Scenario: Structured parse failure remains observable

  • WHEN Executor output parsing fails
  • THEN orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
  • AND the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict

Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures

The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.

Scenario: schema rule rejects removed fields

  • WHEN Executor structured output contains diagnosis_summary or user_facing_answer
  • THEN gatekeeper_result.status SHALL be fail
  • AND gatekeeper_result.failed_rules SHALL contain schema.executor_v2

Scenario: schema rule rejects missing evidence bindings

  • WHEN a confirmed claim has no evidence_bindings
  • THEN gatekeeper_result.status SHALL be fail
  • AND gatekeeper_result.failed_rules SHALL contain schema.executor_v2

Scenario: invocation rule rejects fabricated invocation ids

  • WHEN a claim evidence binding references an invocation id absent from current-run tool_invocation rows
  • THEN gatekeeper_result.status SHALL be fail
  • AND gatekeeper_result.failed_rules SHALL contain evidence.invocation_ref

Scenario: invocation rule rejects tool name mismatch

  • WHEN a claim evidence binding references an existing current-run invocation id
  • AND binding tool_name does not match persisted invocation tool_name
  • THEN gatekeeper_result.status SHALL be fail
  • AND gatekeeper_result.failed_rules SHALL contain evidence.invocation_ref

Scenario: valid structured output passes initial gatekeeper rules

  • WHEN Executor emits legal executor_evidence_v2
  • AND each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
  • THEN gatekeeper_result.status SHALL be pass
  • AND gatekeeper_result.failed_rules SHALL be empty

Scenario: gatekeeper reject bypasses Verifier

  • WHEN normalized Gatekeeper status is REJECT
  • THEN the Graph SHALL route directly to pre-verification Fallback
  • AND Verifier SHALL NOT execute

Scenario: gatekeeper low confidence is bounded

  • WHEN normalized Gatekeeper status is LOW_CONFID with at least one passed binding
  • THEN verified input SHALL contain only passed bindings
  • AND effective verdict SHALL NOT exceed LOW_CONFID

Requirement: Verifier SHALL use verified claim-local evidence for derivability

Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.

Scenario: Verified excerpt supports direct observation

  • WHEN verdict ceiling is PASS
  • AND a claim's verified evidence matched text directly contains the claim's concrete facts
  • THEN Verifier MAY classify that claim as direct_observation

Scenario: Verified evidence is complete Verifier context

  • WHEN verified claims and evidence are available
  • THEN Verifier SHALL use them as its evidence context
  • AND it SHALL NOT require or request a complete tool trace summary
  • AND it SHALL NOT read raw Executor or unreferenced tool material

Requirement: Gatekeeper severity SHALL constrain effective verdict

Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.

Scenario: Reject severity bypasses Verifier

  • WHEN gatekeeper_result.severity=reject
  • THEN the Graph SHALL route to pre-verification Fallback without invoking Verifier
  • AND it SHALL NOT fabricate an effective diagnostic verdict

Scenario: Low confidence severity prevents PASS

  • WHEN gatekeeper_result.severity=low_confid
  • AND the Verifier model returns verdict=PASS
  • THEN deterministic effective-verdict calculation SHALL downgrade the result
  • AND effective verdict SHALL be LOW_CONFID

Scenario: Gatekeeper audit includes severity

  • WHEN Graph verifier evaluation is persisted
  • THEN diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result SHALL include available status, severity, checked_bindings, failed_rules, warnings, and errors