Files

6.1 KiB

ADDED Requirements

Requirement: Executor evidence bindings SHALL support precise evidence references

Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.

Scenario: Precise evidence binding contains invocation path and excerpt

  • WHEN Executor binds evidence to a claim
  • THEN the binding SHOULD include singular source_invocation_id
  • AND the binding SHOULD include raw_path
  • AND the binding SHALL include tool_name and evidence_excerpt
  • AND the raw_path SHALL be interpreted relative to the referenced tool invocation's retrieval_details.evidence_refs

Scenario: Legacy plural invocation ids remain compatibility only

  • WHEN Executor emits legacy source_invocation_ids
  • THEN the system MAY read them for compatibility
  • AND they SHALL NOT be sufficient for a precise Gatekeeper pass without raw_path

Requirement: Gatekeeper SHALL validate evidence reference fidelity

Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.

Scenario: Valid precise binding passes

  • WHEN a binding's source_invocation_id exists in the current session
  • AND the binding's tool_name matches the persisted invocation
  • AND the binding's raw_path exists in retrieval_details.evidence_refs
  • AND the binding's evidence_excerpt is supported by the matching evidence ref text
  • THEN Gatekeeper SHALL return status=pass
  • AND Gatekeeper SHALL return severity=none

Scenario: Missing raw path is low confidence

  • WHEN a binding references an existing invocation
  • AND the binding omits raw_path
  • THEN Gatekeeper SHALL return status=fail
  • AND Gatekeeper SHALL return severity=low_confid
  • AND the effective verifier result SHALL NOT be PASS

Scenario: Old invocation without evidence refs is low confidence

  • WHEN a binding references an existing invocation
  • AND the invocation does not contain retrieval_details.evidence_refs
  • THEN Gatekeeper SHALL return status=fail
  • AND Gatekeeper SHALL return severity=low_confid
  • AND the effective verifier result SHALL NOT be PASS

Scenario: Unknown raw path is rejected

  • WHEN a binding references an existing invocation
  • AND the binding's raw_path is absent from that invocation's retrieval_details.evidence_refs
  • THEN Gatekeeper SHALL return status=fail
  • AND Gatekeeper SHALL return severity=reject
  • AND failed_rules SHALL include evidence.raw_path

Scenario: Mismatched excerpt is rejected

  • WHEN a binding references an existing invocation and raw path
  • AND the binding's evidence_excerpt is not supported by the matching system-side evidence ref text
  • THEN Gatekeeper SHALL return status=fail
  • AND Gatekeeper SHALL return severity=reject
  • AND failed_rules SHALL include evidence.excerpt_mismatch

Requirement: Verifier SHALL use verified claim-local evidence for derivability

Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.

Scenario: Verified excerpt supports direct observation

  • WHEN gatekeeper_result.severity=none
  • AND a claim's verified evidence excerpts directly contain the claim's concrete facts
  • THEN Verifier MAY classify that claim as direct_observation

Scenario: Tool trace summary is navigation context

  • WHEN executor_structured_output.claims[].evidence_bindings are available
  • THEN Verifier SHALL use tool_trace_summary as navigation and audit context
  • AND it SHALL NOT require tool_trace_summary.output_summary to contain every fact already present in verified claim-local evidence

Requirement: Gatekeeper severity SHALL constrain effective verdict

Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.

Scenario: Reject severity prevents PASS

  • WHEN gatekeeper_result.severity=reject
  • AND the Verifier model returns verdict=PASS
  • THEN ChatService SHALL downgrade the effective verdict
  • AND the effective verdict SHALL be REJECT

Scenario: Low confidence severity prevents PASS

  • WHEN gatekeeper_result.severity=low_confid
  • AND the Verifier model returns verdict=PASS
  • THEN ChatService SHALL downgrade the effective verdict
  • AND the effective verdict SHALL be LOW_CONFID

Scenario: Gatekeeper audit includes severity

  • WHEN verifier evaluation is persisted
  • THEN diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result SHALL include status, severity, checked_bindings, failed_rules, warnings, and errors

Requirement: Verifier input hook SHALL only perform narrow compatibility backfill

The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.

Scenario: Unique invocation candidate may be backfilled

  • WHEN an evidence binding omits source_invocation_id
  • AND exactly one current-session invocation exists for the binding's tool_name
  • THEN the hook MAY backfill source_invocation_id
  • AND it SHALL add a Gatekeeper warning describing the auto-backfill

Scenario: Raw path is never backfilled

  • WHEN an evidence binding omits raw_path
  • THEN the hook SHALL NOT synthesize raw_path
  • AND Gatekeeper SHALL treat the binding as not precise enough to pass

Requirement: Executor prompt SHALL constrain narrow-scope over-expansion

The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.

Scenario: Narrow scope produces minimal observation claims

  • WHEN the user asks to confirm one specific service, alert, or symptom
  • THEN Executor SHOULD output the minimum necessary claims, normally one and at most two
  • AND those claims SHALL be observation or negative_observation unless current-session evidence proves more
  • AND Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts