6.1 KiB
6.1 KiB
ADDED Requirements
Requirement: Executor evidence bindings SHALL support precise evidence references
Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.
Scenario: Precise evidence binding contains invocation path and excerpt
- WHEN Executor binds evidence to a claim
- THEN the binding SHOULD include singular
source_invocation_id - AND the binding SHOULD include
raw_path - AND the binding SHALL include
tool_nameandevidence_excerpt - AND the
raw_pathSHALL be interpreted relative to the referenced tool invocation'sretrieval_details.evidence_refs
Scenario: Legacy plural invocation ids remain compatibility only
- WHEN Executor emits legacy
source_invocation_ids - THEN the system MAY read them for compatibility
- AND they SHALL NOT be sufficient for a precise Gatekeeper pass without
raw_path
Requirement: Gatekeeper SHALL validate evidence reference fidelity
Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.
Scenario: Valid precise binding passes
- WHEN a binding's
source_invocation_idexists in the current session - AND the binding's
tool_namematches the persisted invocation - AND the binding's
raw_pathexists inretrieval_details.evidence_refs - AND the binding's
evidence_excerptis supported by the matching evidence ref text - THEN Gatekeeper SHALL return
status=pass - AND Gatekeeper SHALL return
severity=none
Scenario: Missing raw path is low confidence
- WHEN a binding references an existing invocation
- AND the binding omits
raw_path - THEN Gatekeeper SHALL return
status=fail - AND Gatekeeper SHALL return
severity=low_confid - AND the effective verifier result SHALL NOT be
PASS
Scenario: Old invocation without evidence refs is low confidence
- WHEN a binding references an existing invocation
- AND the invocation does not contain
retrieval_details.evidence_refs - THEN Gatekeeper SHALL return
status=fail - AND Gatekeeper SHALL return
severity=low_confid - AND the effective verifier result SHALL NOT be
PASS
Scenario: Unknown raw path is rejected
- WHEN a binding references an existing invocation
- AND the binding's
raw_pathis absent from that invocation'sretrieval_details.evidence_refs - THEN Gatekeeper SHALL return
status=fail - AND Gatekeeper SHALL return
severity=reject - AND
failed_rulesSHALL includeevidence.raw_path
Scenario: Mismatched excerpt is rejected
- WHEN a binding references an existing invocation and raw path
- AND the binding's
evidence_excerptis not supported by the matching system-side evidence ref text - THEN Gatekeeper SHALL return
status=fail - AND Gatekeeper SHALL return
severity=reject - AND
failed_rulesSHALL includeevidence.excerpt_mismatch
Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
Scenario: Verified excerpt supports direct observation
- WHEN
gatekeeper_result.severity=none - AND a claim's verified evidence excerpts directly contain the claim's concrete facts
- THEN Verifier MAY classify that claim as
direct_observation
Scenario: Tool trace summary is navigation context
- WHEN
executor_structured_output.claims[].evidence_bindingsare available - THEN Verifier SHALL use
tool_trace_summaryas navigation and audit context - AND it SHALL NOT require
tool_trace_summary.output_summaryto contain every fact already present in verified claim-local evidence
Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
Scenario: Reject severity prevents PASS
- WHEN
gatekeeper_result.severity=reject - AND the Verifier model returns
verdict=PASS - THEN ChatService SHALL downgrade the effective verdict
- AND the effective verdict SHALL be
REJECT
Scenario: Low confidence severity prevents PASS
- WHEN
gatekeeper_result.severity=low_confid - AND the Verifier model returns
verdict=PASS - THEN ChatService SHALL downgrade the effective verdict
- AND the effective verdict SHALL be
LOW_CONFID
Scenario: Gatekeeper audit includes severity
- WHEN verifier evaluation is persisted
- THEN
diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_resultSHALL includestatus,severity,checked_bindings,failed_rules,warnings, anderrors
Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
Scenario: Unique invocation candidate may be backfilled
- WHEN an evidence binding omits
source_invocation_id - AND exactly one current-session invocation exists for the binding's
tool_name - THEN the hook MAY backfill
source_invocation_id - AND it SHALL add a Gatekeeper warning describing the auto-backfill
Scenario: Raw path is never backfilled
- WHEN an evidence binding omits
raw_path - THEN the hook SHALL NOT synthesize
raw_path - AND Gatekeeper SHALL treat the binding as not precise enough to pass
Requirement: Executor prompt SHALL constrain narrow-scope over-expansion
The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.
Scenario: Narrow scope produces minimal observation claims
- WHEN the user asks to confirm one specific service, alert, or symptom
- THEN Executor SHOULD output the minimum necessary claims, normally one and at most two
- AND those claims SHALL be
observationornegative_observationunless current-session evidence proves more - AND Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts