feat(graph): cut over chat diagnosis stategraph
This commit is contained in:
@@ -4,41 +4,41 @@
|
||||
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
|
||||
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
|
||||
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
|
||||
- **AND** at least one critical claim has direct observation
|
||||
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
|
||||
- **AND** `gatekeeper_result.status` is not `fail`
|
||||
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
- **AND** verdict ceiling is PASS
|
||||
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **WHEN** no critical claim contradicts verified evidence
|
||||
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only inference support
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **WHEN** no critical claim contradicts verified evidence
|
||||
- **AND** all critical claims are only `reasonable_inference`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
|
||||
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
|
||||
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
|
||||
- **THEN** the Verifier SHALL output model verdict="REJECT"
|
||||
|
||||
#### Scenario: Structured Executor claims are verified first
|
||||
- **WHEN** `executor_structured_output.claims` is present and valid
|
||||
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
|
||||
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
|
||||
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
#### Scenario: Verified structured claims are the only verification target
|
||||
- **WHEN** `verified_executor_output.claims` is present
|
||||
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
|
||||
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
|
||||
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
|
||||
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
|
||||
|
||||
#### Scenario: Malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** Executor does not return parseable structured output
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
#### Scenario: Executor output is invalid
|
||||
- **WHEN** Executor does not return a legal structured contract
|
||||
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
|
||||
@@ -64,11 +64,11 @@ The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
|
||||
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
|
||||
|
||||
#### Scenario: claim checks are mapped to legacy facts
|
||||
- **WHEN** Verifier output contains `claim_checks`
|
||||
- **THEN** ChatService SHALL derive compatibility `facts_checked`
|
||||
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
|
||||
- **AND** `direct_observation` SHALL map to `direct_evidence`
|
||||
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
|
||||
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
|
||||
@@ -89,73 +89,86 @@ The groundedness score SHALL be computed from critical fact classifications inst
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
|
||||
|
||||
#### Scenario: PASS -> Composer output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
#### Scenario: PASS routes to Composer
|
||||
- **WHEN** Verifier completes with effective verdict="PASS"
|
||||
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
|
||||
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
|
||||
- **THEN** the Graph SHALL route to Composer without another Planner cycle
|
||||
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
|
||||
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
|
||||
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
|
||||
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
|
||||
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
|
||||
- **WHEN** Verifier completes with effective verdict="REJECT"
|
||||
- **THEN** the Graph SHALL NOT start an evidence supplementation round
|
||||
- **AND** it SHALL route to Composer-safe output
|
||||
|
||||
#### Scenario: REJECT -> degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
#### Scenario: REJECT produces bounded output
|
||||
- **WHEN** effective verdict is REJECT
|
||||
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
|
||||
- **AND** it SHALL NOT pass through raw Executor answer
|
||||
- **AND** it SHALL NOT include an unsupported root-cause conclusion
|
||||
|
||||
#### Scenario: Verifier execution fails
|
||||
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
|
||||
- **THEN** the Graph SHALL route to pre-verification Fallback
|
||||
- **AND** no execution status string SHALL be used as model or effective verdict
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||
- **WHEN** a completed Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
|
||||
|
||||
#### Scenario: composer output written to self_evaluation
|
||||
- **WHEN** final answer composition completes
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
|
||||
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
|
||||
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
|
||||
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
|
||||
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
- **WHEN** Verifier completes
|
||||
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
|
||||
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
|
||||
|
||||
#### Scenario: pre-verification fallback is persisted
|
||||
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
|
||||
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
|
||||
|
||||
#### Scenario: gatekeeper result written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
|
||||
- **WHEN** Graph result mapping persists available Gatekeeper state
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
|
||||
|
||||
#### Scenario: prompt audit written to verifier evaluation
|
||||
- **WHEN** the Chat verifier evaluation is persisted
|
||||
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
|
||||
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
|
||||
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
|
||||
- **WHEN** a complex Chat Graph result is persisted
|
||||
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
|
||||
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
|
||||
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
|
||||
- **AND** full Prompt text SHALL NOT be persisted
|
||||
|
||||
#### Scenario: prompt audit available on fallback paths
|
||||
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
|
||||
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
|
||||
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||
|
||||
#### Scenario: evaluation payload is inspected
|
||||
- **WHEN** Graph verifier evaluation is persisted
|
||||
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
|
||||
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
@@ -175,88 +188,51 @@ The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation ch
|
||||
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
|
||||
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
- **WHEN** the Verifier Graph Node starts
|
||||
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
|
||||
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
|
||||
|
||||
#### Scenario: Verifier remains isolated from intermediate reasoning
|
||||
- **WHEN** `executor_structured_output` is added to the verifier input
|
||||
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
|
||||
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, gatekeeper result, and retry context
|
||||
#### Scenario: Verifier remains isolated from intermediate and raw material
|
||||
- **WHEN** the Verifier input is serialized
|
||||
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
|
||||
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
|
||||
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
|
||||
#### Scenario: only passed bindings are available
|
||||
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
|
||||
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
|
||||
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
|
||||
|
||||
#### Scenario: tool trace summary preserves invocation references
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** each summary item SHALL include a stable `trace_ref`
|
||||
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
|
||||
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
|
||||
#### Scenario: verified evidence preserves precise references
|
||||
- **WHEN** the system prepares Verifier input
|
||||
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
|
||||
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
|
||||
|
||||
#### Scenario: only evidence-bearing tools included
|
||||
- **WHEN** the system generates `tool_trace_summary`
|
||||
- **THEN** it SHALL include only evidence-bearing tool invocations
|
||||
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
|
||||
#### Scenario: gatekeeper audit and ceiling are available
|
||||
- **WHEN** the system prepares Verifier input
|
||||
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
|
||||
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
|
||||
|
||||
#### Scenario: failed evidence calls preserved as evidence gaps
|
||||
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
|
||||
- **THEN** the summary SHALL still include that invocation
|
||||
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
|
||||
|
||||
#### Scenario: repeated tool calls may be compacted
|
||||
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
|
||||
- **THEN** the system MAY compact them into a merged summary entry
|
||||
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
|
||||
|
||||
#### Scenario: raw outputs not passed through in full
|
||||
- **WHEN** a tool invocation returns large raw content
|
||||
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
|
||||
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
|
||||
|
||||
#### Scenario: MessagesModelHook used only for noise reduction
|
||||
- **WHEN** a MessagesModelHook is used for the Verifier
|
||||
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
|
||||
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
|
||||
|
||||
#### Scenario: gatekeeper result available to Verifier
|
||||
- **WHEN** the system prepares verifier inputs from Executor output
|
||||
- **THEN** the payload SHALL include `gatekeeper_result`
|
||||
- **AND** `gatekeeper_result.status` SHALL be one of `pass`, `warn`, or `fail`
|
||||
- **AND** `gatekeeper_result` SHALL include `failed_rules`, `warnings`, and `errors`
|
||||
|
||||
#### Scenario: structured claims are the primary verification target
|
||||
- **WHEN** `executor_output_parse_status.status` is `valid`
|
||||
- **AND** `executor_structured_output.claims` is available
|
||||
- **THEN** Verifier SHALL verify each claim through `claim_checks`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
|
||||
#### Scenario: malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
#### Scenario: technical retry occurs
|
||||
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
|
||||
- **THEN** the second attempt SHALL receive byte-identical serialized input
|
||||
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
|
||||
|
||||
#### Scenario: facts_checked contains evidence refs
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL include `evidence_refs`
|
||||
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
|
||||
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
|
||||
#### Scenario: claim checks contain evidence refs
|
||||
- **WHEN** the Verifier emits `claim_checks`
|
||||
- **THEN** each check SHALL include an `evidence_refs` array
|
||||
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
|
||||
- **AND** it SHALL NOT reference a failed or unverified binding
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **WHEN** Graph result mapping persists verifier evaluation
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
|
||||
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
@@ -346,21 +322,21 @@ The system SHALL preserve a readable Chinese answer for normal Chat users even w
|
||||
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
|
||||
|
||||
### Requirement: Structured Executor output SHALL degrade safely
|
||||
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
|
||||
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
|
||||
|
||||
#### Scenario: Malformed Executor JSON is captured
|
||||
#### Scenario: Malformed Executor JSON is classified
|
||||
- **WHEN** Executor returns malformed JSON or text outside the expected contract
|
||||
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
|
||||
- **AND** it SHALL mark `executor_output_parse_status` as failed
|
||||
- **AND** Verifier SHALL use natural-language fallback behavior
|
||||
- **THEN** Executor Node SHALL set INVALID_OUTPUT
|
||||
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
|
||||
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
|
||||
|
||||
#### Scenario: Structured parse failure remains observable
|
||||
- **WHEN** Executor output parsing fails
|
||||
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
|
||||
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
|
||||
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
|
||||
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
|
||||
|
||||
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
|
||||
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
|
||||
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
|
||||
|
||||
#### Scenario: schema rule rejects removed fields
|
||||
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
|
||||
@@ -373,32 +349,31 @@ The system SHALL run deterministic Gatekeeper checks after Executor output parsi
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
|
||||
|
||||
#### Scenario: invocation rule rejects fabricated invocation ids
|
||||
- **WHEN** a claim evidence binding references a `source_invocation_ids` value that is not present in current-session `tool_invocation` rows
|
||||
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||
|
||||
#### Scenario: invocation rule rejects tool name mismatch
|
||||
- **WHEN** a claim evidence binding references an existing invocation id
|
||||
- **AND** the binding `tool_name` does not match the invocation's persisted `tool_name`
|
||||
- **WHEN** a claim evidence binding references an existing current-run invocation id
|
||||
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||
|
||||
#### Scenario: valid structured output passes initial gatekeeper rules
|
||||
- **WHEN** Executor emits `executor_evidence_v2`
|
||||
- **AND** each claim has evidence bindings pointing to current-session invocations with matching tool names
|
||||
- **WHEN** Executor emits legal `executor_evidence_v2`
|
||||
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `pass`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
|
||||
|
||||
#### Scenario: gatekeeper fail prevents PASS
|
||||
- **WHEN** `gatekeeper_result.status` is `fail`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||
- **AND** the effective verdict SHALL NOT be `PASS`
|
||||
#### Scenario: gatekeeper reject bypasses Verifier
|
||||
- **WHEN** normalized Gatekeeper status is REJECT
|
||||
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||
- **AND** Verifier SHALL NOT execute
|
||||
|
||||
#### Scenario: invocation reference failure downgrades to reject
|
||||
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
|
||||
#### Scenario: gatekeeper low confidence is bounded
|
||||
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
|
||||
- **THEN** verified input SHALL contain only passed bindings
|
||||
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
|
||||
|
||||
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
|
||||
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
|
||||
@@ -489,36 +464,36 @@ Gatekeeper SHALL validate that Executor evidence bindings point to real current-
|
||||
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
|
||||
|
||||
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
|
||||
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
|
||||
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
|
||||
|
||||
#### Scenario: Verified excerpt supports direct observation
|
||||
- **WHEN** `gatekeeper_result.severity=none`
|
||||
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
|
||||
- **WHEN** verdict ceiling is PASS
|
||||
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
|
||||
- **THEN** Verifier MAY classify that claim as `direct_observation`
|
||||
|
||||
#### Scenario: Tool trace summary is navigation context
|
||||
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
|
||||
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
|
||||
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
|
||||
#### Scenario: Verified evidence is complete Verifier context
|
||||
- **WHEN** verified claims and evidence are available
|
||||
- **THEN** Verifier SHALL use them as its evidence context
|
||||
- **AND** it SHALL NOT require or request a complete tool trace summary
|
||||
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
|
||||
|
||||
### Requirement: Gatekeeper severity SHALL constrain effective verdict
|
||||
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
|
||||
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
|
||||
|
||||
#### Scenario: Reject severity prevents PASS
|
||||
#### Scenario: Reject severity bypasses Verifier
|
||||
- **WHEN** `gatekeeper_result.severity=reject`
|
||||
- **AND** the Verifier model returns `verdict=PASS`
|
||||
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||
- **AND** the effective verdict SHALL be `REJECT`
|
||||
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
|
||||
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
|
||||
|
||||
#### Scenario: Low confidence severity prevents PASS
|
||||
- **WHEN** `gatekeeper_result.severity=low_confid`
|
||||
- **AND** the Verifier model returns `verdict=PASS`
|
||||
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
|
||||
- **AND** effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
#### Scenario: Gatekeeper audit includes severity
|
||||
- **WHEN** verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||
- **WHEN** Graph verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||
|
||||
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
|
||||
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
|
||||
|
||||
Reference in New Issue
Block a user