## MODIFIED Requirements ### Requirement: Verifier SHALL consume explicit verification inputs The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history. #### Scenario: explicit input blocks available to Verifier - **WHEN** the Verifier starts - **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs - **AND** `retry_context` SHALL be provided on the second round only - **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output` - **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure - **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise #### Scenario: Verifier remains isolated from intermediate reasoning - **WHEN** `executor_structured_output` is added to the verifier input - **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning - **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context ### Requirement: Verifier SHALL fact-check Executor answers The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict. #### Scenario: Structured Executor claims are verified first - **WHEN** `executor_structured_output.claims` is present and valid - **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` - **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available - **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence` #### Scenario: Extra confirmed-sounding answer facts are still checked - **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims` - **THEN** Verifier SHALL add those facts to `facts_checked` - **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix #### Scenario: Natural-language fallback remains available - **WHEN** Executor does not return parseable structured output - **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer` - **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules ## ADDED Requirements ### Requirement: Executor SHALL output an evidence-attribution contract The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information. #### Scenario: Executor final output contains required top-level fields - **WHEN** Executor completes a Chat diagnosis step - **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer` - **AND** the output SHOULD be parseable as one JSON object without Markdown fences #### Scenario: Confirmed claims carry evidence bindings - **WHEN** Executor emits an item under `claims` - **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings` - **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none` - **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding #### Scenario: Evidence bindings support multiple tool types - **WHEN** Executor binds evidence to a claim - **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt` - **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces - **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id` #### Scenario: Unsupported conclusions are not confirmed claims - **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence - **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info` - **AND** Executor SHALL NOT present it as a confirmed claim #### Scenario: Runbook and skill guidance do not become incident facts - **WHEN** Executor uses runbook, skill, or historical-case guidance - **THEN** the guidance MAY influence `recommended_actions` - **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it ### Requirement: User-facing Chat answers SHALL remain readable Chinese The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract. #### Scenario: User-facing answer is available - **WHEN** Executor emits structured output - **THEN** `user_facing_answer` SHALL be written in Chinese - **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object #### Scenario: Machine contract remains available for trace inspection - **WHEN** the Chat trace or verifier evaluation is inspected - **THEN** the structured Executor contract MAY be shown for debugging or audit - **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default ### Requirement: Structured Executor output SHALL degrade safely The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow. #### Scenario: Malformed Executor JSON is captured - **WHEN** Executor returns malformed JSON or text outside the expected contract - **THEN** Chat runtime SHALL preserve the raw `executor_final_answer` - **AND** it SHALL mark `executor_output_parse_status` as failed - **AND** Verifier SHALL use natural-language fallback behavior #### Scenario: Structured parse failure remains observable - **WHEN** Executor output parsing fails - **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible - **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract