6.2 KiB
6.2 KiB
MODIFIED Requirements
Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
Scenario: explicit input blocks available to Verifier
- WHEN the Verifier starts
- THEN the system SHALL provide
original_query,executor_final_answer, andtool_trace_summaryas explicit inputs - AND
retry_contextSHALL be provided on the second round only - AND when Executor returns a valid evidence-attribution contract, the system SHALL provide
executor_structured_output - AND when Executor output parsing fails, the system SHALL provide an
executor_output_parse_statusthat indicates the failure - AND message filtering MAY be used only to remove intermediate reasoning or unrelated noise
Scenario: Verifier remains isolated from intermediate reasoning
- WHEN
executor_structured_outputis added to the verifier input - THEN the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- AND the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
Scenario: Structured Executor claims are verified first
- WHEN
executor_structured_output.claimsis present and valid - THEN Verifier SHALL verify each structured claim against
tool_trace_summary - AND each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- AND a claim with fabricated or missing evidence references SHALL NOT be classified as
direct_evidence
Scenario: Extra confirmed-sounding answer facts are still checked
- WHEN
executor_structured_output.user_facing_answercontains confirmed-sounding facts that are absent fromexecutor_structured_output.claims - THEN Verifier SHALL add those facts to
facts_checked - AND unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
Scenario: Natural-language fallback remains available
- WHEN Executor does not return parseable structured output
- THEN Verifier SHALL fall back to extracting facts from
executor_final_answer - AND the final verdict SHALL still follow the existing groundedness and evidence classification rules
ADDED Requirements
Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
Scenario: Executor final output contains required top-level fields
- WHEN Executor completes a Chat diagnosis step
- THEN its final output SHALL contain
answer_version,diagnosis_summary,claims,hypotheses,recommended_actions,missing_info, anduser_facing_answer - AND the output SHOULD be parseable as one JSON object without Markdown fences
Scenario: Confirmed claims carry evidence bindings
- WHEN Executor emits an item under
claims - THEN the item SHALL include
claim_id,claim_type,claim_text,support_level, andevidence_bindings - AND
support_levelSHALL be one ofdirect,indirect, ornone - AND claims with
support_level=directorsupport_level=indirectSHALL include at least one evidence binding
Scenario: Evidence bindings support multiple tool types
- WHEN Executor binds evidence to a claim
- THEN each binding SHALL include
source_type,source_id,tool_name,source_invocation_ids, andevidence_excerpt - AND the binding SHALL be able to reference
lookup_knowledge,query_logs,query_metrics, or other evidence-bearing tool traces - AND the binding SHALL NOT rely only on a RAG-specific
chunk_id
Scenario: Unsupported conclusions are not confirmed claims
- WHEN a possible root cause, detail, or remediation lacks current-session tool evidence
- THEN Executor SHALL place it under
hypotheses,recommended_actions, ormissing_info - AND Executor SHALL NOT present it as a confirmed claim
Scenario: Runbook and skill guidance do not become incident facts
- WHEN Executor uses runbook, skill, or historical-case guidance
- THEN the guidance MAY influence
recommended_actions - AND the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
Scenario: User-facing answer is available
- WHEN Executor emits structured output
- THEN
user_facing_answerSHALL be written in Chinese - AND it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
Scenario: Machine contract remains available for trace inspection
- WHEN the Chat trace or verifier evaluation is inspected
- THEN the structured Executor contract MAY be shown for debugging or audit
- AND normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
Requirement: Structured Executor output SHALL degrade safely
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
Scenario: Malformed Executor JSON is captured
- WHEN Executor returns malformed JSON or text outside the expected contract
- THEN Chat runtime SHALL preserve the raw
executor_final_answer - AND it SHALL mark
executor_output_parse_statusas failed - AND Verifier SHALL use natural-language fallback behavior
Scenario: Structured parse failure remains observable
- WHEN Executor output parsing fails
- THEN the verifier evaluation or trace snapshot SHALL make the parse failure visible
- AND the failure SHALL NOT be silently treated as a successful evidence-attribution contract