Files
SuperBizAgent-java/openspec/specs/chat-verifier-agent/spec.md
T

312 lines
18 KiB
Markdown

# chat-verifier-agent Specification
## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
## Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
- **AND** at least one critical fact has direct evidence
- **AND** no critical fact is contradicted
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical fact contradicts the tool results
- **AND** some critical facts have no supporting evidence
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only indirect support
- **WHEN** no critical fact contradicts the tool results
- **AND** all critical facts are only indirectly supported
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
- **THEN** the Verifier SHALL output verdict="REJECT"
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence`
#### Scenario: Extra confirmed-sounding answer facts are still checked
- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims`
- **THEN** Verifier SHALL add those facts to `facts_checked`
- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
#### Scenario: Natural-language fallback remains available
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer`
- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules
### Requirement: Verifier SHALL output structured JSON
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
#### Scenario: Output format validation
- **WHEN** the Verifier completes its analysis
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
#### Scenario: strict schema output
- **WHEN** the Verifier returns its result
- **THEN** it SHALL output exactly one JSON object
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
- **AND** the JSON object SHALL include `critical_fact_count`
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
### Requirement: facts_checked SHALL use a fixed classification set
Each checked fact SHALL be labeled using a fixed evidence classification.
#### Scenario: fact classification values
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
### Requirement: groundedness_score SHALL be derived from fact classifications
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
#### Scenario: contradicted fact forces reject
- **WHEN** any critical fact is labeled `contradicted`
- **THEN** the Verifier SHALL output verdict="REJECT"
- **AND** groundedness_score SHALL be `0.0`
#### Scenario: score derived from supported facts
- **WHEN** no critical fact is contradicted
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
#### Scenario: PASS → direct output
- **WHEN** Verifier outputs verdict="PASS"
- **THEN** the system SHALL output the Executor's answer directly
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL NOT start a retry round for evidence补充
- **AND** it SHALL produce a degraded output directly
#### Scenario: REJECT → degraded output
- **WHEN** Verifier outputs verdict="REJECT"
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
### Requirement: User-facing verifier outputs SHALL follow fixed templates
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
#### Scenario: LOW_CONFID uses disclaimer template
- **WHEN** the final verdict is `LOW_CONFID`
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
#### Scenario: REJECT uses degraded template
- **WHEN** the final verdict is `REJECT`
- **THEN** the user-facing response SHALL use a degraded template
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: verdict written to self_evaluation
- **WHEN** the Verifier produces a verdict
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
#### Scenario: rule evaluation stored separately
- **WHEN** the rule-based evidence scoring completes
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
- **AND** existing `verifier_evaluation` data SHALL be preserved
#### Scenario: verifier evaluation stored separately
- **WHEN** the Verifier completes
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
- **AND** existing `rule_evaluation` data SHALL be preserved
#### Scenario: no whole-object overwrite after initialization
- **WHEN** either evaluation channel updates `self_evaluation`
- **THEN** the implementation SHALL use read-modify-write semantics
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
#### Scenario: Verifier remains isolated from intermediate reasoning
- **WHEN** `executor_structured_output` is added to the verifier input
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
#### Scenario: tool trace summary derived from tool facts
- **WHEN** the system prepares verifier inputs
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
#### Scenario: tool trace summary preserves invocation references
- **WHEN** the system prepares verifier inputs
- **THEN** each summary item SHALL include a stable `trace_ref`
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
#### Scenario: only evidence-bearing tools included
- **WHEN** the system generates `tool_trace_summary`
- **THEN** it SHALL include only evidence-bearing tool invocations
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
#### Scenario: failed evidence calls preserved as evidence gaps
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
- **THEN** the summary SHALL still include that invocation
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
#### Scenario: repeated tool calls may be compacted
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
- **THEN** the system MAY compact them into a merged summary entry
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
#### Scenario: raw outputs not passed through in full
- **WHEN** a tool invocation returns large raw content
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
#### Scenario: MessagesModelHook used only for noise reduction
- **WHEN** a MessagesModelHook is used for the Verifier
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
### Requirement: Verifier facts SHALL be auditable
Verifier facts SHALL be linkable to the evidence summaries used during verification.
#### Scenario: facts_checked contains evidence refs
- **WHEN** the Verifier emits `facts_checked`
- **THEN** each fact SHALL include `evidence_refs`
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** the ChatService persists `verifier_evaluation`
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
#### Scenario: Log evidence contains a concrete matching message
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
#### Scenario: Metrics evidence contains concrete alert fields
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
#### Scenario: Summary remains bounded
- **WHEN** a tool output is large
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
#### Scenario: LOW_CONFID with critical evidence gaps
- **WHEN** Verifier labels critical facts as `no_evidence`
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor final output contains required top-level fields
- **WHEN** Executor completes a Chat diagnosis step
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: User-facing answer is available
- **WHEN** Executor emits structured output
- **THEN** `user_facing_answer` SHALL be written in Chinese
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
### Requirement: Structured Executor output SHALL degrade safely
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
#### Scenario: Malformed Executor JSON is captured
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
- **AND** it SHALL mark `executor_output_parse_status` as failed
- **AND** Verifier SHALL use natural-language fallback behavior
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract