feat(agent): add executor evidence v2 contract

This commit is contained in:
aruo
2026-07-08 01:37:15 +08:00
parent a6afbfaa9d
commit 050cbc8fee
21 changed files with 2707 additions and 20 deletions
@@ -0,0 +1,50 @@
## MODIFIED Requirements
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default