diff --git a/devflow/index.md b/devflow/index.md index b75a82e..9dfd555 100644 --- a/devflow/index.md +++ b/devflow/index.md @@ -5,7 +5,7 @@ | 日期 | slug | 领域 | 关键词 | 状态 | |---|---|---|---|---| | 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented | -| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/executor-evidence-output-contract | proposed | +| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived | | 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived | | 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived | | 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived | diff --git a/openspec/changes/executor-evidence-output-contract/design.md b/openspec/changes/archive/2026-07-07-executor-evidence-output-contract/design.md similarity index 100% rename from openspec/changes/executor-evidence-output-contract/design.md rename to openspec/changes/archive/2026-07-07-executor-evidence-output-contract/design.md diff --git a/openspec/changes/executor-evidence-output-contract/proposal.md b/openspec/changes/archive/2026-07-07-executor-evidence-output-contract/proposal.md similarity index 100% rename from openspec/changes/executor-evidence-output-contract/proposal.md rename to openspec/changes/archive/2026-07-07-executor-evidence-output-contract/proposal.md diff --git a/openspec/changes/executor-evidence-output-contract/specs/chat-verifier-agent/spec.md b/openspec/changes/archive/2026-07-07-executor-evidence-output-contract/specs/chat-verifier-agent/spec.md similarity index 100% rename from openspec/changes/executor-evidence-output-contract/specs/chat-verifier-agent/spec.md rename to openspec/changes/archive/2026-07-07-executor-evidence-output-contract/specs/chat-verifier-agent/spec.md diff --git a/openspec/changes/executor-evidence-output-contract/tasks.md b/openspec/changes/archive/2026-07-07-executor-evidence-output-contract/tasks.md similarity index 100% rename from openspec/changes/executor-evidence-output-contract/tasks.md rename to openspec/changes/archive/2026-07-07-executor-evidence-output-contract/tasks.md diff --git a/openspec/specs/chat-verifier-agent/spec.md b/openspec/specs/chat-verifier-agent/spec.md index 8c5e298..06643b9 100644 --- a/openspec/specs/chat-verifier-agent/spec.md +++ b/openspec/specs/chat-verifier-agent/spec.md @@ -27,6 +27,22 @@ The system SHALL have a Verifier Agent that reads the Executor's answer and the - **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence - **THEN** the Verifier SHALL output verdict="REJECT" +#### Scenario: Structured Executor claims are verified first +- **WHEN** `executor_structured_output.claims` is present and valid +- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` +- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available +- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence` + +#### Scenario: Extra confirmed-sounding answer facts are still checked +- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims` +- **THEN** Verifier SHALL add those facts to `facts_checked` +- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix + +#### Scenario: Natural-language fallback remains available +- **WHEN** Executor does not return parseable structured output +- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer` +- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules + ### Requirement: Verifier SHALL output structured JSON The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale. @@ -138,8 +154,15 @@ The Verifier SHALL receive explicit verification inputs rather than inferring th - **WHEN** the Verifier starts - **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs - **AND** `retry_context` SHALL be provided on the second round only +- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output` +- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure - **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise +#### Scenario: Verifier remains isolated from intermediate reasoning +- **WHEN** `executor_structured_output` is added to the verifier input +- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning +- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context + #### Scenario: tool trace summary derived from tool facts - **WHEN** the system prepares verifier inputs - **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts @@ -230,3 +253,59 @@ When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate co - **THEN** the final user-facing response SHALL identify those gaps from verifier output - **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions +### Requirement: Executor SHALL output an evidence-attribution contract +The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information. + +#### Scenario: Executor final output contains required top-level fields +- **WHEN** Executor completes a Chat diagnosis step +- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer` +- **AND** the output SHOULD be parseable as one JSON object without Markdown fences + +#### Scenario: Confirmed claims carry evidence bindings +- **WHEN** Executor emits an item under `claims` +- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings` +- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none` +- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding + +#### Scenario: Evidence bindings support multiple tool types +- **WHEN** Executor binds evidence to a claim +- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt` +- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces +- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id` + +#### Scenario: Unsupported conclusions are not confirmed claims +- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence +- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info` +- **AND** Executor SHALL NOT present it as a confirmed claim + +#### Scenario: Runbook and skill guidance do not become incident facts +- **WHEN** Executor uses runbook, skill, or historical-case guidance +- **THEN** the guidance MAY influence `recommended_actions` +- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it + +### Requirement: User-facing Chat answers SHALL remain readable Chinese +The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract. + +#### Scenario: User-facing answer is available +- **WHEN** Executor emits structured output +- **THEN** `user_facing_answer` SHALL be written in Chinese +- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object + +#### Scenario: Machine contract remains available for trace inspection +- **WHEN** the Chat trace or verifier evaluation is inspected +- **THEN** the structured Executor contract MAY be shown for debugging or audit +- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default + +### Requirement: Structured Executor output SHALL degrade safely +The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow. + +#### Scenario: Malformed Executor JSON is captured +- **WHEN** Executor returns malformed JSON or text outside the expected contract +- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer` +- **AND** it SHALL mark `executor_output_parse_status` as failed +- **AND** Verifier SHALL use natural-language fallback behavior + +#### Scenario: Structured parse failure remains observable +- **WHEN** Executor output parsing fails +- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible +- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract