chore(openspec): archive executor evidence output contract

This commit is contained in:
zhuyongxin
2026-07-07 19:04:24 +08:00
parent 04eb50e2b4
commit 3b62a8940c
6 changed files with 80 additions and 1 deletions
@@ -0,0 +1,161 @@
## Context
Current Chat flow explicitly runs `Planner -> Executor -> Verifier`. `VerifierInputHook` builds a structured verifier payload containing:
- `original_query`
- `executor_final_answer`
- `tool_trace_summary`
- `retry_context`
This satisfies the earlier design goal that Verifier should not see Planner/Executor intermediate reasoning. However, Executor's answer is still natural language. Verifier must infer claims from prose, then compare those claims to tool traces. Recent low-confidence sessions show that Executor often converts weak or unrelated evidence into confirmed conclusions before Verifier sees it.
The fix should move evidence attribution earlier: Executor must state which claims are confirmed, which evidence supports them, and which points remain hypotheses or gaps.
## Goals / Non-Goals
**Goals:**
- Make Executor final output machine-checkable.
- Separate confirmed claims, hypotheses, recommended actions, and missing information.
- Make every confirmed claim bind to concrete tool evidence identifiers or excerpts.
- Let Verifier consume structured claims directly instead of reconstructing all facts from prose.
- Preserve Verifier input isolation from intermediate reasoning.
- Preserve existing behavior when Executor output is malformed by falling back to natural-language verification and low-confidence handling.
**Non-Goals:**
- Do not add database tables or columns.
- Do not change evidence tool method signatures.
- Do not let Verifier call tools.
- Do not require full raw tool outputs in verifier input.
- Do not make user-facing answers become raw JSON-only unless the product layer explicitly chooses that rendering.
## Decisions
### Decision: Executor emits an evidence-attribution contract
Executor final output SHALL be a single JSON object. The recommended contract is:
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2 sentence summary using only supported facts",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "root_cause",
"claim_text": "The payment-service connection pool is saturated",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"evidence_excerpt": "active connections reached max pool size"
}
]
}
],
"hypotheses": [
{
"hypothesis_text": "A connection leak may be contributing",
"basis": "metrics show saturation, but no leak evidence was returned",
"needed_evidence": ["connection lifetime metrics", "leak detection logs"]
}
],
"recommended_actions": [
{
"action_text": "Check HikariCP active/idle/pending connection metrics",
"reason": "Needed to confirm pool saturation scope",
"evidence_bindings": []
}
],
"missing_info": [
"No log evidence confirming a connection leak was returned"
],
"user_facing_answer": "Chinese answer rendered from the same confirmed claims, hypotheses, actions, and gaps"
}
```
Rationale: this keeps the machine contract explicit while still allowing the product to return a readable Chinese answer.
### Decision: Evidence binding uses generic source fields
`chunk_id` alone is too specific to `lookup_knowledge`. The binding shape SHALL support all evidence-bearing tools:
- `source_type`: `tool_trace`, `lookup_evidence_block`, or another stable source family
- `source_id`: `tool_trace_summary.trace_ref`, evidence block id, or equivalent stable id
- `tool_name`: evidence tool name when available
- `source_invocation_ids`: persisted `tool_invocation` ids when available
- `evidence_excerpt`: bounded excerpt copied or summarized from tool evidence
Rationale: `query_logs`, `query_metrics`, and `lookup_knowledge` expose evidence differently. A generic binding prevents the prompt contract from overfitting to RAG chunks.
### Decision: Confirmed claims are stricter than hypotheses and actions
Confirmed `claims` SHALL contain only facts supported by current-session evidence. Runbook instructions, skill guidance, and historical case patterns SHALL NOT appear as confirmed claims unless the current tool trace supports that exact fact.
Unsupported but useful diagnostic ideas SHALL be placed under `hypotheses` or `recommended_actions`.
Rationale: this directly addresses the observed hallucination: turning plausible patterns into current incident facts.
### Decision: Verifier prefers structured claims but keeps fallback
Verifier prompt SHALL use this order:
1. If `executor_structured_output.claims` is valid, verify each claim directly.
2. Also scan `user_facing_answer` for extra confirmed-sounding facts not present in `claims`; mark them as facts to check.
3. If the structured output is missing or invalid, fall back to the existing natural-language extraction from `executor_final_answer`.
Invalid structured output SHALL NOT crash the flow. It SHOULD produce a fallback verifier decision and make malformed structure visible in `verifier_evaluation`.
Rationale: the new contract should improve precision without making runtime brittle.
### Decision: Verifier still does not see intermediate reasoning
The verifier payload MAY add:
- `executor_structured_output`
- `executor_output_parse_status`
It SHALL continue to exclude Planner reasoning, Executor intermediate reasoning, and raw conversation noise.
Rationale: this preserves the original verifier design: verify the final answer and tool traces, not hidden reasoning.
### Decision: User-facing output remains Chinese
For Chat user-facing pages and answers, the rendered answer SHALL be Chinese. If the Executor emits JSON, either:
- `user_facing_answer` is returned to the user after verifier routing, or
- ChatService renders a Chinese answer from the structured contract.
The raw machine contract may remain visible in trace/debug views, but the normal user answer should not become an English/JSON-only artifact.
Rationale: this aligns with current UI/product language requirements while keeping machine-checkable evidence.
## Risks / Trade-offs
- [Risk] LLM may emit malformed JSON. Mitigation: parse-status fallback and verifier natural-language fallback.
- [Risk] Executor may put unsupported facts in `user_facing_answer` but omit them from `claims`. Mitigation: Verifier scans `user_facing_answer` for extra confirmed-sounding facts.
- [Risk] Evidence excerpts may be fabricated. Mitigation: Verifier checks binding ids against `tool_trace_summary` and treats missing refs as no evidence.
- [Risk] More verbose output increases tokens. Mitigation: keep excerpts bounded and put detailed raw evidence only in tool traces.
- [Risk] Prompt-only enforcement is soft. Mitigation: add focused tests/eval fixtures and later consider code-level schema validation.
## Migration Plan
1. Update `chat-executor-prompt.md` with the evidence-attribution JSON contract.
2. Add parsing in Chat runtime for Executor output:
- valid JSON -> preserve parsed `executor_structured_output`
- invalid JSON -> mark parse status and keep raw `executor_final_answer`
3. Update verifier payload assembly to include `executor_structured_output` and parse status.
4. Update `chat-verifier-prompt.md` to prefer structured claims and verify extra facts in `user_facing_answer`.
5. Persist parsed or raw structured output in existing trace/evaluation snapshots without schema changes.
6. Add focused tests and regression fixtures for unsupported confirmed claims.
7. Validate OpenSpec and run targeted tests.
## Open Questions
- Should `user_facing_answer` be mandatory in Executor JSON, or should ChatService render it from structured fields?
- Should malformed Executor JSON force `LOW_CONFID`, or should Verifier decide based on natural-language fallback alone?
- Should `support_level` allow only `direct`, `indirect`, and `none`, or should it also include `contradicted` for Executor self-reporting?
@@ -0,0 +1,42 @@
# Proposal: executor-evidence-output-contract
## Why
Recent Chat diagnosis sessions frequently end as `LOW_CONFID` even though evidence tools were called successfully. The main failure mode is not missing tool execution; it is that Executor blends tool evidence, runbook/reference patterns, and model inference into one natural-language answer, then presents unsupported details as confirmed incident facts.
Verifier currently receives `original_query`, `executor_final_answer`, `tool_trace_summary`, and optional `retry_context`. It can catch unsupported claims, but it must first infer facts from unstructured prose. That makes the quality gate reactive and noisy: unsupported claims are detected after the answer has already been shaped as a confident story.
## What Changes
- Update the Chat Executor prompt so its final output follows a strict evidence-attribution JSON contract.
- Require confirmed claims to carry explicit evidence bindings and require unsupported items to be placed in `hypotheses`, `missing_info`, or `recommended_actions` instead of confirmed conclusions.
- Extend the verifier input contract with a parsed or raw `executor_structured_output` field while preserving the existing `executor_final_answer` field for fallback compatibility.
- Update the Chat Verifier prompt so it prefers Executor-provided structured claims over natural-language fact extraction.
- Keep Verifier isolated from Planner/Executor intermediate reasoning; it still receives only the original query, Executor final output, structured claim contract, tool trace summary, and retry context.
- Do not change evidence tool signatures, database schema, or the `Planner -> Executor -> Verifier` topology.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: Chat verification now consumes an Executor evidence-attribution contract when available.
## Impact
- Affected prompts: `chat-executor-prompt.md`, `chat-verifier-prompt.md`.
- Affected integration: `VerifierInputHook` or equivalent verifier payload assembly must include `executor_structured_output` when Executor returns valid structured JSON.
- Affected parsing: `ChatService` may parse Executor output to separate machine contract from user-facing text, with fallback when parsing fails.
- Affected tests/eval: prompt contract tests, verifier input assembly tests, and unsupported-claim regression fixtures.
- No database schema change is required. Structured Executor output can be persisted in existing step output fields and verifier snapshots.
## Out Of Scope
- Loosening Verifier scoring or `PASS` criteria.
- Raising `verifier.low-confidence-threshold` to hide unsupported claims.
- Treating runbook, skill, or historical case text as current incident evidence unless it was returned as evidence for the current query and bound explicitly.
- Adding a new verifier tool or allowing Verifier to perform retrieval.
- Replacing the existing tool trace summary contract.
@@ -0,0 +1,95 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier starts
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
- **AND** `retry_context` SHALL be provided on the second round only
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
#### Scenario: Verifier remains isolated from intermediate reasoning
- **WHEN** `executor_structured_output` is added to the verifier input
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
#### Scenario: Structured Executor claims are verified first
- **WHEN** `executor_structured_output.claims` is present and valid
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary`
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence`
#### Scenario: Extra confirmed-sounding answer facts are still checked
- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims`
- **THEN** Verifier SHALL add those facts to `facts_checked`
- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
#### Scenario: Natural-language fallback remains available
- **WHEN** Executor does not return parseable structured output
- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer`
- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules
## ADDED Requirements
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor final output contains required top-level fields
- **WHEN** Executor completes a Chat diagnosis step
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: User-facing answer is available
- **WHEN** Executor emits structured output
- **THEN** `user_facing_answer` SHALL be written in Chinese
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
### Requirement: Structured Executor output SHALL degrade safely
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
#### Scenario: Malformed Executor JSON is captured
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
- **AND** it SHALL mark `executor_output_parse_status` as failed
- **AND** Verifier SHALL use natural-language fallback behavior
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
@@ -0,0 +1,31 @@
## 1. Executor Contract
- [x] 1.1 Update `chat-executor-prompt.md` with the evidence-attribution JSON contract.
- [x] 1.2 Require confirmed claims, hypotheses, recommended actions, missing information, and Chinese `user_facing_answer`.
- [x] 1.3 Add prompt-level constraints that runbook/skill/history guidance cannot become current incident facts without current-session evidence.
## 2. Runtime Parsing And Verifier Input
- [x] 2.1 Parse Executor final output into `executor_structured_output` when it is valid JSON.
- [x] 2.2 Add `executor_output_parse_status` for valid, missing, and malformed outputs.
- [x] 2.3 Pass `executor_structured_output` and parse status into the verifier payload while preserving existing `executor_final_answer`.
- [x] 2.4 Persist or expose the structured output through existing trace/evaluation snapshots without a schema change.
## 3. Verifier Prompt
- [x] 3.1 Update `chat-verifier-prompt.md` to verify structured claims before natural-language extraction.
- [x] 3.2 Add a verifier rule to scan `user_facing_answer` for extra confirmed-sounding facts not present in structured `claims`.
- [x] 3.3 Keep the existing natural-language fallback for malformed or absent structured output.
## 4. Tests And Evaluation
- [x] 4.1 Add focused tests for Executor output parsing and parse-status fallback.
- [x] 4.2 Add focused tests for verifier payload assembly with structured Executor output.
- [x] 4.3 Add prompt/fixture regression coverage for unsupported claims being placed outside confirmed `claims`.
- [x] 4.4 Add or update eval fixture checks that confirmed claims must have evidence bindings.
## 5. Verification
- [x] 5.1 Run targeted unit tests for the parsing and verifier-input changes.
- [x] 5.2 Run relevant Chat verifier/eval regression tests.
- [x] 5.3 Run OpenSpec validation for `executor-evidence-output-contract`.