feat(agent): add executor evidence output contract
This commit is contained in:
@@ -0,0 +1,161 @@
|
||||
## Context
|
||||
|
||||
Current Chat flow explicitly runs `Planner -> Executor -> Verifier`. `VerifierInputHook` builds a structured verifier payload containing:
|
||||
|
||||
- `original_query`
|
||||
- `executor_final_answer`
|
||||
- `tool_trace_summary`
|
||||
- `retry_context`
|
||||
|
||||
This satisfies the earlier design goal that Verifier should not see Planner/Executor intermediate reasoning. However, Executor's answer is still natural language. Verifier must infer claims from prose, then compare those claims to tool traces. Recent low-confidence sessions show that Executor often converts weak or unrelated evidence into confirmed conclusions before Verifier sees it.
|
||||
|
||||
The fix should move evidence attribution earlier: Executor must state which claims are confirmed, which evidence supports them, and which points remain hypotheses or gaps.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make Executor final output machine-checkable.
|
||||
- Separate confirmed claims, hypotheses, recommended actions, and missing information.
|
||||
- Make every confirmed claim bind to concrete tool evidence identifiers or excerpts.
|
||||
- Let Verifier consume structured claims directly instead of reconstructing all facts from prose.
|
||||
- Preserve Verifier input isolation from intermediate reasoning.
|
||||
- Preserve existing behavior when Executor output is malformed by falling back to natural-language verification and low-confidence handling.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add database tables or columns.
|
||||
- Do not change evidence tool method signatures.
|
||||
- Do not let Verifier call tools.
|
||||
- Do not require full raw tool outputs in verifier input.
|
||||
- Do not make user-facing answers become raw JSON-only unless the product layer explicitly chooses that rendering.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision: Executor emits an evidence-attribution contract
|
||||
|
||||
Executor final output SHALL be a single JSON object. The recommended contract is:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2 sentence summary using only supported facts",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "The payment-service connection pool is saturated",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "trace-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_ids": [101],
|
||||
"evidence_excerpt": "active connections reached max pool size"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "A connection leak may be contributing",
|
||||
"basis": "metrics show saturation, but no leak evidence was returned",
|
||||
"needed_evidence": ["connection lifetime metrics", "leak detection logs"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "Check HikariCP active/idle/pending connection metrics",
|
||||
"reason": "Needed to confirm pool saturation scope",
|
||||
"evidence_bindings": []
|
||||
}
|
||||
],
|
||||
"missing_info": [
|
||||
"No log evidence confirming a connection leak was returned"
|
||||
],
|
||||
"user_facing_answer": "Chinese answer rendered from the same confirmed claims, hypotheses, actions, and gaps"
|
||||
}
|
||||
```
|
||||
|
||||
Rationale: this keeps the machine contract explicit while still allowing the product to return a readable Chinese answer.
|
||||
|
||||
### Decision: Evidence binding uses generic source fields
|
||||
|
||||
`chunk_id` alone is too specific to `lookup_knowledge`. The binding shape SHALL support all evidence-bearing tools:
|
||||
|
||||
- `source_type`: `tool_trace`, `lookup_evidence_block`, or another stable source family
|
||||
- `source_id`: `tool_trace_summary.trace_ref`, evidence block id, or equivalent stable id
|
||||
- `tool_name`: evidence tool name when available
|
||||
- `source_invocation_ids`: persisted `tool_invocation` ids when available
|
||||
- `evidence_excerpt`: bounded excerpt copied or summarized from tool evidence
|
||||
|
||||
Rationale: `query_logs`, `query_metrics`, and `lookup_knowledge` expose evidence differently. A generic binding prevents the prompt contract from overfitting to RAG chunks.
|
||||
|
||||
### Decision: Confirmed claims are stricter than hypotheses and actions
|
||||
|
||||
Confirmed `claims` SHALL contain only facts supported by current-session evidence. Runbook instructions, skill guidance, and historical case patterns SHALL NOT appear as confirmed claims unless the current tool trace supports that exact fact.
|
||||
|
||||
Unsupported but useful diagnostic ideas SHALL be placed under `hypotheses` or `recommended_actions`.
|
||||
|
||||
Rationale: this directly addresses the observed hallucination: turning plausible patterns into current incident facts.
|
||||
|
||||
### Decision: Verifier prefers structured claims but keeps fallback
|
||||
|
||||
Verifier prompt SHALL use this order:
|
||||
|
||||
1. If `executor_structured_output.claims` is valid, verify each claim directly.
|
||||
2. Also scan `user_facing_answer` for extra confirmed-sounding facts not present in `claims`; mark them as facts to check.
|
||||
3. If the structured output is missing or invalid, fall back to the existing natural-language extraction from `executor_final_answer`.
|
||||
|
||||
Invalid structured output SHALL NOT crash the flow. It SHOULD produce a fallback verifier decision and make malformed structure visible in `verifier_evaluation`.
|
||||
|
||||
Rationale: the new contract should improve precision without making runtime brittle.
|
||||
|
||||
### Decision: Verifier still does not see intermediate reasoning
|
||||
|
||||
The verifier payload MAY add:
|
||||
|
||||
- `executor_structured_output`
|
||||
- `executor_output_parse_status`
|
||||
|
||||
It SHALL continue to exclude Planner reasoning, Executor intermediate reasoning, and raw conversation noise.
|
||||
|
||||
Rationale: this preserves the original verifier design: verify the final answer and tool traces, not hidden reasoning.
|
||||
|
||||
### Decision: User-facing output remains Chinese
|
||||
|
||||
For Chat user-facing pages and answers, the rendered answer SHALL be Chinese. If the Executor emits JSON, either:
|
||||
|
||||
- `user_facing_answer` is returned to the user after verifier routing, or
|
||||
- ChatService renders a Chinese answer from the structured contract.
|
||||
|
||||
The raw machine contract may remain visible in trace/debug views, but the normal user answer should not become an English/JSON-only artifact.
|
||||
|
||||
Rationale: this aligns with current UI/product language requirements while keeping machine-checkable evidence.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] LLM may emit malformed JSON. Mitigation: parse-status fallback and verifier natural-language fallback.
|
||||
- [Risk] Executor may put unsupported facts in `user_facing_answer` but omit them from `claims`. Mitigation: Verifier scans `user_facing_answer` for extra confirmed-sounding facts.
|
||||
- [Risk] Evidence excerpts may be fabricated. Mitigation: Verifier checks binding ids against `tool_trace_summary` and treats missing refs as no evidence.
|
||||
- [Risk] More verbose output increases tokens. Mitigation: keep excerpts bounded and put detailed raw evidence only in tool traces.
|
||||
- [Risk] Prompt-only enforcement is soft. Mitigation: add focused tests/eval fixtures and later consider code-level schema validation.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Update `chat-executor-prompt.md` with the evidence-attribution JSON contract.
|
||||
2. Add parsing in Chat runtime for Executor output:
|
||||
- valid JSON -> preserve parsed `executor_structured_output`
|
||||
- invalid JSON -> mark parse status and keep raw `executor_final_answer`
|
||||
3. Update verifier payload assembly to include `executor_structured_output` and parse status.
|
||||
4. Update `chat-verifier-prompt.md` to prefer structured claims and verify extra facts in `user_facing_answer`.
|
||||
5. Persist parsed or raw structured output in existing trace/evaluation snapshots without schema changes.
|
||||
6. Add focused tests and regression fixtures for unsupported confirmed claims.
|
||||
7. Validate OpenSpec and run targeted tests.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should `user_facing_answer` be mandatory in Executor JSON, or should ChatService render it from structured fields?
|
||||
- Should malformed Executor JSON force `LOW_CONFID`, or should Verifier decide based on natural-language fallback alone?
|
||||
- Should `support_level` allow only `direct`, `indirect`, and `none`, or should it also include `contradicted` for Executor self-reporting?
|
||||
@@ -0,0 +1,42 @@
|
||||
# Proposal: executor-evidence-output-contract
|
||||
|
||||
## Why
|
||||
|
||||
Recent Chat diagnosis sessions frequently end as `LOW_CONFID` even though evidence tools were called successfully. The main failure mode is not missing tool execution; it is that Executor blends tool evidence, runbook/reference patterns, and model inference into one natural-language answer, then presents unsupported details as confirmed incident facts.
|
||||
|
||||
Verifier currently receives `original_query`, `executor_final_answer`, `tool_trace_summary`, and optional `retry_context`. It can catch unsupported claims, but it must first infer facts from unstructured prose. That makes the quality gate reactive and noisy: unsupported claims are detected after the answer has already been shaped as a confident story.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Update the Chat Executor prompt so its final output follows a strict evidence-attribution JSON contract.
|
||||
- Require confirmed claims to carry explicit evidence bindings and require unsupported items to be placed in `hypotheses`, `missing_info`, or `recommended_actions` instead of confirmed conclusions.
|
||||
- Extend the verifier input contract with a parsed or raw `executor_structured_output` field while preserving the existing `executor_final_answer` field for fallback compatibility.
|
||||
- Update the Chat Verifier prompt so it prefers Executor-provided structured claims over natural-language fact extraction.
|
||||
- Keep Verifier isolated from Planner/Executor intermediate reasoning; it still receives only the original query, Executor final output, structured claim contract, tool trace summary, and retry context.
|
||||
- Do not change evidence tool signatures, database schema, or the `Planner -> Executor -> Verifier` topology.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `chat-verifier-agent`: Chat verification now consumes an Executor evidence-attribution contract when available.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected prompts: `chat-executor-prompt.md`, `chat-verifier-prompt.md`.
|
||||
- Affected integration: `VerifierInputHook` or equivalent verifier payload assembly must include `executor_structured_output` when Executor returns valid structured JSON.
|
||||
- Affected parsing: `ChatService` may parse Executor output to separate machine contract from user-facing text, with fallback when parsing fails.
|
||||
- Affected tests/eval: prompt contract tests, verifier input assembly tests, and unsupported-claim regression fixtures.
|
||||
- No database schema change is required. Structured Executor output can be persisted in existing step output fields and verifier snapshots.
|
||||
|
||||
## Out Of Scope
|
||||
|
||||
- Loosening Verifier scoring or `PASS` criteria.
|
||||
- Raising `verifier.low-confidence-threshold` to hide unsupported claims.
|
||||
- Treating runbook, skill, or historical case text as current incident evidence unless it was returned as evidence for the current query and bound explicitly.
|
||||
- Adding a new verifier tool or allowing Verifier to perform retrieval.
|
||||
- Replacing the existing tool trace summary contract.
|
||||
@@ -0,0 +1,95 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** when Executor returns a valid evidence-attribution contract, the system SHALL provide `executor_structured_output`
|
||||
- **AND** when Executor output parsing fails, the system SHALL provide an `executor_output_parse_status` that indicates the failure
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
|
||||
#### Scenario: Verifier remains isolated from intermediate reasoning
|
||||
- **WHEN** `executor_structured_output` is added to the verifier input
|
||||
- **THEN** the input SHALL still exclude Planner reasoning and Executor intermediate reasoning
|
||||
- **AND** the input SHALL be limited to the original query, final Executor output, parsed Executor evidence contract, tool trace summary, and retry context
|
||||
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
|
||||
#### Scenario: Structured Executor claims are verified first
|
||||
- **WHEN** `executor_structured_output.claims` is present and valid
|
||||
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary`
|
||||
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
|
||||
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_evidence`
|
||||
|
||||
#### Scenario: Extra confirmed-sounding answer facts are still checked
|
||||
- **WHEN** `executor_structured_output.user_facing_answer` contains confirmed-sounding facts that are absent from `executor_structured_output.claims`
|
||||
- **THEN** Verifier SHALL add those facts to `facts_checked`
|
||||
- **AND** unsupported extra facts SHALL lower the verdict according to the existing verdict matrix
|
||||
|
||||
#### Scenario: Natural-language fallback remains available
|
||||
- **WHEN** Executor does not return parseable structured output
|
||||
- **THEN** Verifier SHALL fall back to extracting facts from `executor_final_answer`
|
||||
- **AND** the final verdict SHALL still follow the existing groundedness and evidence classification rules
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Executor SHALL output an evidence-attribution contract
|
||||
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
|
||||
|
||||
#### Scenario: Executor final output contains required top-level fields
|
||||
- **WHEN** Executor completes a Chat diagnosis step
|
||||
- **THEN** its final output SHALL contain `answer_version`, `diagnosis_summary`, `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and `user_facing_answer`
|
||||
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
|
||||
|
||||
#### Scenario: Confirmed claims carry evidence bindings
|
||||
- **WHEN** Executor emits an item under `claims`
|
||||
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
|
||||
- **AND** `support_level` SHALL be one of `direct`, `indirect`, or `none`
|
||||
- **AND** claims with `support_level=direct` or `support_level=indirect` SHALL include at least one evidence binding
|
||||
|
||||
#### Scenario: Evidence bindings support multiple tool types
|
||||
- **WHEN** Executor binds evidence to a claim
|
||||
- **THEN** each binding SHALL include `source_type`, `source_id`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
|
||||
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
|
||||
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
|
||||
|
||||
#### Scenario: Unsupported conclusions are not confirmed claims
|
||||
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
|
||||
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
|
||||
- **AND** Executor SHALL NOT present it as a confirmed claim
|
||||
|
||||
#### Scenario: Runbook and skill guidance do not become incident facts
|
||||
- **WHEN** Executor uses runbook, skill, or historical-case guidance
|
||||
- **THEN** the guidance MAY influence `recommended_actions`
|
||||
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
|
||||
|
||||
### Requirement: User-facing Chat answers SHALL remain readable Chinese
|
||||
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
|
||||
|
||||
#### Scenario: User-facing answer is available
|
||||
- **WHEN** Executor emits structured output
|
||||
- **THEN** `user_facing_answer` SHALL be written in Chinese
|
||||
- **AND** it SHALL be consistent with the confirmed claims, hypotheses, recommended actions, and missing information in the same JSON object
|
||||
|
||||
#### Scenario: Machine contract remains available for trace inspection
|
||||
- **WHEN** the Chat trace or verifier evaluation is inspected
|
||||
- **THEN** the structured Executor contract MAY be shown for debugging or audit
|
||||
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
|
||||
|
||||
### Requirement: Structured Executor output SHALL degrade safely
|
||||
The system SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow.
|
||||
|
||||
#### Scenario: Malformed Executor JSON is captured
|
||||
- **WHEN** Executor returns malformed JSON or text outside the expected contract
|
||||
- **THEN** Chat runtime SHALL preserve the raw `executor_final_answer`
|
||||
- **AND** it SHALL mark `executor_output_parse_status` as failed
|
||||
- **AND** Verifier SHALL use natural-language fallback behavior
|
||||
|
||||
#### Scenario: Structured parse failure remains observable
|
||||
- **WHEN** Executor output parsing fails
|
||||
- **THEN** the verifier evaluation or trace snapshot SHALL make the parse failure visible
|
||||
- **AND** the failure SHALL NOT be silently treated as a successful evidence-attribution contract
|
||||
@@ -0,0 +1,31 @@
|
||||
## 1. Executor Contract
|
||||
|
||||
- [x] 1.1 Update `chat-executor-prompt.md` with the evidence-attribution JSON contract.
|
||||
- [x] 1.2 Require confirmed claims, hypotheses, recommended actions, missing information, and Chinese `user_facing_answer`.
|
||||
- [x] 1.3 Add prompt-level constraints that runbook/skill/history guidance cannot become current incident facts without current-session evidence.
|
||||
|
||||
## 2. Runtime Parsing And Verifier Input
|
||||
|
||||
- [x] 2.1 Parse Executor final output into `executor_structured_output` when it is valid JSON.
|
||||
- [x] 2.2 Add `executor_output_parse_status` for valid, missing, and malformed outputs.
|
||||
- [x] 2.3 Pass `executor_structured_output` and parse status into the verifier payload while preserving existing `executor_final_answer`.
|
||||
- [x] 2.4 Persist or expose the structured output through existing trace/evaluation snapshots without a schema change.
|
||||
|
||||
## 3. Verifier Prompt
|
||||
|
||||
- [x] 3.1 Update `chat-verifier-prompt.md` to verify structured claims before natural-language extraction.
|
||||
- [x] 3.2 Add a verifier rule to scan `user_facing_answer` for extra confirmed-sounding facts not present in structured `claims`.
|
||||
- [x] 3.3 Keep the existing natural-language fallback for malformed or absent structured output.
|
||||
|
||||
## 4. Tests And Evaluation
|
||||
|
||||
- [x] 4.1 Add focused tests for Executor output parsing and parse-status fallback.
|
||||
- [x] 4.2 Add focused tests for verifier payload assembly with structured Executor output.
|
||||
- [x] 4.3 Add prompt/fixture regression coverage for unsupported claims being placed outside confirmed `claims`.
|
||||
- [x] 4.4 Add or update eval fixture checks that confirmed claims must have evidence bindings.
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- [x] 5.1 Run targeted unit tests for the parsing and verifier-input changes.
|
||||
- [x] 5.2 Run relevant Chat verifier/eval regression tests.
|
||||
- [x] 5.3 Run OpenSpec validation for `executor-evidence-output-contract`.
|
||||
Reference in New Issue
Block a user