feat(agent): add executor evidence v2 contract

This commit is contained in:
aruo
2026-07-08 01:37:15 +08:00
parent a6afbfaa9d
commit 050cbc8fee
21 changed files with 2707 additions and 20 deletions
@@ -0,0 +1 @@
archive-ready
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-07
@@ -0,0 +1,112 @@
# Decisions: executor-v2-output-contract
## sm-flow Progress
### Clarify
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
Slug: `executor-v2-output-contract`
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
### Context
Relevant devflow history:
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
Conflict with historical decision:
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
Current code shape:
- `chat-executor-prompt.md` defines the V1 Executor output contract.
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
### Grill
Question pool:
| Question | Mode | Resolution |
|---|---|---|
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
### Specify
OpenSpec artifacts:
- `proposal.md`: scope and compatibility boundary for stage one.
- `design.md`: V2 Executor contract and temporary rendering strategy.
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
- `tasks.md`: executable implementation and verification checklist.
### Audit
Architecture risk summary:
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
- No database schema or evidence tool signature changes are required.
Cross-artifact alignment:
| Source | Target | Status |
|---|---|---|
| issue background / stage one | proposal | aligned |
| proposal scope / non-goals | design | aligned |
| design contract and rendering bridge | specs | aligned |
| specs observable behavior | tasks | aligned |
Interface impact:
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
### Commit
Commit gate result: passed.
- `proposal.md` exists and explains why this phase is needed.
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
- `tasks.md` contains executable implementation and verification tasks.
- `cmd /c openspec validate executor-v2-output-contract` passed.
- No unresolved user-interview questions remain for this stage.
### Apply
Implementation summary:
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
- Preserved V1 `user_facing_answer` extraction for compatibility.
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
Verification:
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
- `cmd /c openspec validate executor-v2-output-contract` passed.
Known limitations:
- Gatekeeper is not implemented in this phase.
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
- The V2 renderer is temporary and should be replaced by Composer in a later phase.
@@ -0,0 +1,112 @@
## Context
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
- `diagnosis_summary`
- `user_facing_answer`
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
This stage implements only the first slice of the V2 migration:
```text
Executor V2 output contract
-> existing VerifierInputHook parsing
-> existing Verifier
-> temporary ChatService structured renderer
```
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
## Goals / Non-Goals
Goals:
- Make Chat Executor emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
- Preserve evidence-binding requirements for confirmed claims.
- Prevent PASS routing from returning raw JSON to normal Chat users.
Non-goals:
- No Gatekeeper implementation.
- No Verifier prompt rewrite to `claim_checks`.
- No Composer agent.
- No database schema changes.
- No Planner changes.
- No evidence tool signature changes.
- No retry behavior changes.
## Executor V2 Contract
Executor final output SHALL be one JSON object:
```json
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "payment-service 出现请求超时日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_logs",
"source_invocation_ids": [394],
"evidence_excerpt": "request timeout"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
}
```
Removed fields:
- `diagnosis_summary`
- `user_facing_answer`
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
## Temporary Rendering Strategy
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
When Verifier returns `PASS`:
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
- `claims[].claim_text`
- `hypotheses[].hypothesis_text`
- `missing_info[]`
- `recommended_actions[].action_text` and `reason`
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
## Parser Boundary
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
## Interface Impact
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
## Risks / Mitigations
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
- Mitigation: add temporary V2 renderer in `ChatService`.
- Risk: current Verifier prompt still mentions `user_facing_answer`.
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
- Risk: tests assume V1 fields.
- Mitigation: update/add focused tests for V2 output without final-expression fields.
@@ -0,0 +1,33 @@
## Why
The current Chat Executor evidence contract still mixes diagnostic material with final user-facing prose through `diagnosis_summary` and `user_facing_answer`. This keeps the Executor in a "diagnose and narrate" mode, so unsupported details can be smuggled into the final answer before later Gatekeeper, Verifier V2, and Composer phases exist.
This first phase narrows Executor to structured diagnostic material only and adds a temporary safe rendering path so normal Chat responses do not expose raw Executor JSON while later phases are implemented.
## What Changes
- **BREAKING internal Agent contract**: Chat Executor output changes from `executor_evidence_v1` to `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from the Chat Executor final JSON contract.
- Keep the existing structured arrays: `claims`, `hypotheses`, `recommended_actions`, and `missing_info`.
- Preserve `claims[].evidence_bindings` and current-session evidence attribution rules.
- Adjust runtime final-answer handling so a PASS result with V2 Executor output is rendered into readable Chinese from structured fields instead of returning raw JSON.
- Keep Planner, Verifier, Gatekeeper, retry behavior, database schema, and tool signatures unchanged in this phase.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `chat-verifier-agent`: The Executor evidence-attribution contract is tightened so V2 structured output no longer contains final-expression fields. Verifier still receives `executor_final_answer` as raw text and `executor_structured_output` when parseable.
## Impact
- Affected prompt: `src/main/resources/prompts/chat-executor-prompt.md`.
- Affected runtime: `ChatService` PASS answer extraction/rendering for Executor V2.
- Affected parser boundary: `VerifierInputHook` should continue parsing JSON but must not treat schema validation as its own responsibility in this phase.
- Affected tests: ChatService sequential flow tests and VerifierInputHook parsing tests for V2 output without `user_facing_answer`.
- Interface impact: L4 for internal Agent output contract because fields are removed from Executor JSON; external HTTP/chat answer behavior must remain readable Chinese and must not expose raw JSON.
@@ -0,0 +1,50 @@
## MODIFIED Requirements
### Requirement: Executor SHALL output an evidence-attribution contract
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
#### Scenario: Executor V2 final output contains only structured diagnostic fields
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
- **AND** the output SHALL NOT contain `diagnosis_summary`
- **AND** the output SHALL NOT contain `user_facing_answer`
#### Scenario: Confirmed claims carry evidence bindings
- **WHEN** Executor emits an item under `claims`
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
- **AND** `support_level` SHALL be one of `direct` or `indirect`
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
#### Scenario: Evidence bindings support multiple tool types
- **WHEN** Executor binds evidence to a claim
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
- **AND** the binding MAY include `source_id`
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
#### Scenario: Unsupported conclusions are not confirmed claims
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
- **AND** Executor SHALL NOT present it as a confirmed claim
#### Scenario: Runbook and skill guidance do not become incident facts
- **WHEN** Executor uses runbook, skill, or historical-case guidance
- **THEN** the guidance MAY influence `recommended_actions`
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
### Requirement: User-facing Chat answers SHALL remain readable Chinese
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
#### Scenario: V2 machine contract is not exposed as normal user answer
- **WHEN** Executor emits `executor_evidence_v2`
- **AND** Verifier returns `PASS`
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
- **AND** normal user output SHALL NOT be the raw Executor JSON object
#### Scenario: Machine contract remains available for trace inspection
- **WHEN** the Chat trace or verifier evaluation is inspected
- **THEN** the structured Executor contract MAY be shown for debugging or audit
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
@@ -0,0 +1,24 @@
## 1. Executor V2 Prompt
- [x] 1.1 Update `src/main/resources/prompts/chat-executor-prompt.md` so the final contract uses `answer_version="executor_evidence_v2"`.
- [x] 1.2 Remove `diagnosis_summary` and `user_facing_answer` from the required Executor output schema and validation rules.
- [x] 1.3 Keep claims, hypotheses, recommended actions, missing information, and evidence-binding rules.
- [x] 1.4 Keep Planner and tool-use behavior unchanged.
## 2. Runtime Rendering Safety
- [x] 2.1 Update `ChatService` PASS handling so V2 structured output is rendered into readable Chinese instead of raw JSON.
- [x] 2.2 Preserve V1 `user_facing_answer` extraction for compatibility.
- [x] 2.3 Ensure fallback behavior does not expose raw Executor JSON when structured rendering fails.
## 3. Parser Compatibility
- [x] 3.1 Keep `VerifierInputHook` parse-only behavior compatible with V2 output.
- [x] 3.2 Add or update tests proving V2 output without `user_facing_answer` parses into `executor_structured_output`.
## 4. Tests And Verification
- [x] 4.1 Add or update ChatService tests for PASS with `executor_evidence_v2`.
- [x] 4.2 Add or update tests proving final user answer does not contain raw JSON contract text.
- [x] 4.3 Run targeted tests for ChatService and VerifierInputHook.
- [x] 4.4 Validate this OpenSpec change.