feat(agent): add executor evidence v2 contract
This commit is contained in:
@@ -0,0 +1 @@
|
||||
archive-ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-07
|
||||
@@ -0,0 +1,112 @@
|
||||
# Decisions: executor-v2-output-contract
|
||||
|
||||
## sm-flow Progress
|
||||
|
||||
### Clarify
|
||||
|
||||
Entry summary: implement stage one of `Executor Structured Output V2`: narrow Chat Executor output to structured diagnostic material and prevent raw JSON from leaking to users before later Gatekeeper/Verifier/Composer phases.
|
||||
|
||||
Slug: `executor-v2-output-contract`
|
||||
|
||||
Scale: complex overall program, but this change is the first vertical stage. It is still treated with full sm-flow gates because it changes an internal Agent output contract and must be archived before the next phase.
|
||||
|
||||
### Context
|
||||
|
||||
Relevant devflow history:
|
||||
|
||||
- `chat-verifier-agent`: Verifier is isolated from Planner/Executor intermediate reasoning and consumes explicit verification inputs.
|
||||
- `evidence-trace-hardening`: evidence-bearing tool traces are persisted and summarized through `ToolTraceSummaryService`.
|
||||
- `executor-evidence-output-contract`: V1 introduced `executor_evidence_v1` with `diagnosis_summary`, structured claims, and `user_facing_answer`.
|
||||
|
||||
Conflict with historical decision:
|
||||
|
||||
- Previous `executor-evidence-output-contract` deliberately kept `user_facing_answer` in Executor output.
|
||||
- New V2 design deliberately removes it so Composer becomes the only final-expression layer in a later phase.
|
||||
- For this stage, code must bridge the gap by rendering a safe temporary Chinese answer from V2 structured fields; it must not restore Executor `user_facing_answer`.
|
||||
|
||||
Current code shape:
|
||||
|
||||
- `chat-executor-prompt.md` defines the V1 Executor output contract.
|
||||
- `VerifierInputHook` parses Executor JSON and sets `executor_structured_output`.
|
||||
- `ChatService.extractUserFacingAnswer(...)` currently reads `user_facing_answer` on PASS.
|
||||
- If no replacement is added, PASS may expose raw Executor JSON after V2 removes `user_facing_answer`.
|
||||
|
||||
### Grill
|
||||
|
||||
Question pool:
|
||||
|
||||
| Question | Mode | Resolution |
|
||||
|---|---|---|
|
||||
| Does stage one include Gatekeeper? | evidence-driven | No. The issue splits Gatekeeper into stage two. |
|
||||
| Does stage one change Planner? | evidence-driven | No. Planner is explicitly out of scope. |
|
||||
| Can `user_facing_answer` remain temporarily in Executor? | evidence-driven | No. The V2 design requires removing it in stage one. |
|
||||
| How do users get readable output before Composer exists? | evidence-driven | ChatService must use a temporary structured renderer for V2 PASS output. |
|
||||
| Is the internal Agent contract breaking? | evidence-driven | Yes. Removing fields from Executor JSON is internal L4, but external Chat answer behavior remains readable. |
|
||||
|
||||
No user-interview questions are open for stage one because the user already approved the staged design and asked for automatic phased implementation; decision questions should pause only if implementation reveals a new product trade-off.
|
||||
|
||||
### Specify
|
||||
|
||||
OpenSpec artifacts:
|
||||
|
||||
- `proposal.md`: scope and compatibility boundary for stage one.
|
||||
- `design.md`: V2 Executor contract and temporary rendering strategy.
|
||||
- `specs/chat-verifier-agent/spec.md`: delta requirements for the Executor contract.
|
||||
- `tasks.md`: executable implementation and verification checklist.
|
||||
|
||||
### Audit
|
||||
|
||||
Architecture risk summary:
|
||||
|
||||
- The first-stage change deliberately breaks the internal Executor JSON contract by removing `diagnosis_summary` and `user_facing_answer`.
|
||||
- External Chat answers must remain readable Chinese, so `ChatService` needs a temporary V2 renderer before Composer exists.
|
||||
- `VerifierInputHook` should remain parse-only; full schema/evidence validation is deferred to the Gatekeeper stage.
|
||||
- No database schema or evidence tool signature changes are required.
|
||||
|
||||
Cross-artifact alignment:
|
||||
|
||||
| Source | Target | Status |
|
||||
|---|---|---|
|
||||
| issue background / stage one | proposal | aligned |
|
||||
| proposal scope / non-goals | design | aligned |
|
||||
| design contract and rendering bridge | specs | aligned |
|
||||
| specs observable behavior | tasks | aligned |
|
||||
|
||||
Interface impact:
|
||||
|
||||
- Internal Agent output contract: L4, because `diagnosis_summary` and `user_facing_answer` are removed.
|
||||
- Verifier payload: L2, because `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
|
||||
- External Chat/API answer: intended compatible behavior; users must still receive readable Chinese rather than raw JSON.
|
||||
|
||||
### Commit
|
||||
|
||||
Commit gate result: passed.
|
||||
|
||||
- `proposal.md` exists and explains why this phase is needed.
|
||||
- `design.md` records the V2 contract, temporary rendering strategy, non-goals, and interface impact.
|
||||
- `specs/chat-verifier-agent/spec.md` expresses observable behavior for Executor V2 and user-facing rendering safety.
|
||||
- `tasks.md` contains executable implementation and verification tasks.
|
||||
- `cmd /c openspec validate executor-v2-output-contract` passed.
|
||||
- No unresolved user-interview questions remain for this stage.
|
||||
|
||||
### Apply
|
||||
|
||||
Implementation summary:
|
||||
|
||||
- Updated `chat-executor-prompt.md` to require `answer_version="executor_evidence_v2"`.
|
||||
- Removed `diagnosis_summary` and `user_facing_answer` from the Executor final output schema and output validation rules.
|
||||
- Added a temporary `ChatService` structured renderer for PASS + `executor_evidence_v2` so normal users receive readable Chinese instead of raw JSON.
|
||||
- Preserved V1 `user_facing_answer` extraction for compatibility.
|
||||
- Kept `VerifierInputHook` parse-only behavior compatible with V2 output.
|
||||
- Adjusted `chat-verifier-prompt.md` wording so `user_facing_answer` is treated as a compatibility field, not a V2 required field.
|
||||
|
||||
Verification:
|
||||
|
||||
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
|
||||
- `cmd /c openspec validate executor-v2-output-contract` passed.
|
||||
|
||||
Known limitations:
|
||||
|
||||
- Gatekeeper is not implemented in this phase.
|
||||
- Verifier still outputs `facts_checked`; `claim_checks` belongs to a later phase.
|
||||
- The V2 renderer is temporary and should be replaced by Composer in a later phase.
|
||||
@@ -0,0 +1,112 @@
|
||||
## Context
|
||||
|
||||
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
|
||||
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
|
||||
|
||||
This stage implements only the first slice of the V2 migration:
|
||||
|
||||
```text
|
||||
Executor V2 output contract
|
||||
-> existing VerifierInputHook parsing
|
||||
-> existing Verifier
|
||||
-> temporary ChatService structured renderer
|
||||
```
|
||||
|
||||
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
Goals:
|
||||
|
||||
- Make Chat Executor emit `executor_evidence_v2`.
|
||||
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
|
||||
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
|
||||
- Preserve evidence-binding requirements for confirmed claims.
|
||||
- Prevent PASS routing from returning raw JSON to normal Chat users.
|
||||
|
||||
Non-goals:
|
||||
|
||||
- No Gatekeeper implementation.
|
||||
- No Verifier prompt rewrite to `claim_checks`.
|
||||
- No Composer agent.
|
||||
- No database schema changes.
|
||||
- No Planner changes.
|
||||
- No evidence tool signature changes.
|
||||
- No retry behavior changes.
|
||||
|
||||
## Executor V2 Contract
|
||||
|
||||
Executor final output SHALL be one JSON object:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "payment-service 出现请求超时日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_ids": [394],
|
||||
"evidence_excerpt": "request timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
Removed fields:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
|
||||
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
|
||||
|
||||
## Temporary Rendering Strategy
|
||||
|
||||
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
|
||||
|
||||
When Verifier returns `PASS`:
|
||||
|
||||
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
|
||||
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
|
||||
- `claims[].claim_text`
|
||||
- `hypotheses[].hypothesis_text`
|
||||
- `missing_info[]`
|
||||
- `recommended_actions[].action_text` and `reason`
|
||||
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
|
||||
|
||||
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
|
||||
|
||||
## Parser Boundary
|
||||
|
||||
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
|
||||
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
|
||||
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
|
||||
|
||||
## Risks / Mitigations
|
||||
|
||||
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
|
||||
- Mitigation: add temporary V2 renderer in `ChatService`.
|
||||
- Risk: current Verifier prompt still mentions `user_facing_answer`.
|
||||
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
|
||||
- Risk: tests assume V1 fields.
|
||||
- Mitigation: update/add focused tests for V2 output without final-expression fields.
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
## Why
|
||||
|
||||
The current Chat Executor evidence contract still mixes diagnostic material with final user-facing prose through `diagnosis_summary` and `user_facing_answer`. This keeps the Executor in a "diagnose and narrate" mode, so unsupported details can be smuggled into the final answer before later Gatekeeper, Verifier V2, and Composer phases exist.
|
||||
|
||||
This first phase narrows Executor to structured diagnostic material only and adds a temporary safe rendering path so normal Chat responses do not expose raw Executor JSON while later phases are implemented.
|
||||
|
||||
## What Changes
|
||||
|
||||
- **BREAKING internal Agent contract**: Chat Executor output changes from `executor_evidence_v1` to `executor_evidence_v2`.
|
||||
- Remove `diagnosis_summary` and `user_facing_answer` from the Chat Executor final JSON contract.
|
||||
- Keep the existing structured arrays: `claims`, `hypotheses`, `recommended_actions`, and `missing_info`.
|
||||
- Preserve `claims[].evidence_bindings` and current-session evidence attribution rules.
|
||||
- Adjust runtime final-answer handling so a PASS result with V2 Executor output is rendered into readable Chinese from structured fields instead of returning raw JSON.
|
||||
- Keep Planner, Verifier, Gatekeeper, retry behavior, database schema, and tool signatures unchanged in this phase.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `chat-verifier-agent`: The Executor evidence-attribution contract is tightened so V2 structured output no longer contains final-expression fields. Verifier still receives `executor_final_answer` as raw text and `executor_structured_output` when parseable.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected prompt: `src/main/resources/prompts/chat-executor-prompt.md`.
|
||||
- Affected runtime: `ChatService` PASS answer extraction/rendering for Executor V2.
|
||||
- Affected parser boundary: `VerifierInputHook` should continue parsing JSON but must not treat schema validation as its own responsibility in this phase.
|
||||
- Affected tests: ChatService sequential flow tests and VerifierInputHook parsing tests for V2 output without `user_facing_answer`.
|
||||
- Interface impact: L4 for internal Agent output contract because fields are removed from Executor JSON; external HTTP/chat answer behavior must remain readable Chinese and must not expose raw JSON.
|
||||
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Executor SHALL output an evidence-attribution contract
|
||||
The Chat Executor SHALL produce a machine-checkable final output that separates confirmed claims from hypotheses, recommendations, and missing information.
|
||||
|
||||
#### Scenario: Executor V2 final output contains only structured diagnostic fields
|
||||
- **WHEN** Executor completes a Chat diagnosis step under the V2 contract
|
||||
- **THEN** its final output SHALL contain `answer_version`, `claims`, `hypotheses`, `recommended_actions`, and `missing_info`
|
||||
- **AND** `answer_version` SHALL equal `executor_evidence_v2`
|
||||
- **AND** the output SHOULD be parseable as one JSON object without Markdown fences
|
||||
- **AND** the output SHALL NOT contain `diagnosis_summary`
|
||||
- **AND** the output SHALL NOT contain `user_facing_answer`
|
||||
|
||||
#### Scenario: Confirmed claims carry evidence bindings
|
||||
- **WHEN** Executor emits an item under `claims`
|
||||
- **THEN** the item SHALL include `claim_id`, `claim_type`, `claim_text`, `support_level`, and `evidence_bindings`
|
||||
- **AND** `support_level` SHALL be one of `direct` or `indirect`
|
||||
- **AND** `evidence_bindings` SHALL contain at least one evidence binding
|
||||
|
||||
#### Scenario: Evidence bindings support multiple tool types
|
||||
- **WHEN** Executor binds evidence to a claim
|
||||
- **THEN** each binding SHALL include `source_type`, `tool_name`, `source_invocation_ids`, and `evidence_excerpt`
|
||||
- **AND** the binding MAY include `source_id`
|
||||
- **AND** the binding SHALL be able to reference `lookup_knowledge`, `query_logs`, `query_metrics`, or other evidence-bearing tool traces
|
||||
- **AND** the binding SHALL NOT rely only on a RAG-specific `chunk_id`
|
||||
|
||||
#### Scenario: Unsupported conclusions are not confirmed claims
|
||||
- **WHEN** a possible root cause, detail, or remediation lacks current-session tool evidence
|
||||
- **THEN** Executor SHALL place it under `hypotheses`, `recommended_actions`, or `missing_info`
|
||||
- **AND** Executor SHALL NOT present it as a confirmed claim
|
||||
|
||||
#### Scenario: Runbook and skill guidance do not become incident facts
|
||||
- **WHEN** Executor uses runbook, skill, or historical-case guidance
|
||||
- **THEN** the guidance MAY influence `recommended_actions`
|
||||
- **AND** the guidance SHALL NOT be emitted as a current incident fact unless current-session tool evidence supports it
|
||||
|
||||
### Requirement: User-facing Chat answers SHALL remain readable Chinese
|
||||
The system SHALL preserve a readable Chinese answer for normal Chat users even when Executor emits a machine-checkable contract.
|
||||
|
||||
#### Scenario: V2 machine contract is not exposed as normal user answer
|
||||
- **WHEN** Executor emits `executor_evidence_v2`
|
||||
- **AND** Verifier returns `PASS`
|
||||
- **THEN** normal user output SHALL be rendered as readable Chinese from the structured contract or a safe fallback template
|
||||
- **AND** normal user output SHALL NOT be the raw Executor JSON object
|
||||
|
||||
#### Scenario: Machine contract remains available for trace inspection
|
||||
- **WHEN** the Chat trace or verifier evaluation is inspected
|
||||
- **THEN** the structured Executor contract MAY be shown for debugging or audit
|
||||
- **AND** normal user output SHALL use the existing verifier-routed display path rather than exposing raw JSON by default
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
## 1. Executor V2 Prompt
|
||||
|
||||
- [x] 1.1 Update `src/main/resources/prompts/chat-executor-prompt.md` so the final contract uses `answer_version="executor_evidence_v2"`.
|
||||
- [x] 1.2 Remove `diagnosis_summary` and `user_facing_answer` from the required Executor output schema and validation rules.
|
||||
- [x] 1.3 Keep claims, hypotheses, recommended actions, missing information, and evidence-binding rules.
|
||||
- [x] 1.4 Keep Planner and tool-use behavior unchanged.
|
||||
|
||||
## 2. Runtime Rendering Safety
|
||||
|
||||
- [x] 2.1 Update `ChatService` PASS handling so V2 structured output is rendered into readable Chinese instead of raw JSON.
|
||||
- [x] 2.2 Preserve V1 `user_facing_answer` extraction for compatibility.
|
||||
- [x] 2.3 Ensure fallback behavior does not expose raw Executor JSON when structured rendering fails.
|
||||
|
||||
## 3. Parser Compatibility
|
||||
|
||||
- [x] 3.1 Keep `VerifierInputHook` parse-only behavior compatible with V2 output.
|
||||
- [x] 3.2 Add or update tests proving V2 output without `user_facing_answer` parses into `executor_structured_output`.
|
||||
|
||||
## 4. Tests And Verification
|
||||
|
||||
- [x] 4.1 Add or update ChatService tests for PASS with `executor_evidence_v2`.
|
||||
- [x] 4.2 Add or update tests proving final user answer does not contain raw JSON contract text.
|
||||
- [x] 4.3 Run targeted tests for ChatService and VerifierInputHook.
|
||||
- [x] 4.4 Validate this OpenSpec change.
|
||||
Reference in New Issue
Block a user