113 lines
3.9 KiB
Markdown
113 lines
3.9 KiB
Markdown
## Context
|
|
|
|
The prior `executor_evidence_v1` contract made Executor responsible for both evidence attribution and final answer wording:
|
|
|
|
- `diagnosis_summary`
|
|
- `user_facing_answer`
|
|
|
|
That shape helped the first Verifier integration remain readable, but it also preserved the original problem: Executor can write unsupported or over-confident natural-language conclusions before the quality gate is complete.
|
|
|
|
This stage implements only the first slice of the V2 migration:
|
|
|
|
```text
|
|
Executor V2 output contract
|
|
-> existing VerifierInputHook parsing
|
|
-> existing Verifier
|
|
-> temporary ChatService structured renderer
|
|
```
|
|
|
|
Gatekeeper, Verifier V2 `claim_checks`, and Composer are later phases.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
Goals:
|
|
|
|
- Make Chat Executor emit `executor_evidence_v2`.
|
|
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
|
|
- Keep confirmed claims, hypotheses, recommended actions, and missing information as structured fields.
|
|
- Preserve evidence-binding requirements for confirmed claims.
|
|
- Prevent PASS routing from returning raw JSON to normal Chat users.
|
|
|
|
Non-goals:
|
|
|
|
- No Gatekeeper implementation.
|
|
- No Verifier prompt rewrite to `claim_checks`.
|
|
- No Composer agent.
|
|
- No database schema changes.
|
|
- No Planner changes.
|
|
- No evidence tool signature changes.
|
|
- No retry behavior changes.
|
|
|
|
## Executor V2 Contract
|
|
|
|
Executor final output SHALL be one JSON object:
|
|
|
|
```json
|
|
{
|
|
"answer_version": "executor_evidence_v2",
|
|
"claims": [
|
|
{
|
|
"claim_id": "claim-1",
|
|
"claim_type": "symptom",
|
|
"claim_text": "payment-service 出现请求超时日志。",
|
|
"support_level": "direct",
|
|
"evidence_bindings": [
|
|
{
|
|
"source_type": "tool_trace",
|
|
"source_id": "trace-1",
|
|
"tool_name": "query_logs",
|
|
"source_invocation_ids": [394],
|
|
"evidence_excerpt": "request timeout"
|
|
}
|
|
]
|
|
}
|
|
],
|
|
"hypotheses": [],
|
|
"recommended_actions": [],
|
|
"missing_info": []
|
|
}
|
|
```
|
|
|
|
Removed fields:
|
|
|
|
- `diagnosis_summary`
|
|
- `user_facing_answer`
|
|
|
|
`hypotheses`, `recommended_actions`, and `missing_info` SHOULD be present as arrays. They may be empty.
|
|
|
|
## Temporary Rendering Strategy
|
|
|
|
Before Composer exists, `ChatService` needs a safe PASS fallback for V2 output.
|
|
|
|
When Verifier returns `PASS`:
|
|
|
|
1. If Executor output has `user_facing_answer`, keep the existing V1 behavior.
|
|
2. Else, if Executor output is `executor_evidence_v2`, render a readable Chinese answer from:
|
|
- `claims[].claim_text`
|
|
- `hypotheses[].hypothesis_text`
|
|
- `missing_info[]`
|
|
- `recommended_actions[].action_text` and `reason`
|
|
3. If structured rendering fails, fall back to the existing low-confidence/degraded style rather than returning raw JSON.
|
|
|
|
The temporary renderer is not a Composer replacement. It is only a safety bridge until the Composer phase.
|
|
|
|
## Parser Boundary
|
|
|
|
`VerifierInputHook` may continue parsing raw Executor output into `executor_structured_output` when it is a JSON object. In this phase, it should not enforce the full V2 schema. Schema and evidence-reference validation belong to the later Gatekeeper phase.
|
|
|
|
## Interface Impact
|
|
|
|
- Internal Agent output contract: L4, because two fields are removed from Executor JSON.
|
|
- Verifier payload: L2, because existing `executor_final_answer` remains raw text and `executor_structured_output` remains optional.
|
|
- External Chat/API answer: intended compatible behavior; users still receive readable Chinese, not raw JSON.
|
|
|
|
## Risks / Mitigations
|
|
|
|
- Risk: existing PASS path exposes raw JSON because `user_facing_answer` is gone.
|
|
- Mitigation: add temporary V2 renderer in `ChatService`.
|
|
- Risk: current Verifier prompt still mentions `user_facing_answer`.
|
|
- Mitigation: stage one keeps Verifier behavior compatible; it should verify `claims` when structured output is valid and simply find no extra `user_facing_answer`.
|
|
- Risk: tests assume V1 fields.
|
|
- Mitigation: update/add focused tests for V2 output without final-expression fields.
|
|
|