feat(agent): add executor evidence output contract

This commit is contained in:
zhuyongxin
2026-07-07 19:02:02 +08:00
parent 7f2e47ca38
commit 04eb50e2b4
19 changed files with 986 additions and 25 deletions
@@ -0,0 +1,16 @@
# Acceptance: executor-evidence-output-contract
## Draft Acceptance
- [x] Issue exists: `mvp/issues/executor-evidence-attribution-hallucination.md`.
- [x] OpenSpec change artifacts exist.
- [x] devflow tracking files exist.
- [x] OpenSpec validation passes.
- [x] Implementation updates Chat Executor prompt.
- [x] Implementation passes structured Executor output to Verifier.
- [x] Verifier prefers structured claims and still falls back safely.
- [x] Focused tests cover parsing, payload assembly, and unsupported confirmed claims.
## Notes
This project is currently in proposal/design stage. Runtime code is intentionally not changed yet.
@@ -0,0 +1,23 @@
# Brief: executor-evidence-output-contract
## Summary
Executor currently returns natural-language diagnosis answers that may mix confirmed evidence, runbook guidance, historical patterns, and unsupported inference. Verifier catches many unsupported facts, but only after extracting claims from prose.
This project defines a structured Executor evidence-attribution contract and updates the Verifier input/verification path to consume it.
## Goal
Make Chat Executor output machine-checkable so confirmed claims are explicitly bound to current-session evidence, while hypotheses and evidence gaps remain visibly separate.
## Scope
- Chat Executor prompt contract.
- Executor structured output parsing.
- Verifier payload extension.
- Chat Verifier prompt behavior.
- Focused tests/eval fixtures.
## Related OpenSpec
`openspec/changes/executor-evidence-output-contract/`
@@ -0,0 +1,71 @@
# Decisions: executor-evidence-output-contract
## sm-flow Progress
### Clarify
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
Slug: `executor-evidence-output-contract`
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
### Context
Relevant history:
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
Current code shape:
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Verifier prompt currently extracts facts from `executor_final_answer`.
### Grill
Question: Should Executor output only JSON or JSON plus readable answer?
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
Question: Should evidence binding use `chunk_id`?
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
Question: Should Verifier trust Executor-provided claims completely?
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
Question: What happens when Executor JSON is malformed?
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope
- `design.md`: contract, verifier behavior, risks
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
- `tasks.md`: implementation checklist
### Audit
Cross-artifact alignment:
- Issue describes evidence attribution hallucination.
- Proposal scopes the fix to Executor output and Verifier consumption.
- Design preserves existing verifier isolation.
- Spec adds observable behavior without changing database schema.
- Tasks remain implementation-oriented and unchecked.
Interface impact:
- Prompt/output contract: L2 internal Agent contract change.
- Verifier payload: L2 internal structured input extension.
- Database schema: no change.
- External HTTP API: no intended change.
@@ -0,0 +1,17 @@
# Evidence: executor-evidence-output-contract
## Repository Evidence
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
## Runtime Evidence From Recent Sessions
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
## Design Evidence
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.