# Decisions: executor-evidence-output-contract ## sm-flow Progress ### Clarify Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls. Slug: `executor-evidence-output-contract` Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change. ### Context Relevant history: - `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`. - `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`. - `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics. - `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence. Current code shape: - `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt. - `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change. - `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`. - Verifier prompt currently extracts facts from `executor_final_answer`. ### Grill Question: Should Executor output only JSON or JSON plus readable answer? Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims. Question: Should evidence binding use `chunk_id`? Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks. Question: Should Verifier trust Executor-provided claims completely? Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`. Question: What happens when Executor JSON is malformed? Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification. ### Specify OpenSpec artifacts: - `proposal.md`: why and scope - `design.md`: contract, verifier behavior, risks - `specs/chat-verifier-agent/spec.md`: modified and added requirements - `tasks.md`: implementation checklist ### Audit Cross-artifact alignment: - Issue describes evidence attribution hallucination. - Proposal scopes the fix to Executor output and Verifier consumption. - Design preserves existing verifier isolation. - Spec adds observable behavior without changing database schema. - Tasks remain implementation-oriented and unchecked. Interface impact: - Prompt/output contract: L2 internal Agent contract change. - Verifier payload: L2 internal structured input extension. - Database schema: no change. - External HTTP API: no intended change.