3.0 KiB
Decisions: executor-evidence-output-contract
sm-flow Progress
Clarify
Entry summary: recent Chat diagnosis sessions are LOW_CONFID because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
Slug: executor-evidence-output-contract
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
Context
Relevant history:
executor-action-memory-relevance: Executor already has retrieval quality constraints and should avoid repeatedlookup_knowledge.chat-verifier-agent: Verifier should not see intermediate reasoning; it receives explicitoriginal_query,executor_final_answer,tool_trace_summary, andretry_context.evidence-trace-hardening: evidence-bearing tools persist stable traces and no-evidence semantics.modular-rag-pipeline:lookup_knowledgeexposes evidence blocks and context packs; L0 hints are not fact evidence.
Current code shape:
src/main/resources/prompts/chat-executor-prompt.mdis the Chat Executor prompt.src/main/resources/prompts/executor-prompt.mdis for the AiOps flow and is not the target of this Chat change.VerifierInputHookcurrently builds a payload withoriginal_query,executor_final_answer,tool_trace_summary, andretry_context.- Verifier prompt currently extracts facts from
executor_final_answer.
Grill
Question: Should Executor output only JSON or JSON plus readable answer?
Decision: use one JSON object containing both machine fields and user_facing_answer. This avoids losing a readable Chinese answer while giving Verifier structured claims.
Question: Should evidence binding use chunk_id?
Decision: no. Use generic binding fields because query_logs and query_metrics do not naturally expose RAG chunks.
Question: Should Verifier trust Executor-provided claims completely?
Decision: no. Verifier should verify structured claims first, then scan user_facing_answer for extra confirmed-sounding facts omitted from claims.
Question: What happens when Executor JSON is malformed?
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
Specify
OpenSpec artifacts:
proposal.md: why and scopedesign.md: contract, verifier behavior, risksspecs/chat-verifier-agent/spec.md: modified and added requirementstasks.md: implementation checklist
Audit
Cross-artifact alignment:
- Issue describes evidence attribution hallucination.
- Proposal scopes the fix to Executor output and Verifier consumption.
- Design preserves existing verifier isolation.
- Spec adds observable behavior without changing database schema.
- Tasks remain implementation-oriented and unchecked.
Interface impact:
- Prompt/output contract: L2 internal Agent contract change.
- Verifier payload: L2 internal structured input extension.
- Database schema: no change.
- External HTTP API: no intended change.