Files
SuperBizAgent-java/devflow/projects/2026-07-07-executor-evidence-output-contract/decisions.md
T

3.0 KiB

Decisions: executor-evidence-output-contract

sm-flow Progress

Clarify

Entry summary: recent Chat diagnosis sessions are LOW_CONFID because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.

Slug: executor-evidence-output-contract

Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.

Context

Relevant history:

  • executor-action-memory-relevance: Executor already has retrieval quality constraints and should avoid repeated lookup_knowledge.
  • chat-verifier-agent: Verifier should not see intermediate reasoning; it receives explicit original_query, executor_final_answer, tool_trace_summary, and retry_context.
  • evidence-trace-hardening: evidence-bearing tools persist stable traces and no-evidence semantics.
  • modular-rag-pipeline: lookup_knowledge exposes evidence blocks and context packs; L0 hints are not fact evidence.

Current code shape:

  • src/main/resources/prompts/chat-executor-prompt.md is the Chat Executor prompt.
  • src/main/resources/prompts/executor-prompt.md is for the AiOps flow and is not the target of this Chat change.
  • VerifierInputHook currently builds a payload with original_query, executor_final_answer, tool_trace_summary, and retry_context.
  • Verifier prompt currently extracts facts from executor_final_answer.

Grill

Question: Should Executor output only JSON or JSON plus readable answer?

Decision: use one JSON object containing both machine fields and user_facing_answer. This avoids losing a readable Chinese answer while giving Verifier structured claims.

Question: Should evidence binding use chunk_id?

Decision: no. Use generic binding fields because query_logs and query_metrics do not naturally expose RAG chunks.

Question: Should Verifier trust Executor-provided claims completely?

Decision: no. Verifier should verify structured claims first, then scan user_facing_answer for extra confirmed-sounding facts omitted from claims.

Question: What happens when Executor JSON is malformed?

Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.

Specify

OpenSpec artifacts:

  • proposal.md: why and scope
  • design.md: contract, verifier behavior, risks
  • specs/chat-verifier-agent/spec.md: modified and added requirements
  • tasks.md: implementation checklist

Audit

Cross-artifact alignment:

  • Issue describes evidence attribution hallucination.
  • Proposal scopes the fix to Executor output and Verifier consumption.
  • Design preserves existing verifier isolation.
  • Spec adds observable behavior without changing database schema.
  • Tasks remain implementation-oriented and unchecked.

Interface impact:

  • Prompt/output contract: L2 internal Agent contract change.
  • Verifier payload: L2 internal structured input extension.
  • Database schema: no change.
  • External HTTP API: no intended change.