1.3 KiB
1.3 KiB
Evidence: executor-evidence-output-contract
Repository Evidence
chat-executor-prompt.mdcurrently requires using real tool data but does not require a structured evidence-attribution output.chat-verifier-prompt.mdcurrently extracts facts fromexecutor_final_answerprose and compares them withtool_trace_summary.VerifierInputHookcurrently providesoriginal_query,executor_final_answer,tool_trace_summary, andretry_context.openspec/specs/chat-verifier-agent/spec.mdalready requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.openspec/specs/evidence-trace-hardening/spec.mdalready distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
Runtime Evidence From Recent Sessions
Recent MySQL inspection showed repeated LOW_CONFID verifier results with many no_evidence facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
Design Evidence
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.