18 lines
1.3 KiB
Markdown
18 lines
1.3 KiB
Markdown
# Evidence: executor-evidence-output-contract
|
|
|
|
## Repository Evidence
|
|
|
|
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
|
|
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
|
|
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
|
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
|
|
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
|
|
|
|
## Runtime Evidence From Recent Sessions
|
|
|
|
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
|
|
|
|
## Design Evidence
|
|
|
|
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.
|