# verifier-evidence-reference-fidelity ## Problem Recent end-to-end checks show that some narrow diagnosis questions still become `LOW_CONFID` even when the raw tool output and Executor evidence excerpts contain enough concrete evidence. The failure is caused by evidence being compressed or lost before Verifier reasoning, plus mock log no-hit behavior that can return `generic-service` placeholder logs. Current weak points: - Verifier still relies too heavily on `tool_trace_summary.output_summary`. - Executor evidence bindings identify tool invocations but do not precisely locate evidence inside the invocation output. - Gatekeeper validates invocation IDs and tool names, but does not yet verify `raw_path` and excerpt fidelity. - `query_logs` mock data cannot reliably produce a positive HikariCP connection-pool exhaustion path and can pollute no-hit results with placeholder logs. - Narrow-scope Executor answers can still over-expand into unrelated claims. ## Proposed Change Introduce a claim-local evidence reference protocol: ```text tool raw output -> ToolInvocationRecorder stores retrieval_details.evidence_refs -> Executor outputs claim + source_invocation_id + raw_path + evidence_excerpt -> Gatekeeper verifies that the reference is real -> Verifier judges whether verified evidence can derive the claim -> Composer only expresses Verifier-allowed material ``` The first implementation keeps the orchestration unchanged. Gatekeeper remains in the Verifier input hook path. Planner `scope_contract` is out of scope for this change. ## Scope - Add minimal `retrieval_details.evidence_refs` extraction for `query_metrics`, `query_logs`, and `lookup_knowledge`. - Tighten Executor evidence bindings to prefer singular `source_invocation_id`, stable `raw_path`, and `evidence_excerpt`. - Extend Gatekeeper to validate `source_invocation_id + raw_path + evidence_excerpt`. - Add Gatekeeper severity: `none`, `low_confid`, `reject`. - Make Verifier consume verified `evidence_excerpt` as the primary claim-local evidence. - Tighten `VerifierInputHook` auto-backfill: only unique invocation candidate, never `raw_path`, and no `PASS` without a precise reference. - Fix HikariCP mock log matching and no-hit behavior. - Update Executor and Verifier prompts for narrow-scope and derivability behavior. - Add focused tests and end-to-end checks for the minimum acceptance matrix. ## Non-goals - Do not change Planner output. - Do not implement Planner `scope_contract`. - Do not add new database tables. - Do not implement a general JSONPath engine. - Do not make `tool_trace_summary` the primary evidence source again. - Do not allow Executor `diagnosis_summary` or `user_facing_answer` to re-enter the V2 contract. ## Context Constraints - `tool_invocation.retrieval_details` is the preferred place for tool-specific structured details. - `DiagnosisSession.selfEvaluation.verifier_evaluation.gatekeeper_result` is the existing audit container and must be preserved. - `read_skill` / runbook guidance is not incident evidence. - Existing Composer routing must keep raw Executor JSON out of normal user answers. ## Risks - Existing tests or prompts may still assume plural `source_invocation_ids`. - Some old invocations will not have `evidence_refs`; those must downgrade to `LOW_CONFID`, not `PASS`. - Similarity checks must tolerate formatting changes without accepting unrelated text.