# Tasks ## 1. Evidence reference extraction - [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_metrics` alert arrays. - [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_logs` log arrays. - [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `lookup_knowledge` evidence blocks. - [x] Add focused recorder tests for `$.alerts[i]`, `$.logs[i]`, and `$.evidence_blocks[i]`. Acceptance: - Persisted `retrieval_details` contains minimal `raw_path` and `text`. - Extraction does not infer root cause or diagnosis. ## 2. Gatekeeper reference fidelity - [x] Extend `ExecutorGatekeeperService` to read singular `source_invocation_id`, `raw_path`, and `evidence_excerpt`. - [x] Keep legacy `source_invocation_ids` compatibility where needed, but require `raw_path` for precise pass. - [x] Validate invocation existence, session ownership, tool name, `raw_path`, and excerpt similarity. - [x] Add `severity` and checked binding details to Gatekeeper output. - [x] Add tests for valid reference, missing `raw_path`, missing `evidence_refs`, unknown `raw_path`, mismatched excerpt, fabricated invocation ID, and tool mismatch. Acceptance: - Valid precise references pass. - Missing precision downgrades to `LOW_CONFID`. - Fabricated or mismatched references become `REJECT`. ## 3. Verifier input hook and prompts - [x] Tighten `VerifierInputHook` auto-backfill to only unique single invocation candidates. - [x] Prevent auto-filled bindings without `raw_path` from passing Gatekeeper. - [x] Add auto-backfill warnings into Gatekeeper/audit output. - [x] Update `chat-executor-prompt.md` with evidence-reference and narrow-scope constraints. - [x] Update `chat-verifier-prompt.md` so verified excerpts are the primary derivability evidence. - [x] Add focused hook and prompt-sensitive tests where practical. Acceptance: - Hook no longer bulk-fills invocation IDs. - Verifier can see `gatekeeper_result.severity`. - Executor is instructed to output precise references and avoid over-expansion. ## 4. Effective verdict and audit persistence - [x] Ensure `severity=reject` prevents effective `PASS` and maps to `REJECT` when appropriate. - [x] Ensure `severity=low_confid` prevents effective `PASS` and maps to `LOW_CONFID`. - [x] Ensure `gatekeeper_result` with `severity` is persisted under `verifier_evaluation`. - [x] Add ChatService or integration tests for effective verdict guardrails. Acceptance: - Gatekeeper fail cannot become final PASS. - Audit JSON contains the minimum Gatekeeper fields. ## 5. HikariCP mock quality - [x] Add positive HikariCP mock logs for `order-service`. - [x] Support HikariCP synonym matching. - [x] Remove `generic-service` placeholder evidence for no-hit cases. - [x] Add tests for HikariCP positive and negative no-hit behavior. Acceptance: - Positive HikariCP query returns order-service logs. - Negative HikariCP query returns `logs=[]` and `evidence_status=no_evidence`. ## 6. Verification - [x] Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior. - [x] Run `openspec validate verifier-evidence-reference-fidelity --strict`. - [x] Run `openspec validate --specs`. - [x] Start the Java project using `mvn spring-boot:run`. - [x] Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim. - [x] Query MySQL audit data with `scripts/query_mysql.py` to confirm persisted `gatekeeper_result`. Acceptance: - Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.