Files

80 lines
3.6 KiB
Markdown

# Tasks
## 1. Evidence reference extraction
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_metrics` alert arrays.
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_logs` log arrays.
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `lookup_knowledge` evidence blocks.
- [x] Add focused recorder tests for `$.alerts[i]`, `$.logs[i]`, and `$.evidence_blocks[i]`.
Acceptance:
- Persisted `retrieval_details` contains minimal `raw_path` and `text`.
- Extraction does not infer root cause or diagnosis.
## 2. Gatekeeper reference fidelity
- [x] Extend `ExecutorGatekeeperService` to read singular `source_invocation_id`, `raw_path`, and `evidence_excerpt`.
- [x] Keep legacy `source_invocation_ids` compatibility where needed, but require `raw_path` for precise pass.
- [x] Validate invocation existence, session ownership, tool name, `raw_path`, and excerpt similarity.
- [x] Add `severity` and checked binding details to Gatekeeper output.
- [x] Add tests for valid reference, missing `raw_path`, missing `evidence_refs`, unknown `raw_path`, mismatched excerpt, fabricated invocation ID, and tool mismatch.
Acceptance:
- Valid precise references pass.
- Missing precision downgrades to `LOW_CONFID`.
- Fabricated or mismatched references become `REJECT`.
## 3. Verifier input hook and prompts
- [x] Tighten `VerifierInputHook` auto-backfill to only unique single invocation candidates.
- [x] Prevent auto-filled bindings without `raw_path` from passing Gatekeeper.
- [x] Add auto-backfill warnings into Gatekeeper/audit output.
- [x] Update `chat-executor-prompt.md` with evidence-reference and narrow-scope constraints.
- [x] Update `chat-verifier-prompt.md` so verified excerpts are the primary derivability evidence.
- [x] Add focused hook and prompt-sensitive tests where practical.
Acceptance:
- Hook no longer bulk-fills invocation IDs.
- Verifier can see `gatekeeper_result.severity`.
- Executor is instructed to output precise references and avoid over-expansion.
## 4. Effective verdict and audit persistence
- [x] Ensure `severity=reject` prevents effective `PASS` and maps to `REJECT` when appropriate.
- [x] Ensure `severity=low_confid` prevents effective `PASS` and maps to `LOW_CONFID`.
- [x] Ensure `gatekeeper_result` with `severity` is persisted under `verifier_evaluation`.
- [x] Add ChatService or integration tests for effective verdict guardrails.
Acceptance:
- Gatekeeper fail cannot become final PASS.
- Audit JSON contains the minimum Gatekeeper fields.
## 5. HikariCP mock quality
- [x] Add positive HikariCP mock logs for `order-service`.
- [x] Support HikariCP synonym matching.
- [x] Remove `generic-service` placeholder evidence for no-hit cases.
- [x] Add tests for HikariCP positive and negative no-hit behavior.
Acceptance:
- Positive HikariCP query returns order-service logs.
- Negative HikariCP query returns `logs=[]` and `evidence_status=no_evidence`.
## 6. Verification
- [x] Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior.
- [x] Run `openspec validate verifier-evidence-reference-fidelity --strict`.
- [x] Run `openspec validate --specs`.
- [x] Start the Java project using `mvn spring-boot:run`.
- [x] Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim.
- [x] Query MySQL audit data with `scripts/query_mysql.py` to confirm persisted `gatekeeper_result`.
Acceptance:
- Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.