80 lines
3.6 KiB
Markdown
80 lines
3.6 KiB
Markdown
# Tasks
|
|
|
|
## 1. Evidence reference extraction
|
|
|
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_metrics` alert arrays.
|
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_logs` log arrays.
|
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `lookup_knowledge` evidence blocks.
|
|
- [x] Add focused recorder tests for `$.alerts[i]`, `$.logs[i]`, and `$.evidence_blocks[i]`.
|
|
|
|
Acceptance:
|
|
|
|
- Persisted `retrieval_details` contains minimal `raw_path` and `text`.
|
|
- Extraction does not infer root cause or diagnosis.
|
|
|
|
## 2. Gatekeeper reference fidelity
|
|
|
|
- [x] Extend `ExecutorGatekeeperService` to read singular `source_invocation_id`, `raw_path`, and `evidence_excerpt`.
|
|
- [x] Keep legacy `source_invocation_ids` compatibility where needed, but require `raw_path` for precise pass.
|
|
- [x] Validate invocation existence, session ownership, tool name, `raw_path`, and excerpt similarity.
|
|
- [x] Add `severity` and checked binding details to Gatekeeper output.
|
|
- [x] Add tests for valid reference, missing `raw_path`, missing `evidence_refs`, unknown `raw_path`, mismatched excerpt, fabricated invocation ID, and tool mismatch.
|
|
|
|
Acceptance:
|
|
|
|
- Valid precise references pass.
|
|
- Missing precision downgrades to `LOW_CONFID`.
|
|
- Fabricated or mismatched references become `REJECT`.
|
|
|
|
## 3. Verifier input hook and prompts
|
|
|
|
- [x] Tighten `VerifierInputHook` auto-backfill to only unique single invocation candidates.
|
|
- [x] Prevent auto-filled bindings without `raw_path` from passing Gatekeeper.
|
|
- [x] Add auto-backfill warnings into Gatekeeper/audit output.
|
|
- [x] Update `chat-executor-prompt.md` with evidence-reference and narrow-scope constraints.
|
|
- [x] Update `chat-verifier-prompt.md` so verified excerpts are the primary derivability evidence.
|
|
- [x] Add focused hook and prompt-sensitive tests where practical.
|
|
|
|
Acceptance:
|
|
|
|
- Hook no longer bulk-fills invocation IDs.
|
|
- Verifier can see `gatekeeper_result.severity`.
|
|
- Executor is instructed to output precise references and avoid over-expansion.
|
|
|
|
## 4. Effective verdict and audit persistence
|
|
|
|
- [x] Ensure `severity=reject` prevents effective `PASS` and maps to `REJECT` when appropriate.
|
|
- [x] Ensure `severity=low_confid` prevents effective `PASS` and maps to `LOW_CONFID`.
|
|
- [x] Ensure `gatekeeper_result` with `severity` is persisted under `verifier_evaluation`.
|
|
- [x] Add ChatService or integration tests for effective verdict guardrails.
|
|
|
|
Acceptance:
|
|
|
|
- Gatekeeper fail cannot become final PASS.
|
|
- Audit JSON contains the minimum Gatekeeper fields.
|
|
|
|
## 5. HikariCP mock quality
|
|
|
|
- [x] Add positive HikariCP mock logs for `order-service`.
|
|
- [x] Support HikariCP synonym matching.
|
|
- [x] Remove `generic-service` placeholder evidence for no-hit cases.
|
|
- [x] Add tests for HikariCP positive and negative no-hit behavior.
|
|
|
|
Acceptance:
|
|
|
|
- Positive HikariCP query returns order-service logs.
|
|
- Negative HikariCP query returns `logs=[]` and `evidence_status=no_evidence`.
|
|
|
|
## 6. Verification
|
|
|
|
- [x] Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior.
|
|
- [x] Run `openspec validate verifier-evidence-reference-fidelity --strict`.
|
|
- [x] Run `openspec validate --specs`.
|
|
- [x] Start the Java project using `mvn spring-boot:run`.
|
|
- [x] Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim.
|
|
- [x] Query MySQL audit data with `scripts/query_mysql.py` to confirm persisted `gatekeeper_result`.
|
|
|
|
Acceptance:
|
|
|
|
- Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.
|