3.6 KiB
3.6 KiB
Tasks
1. Evidence reference extraction
- Add
evidence_refsextraction inToolInvocationRecorderforquery_metricsalert arrays. - Add
evidence_refsextraction inToolInvocationRecorderforquery_logslog arrays. - Add
evidence_refsextraction inToolInvocationRecorderforlookup_knowledgeevidence blocks. - Add focused recorder tests for
$.alerts[i],$.logs[i], and$.evidence_blocks[i].
Acceptance:
- Persisted
retrieval_detailscontains minimalraw_pathandtext. - Extraction does not infer root cause or diagnosis.
2. Gatekeeper reference fidelity
- Extend
ExecutorGatekeeperServiceto read singularsource_invocation_id,raw_path, andevidence_excerpt. - Keep legacy
source_invocation_idscompatibility where needed, but requireraw_pathfor precise pass. - Validate invocation existence, session ownership, tool name,
raw_path, and excerpt similarity. - Add
severityand checked binding details to Gatekeeper output. - Add tests for valid reference, missing
raw_path, missingevidence_refs, unknownraw_path, mismatched excerpt, fabricated invocation ID, and tool mismatch.
Acceptance:
- Valid precise references pass.
- Missing precision downgrades to
LOW_CONFID. - Fabricated or mismatched references become
REJECT.
3. Verifier input hook and prompts
- Tighten
VerifierInputHookauto-backfill to only unique single invocation candidates. - Prevent auto-filled bindings without
raw_pathfrom passing Gatekeeper. - Add auto-backfill warnings into Gatekeeper/audit output.
- Update
chat-executor-prompt.mdwith evidence-reference and narrow-scope constraints. - Update
chat-verifier-prompt.mdso verified excerpts are the primary derivability evidence. - Add focused hook and prompt-sensitive tests where practical.
Acceptance:
- Hook no longer bulk-fills invocation IDs.
- Verifier can see
gatekeeper_result.severity. - Executor is instructed to output precise references and avoid over-expansion.
4. Effective verdict and audit persistence
- Ensure
severity=rejectprevents effectivePASSand maps toREJECTwhen appropriate. - Ensure
severity=low_confidprevents effectivePASSand maps toLOW_CONFID. - Ensure
gatekeeper_resultwithseverityis persisted underverifier_evaluation. - Add ChatService or integration tests for effective verdict guardrails.
Acceptance:
- Gatekeeper fail cannot become final PASS.
- Audit JSON contains the minimum Gatekeeper fields.
5. HikariCP mock quality
- Add positive HikariCP mock logs for
order-service. - Support HikariCP synonym matching.
- Remove
generic-serviceplaceholder evidence for no-hit cases. - Add tests for HikariCP positive and negative no-hit behavior.
Acceptance:
- Positive HikariCP query returns order-service logs.
- Negative HikariCP query returns
logs=[]andevidence_status=no_evidence.
6. Verification
- Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior.
- Run
openspec validate verifier-evidence-reference-fidelity --strict. - Run
openspec validate --specs. - Start the Java project using
mvn spring-boot:run. - Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim.
- Query MySQL audit data with
scripts/query_mysql.pyto confirm persistedgatekeeper_result.
Acceptance:
- Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.