# Design ## Data Flow ```text Evidence tool result -> ToolInvocationRecorder persists tool_invocation.retrieval_details.evidence_refs -> Executor emits claim-local evidence_bindings -> VerifierInputHook preserves structured payload and runs Gatekeeper -> ExecutorGatekeeperService validates reference authenticity -> Verifier checks derivability from verified excerpts -> ChatService / Composer enforces effective verdict and safe final answer ``` ## Evidence Reference Contract `tool_invocation.retrieval_details.evidence_refs` is an array of minimal evidence references: ```json { "evidence_refs": [ { "raw_path": "$.alerts[0]", "text": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m" } ] } ``` Supported `raw_path` formats in this change: - `$.alerts[i]` for `query_metrics` - `$.logs[i]` for `query_logs` - `$.evidence_blocks[i]` for `lookup_knowledge` Unsupported in this change: - Deep JSONPath such as `$.alerts[1].description` - Nested aliases such as `$.retrieval_details.evidence_blocks[0]` - Filter expressions - Tool-specific metadata beyond `raw_path` and `text` ## Executor Binding Contract Executor V2 claim bindings should use: ```json { "tool_name": "query_metrics", "source_invocation_id": 12345, "raw_path": "$.alerts[0]", "evidence_excerpt": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m" } ``` Compatibility: - Legacy `source_invocation_ids` may still be read for transition. - New precise validation requires singular `source_invocation_id` plus `raw_path`. - Missing `raw_path` is `LOW_CONFID`, not `PASS`. ## Gatekeeper Severity Gatekeeper output includes: ```json { "status": "fail", "severity": "reject", "checked_bindings": [], "failed_rules": [], "warnings": [], "errors": [] } ``` Severity mapping: - `status=pass`, `severity=none`: all checked bindings are authentic. - `status=fail`, `severity=low_confid`: evidence is missing or incomplete but not fabricated. - `status=fail`, `severity=reject`: Executor cites fabricated or mismatched evidence. `reject` cases: - Invocation ID does not exist. - Invocation belongs to another session. - Tool name mismatches persisted invocation. - `raw_path` is not present in `retrieval_details.evidence_refs`. - `evidence_excerpt` clearly mismatches the system-side `text`. `low_confid` cases: - Confirmed claim has empty `evidence_bindings`. - `source_invocation_id` exists but `raw_path` is missing. - Old invocation lacks `evidence_refs`. - `evidence_excerpt` is too short or generic to compare. - Output is incomplete without concrete fabricated IDs or paths. ## Verifier Behavior Verifier should treat `gatekeeper_result.severity` as a hard boundary: - `reject`: effective result cannot be `PASS`; fabricated-reference cases should become `REJECT`. - `low_confid`: effective result cannot be `PASS`; missing-reference cases should become `LOW_CONFID`. - `none`: Verifier judges whether claim text is derivable from verified excerpts. `tool_trace_summary` remains available as navigation and audit context, but claim-local verified excerpts are the primary evidence for derivability. ## Hook Placement Gatekeeper remains in the Verifier hook path for this change. Retry or rollback into Executor is not implemented in this stage. Existing ad hoc verifier-hook validation should be replaced by Gatekeeper output. The hook may normalize compatibility fields, but it should not independently decide pass/fail outside Gatekeeper semantics. ## Auto-backfill `VerifierInputHook` may only auto-fill a missing `source_invocation_id` when exactly one invocation candidate exists for the binding's `tool_name`. Rules: - Never bulk-fill multiple IDs. - Never auto-fill `raw_path`. - Add a warning when auto-fill occurs. - Auto-filled binding without `raw_path` must remain `LOW_CONFID`. ## Prompt Constraints Executor prompt changes are prompt-first, contract-later: - Executor is an evidence collector and micro-fact extractor. - Narrow-scope questions should usually produce one claim and at most two claims. - Do not limit evidence binding count. - Emit observation or negative observation, not root-cause certainty, for narrow confirmation questions. - Runbook and skill content cannot become current incident facts. - Recommended actions, if any, are evidence-collection next steps, not remediation actions. ## HikariCP Mock Behavior `query_logs` should support positive mock hits for: - `HikariCP` - `HikariPool` - `connection pool` - `数据库连接池` - `连接池耗尽` - `active=50/50` - `waiting` - `request timed out after 30000ms` - `order-service` Positive output should include order-service HikariCP log records. No-hit should return `logs=[]` and `evidence_status=no_evidence`, not `generic-service` placeholder logs. ## Audit Gatekeeper result must be persisted under: ```text diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result ``` The minimum persisted fields are: - `status` - `severity` - `checked_bindings` - `failed_rules` - `warnings` - `errors` ## Interface Impact Impact level: L2 internal contract change. The change affects internal Agent payload JSON, persisted audit JSON, prompts, and test fixtures. It does not add public HTTP endpoints or new database tables.