Files
SuperBizAgent-java/openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity/decisions.md
T

74 lines
4.9 KiB
Markdown

# Decisions
## Discover Context
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`.
- Related existing specs: `chat-verifier-agent`, `evidence-trace-hardening`.
- Related devflow records: `executor-gatekeeper-hook`, `executor-verifier-claim-checks`, `executor-composer-final-answer`, `evidence-trace-hardening`.
- Current repo instruction requested semantic code search and LSP confirmation before code changes; those tools are not exposed in this environment, so implementation will use `rg`, direct code reading, and focused tests as fallback evidence.
## Classification
- sm-flow scale: `standard`.
- Reason: internal contract change across recorder, hook, Gatekeeper, verifier prompt, mock tools, tests, and E2E validation.
## Question Pool
| Question | Mode | Resolution |
|---|---|---|
| Should Planner change or add `scope_contract`? | user-interview, already resolved in issue | No. This change is prompt-first and does not alter Planner. |
| Should Gatekeeper move to Executor hook for retry? | user-interview, already resolved in issue | No. Gatekeeper remains in Verifier hook path for this version. |
| Should new DB tables be added for evidence refs or audit? | user-interview, already resolved in issue | No. Use `tool_invocation.retrieval_details.evidence_refs` and existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
| Should `tool_trace_summary` remain the primary evidence source? | user-interview, already resolved in issue | No. It becomes navigation/audit context; verified claim-local excerpts become primary evidence for derivability. |
| Should full JSONPath be supported? | user-interview, already resolved in issue | No. Only stable raw paths listed in design are supported. |
## Cross-artifact Alignment
| Check | Status |
|---|---|
| Issue background and target -> proposal | aligned |
| Proposal scope and non-goals -> design | aligned |
| Design protocol and risks -> specs/tasks | aligned |
| Specs observable behavior -> tasks acceptance | aligned |
## Architecture Audit
The change keeps the existing agent orchestration and only hardens the evidence payload between Executor, Gatekeeper, and Verifier. The highest coupling risk is transition compatibility from plural `source_invocation_ids` to singular `source_invocation_id`; the implementation must accept old data but only allow precise new bindings to pass. No new DB table is introduced, reducing migration risk. Verifier prompt and effective verdict guardrails must agree on `gatekeeper_result.severity`, otherwise the model could still produce a PASS that runtime later must downgrade.
## Commit Gate
Passed on 2026-07-08.
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
- File completeness: proposal, design, specs, tasks present.
- Consistency: proposal -> design -> specs -> tasks aligned.
- Apply authorization: user requested continuing implementation and fixing autonomously; proceed to Apply.
## Apply Verification
Focused tests:
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`: passed, 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`: passed, 9 tests after adding the old-invocation-without-evidence-refs case.
OpenSpec:
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
End-to-end sessions:
| Case | Session | Result | Gatekeeper |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | `PASS`; answer confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | `LOW_CONFID`; no `generic-service` pollution; no positive inventory-service confirmation | `fail / low_confid` due incomplete negative evidence references |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | `PASS`; answer confirmed HighMemoryUsage 91% without confirming memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | `PASS`; answer confirmed SlowResponse and slow request logs without DB-pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | `PASS`; answer only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
Residual observation:
- HikariCP negative no longer returns `generic-service`, and no-hit rows persist `evidence_status=no_evidence`.
- The model still issued an extra broad HikariCP query without the `inventory-service` filter and used order-service as a context claim. This is a remaining narrow-scope behavior issue, not a mock data pollution issue. It is acceptable for this change because the final verdict did not become a false positive for inventory-service.