## Context The current MVP already has the core pieces required for traceable agent execution: - `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation` - `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder` - `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries - `ChatService` already contains fallback behavior for missing or invalid `verifier_output` The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions: - `LookupKnowledgeTool` builds `ToolInvocation` rows itself - the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)` - “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools - degraded output behavior exists in code but is only lightly covered by tests For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow. ## Goals / Non-Goals **Goals:** - Centralize the common persistence contract for evidence-bearing tools. - Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created. - Define stable summarization semantics for: - successful evidence - no-hit / no-usable-evidence - deduped retrievals - failed evidence queries - Make `ChatService` fallback and degraded-output paths testable as explicit product behavior. - Keep the scope small enough to unblock the next P1-B evaluation harness. **Non-Goals:** - No new table, column, or Flyway migration. - No new public API. - No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`. - No attempt to redesign planner/executor routing. - No full offline runtime or end-to-end benchmark harness in this slice. ## Decisions | Decision | Choice | Alternative Considered | Rationale | |---|---|---|---| | Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. | | `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. | | No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. | | Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. | | Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. | ## Risks / Trade-offs - [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output. - [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload. - [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots. - [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic. ## Migration Plan - No deployment migration is required beyond shipping the code changes. - Existing `tool_invocation` rows remain valid because this change reuses the same schema. - Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is. ## Open Questions - Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.