Files
SuperBizAgent-java/openspec/changes/archive/2026-07-04-evidence-trace-hardening/design.md
T
2026-07-04 22:36:30 +08:00

4.6 KiB

Context

The current MVP already has the core pieces required for traceable agent execution:

  • LookupKnowledgeTool writes rich retrieval metadata into tool_invocation
  • QueryLogsTools and QueryMetricsTools use ToolInvocationRecorder
  • ToolTraceSummaryService turns persisted rows into verifier-facing evidence summaries
  • ChatService already contains fallback behavior for missing or invalid verifier_output

The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:

  • LookupKnowledgeTool builds ToolInvocation rows itself
  • the other evidence tools use ToolInvocationRecorder.recordEvidenceTool(...)
  • “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
  • degraded output behavior exists in code but is only lightly covered by tests

For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.

Goals / Non-Goals

Goals:

  • Centralize the common persistence contract for evidence-bearing tools.
  • Preserve lookup_knowledge-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
  • Define stable summarization semantics for:
    • successful evidence
    • no-hit / no-usable-evidence
    • deduped retrievals
    • failed evidence queries
  • Make ChatService fallback and degraded-output paths testable as explicit product behavior.
  • Keep the scope small enough to unblock the next P1-B evaluation harness.

Non-Goals:

  • No new table, column, or Flyway migration.
  • No new public API.
  • No new verifier verdict type beyond PASS / LOW_CONFID / REJECT.
  • No attempt to redesign planner/executor routing.
  • No full offline runtime or end-to-end benchmark harness in this slice.

Decisions

Decision Choice Alternative Considered Rationale
Evidence persistence ownership Keep ToolInvocationRecorder as the single common entry point Let each tool continue building ToolInvocation rows ad hoc The recorder already exists and is the right seam for contract hardening.
lookup_knowledge integration style Add a richer recorder entry path for retrieval-aware calls Force lookup_knowledge into the same minimal method used by logs/metrics lookup_knowledge carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured.
No-evidence semantics Distinguish failed calls from successful calls that yield no usable evidence Collapse all non-successful evidence into one bucket Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”.
Degraded-path hardening Add focused unit tests around verifier fallback and output shaping Rely on runtime demo only Interview value comes from proving the system fails predictably, not just that the happy path ran once.
Scope boundary Keep changes additive and contract-oriented Expand into P1-B evaluation harness in the same change This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work.

Risks / Trade-offs

  • [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
  • [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
  • [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
  • [Risk] lookup_knowledge dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve dedupReason and treat dedup as a first-class no-new-evidence case in summary logic.

Migration Plan

  • No deployment migration is required beyond shipping the code changes.
  • Existing tool_invocation rows remain valid because this change reuses the same schema.
  • Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.

Open Questions

  • Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.