4.6 KiB
4.6 KiB
Context
The current MVP already has the core pieces required for traceable agent execution:
LookupKnowledgeToolwrites rich retrieval metadata intotool_invocationQueryLogsToolsandQueryMetricsToolsuseToolInvocationRecorderToolTraceSummaryServiceturns persisted rows into verifier-facing evidence summariesChatServicealready contains fallback behavior for missing or invalidverifier_output
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
LookupKnowledgeToolbuildsToolInvocationrows itself- the other evidence tools use
ToolInvocationRecorder.recordEvidenceTool(...) - “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
- degraded output behavior exists in code but is only lightly covered by tests
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
Goals / Non-Goals
Goals:
- Centralize the common persistence contract for evidence-bearing tools.
- Preserve
lookup_knowledge-specific retrieval fields while removing ad hoc duplication in how evidence rows are created. - Define stable summarization semantics for:
- successful evidence
- no-hit / no-usable-evidence
- deduped retrievals
- failed evidence queries
- Make
ChatServicefallback and degraded-output paths testable as explicit product behavior. - Keep the scope small enough to unblock the next P1-B evaluation harness.
Non-Goals:
- No new table, column, or Flyway migration.
- No new public API.
- No new verifier verdict type beyond
PASS/LOW_CONFID/REJECT. - No attempt to redesign planner/executor routing.
- No full offline runtime or end-to-end benchmark harness in this slice.
Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evidence persistence ownership | Keep ToolInvocationRecorder as the single common entry point |
Let each tool continue building ToolInvocation rows ad hoc |
The recorder already exists and is the right seam for contract hardening. |
lookup_knowledge integration style |
Add a richer recorder entry path for retrieval-aware calls | Force lookup_knowledge into the same minimal method used by logs/metrics |
lookup_knowledge carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
Risks / Trade-offs
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
- [Risk]
lookup_knowledgededup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preservededupReasonand treat dedup as a first-class no-new-evidence case in summary logic.
Migration Plan
- No deployment migration is required beyond shipping the code changes.
- Existing
tool_invocationrows remain valid because this change reuses the same schema. - Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
Open Questions
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.