Harden evidence trace semantics
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
## Context
|
||||
|
||||
The current MVP already has the core pieces required for traceable agent execution:
|
||||
|
||||
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
|
||||
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
|
||||
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
|
||||
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
|
||||
|
||||
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
|
||||
|
||||
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
|
||||
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
|
||||
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
|
||||
- degraded output behavior exists in code but is only lightly covered by tests
|
||||
|
||||
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Centralize the common persistence contract for evidence-bearing tools.
|
||||
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
|
||||
- Define stable summarization semantics for:
|
||||
- successful evidence
|
||||
- no-hit / no-usable-evidence
|
||||
- deduped retrievals
|
||||
- failed evidence queries
|
||||
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
|
||||
- Keep the scope small enough to unblock the next P1-B evaluation harness.
|
||||
|
||||
**Non-Goals:**
|
||||
- No new table, column, or Flyway migration.
|
||||
- No new public API.
|
||||
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
|
||||
- No attempt to redesign planner/executor routing.
|
||||
- No full offline runtime or end-to-end benchmark harness in this slice.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
|
||||
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
|
||||
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
|
||||
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
|
||||
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
|
||||
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
|
||||
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
|
||||
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required beyond shipping the code changes.
|
||||
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
|
||||
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
|
||||
Reference in New Issue
Block a user