Files
2026-07-04 22:36:30 +08:00

65 lines
4.6 KiB
Markdown

## Context
The current MVP already has the core pieces required for traceable agent execution:
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
- degraded output behavior exists in code but is only lightly covered by tests
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
## Goals / Non-Goals
**Goals:**
- Centralize the common persistence contract for evidence-bearing tools.
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
- Define stable summarization semantics for:
- successful evidence
- no-hit / no-usable-evidence
- deduped retrievals
- failed evidence queries
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
- Keep the scope small enough to unblock the next P1-B evaluation harness.
**Non-Goals:**
- No new table, column, or Flyway migration.
- No new public API.
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
- No attempt to redesign planner/executor routing.
- No full offline runtime or end-to-end benchmark harness in this slice.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
## Risks / Trade-offs
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
## Migration Plan
- No deployment migration is required beyond shipping the code changes.
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
## Open Questions
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.