2.0 KiB
2.0 KiB
Why
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
What Changes
- Standardize the persisted evidence-tool contract across
lookup_knowledge,query_logs, andquery_metrics. - Align how evidence tools represent success, no-hit, deduped, and failed calls so
ToolTraceSummaryServicecan summarize them consistently. - Harden
ChatServicefallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable. - Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
Capabilities
New Capabilities
evidence-trace-hardening: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
Modified Capabilities
chat-verifier-agent: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
Impact
- Affected code:
ToolInvocationRecorder,LookupKnowledgeTool,QueryLogsTools,QueryMetricsTools,ToolTraceSummaryService,ChatService, and focused test classes. - Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
- Affected APIs: none. No new endpoint or schema is introduced.
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.