Harden evidence trace semantics
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,64 @@
|
||||
## Context
|
||||
|
||||
The current MVP already has the core pieces required for traceable agent execution:
|
||||
|
||||
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
|
||||
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
|
||||
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
|
||||
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
|
||||
|
||||
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
|
||||
|
||||
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
|
||||
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
|
||||
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
|
||||
- degraded output behavior exists in code but is only lightly covered by tests
|
||||
|
||||
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Centralize the common persistence contract for evidence-bearing tools.
|
||||
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
|
||||
- Define stable summarization semantics for:
|
||||
- successful evidence
|
||||
- no-hit / no-usable-evidence
|
||||
- deduped retrievals
|
||||
- failed evidence queries
|
||||
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
|
||||
- Keep the scope small enough to unblock the next P1-B evaluation harness.
|
||||
|
||||
**Non-Goals:**
|
||||
- No new table, column, or Flyway migration.
|
||||
- No new public API.
|
||||
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
|
||||
- No attempt to redesign planner/executor routing.
|
||||
- No full offline runtime or end-to-end benchmark harness in this slice.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
|
||||
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
|
||||
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
|
||||
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
|
||||
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
|
||||
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
|
||||
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
|
||||
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required beyond shipping the code changes.
|
||||
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
|
||||
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
|
||||
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
|
||||
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
|
||||
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
|
||||
|
||||
### Modified Capabilities
|
||||
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
|
||||
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
|
||||
- Affected APIs: none. No new endpoint or schema is introduced.
|
||||
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
+82
@@ -0,0 +1,82 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Evidence Persistence Contract
|
||||
|
||||
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
|
||||
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
|
||||
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
|
||||
|
||||
## 2. Verifier-Facing Summary Semantics
|
||||
|
||||
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
|
||||
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
|
||||
|
||||
## 3. Chat Degraded Paths
|
||||
|
||||
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
|
||||
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
|
||||
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
|
||||
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -190,3 +190,17 @@ Verifier facts SHALL be linkable to the evidence summaries used during verificat
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
# evidence-trace-hardening Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
|
||||
Reference in New Issue
Block a user