Harden evidence trace semantics

This commit is contained in:
aruo
2026-07-04 22:36:30 +08:00
parent 246c99b954
commit dc6cd32a67
24 changed files with 1096 additions and 141 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,64 @@
## Context
The current MVP already has the core pieces required for traceable agent execution:
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
- degraded output behavior exists in code but is only lightly covered by tests
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
## Goals / Non-Goals
**Goals:**
- Centralize the common persistence contract for evidence-bearing tools.
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
- Define stable summarization semantics for:
- successful evidence
- no-hit / no-usable-evidence
- deduped retrievals
- failed evidence queries
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
- Keep the scope small enough to unblock the next P1-B evaluation harness.
**Non-Goals:**
- No new table, column, or Flyway migration.
- No new public API.
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
- No attempt to redesign planner/executor routing.
- No full offline runtime or end-to-end benchmark harness in this slice.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
## Risks / Trade-offs
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
## Migration Plan
- No deployment migration is required beyond shipping the code changes.
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
## Open Questions
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
@@ -0,0 +1,26 @@
## Why
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
## What Changes
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
## Capabilities
### New Capabilities
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
### Modified Capabilities
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
## Impact
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
- Affected APIs: none. No new endpoint or schema is introduced.
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
@@ -0,0 +1,14 @@
## ADDED Requirements
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
#### Scenario: Failed evidence remains a verifier-visible gap
- **WHEN** an evidence-bearing tool invocation fails
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
- **AND** the verifier flow SHALL continue without crashing
#### Scenario: Deduped retrievals do not count as fresh support
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
@@ -0,0 +1,82 @@
## ADDED Requirements
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
#### Scenario: Common evidence fields are always persisted
- **WHEN** an evidence-bearing tool finishes a call
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
- **WHEN** `lookup_knowledge` persists a tool invocation
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
#### Scenario: Tool failure is preserved as failure
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
- **THEN** the persisted row SHALL set `success=false`
- **AND** it SHALL preserve an `error_message` explaining the failure
#### Scenario: No usable evidence is preserved without pretending success
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- **THEN** the persisted contract SHALL preserve that the call completed
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
#### Scenario: Deduped retrieval remains auditable
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
- **THEN** the persisted row SHALL preserve the dedup reason
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
#### Scenario: Failed evidence calls remain visible in the summary
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
- **THEN** the summary SHALL retain them
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
#### Scenario: No-hit and deduped calls do not upgrade evidence level
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
- **AND** their counts SHALL still be reflected in the merged summary entry
#### Scenario: Successful evidence keeps the strongest available support
- **WHEN** multiple rows for the same tool and topic domain are merged
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
### Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
#### Scenario: Missing verifier output falls back to LOW_CONFID
- **WHEN** the verifier step completes without a usable `verifier_output`
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
- **WHEN** the verifier returns malformed or non-parseable JSON
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
- **AND** the fallback SHALL still persist a verifier-evaluation record
#### Scenario: REJECT output hides unverified raw answer text
- **WHEN** the final verifier decision is `REJECT`
- **THEN** the user-facing output SHALL use the degraded template
- **AND** it SHALL NOT pass through the raw executor answer
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
#### Scenario: Evidence recorder contract is tested offline
- **WHEN** the test suite runs the focused recorder tests
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
#### Scenario: Trace summary hardening is tested offline
- **WHEN** the test suite runs the focused trace-summary tests
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
#### Scenario: Verifier fallback behavior is tested offline
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
@@ -0,0 +1,22 @@
## 1. Evidence Persistence Contract
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
## 2. Verifier-Facing Summary Semantics
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
## 3. Chat Degraded Paths
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
## 4. Verification
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
- [x] 4.3 Run compile verification.