## ADDED Requirements ### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`. #### Scenario: Common evidence fields are always persisted - **WHEN** an evidence-bearing tool finishes a call - **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state #### Scenario: Retrieval-aware tools preserve structured retrieval fields - **WHEN** `lookup_knowledge` persists a tool invocation - **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details - **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths ### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes The system SHALL keep failed calls separate from successful calls that return no usable evidence. #### Scenario: Tool failure is preserved as failure - **WHEN** an evidence-bearing tool throws, times out, or returns an execution error - **THEN** the persisted row SHALL set `success=false` - **AND** it SHALL preserve an `error_message` explaining the failure #### Scenario: No usable evidence is preserved without pretending success - **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier - **THEN** the persisted contract SHALL preserve that the call completed - **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support #### Scenario: Deduped retrieval remains auditable - **WHEN** `lookup_knowledge` is blocked by session-level deduplication - **THEN** the persisted row SHALL preserve the dedup reason - **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit ### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes. #### Scenario: Failed evidence calls remain visible in the summary - **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows - **THEN** the summary SHALL retain them - **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap #### Scenario: No-hit and deduped calls do not upgrade evidence level - **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped - **THEN** those rows SHALL NOT be promoted to direct or indirect evidence - **AND** their counts SHALL still be reflected in the merged summary entry #### Scenario: Successful evidence keeps the strongest available support - **WHEN** multiple rows for the same tool and topic domain are merged - **THEN** the summary SHALL preserve the strongest successful evidence level among them - **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability ### Requirement: ChatService SHALL degrade predictably on verifier output failures The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error. #### Scenario: Missing verifier output falls back to LOW_CONFID - **WHEN** the verifier step completes without a usable `verifier_output` - **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision - **AND** the final user-facing output SHALL use the fixed low-confidence protocol #### Scenario: Invalid verifier JSON falls back to LOW_CONFID - **WHEN** the verifier returns malformed or non-parseable JSON - **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision - **AND** the fallback SHALL still persist a verifier-evaluation record #### Scenario: REJECT output hides unverified raw answer text - **WHEN** the final verifier decision is `REJECT` - **THEN** the user-facing output SHALL use the degraded template - **AND** it SHALL NOT pass through the raw executor answer ### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior. #### Scenario: Evidence recorder contract is tested offline - **WHEN** the test suite runs the focused recorder tests - **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure #### Scenario: Trace summary hardening is tested offline - **WHEN** the test suite runs the focused trace-summary tests - **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows #### Scenario: Verifier fallback behavior is tested offline - **WHEN** the test suite runs the focused `ChatService` fallback tests - **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database