# evidence-trace-hardening Specification ## Purpose TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive. ## Requirements ### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`. #### Scenario: Common evidence fields are always persisted - **WHEN** an evidence-bearing tool finishes a call - **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state #### Scenario: Retrieval-aware tools preserve structured retrieval fields - **WHEN** `lookup_knowledge` persists a tool invocation - **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details - **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths ### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes The system SHALL keep failed calls separate from successful calls that return no usable evidence. #### Scenario: Tool failure is preserved as failure - **WHEN** an evidence-bearing tool throws, times out, or returns an execution error - **THEN** the persisted row SHALL set `success=false` - **AND** it SHALL preserve an `error_message` explaining the failure #### Scenario: No usable evidence is preserved without pretending success - **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier - **THEN** the persisted contract SHALL preserve that the call completed - **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support #### Scenario: Deduped retrieval remains auditable - **WHEN** `lookup_knowledge` is blocked by session-level deduplication - **THEN** the persisted row SHALL preserve the dedup reason - **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit ### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes. #### Scenario: Failed evidence calls remain visible in the summary - **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows - **THEN** the summary SHALL retain them - **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap #### Scenario: No-hit and deduped calls do not upgrade evidence level - **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped - **THEN** those rows SHALL NOT be promoted to direct or indirect evidence - **AND** their counts SHALL still be reflected in the merged summary entry #### Scenario: Successful evidence keeps the strongest available support - **WHEN** multiple rows for the same tool and topic domain are merged - **THEN** the summary SHALL preserve the strongest successful evidence level among them - **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability ### Requirement: ChatService SHALL degrade predictably on verifier output failures The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error. #### Scenario: Missing verifier output falls back to LOW_CONFID - **WHEN** the verifier step completes without a usable `verifier_output` - **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision - **AND** the final user-facing output SHALL use the fixed low-confidence protocol #### Scenario: Invalid verifier JSON falls back to LOW_CONFID - **WHEN** the verifier returns malformed or non-parseable JSON - **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision - **AND** the fallback SHALL still persist a verifier-evaluation record #### Scenario: REJECT output hides unverified raw answer text - **WHEN** the final verifier decision is `REJECT` - **THEN** the user-facing output SHALL use the degraded template - **AND** it SHALL NOT pass through the raw executor answer ### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior. #### Scenario: Evidence recorder contract is tested offline - **WHEN** the test suite runs the focused recorder tests - **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure #### Scenario: Trace summary hardening is tested offline - **WHEN** the test suite runs the focused trace-summary tests - **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows #### Scenario: Verifier fallback behavior is tested offline - **WHEN** the test suite runs the focused `ChatService` fallback tests - **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database ### Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation. #### Scenario: Merged log calls include one direct hit - **WHEN** repeated `query_logs` calls for the same topic are merged - **AND** at least one call contains a concrete direct evidence hit - **THEN** the merged summary SHALL include a bounded direct-hit excerpt - **AND** it SHALL preserve all contributing `source_invocation_ids` #### Scenario: Truncated preview still contains direct evidence - **WHEN** a persisted output preview is marked truncated - **AND** the preview contains a concrete direct evidence hit - **THEN** the summary SHALL treat the row as successful evidence - **AND** it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated ### Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts. #### Scenario: Diagnosis session exposes aggregate counts - **WHEN** a diagnosis session completes - **THEN** `diagnosis_session.step_count` SHALL count persisted Agent model steps - **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows - **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation` ### Requirement: Evidence tools SHALL persist minimal evidence refs Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`. #### Scenario: Metrics alerts produce evidence refs - **WHEN** a `query_metrics` invocation returns alert entries - **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert - **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]` - **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available #### Scenario: Logs produce evidence refs - **WHEN** a `query_logs` invocation returns log entries - **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log - **AND** each item SHALL include `raw_path` formatted as `$.logs[i]` - **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available #### Scenario: Knowledge lookup produces evidence refs - **WHEN** a `lookup_knowledge` invocation returns evidence blocks - **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block - **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]` - **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available #### Scenario: Evidence ref extraction does not infer diagnosis - **WHEN** the recorder creates `evidence_refs` - **THEN** it SHALL only copy or format concrete tool output fields - **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions ### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence. #### Scenario: HikariCP positive query returns order-service pool evidence - **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms - **THEN** the tool SHALL return order-service HikariCP-related log entries - **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms` - **AND** it SHALL NOT return `generic-service` placeholder logs #### Scenario: HikariCP no-hit query returns no evidence - **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence - **THEN** the tool SHALL return an empty `logs` array - **AND** it SHALL mark the output as `evidence_status=no_evidence` - **AND** it SHALL NOT return `generic-service` placeholder logs