6.8 KiB
evidence-trace-hardening Specification
Purpose
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
Requirements
Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across lookup_knowledge, query_logs, and query_metrics.
Scenario: Common evidence fields are always persisted
- WHEN an evidence-bearing tool finishes a call
- THEN the persisted
tool_invocationrow SHALL includesession_id,tool_name,input_params,duration_ms,success, and an output preview or explicit no-output state
Scenario: Retrieval-aware tools preserve structured retrieval fields
- WHEN
lookup_knowledgepersists a tool invocation - THEN the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- AND those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
Scenario: Tool failure is preserved as failure
- WHEN an evidence-bearing tool throws, times out, or returns an execution error
- THEN the persisted row SHALL set
success=false - AND it SHALL preserve an
error_messageexplaining the failure
Scenario: No usable evidence is preserved without pretending success
- WHEN an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- THEN the persisted contract SHALL preserve that the call completed
- AND the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
Scenario: Deduped retrieval remains auditable
- WHEN
lookup_knowledgeis blocked by session-level deduplication - THEN the persisted row SHALL preserve the dedup reason
- AND the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
Scenario: Failed evidence calls remain visible in the summary
- WHEN
ToolTraceSummaryServiceprocesses failed evidence-bearing tool rows - THEN the summary SHALL retain them
- AND it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
Scenario: No-hit and deduped calls do not upgrade evidence level
- WHEN
ToolTraceSummaryServiceprocesses rows that returned no usable evidence or were deduped - THEN those rows SHALL NOT be promoted to direct or indirect evidence
- AND their counts SHALL still be reflected in the merged summary entry
Scenario: Successful evidence keeps the strongest available support
- WHEN multiple rows for the same tool and topic domain are merged
- THEN the summary SHALL preserve the strongest successful evidence level among them
- AND it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
Scenario: Missing verifier output falls back to LOW_CONFID
- WHEN the verifier step completes without a usable
verifier_output - THEN
ChatServiceSHALL fall back to aLOW_CONFIDdecision - AND the final user-facing output SHALL use the fixed low-confidence protocol
Scenario: Invalid verifier JSON falls back to LOW_CONFID
- WHEN the verifier returns malformed or non-parseable JSON
- THEN
ChatServiceSHALL fall back to aLOW_CONFIDdecision - AND the fallback SHALL still persist a verifier-evaluation record
Scenario: REJECT output hides unverified raw answer text
- WHEN the final verifier decision is
REJECT - THEN the user-facing output SHALL use the degraded template
- AND it SHALL NOT pass through the raw executor answer
Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
Scenario: Evidence recorder contract is tested offline
- WHEN the test suite runs the focused recorder tests
- THEN it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
Scenario: Trace summary hardening is tested offline
- WHEN the test suite runs the focused trace-summary tests
- THEN it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
Scenario: Verifier fallback behavior is tested offline
- WHEN the test suite runs the focused
ChatServicefallback tests - THEN it SHALL verify the missing-output, invalid-JSON,
LOW_CONFID, andREJECTdegraded paths without requiring a real LLM or database
Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations
When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.
Scenario: Merged log calls include one direct hit
- WHEN repeated
query_logscalls for the same topic are merged - AND at least one call contains a concrete direct evidence hit
- THEN the merged summary SHALL include a bounded direct-hit excerpt
- AND it SHALL preserve all contributing
source_invocation_ids
Scenario: Truncated preview still contains direct evidence
- WHEN a persisted output preview is marked truncated
- AND the preview contains a concrete direct evidence hit
- THEN the summary SHALL treat the row as successful evidence
- AND it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated
Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations
The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.
Scenario: Diagnosis session exposes aggregate counts
- WHEN a diagnosis session completes
- THEN
diagnosis_session.step_countSHALL count persisted Agent model steps - AND
diagnosis_session.tool_call_countSHALL count persisted evidence-tool invocation rows - AND helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating
tool_invocation