9.5 KiB
evidence-trace-hardening Specification
Purpose
TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.
Requirements
Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across lookup_knowledge, query_logs, and query_metrics.
Scenario: Common evidence fields are always persisted
- WHEN an evidence-bearing tool finishes a call
- THEN the persisted
tool_invocationrow SHALL includesession_id,tool_name,input_params,duration_ms,success, and an output preview or explicit no-output state
Scenario: Retrieval-aware tools preserve structured retrieval fields
- WHEN
lookup_knowledgepersists a tool invocation - THEN the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
- AND those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
Scenario: Tool failure is preserved as failure
- WHEN an evidence-bearing tool throws, times out, or returns an execution error
- THEN the persisted row SHALL set
success=false - AND it SHALL preserve an
error_messageexplaining the failure
Scenario: No usable evidence is preserved without pretending success
- WHEN an evidence-bearing tool completes normally but yields no usable evidence for the verifier
- THEN the persisted contract SHALL preserve that the call completed
- AND the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
Scenario: Deduped retrieval remains auditable
- WHEN
lookup_knowledgeis blocked by session-level deduplication - THEN the persisted row SHALL preserve the dedup reason
- AND the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
Scenario: Failed evidence calls remain visible in the summary
- WHEN
ToolTraceSummaryServiceprocesses failed evidence-bearing tool rows - THEN the summary SHALL retain them
- AND it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
Scenario: No-hit and deduped calls do not upgrade evidence level
- WHEN
ToolTraceSummaryServiceprocesses rows that returned no usable evidence or were deduped - THEN those rows SHALL NOT be promoted to direct or indirect evidence
- AND their counts SHALL still be reflected in the merged summary entry
Scenario: Successful evidence keeps the strongest available support
- WHEN multiple rows for the same tool and topic domain are merged
- THEN the summary SHALL preserve the strongest successful evidence level among them
- AND it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
Requirement: ChatService SHALL degrade predictably on verifier output failures
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
Scenario: Missing verifier output falls back to LOW_CONFID
- WHEN the verifier step completes without a usable
verifier_output - THEN
ChatServiceSHALL fall back to aLOW_CONFIDdecision - AND the final user-facing output SHALL use the fixed low-confidence protocol
Scenario: Invalid verifier JSON falls back to LOW_CONFID
- WHEN the verifier returns malformed or non-parseable JSON
- THEN
ChatServiceSHALL fall back to aLOW_CONFIDdecision - AND the fallback SHALL still persist a verifier-evaluation record
Scenario: REJECT output hides unverified raw answer text
- WHEN the final verifier decision is
REJECT - THEN the user-facing output SHALL use the degraded template
- AND it SHALL NOT pass through the raw executor answer
Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
Scenario: Evidence recorder contract is tested offline
- WHEN the test suite runs the focused recorder tests
- THEN it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
Scenario: Trace summary hardening is tested offline
- WHEN the test suite runs the focused trace-summary tests
- THEN it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
Scenario: Verifier fallback behavior is tested offline
- WHEN the test suite runs the focused
ChatServicefallback tests - THEN it SHALL verify the missing-output, invalid-JSON,
LOW_CONFID, andREJECTdegraded paths without requiring a real LLM or database
Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations
When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.
Scenario: Merged log calls include one direct hit
- WHEN repeated
query_logscalls for the same topic are merged - AND at least one call contains a concrete direct evidence hit
- THEN the merged summary SHALL include a bounded direct-hit excerpt
- AND it SHALL preserve all contributing
source_invocation_ids
Scenario: Truncated preview still contains direct evidence
- WHEN a persisted output preview is marked truncated
- AND the preview contains a concrete direct evidence hit
- THEN the summary SHALL treat the row as successful evidence
- AND it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated
Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations
The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.
Scenario: Diagnosis session exposes aggregate counts
- WHEN a diagnosis session completes
- THEN
diagnosis_session.step_countSHALL count persisted Agent model steps - AND
diagnosis_session.tool_call_countSHALL count persisted evidence-tool invocation rows - AND helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating
tool_invocation
Requirement: Evidence tools SHALL persist minimal evidence refs
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in tool_invocation.retrieval_details.evidence_refs.
Scenario: Metrics alerts produce evidence refs
- WHEN a
query_metricsinvocation returns alert entries - THEN the persisted retrieval details SHALL include one
evidence_refsitem per usable alert - AND each item SHALL include
raw_pathformatted as$.alerts[i] - AND each item SHALL include bounded
textcontaining concrete alert facts such as alert name, state, service, current value, and duration when available
Scenario: Logs produce evidence refs
- WHEN a
query_logsinvocation returns log entries - THEN the persisted retrieval details SHALL include one
evidence_refsitem per usable log - AND each item SHALL include
raw_pathformatted as$.logs[i] - AND each item SHALL include bounded
textcontaining concrete log facts such as timestamp, level, service, and message when available
Scenario: Knowledge lookup produces evidence refs
- WHEN a
lookup_knowledgeinvocation returns evidence blocks - THEN the persisted retrieval details SHALL include one
evidence_refsitem per usable evidence block - AND each item SHALL include
raw_pathformatted as$.evidence_blocks[i] - AND each item SHALL include bounded
textcontaining concrete block content, title, or source when available
Scenario: Evidence ref extraction does not infer diagnosis
- WHEN the recorder creates
evidence_refs - THEN it SHALL only copy or format concrete tool output fields
- AND it SHALL NOT infer root cause, remediation, or diagnosis conclusions
Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
Scenario: HikariCP positive query returns order-service pool evidence
- WHEN a
query_logsrequest targetsorder-serviceand HikariCP connection-pool exhaustion terms - THEN the tool SHALL return order-service HikariCP-related log entries
- AND the returned evidence SHALL include concrete terms such as
HikariPool,active=50/50,waiting, orrequest timed out after 30000ms - AND it SHALL NOT return
generic-serviceplaceholder logs
Scenario: HikariCP no-hit query returns no evidence
- WHEN a
query_logsrequest targets a service without matching HikariCP mock evidence - THEN the tool SHALL return an empty
logsarray - AND it SHALL mark the output as
evidence_status=no_evidence - AND it SHALL NOT return
generic-serviceplaceholder logs