Files
SuperBizAgent-java/openspec/specs/evidence-trace-hardening/spec.md
T

9.5 KiB

evidence-trace-hardening Specification

Purpose

TBD - created by archiving change evidence-trace-hardening. Update Purpose after archive.

Requirements

Requirement: Evidence-bearing tools SHALL persist a unified invocation contract

The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across lookup_knowledge, query_logs, and query_metrics.

Scenario: Common evidence fields are always persisted

  • WHEN an evidence-bearing tool finishes a call
  • THEN the persisted tool_invocation row SHALL include session_id, tool_name, input_params, duration_ms, success, and an output preview or explicit no-output state

Scenario: Retrieval-aware tools preserve structured retrieval fields

  • WHEN lookup_knowledge persists a tool invocation
  • THEN the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
  • AND those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths

Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes

The system SHALL keep failed calls separate from successful calls that return no usable evidence.

Scenario: Tool failure is preserved as failure

  • WHEN an evidence-bearing tool throws, times out, or returns an execution error
  • THEN the persisted row SHALL set success=false
  • AND it SHALL preserve an error_message explaining the failure

Scenario: No usable evidence is preserved without pretending success

  • WHEN an evidence-bearing tool completes normally but yields no usable evidence for the verifier
  • THEN the persisted contract SHALL preserve that the call completed
  • AND the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support

Scenario: Deduped retrieval remains auditable

  • WHEN lookup_knowledge is blocked by session-level deduplication
  • THEN the persisted row SHALL preserve the dedup reason
  • AND the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit

Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules

The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.

Scenario: Failed evidence calls remain visible in the summary

  • WHEN ToolTraceSummaryService processes failed evidence-bearing tool rows
  • THEN the summary SHALL retain them
  • AND it SHALL mark them as unsuccessful evidence with an output summary that explains the gap

Scenario: No-hit and deduped calls do not upgrade evidence level

  • WHEN ToolTraceSummaryService processes rows that returned no usable evidence or were deduped
  • THEN those rows SHALL NOT be promoted to direct or indirect evidence
  • AND their counts SHALL still be reflected in the merged summary entry

Scenario: Successful evidence keeps the strongest available support

  • WHEN multiple rows for the same tool and topic domain are merged
  • THEN the summary SHALL preserve the strongest successful evidence level among them
  • AND it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability

Requirement: ChatService SHALL degrade predictably on verifier output failures

The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.

Scenario: Missing verifier output falls back to LOW_CONFID

  • WHEN the verifier step completes without a usable verifier_output
  • THEN ChatService SHALL fall back to a LOW_CONFID decision
  • AND the final user-facing output SHALL use the fixed low-confidence protocol

Scenario: Invalid verifier JSON falls back to LOW_CONFID

  • WHEN the verifier returns malformed or non-parseable JSON
  • THEN ChatService SHALL fall back to a LOW_CONFID decision
  • AND the fallback SHALL still persist a verifier-evaluation record

Scenario: REJECT output hides unverified raw answer text

  • WHEN the final verifier decision is REJECT
  • THEN the user-facing output SHALL use the degraded template
  • AND it SHALL NOT pass through the raw executor answer

Requirement: Evidence-trace hardening SHALL be covered by focused offline tests

The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.

Scenario: Evidence recorder contract is tested offline

  • WHEN the test suite runs the focused recorder tests
  • THEN it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure

Scenario: Trace summary hardening is tested offline

  • WHEN the test suite runs the focused trace-summary tests
  • THEN it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows

Scenario: Verifier fallback behavior is tested offline

  • WHEN the test suite runs the focused ChatService fallback tests
  • THEN it SHALL verify the missing-output, invalid-JSON, LOW_CONFID, and REJECT degraded paths without requiring a real LLM or database

Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations

When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.

Scenario: Merged log calls include one direct hit

  • WHEN repeated query_logs calls for the same topic are merged
  • AND at least one call contains a concrete direct evidence hit
  • THEN the merged summary SHALL include a bounded direct-hit excerpt
  • AND it SHALL preserve all contributing source_invocation_ids

Scenario: Truncated preview still contains direct evidence

  • WHEN a persisted output preview is marked truncated
  • AND the preview contains a concrete direct evidence hit
  • THEN the summary SHALL treat the row as successful evidence
  • AND it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated

Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations

The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.

Scenario: Diagnosis session exposes aggregate counts

  • WHEN a diagnosis session completes
  • THEN diagnosis_session.step_count SHALL count persisted Agent model steps
  • AND diagnosis_session.tool_call_count SHALL count persisted evidence-tool invocation rows
  • AND helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating tool_invocation

Requirement: Evidence tools SHALL persist minimal evidence refs

Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in tool_invocation.retrieval_details.evidence_refs.

Scenario: Metrics alerts produce evidence refs

  • WHEN a query_metrics invocation returns alert entries
  • THEN the persisted retrieval details SHALL include one evidence_refs item per usable alert
  • AND each item SHALL include raw_path formatted as $.alerts[i]
  • AND each item SHALL include bounded text containing concrete alert facts such as alert name, state, service, current value, and duration when available

Scenario: Logs produce evidence refs

  • WHEN a query_logs invocation returns log entries
  • THEN the persisted retrieval details SHALL include one evidence_refs item per usable log
  • AND each item SHALL include raw_path formatted as $.logs[i]
  • AND each item SHALL include bounded text containing concrete log facts such as timestamp, level, service, and message when available

Scenario: Knowledge lookup produces evidence refs

  • WHEN a lookup_knowledge invocation returns evidence blocks
  • THEN the persisted retrieval details SHALL include one evidence_refs item per usable evidence block
  • AND each item SHALL include raw_path formatted as $.evidence_blocks[i]
  • AND each item SHALL include bounded text containing concrete block content, title, or source when available

Scenario: Evidence ref extraction does not infer diagnosis

  • WHEN the recorder creates evidence_refs
  • THEN it SHALL only copy or format concrete tool output fields
  • AND it SHALL NOT infer root cause, remediation, or diagnosis conclusions

Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence

The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.

Scenario: HikariCP positive query returns order-service pool evidence

  • WHEN a query_logs request targets order-service and HikariCP connection-pool exhaustion terms
  • THEN the tool SHALL return order-service HikariCP-related log entries
  • AND the returned evidence SHALL include concrete terms such as HikariPool, active=50/50, waiting, or request timed out after 30000ms
  • AND it SHALL NOT return generic-service placeholder logs

Scenario: HikariCP no-hit query returns no evidence

  • WHEN a query_logs request targets a service without matching HikariCP mock evidence
  • THEN the tool SHALL return an empty logs array
  • AND it SHALL mark the output as evidence_status=no_evidence
  • AND it SHALL NOT return generic-service placeholder logs