Files

10 KiB

session-run-trace-isolation Specification

Purpose

Separate multi-turn conversation metadata from per-execution diagnosis state. sessionId identifies the conversation context, while runId identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.

Requirements

Requirement: Conversation metadata SHALL be separated from diagnosis runs

The system SHALL persist multi-turn conversation metadata in chat_session and one execution's auditable state in diagnosis_run.

Scenario: Valid Chat execution creates session metadata and a run

  • WHEN a valid /api/chat request enters the Chat execution path and the service resolves an effective sessionId
  • THEN the system SHALL ensure a chat_session row exists for the effective sessionId
  • AND it SHALL create a new diagnosis_run row with a unique run_id
  • AND the diagnosis_run.session_id SHALL equal the effective sessionId

Scenario: Invalid Chat request does not create a run

  • WHEN a /api/chat request fails parameter validation before execution
  • THEN the system SHALL NOT create a diagnosis_run

Scenario: Chat session stores metadata only

  • WHEN a Chat request completes
  • THEN chat_session SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
  • AND it SHALL NOT store full conversation message history

Requirement: Chat responses SHALL expose run identity

The system SHALL expose the current execution runId to clients that submit Chat requests.

Scenario: Chat response includes runId

  • WHEN /api/chat returns a successful response
  • THEN the response SHALL include sessionId
  • AND the response SHALL include runId for the created diagnosis run

Scenario: Multi-turn Chat keeps one session and multiple runs

  • WHEN two valid /api/chat requests use the same sessionId
  • THEN the system SHALL preserve multi-turn Redis context for that sessionId
  • AND it SHALL persist two distinct diagnosis_run.run_id values

Requirement: Trace details SHALL be scoped by run

The system SHALL write and read agent_step and tool_invocation rows using run_id as the execution boundary.

Scenario: Agent steps are recorded with runId

  • WHEN an Agent model step is persisted during a diagnosis run
  • THEN the agent_step row SHALL include the current run_id
  • AND it SHALL retain the current session_id

Scenario: Tool invocations are recorded with runId

  • WHEN an evidence tool invocation is persisted during a diagnosis run
  • THEN the tool_invocation row SHALL include the current run_id
  • AND it SHALL retain the current session_id

Scenario: Run metrics count only current run rows

  • WHEN a diagnosis run completes
  • THEN its step and tool counts SHALL be calculated from rows matching that run_id
  • AND rows from other runs in the same sessionId SHALL NOT be counted

Requirement: Trace API SHALL support latest-run and exact-run queries

The system SHALL allow callers to query a diagnosis trace by sessionId alone for compatibility or by sessionId + runId for exact run replay.

Scenario: Trace without runId resolves latest run

  • WHEN a caller requests GET /api/diagnosis/{sessionId}/trace without runId
  • THEN the system SHALL resolve the latest run for that session by diagnosis_run.created_at DESC, id DESC
  • AND the response SHALL include the resolved runId

Scenario: Trace with runId returns exact run

  • WHEN a caller requests GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
  • THEN the system SHALL validate that runId belongs to the path sessionId
  • AND it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
  • AND the session summary SHALL come from chat_session metadata when available, while the run summary SHALL come from diagnosis_run

Scenario: Trace rejects run from another session

  • WHEN a caller requests a runId that belongs to a different sessionId
  • THEN the system SHALL return an error instead of leaking trace data from the other session

Requirement: Session runs SHALL be listable without expanding trace details

The system SHALL provide a lightweight run-list API for a Chat Session.

Scenario: Run list returns summaries

  • WHEN a caller requests GET /api/chat/session/{sessionId}/runs
  • THEN the system SHALL return run summaries from diagnosis_run inside the existing API response wrapper
  • AND each summary SHALL include runId, sessionId, query, status, agentFlow, answerPreview, stepCount, toolCallCount, createdAt, and updatedAt
  • AND the response SHALL NOT expand agent_step or tool_invocation detail rows

Scenario: Run list handles session without runs

  • WHEN a caller requests GET /api/chat/session/{sessionId}/runs for an existing chat_session with no runs
  • THEN the system SHALL return a successful empty list

Scenario: Run list rejects missing session

  • WHEN a caller requests GET /api/chat/session/{sessionId}/runs for a session that does not exist in chat_session or diagnosis_run
  • THEN the system SHALL use the existing not-found/error response behavior

Requirement: Feedback SHALL bind to diagnosis runs

The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.

Scenario: Feedback with runId updates specified run

  • WHEN a feedback request includes sessionId and runId
  • THEN the system SHALL validate that the run belongs to the session
  • AND it SHALL update feedback on that run
  • AND the response SHALL include the actual bound runId
  • AND the response SHALL include fallbackToLatestRun=false

Scenario: Feedback without runId falls back observably

  • WHEN a legacy feedback request includes sessionId but omits runId
  • AND at least one diagnosis_run exists for that session
  • THEN the system SHALL bind feedback to the latest run for that session
  • AND the response SHALL include fallbackToLatestRun=true
  • AND the response SHALL include the actual bound runId

Scenario: Historical feedback without run-backed data remains compatible

  • WHEN a legacy feedback request includes sessionId but omits runId
  • AND no diagnosis_run exists for that session
  • AND a historical diagnosis_session row exists for that session
  • THEN the system MAY bind feedback to the historical session row for migration compatibility
  • AND the response SHALL NOT claim latest-run fallback
  • AND the response MAY omit runId

Scenario: Feedback rejects run from another session

  • WHEN a feedback request includes a runId that belongs to a different sessionId
  • THEN the system SHALL return a failed feedback response instead of updating either run

Scenario: Useful feedback creates case from run

  • WHEN feedback for a run is useful
  • THEN the system SHALL create or reuse a case_library row using that run's query and answer
  • AND new automatic case data SHALL store case_library.diagnosis_id as the run_id

Requirement: AIOps executions SHALL use run isolation

The system SHALL create and expose a diagnosis run for every valid /api/ai_ops execution.

Scenario: AIOps creates run

  • WHEN /api/ai_ops starts a valid execution
  • THEN the system SHALL create a diagnosis_run with agent_flow=AI_OPS
  • AND AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that run_id
  • AND AIOps rule evaluation SHALL be stored under diagnosis_run.self_evaluation.aiops_rule_evaluation for the current run

Scenario: AIOps SSE exposes runId

  • WHEN /api/ai_ops streams response metadata to the caller
  • THEN the stream SHALL send a compatible metadata message before report content
  • AND the SSE event name SHALL remain message
  • AND the message type SHALL be metadata
  • AND the metadata payload SHALL expose the resolved sessionId
  • AND the metadata payload SHALL expose the created runId
  • AND report content SHALL continue to use the existing content message shape

Requirement: Migration SHALL preserve historical trace access

The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old diagnosis_session table.

Scenario: Historical session gets compatibility run

  • WHEN migration runs on an existing diagnosis_session row
  • THEN the system SHALL create a compatible diagnosis_run row for that session
  • AND old agent_step and tool_invocation rows for that session SHALL be backfilled to that run_id when possible

Scenario: Old table is retained

  • WHEN the migration completes
  • THEN the diagnosis_session table SHALL remain available for historical comparison and rollback
  • AND new execution writes SHALL target chat_session and diagnosis_run

Requirement: Demo and Trace UI SHALL support runId

The demo tooling and Trace UI SHALL support minimal run-aware workflows.

Scenario: Demo script queries exact trace

  • WHEN a demo script receives a /api/chat or /api/ai_ops response containing runId
  • THEN it SHALL include runId when querying the Trace API

Scenario: Trace UI honors URL runId

  • WHEN the Trace UI is opened with ?sessionId=...&runId=...
  • THEN it SHALL query GET /api/diagnosis/{sessionId}/trace?runId=...

Requirement: Run isolation SHALL be verified against baselines

The change SHALL verify both runtime behavior and evaluation baseline impact.

Scenario: Multi-turn E2E proves run isolation

  • WHEN E2E verification sends two valid Chat requests with the same sessionId
  • THEN database inspection SHALL show two diagnosis_run rows
  • AND each run SHALL have only its own step and tool rows when queried by run_id
  • AND Redis session metadata SHALL still show multi-turn context continuity

Scenario: Baseline drift is checked

  • WHEN verification is complete
  • THEN the project SHALL run or explicitly evaluate the relevant baseline diff command
  • AND any drift caused by run isolation SHALL be documented as expected or investigated as a regression