12 KiB
session-run-trace-isolation Specification
Purpose
Separate multi-turn conversation metadata from per-execution diagnosis state. sessionId identifies the conversation context, while runId identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
Requirements
Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in chat_session and one execution's auditable state in diagnosis_run.
Scenario: Valid Chat execution creates session metadata and a run
- WHEN a valid
/api/chatrequest enters the Chat execution path and the service resolves an effectivesessionId - THEN the system SHALL ensure a
chat_sessionrow exists for the effectivesessionId - AND it SHALL create a new
diagnosis_runrow with a uniquerun_id - AND the
diagnosis_run.session_idSHALL equal the effectivesessionId
Scenario: Invalid Chat request does not create a run
- WHEN a
/api/chatrequest fails parameter validation before execution - THEN the system SHALL NOT create a
diagnosis_run
Scenario: Chat session stores metadata only
- WHEN a Chat request completes
- THEN
chat_sessionSHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time - AND it SHALL NOT store full conversation message history
Requirement: Chat responses SHALL expose run identity
The system SHALL expose the current execution runId to clients that submit Chat requests.
Scenario: Chat response includes runId
- WHEN
/api/chatreturns a successful response - THEN the response SHALL include
sessionId - AND the response SHALL include
runIdfor the created diagnosis run
Scenario: Multi-turn Chat keeps one session and multiple runs
- WHEN two valid
/api/chatrequests use the samesessionId - THEN the system SHALL preserve multi-turn Redis context for that
sessionId - AND it SHALL persist two distinct
diagnosis_run.run_idvalues
Requirement: Trace details SHALL be scoped by run
The system SHALL write and read agent_step and tool_invocation rows using run_id as the execution boundary.
Scenario: Agent steps are recorded with runId
- WHEN an Agent model step is persisted during a diagnosis run
- THEN the
agent_steprow SHALL include the currentrun_id - AND it SHALL retain the current
session_id
Scenario: Tool invocations are recorded with runId
- WHEN an evidence tool invocation is persisted during a diagnosis run
- THEN the
tool_invocationrow SHALL include the currentrun_id - AND it SHALL retain the current
session_id
Scenario: Run metrics count only current run rows
- WHEN a diagnosis run completes
- THEN its step and tool counts SHALL be calculated from rows matching that
run_id - AND rows from other runs in the same
sessionIdSHALL NOT be counted
Requirement: Trace API SHALL require an exact run
The system SHALL require sessionId + runId for every diagnosis trace query and SHALL NOT infer a latest or historical run.
Scenario: Trace without runId is rejected
- WHEN a caller requests
GET /api/diagnosis/{sessionId}/tracewithoutrunId - THEN request validation SHALL reject the request
- AND the service SHALL NOT infer a run from current or historical data
Scenario: Trace with runId returns exact run
- WHEN a caller requests
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx - THEN the system SHALL validate that
runIdbelongs to the pathsessionId - AND it SHALL return only chat-session metadata, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
- AND it SHALL NOT expose a compatibility
sessionprojection
Scenario: Trace rejects run from another session
- WHEN a caller requests a
runIdthat belongs to a differentsessionId - THEN the system SHALL return an error instead of leaking trace data from the other session
Requirement: Session runs SHALL be listable without expanding trace details
The system SHALL provide a lightweight run-list API for a Chat Session.
Scenario: Run list returns summaries
- WHEN a caller requests
GET /api/chat/session/{sessionId}/runs - THEN the system SHALL return run summaries from
diagnosis_runinside the existing API response wrapper - AND each summary SHALL include
runId,sessionId,query,status,agentFlow,answerPreview,stepCount,toolCallCount,createdAt, andupdatedAt - AND the response SHALL NOT expand
agent_steportool_invocationdetail rows
Scenario: Run list handles session without runs
- WHEN a caller requests
GET /api/chat/session/{sessionId}/runsfor an existingchat_sessionwith no runs - THEN the system SHALL return a successful empty list
Scenario: Run list rejects missing session
- WHEN a caller requests
GET /api/chat/session/{sessionId}/runsfor a session that does not exist inchat_sessionordiagnosis_run - THEN the system SHALL use the existing not-found/error response behavior
Requirement: Feedback SHALL bind to diagnosis runs
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
Scenario: Feedback with runId updates specified run
- WHEN a feedback request includes
sessionIdandrunId - THEN the system SHALL validate that the run belongs to the session
- AND it SHALL update feedback on that run
- AND the response SHALL include the actual bound
runId
Scenario: Feedback without runId is rejected
- WHEN a feedback request includes
sessionIdbut omitsrunId - THEN the system SHALL reject the request
- AND it SHALL NOT bind feedback to any run
Scenario: Feedback rejects run from another session
- WHEN a feedback request includes a
runIdthat belongs to a differentsessionId - THEN the system SHALL return a failed feedback response instead of updating either run
Scenario: Useful feedback creates case from run
- WHEN feedback for a run is
useful - THEN the system SHALL create or reuse a
case_libraryrow using that run's query and answer - AND new automatic case data SHALL store
case_library.diagnosis_idas therun_id
Requirement: AIOps executions SHALL use run isolation
The system SHALL create and expose a diagnosis run for every valid /api/ai_ops execution.
Scenario: AIOps creates run
- WHEN
/api/ai_opsstarts a valid execution - THEN the system SHALL create a
diagnosis_runwithagent_flow=AI_OPS - AND AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that
run_id - AND AIOps rule evaluation SHALL be stored under
diagnosis_run.self_evaluation.aiops_rule_evaluationfor the current run
Scenario: AIOps SSE exposes runId
- WHEN
/api/ai_opsstreams response metadata to the caller - THEN the stream SHALL send a compatible metadata message before report content
- AND the SSE event name SHALL remain
message - AND the message type SHALL be
metadata - AND the metadata payload SHALL expose the resolved
sessionId - AND the metadata payload SHALL expose the created
runId - AND report content SHALL continue to use the existing content message shape
Requirement: Runtime SHALL use only run-based diagnosis storage
The system SHALL use chat_session and diagnosis_run as the only runtime diagnosis model and SHALL NOT read or write diagnosis_session.
Scenario: Runtime components are inspected
- WHEN Trace, Feedback, Evaluation, AIOps persistence, Gatekeeper, and case creation execute
- THEN they SHALL resolve data by
runId - AND no executable entity, repository, service fallback, or test SHALL depend on
diagnosis_session
Requirement: Demo and Trace UI SHALL support runId
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
Scenario: Demo script queries exact trace
- WHEN a demo script receives a
/api/chator/api/ai_opsresponse containingrunId - THEN it SHALL include
runIdwhen querying the Trace API
Scenario: Trace UI honors URL runId
- WHEN the Trace UI is opened with
?sessionId=...&runId=... - THEN it SHALL query
GET /api/diagnosis/{sessionId}/trace?runId=...
Requirement: Run isolation SHALL be verified against baselines
The change SHALL verify both runtime behavior and evaluation baseline impact.
Scenario: Multi-turn E2E proves run isolation
- WHEN E2E verification sends two valid Chat requests with the same
sessionId - THEN database inspection SHALL show two
diagnosis_runrows - AND each run SHALL have only its own step and tool rows when queried by
run_id - AND Redis session metadata SHALL still show multi-turn context continuity
Scenario: Baseline drift is checked
- WHEN verification is complete
- THEN the project SHALL run or explicitly evaluate the relevant baseline diff command
- AND any drift caused by run isolation SHALL be documented as expected or investigated as a regression
Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in diagnosis_run.orchestration_trace.
Scenario: Graph reaches Composer
- WHEN a complex Chat Graph terminates through Composer with a safe answer
- THEN the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- AND the summary SHALL be derived from the current Run's actual orchestration events
Scenario: Graph reaches handled Fallback
- WHEN a complex Chat Graph terminates through deterministic Fallback with a safe answer
- THEN the current DiagnosisRun SHALL store a non-empty orchestration trace with
degraded=true - AND the Run status SHALL be SUCCESS
Scenario: Unhandled execution fails after events exist
- WHEN an unhandled failure occurs after one or more real Graph events are available
- THEN the service SHALL best-effort persist a partial orchestration summary for the current failed run
- AND it SHALL NOT add a node or transition that did not occur
Scenario: Orchestration trace content is inspected
- WHEN orchestration trace JSON is serialized
- THEN it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- AND it SHALL NOT contain data owned by another run
Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as run.orchestrationTrace.
Scenario: Exact StateGraph run trace is queried
- WHEN a caller queries a successful new StateGraph Chat run
- THEN
run.orchestrationTraceSHALL be a non-empty parsed JSON object - AND the response SHALL NOT contain a compatibility
sessionprojection - AND no raw orchestration trace field SHALL be added
Scenario: Non-StateGraph run is queried
- WHEN the selected DiagnosisRun has null orchestration trace
- THEN
run.orchestrationTraceMAY be null - AND the service SHALL NOT synthesize events or read another run's trace
Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named orchestration_trace to diagnosis_run for this change.
Scenario: Migration is applied
- WHEN Flyway applies the stage 3 migration
- THEN existing DiagnosisRun rows SHALL remain valid without backfill
- AND no other table or column SHALL be changed by the stage 3 schema migration