# session-run-trace-isolation Specification ## Purpose Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance. ## Requirements ### Requirement: Conversation metadata SHALL be separated from diagnosis runs The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`. #### Scenario: Valid Chat execution creates session metadata and a run - **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId` - **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId` - **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id` - **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId` #### Scenario: Invalid Chat request does not create a run - **WHEN** a `/api/chat` request fails parameter validation before execution - **THEN** the system SHALL NOT create a `diagnosis_run` #### Scenario: Chat session stores metadata only - **WHEN** a Chat request completes - **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time - **AND** it SHALL NOT store full conversation message history ### Requirement: Chat responses SHALL expose run identity The system SHALL expose the current execution `runId` to clients that submit Chat requests. #### Scenario: Chat response includes runId - **WHEN** `/api/chat` returns a successful response - **THEN** the response SHALL include `sessionId` - **AND** the response SHALL include `runId` for the created diagnosis run #### Scenario: Multi-turn Chat keeps one session and multiple runs - **WHEN** two valid `/api/chat` requests use the same `sessionId` - **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId` - **AND** it SHALL persist two distinct `diagnosis_run.run_id` values ### Requirement: Trace details SHALL be scoped by run The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary. #### Scenario: Agent steps are recorded with runId - **WHEN** an Agent model step is persisted during a diagnosis run - **THEN** the `agent_step` row SHALL include the current `run_id` - **AND** it SHALL retain the current `session_id` #### Scenario: Tool invocations are recorded with runId - **WHEN** an evidence tool invocation is persisted during a diagnosis run - **THEN** the `tool_invocation` row SHALL include the current `run_id` - **AND** it SHALL retain the current `session_id` #### Scenario: Run metrics count only current run rows - **WHEN** a diagnosis run completes - **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id` - **AND** rows from other runs in the same `sessionId` SHALL NOT be counted ### Requirement: Trace API SHALL support latest-run and exact-run queries The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay. #### Scenario: Trace without runId resolves latest run - **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId` - **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC` - **AND** the response SHALL include the resolved `runId` #### Scenario: Trace with runId returns exact run - **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx` - **THEN** the system SHALL validate that `runId` belongs to the path `sessionId` - **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run - **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run` #### Scenario: Trace rejects run from another session - **WHEN** a caller requests a `runId` that belongs to a different `sessionId` - **THEN** the system SHALL return an error instead of leaking trace data from the other session ### Requirement: Session runs SHALL be listable without expanding trace details The system SHALL provide a lightweight run-list API for a Chat Session. #### Scenario: Run list returns summaries - **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` - **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper - **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt` - **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows #### Scenario: Run list handles session without runs - **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs - **THEN** the system SHALL return a successful empty list #### Scenario: Run list rejects missing session - **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run` - **THEN** the system SHALL use the existing not-found/error response behavior ### Requirement: Feedback SHALL bind to diagnosis runs The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session. #### Scenario: Feedback with runId updates specified run - **WHEN** a feedback request includes `sessionId` and `runId` - **THEN** the system SHALL validate that the run belongs to the session - **AND** it SHALL update feedback on that run - **AND** the response SHALL include the actual bound `runId` - **AND** the response SHALL include `fallbackToLatestRun=false` #### Scenario: Feedback without runId falls back observably - **WHEN** a legacy feedback request includes `sessionId` but omits `runId` - **AND** at least one `diagnosis_run` exists for that session - **THEN** the system SHALL bind feedback to the latest run for that session - **AND** the response SHALL include `fallbackToLatestRun=true` - **AND** the response SHALL include the actual bound `runId` #### Scenario: Historical feedback without run-backed data remains compatible - **WHEN** a legacy feedback request includes `sessionId` but omits `runId` - **AND** no `diagnosis_run` exists for that session - **AND** a historical `diagnosis_session` row exists for that session - **THEN** the system MAY bind feedback to the historical session row for migration compatibility - **AND** the response SHALL NOT claim latest-run fallback - **AND** the response MAY omit `runId` #### Scenario: Feedback rejects run from another session - **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId` - **THEN** the system SHALL return a failed feedback response instead of updating either run #### Scenario: Useful feedback creates case from run - **WHEN** feedback for a run is `useful` - **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer - **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id` ### Requirement: AIOps executions SHALL use run isolation The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution. #### Scenario: AIOps creates run - **WHEN** `/api/ai_ops` starts a valid execution - **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS` - **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id` - **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run #### Scenario: AIOps SSE exposes runId - **WHEN** `/api/ai_ops` streams response metadata to the caller - **THEN** the stream SHALL send a compatible metadata message before report content - **AND** the SSE event name SHALL remain `message` - **AND** the message type SHALL be `metadata` - **AND** the metadata payload SHALL expose the resolved `sessionId` - **AND** the metadata payload SHALL expose the created `runId` - **AND** report content SHALL continue to use the existing content message shape ### Requirement: Migration SHALL preserve historical trace access The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table. #### Scenario: Historical session gets compatibility run - **WHEN** migration runs on an existing `diagnosis_session` row - **THEN** the system SHALL create a compatible `diagnosis_run` row for that session - **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible #### Scenario: Old table is retained - **WHEN** the migration completes - **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback - **AND** new execution writes SHALL target `chat_session` and `diagnosis_run` ### Requirement: Demo and Trace UI SHALL support runId The demo tooling and Trace UI SHALL support minimal run-aware workflows. #### Scenario: Demo script queries exact trace - **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId` - **THEN** it SHALL include `runId` when querying the Trace API #### Scenario: Trace UI honors URL runId - **WHEN** the Trace UI is opened with `?sessionId=...&runId=...` - **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...` ### Requirement: Run isolation SHALL be verified against baselines The change SHALL verify both runtime behavior and evaluation baseline impact. #### Scenario: Multi-turn E2E proves run isolation - **WHEN** E2E verification sends two valid Chat requests with the same `sessionId` - **THEN** database inspection SHALL show two `diagnosis_run` rows - **AND** each run SHALL have only its own step and tool rows when queried by `run_id` - **AND** Redis session metadata SHALL still show multi-turn context continuity #### Scenario: Baseline drift is checked - **WHEN** verification is complete - **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command - **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression ### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`. #### Scenario: Graph reaches Composer - **WHEN** a complex Chat Graph terminates through Composer with a safe answer - **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count - **AND** the summary SHALL be derived from the current Run's actual orchestration events #### Scenario: Graph reaches handled Fallback - **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer - **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true` - **AND** the Run status SHALL be SUCCESS #### Scenario: Unhandled execution fails after events exist - **WHEN** an unhandled failure occurs after one or more real Graph events are available - **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run - **AND** it SHALL NOT add a node or transition that did not occur #### Scenario: Orchestration trace content is inspected - **WHEN** orchestration trace JSON is serialized - **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots - **AND** it SHALL NOT contain data owned by another run ### Requirement: Trace API SHALL expose orchestration trace only on the run object The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`. #### Scenario: Exact StateGraph run trace is queried - **WHEN** a caller queries a successful new StateGraph Chat run - **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object - **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field - **AND** no raw orchestration trace field SHALL be added #### Scenario: Historical or non-StateGraph run is queried - **WHEN** the selected DiagnosisRun has null orchestration trace - **THEN** `run.orchestrationTrace` MAY be null - **AND** the service SHALL NOT synthesize historical events or read another run's trace ### Requirement: Orchestration trace migration SHALL be additive and nullable The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change. #### Scenario: Migration is applied - **WHEN** Flyway applies the stage 3 migration - **THEN** existing DiagnosisRun rows SHALL remain valid without backfill - **AND** no other table or column SHALL be changed by the stage 3 schema migration