# Change: Session / Run / Trace Isolation ## Problem The current MVP uses the same `sessionId` for two different concepts: - Redis `SessionContext` keeps multi-turn chat history for prompt context. - MySQL `diagnosis_session`, `agent_step`, and `tool_invocation` persist diagnosis trace data for replay, verification, feedback, and evaluation. End-to-end verification showed that two `/api/chat` calls with the same `sessionId` correctly reuse Redis context, but MySQL trace data is mixed under the same key: - `diagnosis_session.query` is overwritten by the second round. - `agent_step` and `tool_invocation` append rows from both rounds under the same `session_id`. - Trace, verifier/evaluation, and feedback can read cross-round evidence. This makes a trace no longer represent one replayable diagnosis run. ## Proposed Solution Introduce a stable split between conversation state and execution state: - `chat_session`: conversation metadata keyed by `session_id`. - `diagnosis_run`: one execution/run keyed by `run_id`, belonging to a `session_id`. - `agent_step` and `tool_invocation`: keep existing trace detail role, add `run_id` while retaining `session_id` for compatibility and coarse filtering. `runId` becomes an official API field: - `/api/chat` returns `sessionId + runId`. - `/api/ai_ops` emits an SSE-compatible metadata message containing `sessionId` and `runId` before report content. - `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility. - `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`. - Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`. Trace remains an aggregate view of `diagnosis_run + agent_step + tool_invocation`; this change does not introduce a separate `diagnosis_trace` or `trace_event` table. ## Scope - Add Flyway migrations and JPA entities/repositories for `chat_session` and `diagnosis_run`. - Add nullable `run_id` to `agent_step` and `tool_invocation`, backfill historical data, then switch new writes to require run context. - Move new Chat writes from `diagnosis_session` to `chat_session + diagnosis_run`. - Update trace reads to resolve latest run or specified run. - Add lightweight run list API: `GET /api/chat/session/{sessionId}/runs`. - Update feedback and case-library creation to bind new data to `run_id`. - Update AIOps to create and expose `runId` before this change is considered production complete. - Update demo scripts and Trace UI with minimal `runId` support. - Update relevant MVP table and architecture documentation. - Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under `logs/`, database queries via `scripts/query_mysql.py`, and baseline drift checks. ## Non-Goals - Do not add `diagnosis_trace` or `trace_event` in this change. - Do not implement a full run-list UI. - Do not remove the historical `diagnosis_session` table in this change. - Do not change the Redis conversation window strategy. - Do not persist full conversation history in MySQL; `chat_session` stores metadata only. - Do not split historical mixed traces into true historical runs when the original run boundary is unavailable. ## Devflow Context Constraints - `session-storage` established the current trace tables and decided that `sessionId` is propagated through `RunnableConfig.metadata` with `SessionContextHolder` as a fallback for tools. - `confidence-feedback` established that useful feedback creates `case_library` from `DiagnosisSession.answer`, and that `feedback` does not change execution `status`. - `mvp-demo-trace-acceptance` established `GET /api/diagnosis/{sessionId}/trace` as a read-only endpoint and demo scripts as part of the observable story. - `aiops-traceable-diagnosis-entry` established `/api/ai_ops` as a traceable SSE entry point and made `sessionId` visible to callers. - `data-model.md` and `session-trace-lifecycle.md` describe the current model as `diagnosis_session + agent_step + tool_invocation`, and list run id as a known follow-up. - `devflow/glossary/CONTEXT.md` now defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks. ## Interface Impact Level: L4 database/API contract migration with compatibility behavior. - New API response field: `runId`. - New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`. - New API: `GET /api/chat/session/{sessionId}/runs`. - Feedback request gains optional/preferred `runId`. - Database contract changes include new tables and new `run_id` columns. - Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable. ## Risks - Historical data has no true per-round boundary; backfill can only create compatibility runs from existing `diagnosis_session` rows. - Context propagation through Agent hooks and tools is easy to break because it currently combines `RunnableConfig.metadata` and `SessionContextHolder`. - Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted. - Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive. - `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.