5.4 KiB
5.4 KiB
Change: Session / Run / Trace Isolation
Problem
The current MVP uses the same sessionId for two different concepts:
- Redis
SessionContextkeeps multi-turn chat history for prompt context. - MySQL
diagnosis_session,agent_step, andtool_invocationpersist diagnosis trace data for replay, verification, feedback, and evaluation.
End-to-end verification showed that two /api/chat calls with the same sessionId correctly reuse Redis context, but MySQL trace data is mixed under the same key:
diagnosis_session.queryis overwritten by the second round.agent_stepandtool_invocationappend rows from both rounds under the samesession_id.- Trace, verifier/evaluation, and feedback can read cross-round evidence.
This makes a trace no longer represent one replayable diagnosis run.
Proposed Solution
Introduce a stable split between conversation state and execution state:
chat_session: conversation metadata keyed bysession_id.diagnosis_run: one execution/run keyed byrun_id, belonging to asession_id.agent_stepandtool_invocation: keep existing trace detail role, addrun_idwhile retainingsession_idfor compatibility and coarse filtering.
runId becomes an official API field:
/api/chatreturnssessionId + runId./api/ai_opsemits or returnsrunIdin the SSE-compatible protocol.GET /api/diagnosis/{sessionId}/tracedefaults to the latest run for compatibility.GET /api/diagnosis/{sessionId}/trace?runId=run-...returns the specified run after validating it belongs to the pathsessionId.- Feedback prefers
runId; missingrunIdtemporarily falls back to the latest run and returns bothfallbackToLatestRun=trueand the actual boundrunId.
Trace remains an aggregate view of diagnosis_run + agent_step + tool_invocation; this change does not introduce a separate diagnosis_trace or trace_event table.
Scope
- Add Flyway migrations and JPA entities/repositories for
chat_sessionanddiagnosis_run. - Add nullable
run_idtoagent_stepandtool_invocation, backfill historical data, then switch new writes to require run context. - Move new Chat writes from
diagnosis_sessiontochat_session + diagnosis_run. - Update trace reads to resolve latest run or specified run.
- Add lightweight run list API:
GET /api/chat/session/{sessionId}/runs. - Update feedback and case-library creation to bind new data to
run_id. - Update AIOps to create and expose
runIdbefore this change is considered production complete. - Update demo scripts and Trace UI with minimal
runIdsupport. - Update relevant MVP table and architecture documentation.
- Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under
logs/, database queries viascripts/query_mysql.py, and baseline drift checks.
Non-Goals
- Do not add
diagnosis_traceortrace_eventin this change. - Do not implement a full run-list UI.
- Do not remove the historical
diagnosis_sessiontable in this change. - Do not change the Redis conversation window strategy.
- Do not persist full conversation history in MySQL;
chat_sessionstores metadata only. - Do not split historical mixed traces into true historical runs when the original run boundary is unavailable.
Devflow Context Constraints
session-storageestablished the current trace tables and decided thatsessionIdis propagated throughRunnableConfig.metadatawithSessionContextHolderas a fallback for tools.confidence-feedbackestablished that useful feedback createscase_libraryfromDiagnosisSession.answer, and thatfeedbackdoes not change executionstatus.mvp-demo-trace-acceptanceestablishedGET /api/diagnosis/{sessionId}/traceas a read-only endpoint and demo scripts as part of the observable story.aiops-traceable-diagnosis-entryestablished/api/ai_opsas a traceable SSE entry point and madesessionIdvisible to callers.data-model.mdandsession-trace-lifecycle.mddescribe the current model asdiagnosis_session + agent_step + tool_invocation, and list run id as a known follow-up.devflow/glossary/CONTEXT.mdnow defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks.
Interface Impact
Level: L4 database/API contract migration with compatibility behavior.
- New API response field:
runId. - New query parameter:
GET /api/diagnosis/{sessionId}/trace?runId=.... - New API:
GET /api/chat/session/{sessionId}/runs. - Feedback request gains optional/preferred
runId. - Database contract changes include new tables and new
run_idcolumns. - Old callers that only pass
sessionIdremain compatible by binding to latest run, but this fallback must be observable.
Risks
- Historical data has no true per-round boundary; backfill can only create compatibility runs from existing
diagnosis_sessionrows. - Context propagation through Agent hooks and tools is easy to break because it currently combines
RunnableConfig.metadataandSessionContextHolder. - Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
case_library.diagnosis_idbecomes transitional: old data may containsession_id, new data containsrun_id.