Files
SuperBizAgent-java/openspec/changes/session-run-trace-isolation/tasks.md
T

5.7 KiB

1. Schema and Compatibility Foundation

  • 1.1 Add Flyway migration for chat_session and diagnosis_run with indexes for unique session_id, unique run_id, latest-run lookup, and run summary listing.
  • 1.2 Add nullable run_id columns to agent_step and tool_invocation with indexes for run-scoped trace queries.
  • 1.3 Backfill one compatibility diagnosis_run for each existing diagnosis_session row.
  • 1.4 Backfill historical agent_step.run_id and tool_invocation.run_id from the compatibility run for the same session_id.
  • 1.5 Add JPA entities and repositories for ChatSession and DiagnosisRun, including latest-run and run-id lookup methods.
  • 1.6 Add focused migration/repository verification that proves old data remains queryable and run lookup methods work.
  • 1.7 Phase 1 gate: run the smallest relevant test/build check, inspect DB migration behavior, update OpenSpec task status, archive phase evidence, and commit before starting Phase 2.

2. Chat Run Write Path

  • 2.1 Add a unified execution context that carries both sessionId and runId through Chat service, Agent hooks, and tool recording.
  • 2.2 Change valid /api/chat executions to create or update chat_session metadata and create one new diagnosis_run.
  • 2.3 Change AgentLoggingHook to write agent_step.run_id for Chat runs while retaining session_id.
  • 2.4 Change ToolInvocationRecorder and evidence tools to write tool_invocation.run_id for Chat runs while retaining session_id.
  • 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from diagnosis_session to the current diagnosis_run.
  • 2.6 Change ToolTraceSummaryService, ExecutorGatekeeperService, and EvaluationService Chat reads from session-scoped tool rows to run-scoped tool rows.
  • 2.7 Change /api/chat response DTO to include official runId.
  • 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
  • 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with scripts/query_mysql.py, review logs/, update task status, archive phase evidence, and commit before starting Phase 3.

3. Trace Read Path and Run Listing

  • 3.1 Change DiagnosisTraceService to resolve latest run by diagnosis_run.created_at DESC, id DESC when runId is omitted.
  • 3.2 Add exact trace lookup for sessionId + runId, including validation that the run belongs to the session.
  • 3.3 Change trace response DTOs to include resolved runId and run summary fields.
  • 3.4 Add GET /api/chat/session/{sessionId}/runs returning lightweight run summaries without expanding trace detail rows.
  • 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
  • 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
  • 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.

4. Feedback and Case Library Run Binding

  • 4.1 Change feedback request handling to prefer runId and validate run/session ownership.
  • 4.2 Implement legacy feedback fallback to latest run with observable fallbackToLatestRun=true and actual bound runId.
  • 4.3 Change feedback persistence to update diagnosis_run.feedback for new data.
  • 4.4 Change CaseLibraryService to create automatic cases from diagnosis_run.query and diagnosis_run.answer.
  • 4.5 Preserve transitional case-library semantics where old diagnosis_id values may be session_id and new automatic values are run_id.
  • 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
  • 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with scripts/query_mysql.py, update task status, archive phase evidence, and commit before starting Phase 5.

5. AIOps Run Isolation

  • 5.1 Change valid /api/ai_ops executions to create diagnosis_run with agent_flow=AI_OPS.
  • 5.2 Expose runId in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
  • 5.3 Propagate runId through AIOps Agent hooks and tool recording.
  • 5.4 Change AIOps final report, status, counts, and diagnosis_run.self_evaluation.aiops_rule_evaluation writes to the current run.
  • 5.5 Add tests for repeated AIOps executions with the same sessionId and run-scoped rule evaluation.
  • 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.

6. Demo, Trace UI, Documentation, and Verification

  • 6.1 Update demo scripts to read runId from Chat/AIOps responses and pass ?runId=... to Trace API.
  • 6.2 Update Trace UI to accept ?sessionId=...&runId=... and query exact trace when runId is present.
  • 6.3 Update MVP table and architecture docs for chat_session, diagnosis_run, run_id, and transitional case_library.diagnosis_id semantics.
  • 6.4 Run final same-session multi-turn E2E using Maven startup if needed; collect DB evidence through scripts/query_mysql.py and inspect logs/.
  • 6.5 Run or explicitly evaluate the relevant baseline diff command and document whether drift is expected or a regression.
  • 6.6 Final gate: ensure all OpenSpec tasks are checked, no new writes depend on diagnosis_session, phase evidence is archived, final commit is created, and the change is ready for OpenSpec archive.