docs(openspec): archive session run isolation

This commit is contained in:
zhuyongxin
2026-07-10 22:52:39 +08:00
parent f9df94377b
commit 3578709896
21 changed files with 511 additions and 45 deletions
@@ -0,0 +1 @@
archive-ready
@@ -0,0 +1,3 @@
committed: true
change: session-run-trace-isolation
validated: openspec validate session-run-trace-isolation --strict
@@ -0,0 +1,191 @@
# Decisions: session-run-trace-isolation
## sm-flow State
- Checkpoint: Apply / Phase 6 ready
- Scale: complex
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
- Change slug: `session-run-trace-isolation`
## Entry Summary
Problem: the same `sessionId` currently represents both multi-turn conversation context and one persisted diagnosis trace. Multi-turn Chat E2E proved that Redis context behaves correctly, but MySQL trace rows from different rounds are mixed under one `session_id`.
Expected result: split session metadata from per-run execution state, expose `runId` as the official run identifier, and make trace, feedback, evaluation, demo scripts, Trace UI, and AIOps read/write by run.
Known modules: Flyway/JPA entities/repositories, `ChatService`, `ChatController`, `AiOpsService`, `DiagnosisTraceService`, `EvaluationService`, `FeedbackService`, `CaseLibraryService`, `AgentLoggingHook`, `ToolInvocationRecorder`, `SessionContextHolder`, demo scripts, static Trace UI, MVP docs.
## Context Collection
`devflow/index.md`: relevant entries found.
Relevant historical decisions:
- `session-storage`: current trace persistence is `diagnosis_session + agent_step + tool_invocation`; `sessionId` propagation uses `RunnableConfig.metadata` with `SessionContextHolder` fallback for tools; AIOps records child agents only.
- `confidence-feedback`: `FeedbackService` writes feedback, useful feedback creates `case_library`, and feedback must not change execution status.
- `mvp-demo-trace-acceptance`: Trace API is read-only and demo artifacts/scripts are part of acceptance.
- `aiops-traceable-diagnosis-entry`: `/api/ai_ops` is an SSE entry point that accepts optional alert payload and exposes `sessionId`.
- `data-model.md`: current `case_library.diagnosis_id` maps to `diagnosis_session.session_id`; this must be treated as legacy data after the change.
- `session-trace-lifecycle.md`: current docs already list run id as a future enhancement for multi-run sessions.
OpenSpec inputs that must be carried forward:
- Current specs mention `diagnosis_session` directly in trace, verifier, evidence, and demo requirements; new specs must either supersede or preserve compatibility for those requirements.
- `GET /api/diagnosis/{sessionId}/trace` must remain read-only.
- Baseline/eval checks must expect changed counts only when the change is explained by run isolation.
## Question Pool
| # | Question | Mode | Status |
|---|---|---|---|
| Q1 | Should this be an added field on existing `diagnosis_session`, or split tables? | user-interview | confirmed: split `chat_session` and `diagnosis_run`, keep existing trace detail tables |
| Q2 | What is metadata in `chat_session`? | user-interview | confirmed: session directory fields only, not full message body |
| Q3 | Is "message body" the full conversation history? | user-interview | confirmed: full conversation history stays in Redis `SessionContext.messageHistory` for now |
| Q4 | Should `runId` be an official API field? | user-interview | confirmed: yes |
| Q5 | How should trace work without `runId`? | user-interview | confirmed: default to latest run for compatibility |
| Q6 | Should Feedback without `runId` fail or fall back? | user-interview | confirmed: short-term fallback to latest run, long-term may tighten |
| Q7 | Should simple Chat Q&A create a run? | user-interview | confirmed: yes, every valid `/api/chat` execution creates a run |
| Q8 | Should AIOps be included? | user-interview | confirmed: yes, same run isolation semantics |
| Q9 | Should there be a separate trace table? | user-interview | confirmed: no, current `agent_step` and `tool_invocation` are enough for this phase |
| Q10 | What should happen to old `diagnosis_session`? | user-interview | confirmed: keep it for history/rollback, new code stops writing it after migration |
| Q11 | Does the issue need demo/Trace UI support? | user-interview | confirmed: yes, minimal `runId` support |
| Q12 | What extra risks were found by document review? | evidence-driven | reported and patched into ISS-010 |
## Evidence-Driven Findings
- Code evidence: `SessionContext` contains `messageHistory` and `getMessagePairCount()`, supporting the decision that Redis holds hot conversation history while MySQL stores auditable per-run query/answer.
- Code evidence: `CaseLibraryService.createFromSession` currently deduplicates by `session.getSessionId()` and maps answer/query from `DiagnosisSession`; this must change for new run-based data.
- Documentation evidence: `mvp/architecture/data-model.md` states `case_library.diagnosis_id = diagnosis_session.session_id`; this becomes transitional legacy semantics.
- Documentation evidence: existing trace OpenSpec requires `GET /api/diagnosis/{sessionId}/trace` to be read-only; run resolution must preserve that invariant.
- E2E evidence from ISS-010: two Chat rounds with the same `sessionId` resulted in one overwritten `diagnosis_session` row and mixed step/tool rows.
## Confirmed Decisions
- `runId` format: `run-` + full UUID.
- Latest run ordering: `diagnosis_run.created_at DESC, id DESC`, not `updated_at`.
- `chat_session` stores metadata: `session_id`, `status`, `message_pair_count`, `created_at`, `last_active_at`, `expires_at`.
- Per-run long-term audit stores `query` and `answer` in `diagnosis_run`.
- `agent_step` and `tool_invocation` retain `session_id` and add `run_id`.
- Missing feedback `runId` returns `fallbackToLatestRun=true` plus actual bound `runId`.
- Historical mixed data is not split into multiple true runs.
## OpenSpec Backwrite Log
- Created `proposal.md` with problem, scope, non-goals, context constraints, interface impact, and risks.
- ISS-010 patched to clarify phase boundary, migration order, AIOps same-change requirement, feedback fallback response, case-library transitional semantics, and Redis/MySQL TTL boundary.
- Glossary patched to include `SessionContext.messageHistory` and its persistence boundary.
- Created `design.md`, `specs/session-run-trace-isolation/spec.md`, `specs/mvp-demo-trace-acceptance/spec.md`, and `tasks.md`.
- Architecture audit found that `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` also read tool rows by `sessionId`; Phase 2 tasks were updated to cover run-scoped reads.
## Cross-Artifact Alignment
| Check | Result | Evidence |
|---|---|---|
| issue/context -> proposal | aligned | proposal carries E2E problem, split-table solution, compatibility, AIOps, feedback, demo/UI, and baseline scope |
| proposal -> design | aligned | design records data model, API impact, migration plan, rollback, and key decisions |
| design -> specs | aligned | specs cover session/run split, Chat, trace, feedback, AIOps, migration, demo/UI, E2E, and baseline behavior |
| specs -> tasks | aligned | tasks implement schema, Chat write path, trace reads, feedback/case, AIOps, demo/UI/docs, and verification gates |
## Architecture Audit
Input to output chain:
```text
Chat/AIOps request
-> ChatController / AIOps endpoint
-> ChatService / AiOpsService
-> execution context(sessionId, runId)
-> AgentLoggingHook -> agent_step
-> ToolInvocationRecorder -> tool_invocation
-> verifier/gatekeeper/evaluation summary reads
-> diagnosis_run answer/self_evaluation/status/counts
-> Trace API / Feedback / CaseLibrary / Demo / Trace UI
```
Audit conclusions:
- Data ownership is clearer with `chat_session` owning conversation metadata and `diagnosis_run` owning execution state; `agent_step` and `tool_invocation` remain trace details owned by one run.
- The highest coupling risk is execution-context propagation because hooks and tools currently use both `RunnableConfig.metadata` and `SessionContextHolder`.
- The main read-path risk is missing one of the session-scoped consumers (`ToolTraceSummaryService`, `ExecutorGatekeeperService`, `EvaluationService`, trace, feedback, case creation).
- Migration is additive and rollback-friendly until constraints are tightened; historical mixed data must be treated as compatibility data.
- AIOps must complete before final archive because otherwise the system would still have one production entry point with mixed trace semantics.
## Commit Gate Preflight
- Interface impact level: L4.
- Required artifacts: proposal, design, specs, tasks.
- Strict OpenSpec validation: passed with `openspec validate session-run-trace-isolation --strict`.
- `.committed` marker: created after successful commit gate.
## Pre-apply Research
- Reference migrations: `V005__create_session_storage.sql`, `V008__add_answer_to_diagnosis_session.sql`, `V010__add_relevance_level_to_tool_invocation.sql`.
- Entity style: JPA entities use Lombok `@Data`, `@Builder`, `@NoArgsConstructor`, `@AllArgsConstructor`, `@PrePersist`, and `@PreUpdate` where timestamps need maintenance.
- Repository style: Spring Data JPA repository interfaces with derived query methods returning `Optional<T>` or `List<T>`.
- Test style: repository tests use `@DataJpaTest`, `@AutoConfigureTestDatabase(replace = NONE)`, Flyway enabled, and `ddl-auto=validate`.
## Phase 1 Apply Notes
- Implemented additive migration `V011__add_session_run_isolation.sql`.
- Implemented `ChatSession` / `DiagnosisRun` entities and repositories.
- Added nullable `runId` fields to `AgentStep` and `ToolInvocation`.
- Added run-scoped repository methods for step/tool lookup and counts.
- Added `DiagnosisRunRepositoryTest`.
- Verification evidence is recorded in `phase-1-evidence.md`.
- Historical DB inspection found one orphan `tool_invocation` row without a matching `diagnosis_session`; it remains `run_id = NULL` because no reliable compatibility run can be inferred.
## Document Review Follow-up
- Clarified that Chat creates runs for the effective resolved `sessionId`, including requests where the server generates a session id.
- Tightened legacy feedback fallback so `fallbackToLatestRun=true` and the actual bound `runId` are response fields, not log-only evidence.
- Clarified AIOps SSE compatibility: emit a metadata message containing `sessionId` and `runId` before report content while preserving the existing content stream shape.
- Clarified that run/session ownership is enforced by service-layer validation and indexed lookup in this change; database foreign keys are intentionally deferred to preserve compatibility with historical orphan detail rows and rollback.
- Clarified that `chat_session.expires_at` is nullable directory metadata / best-effort TTL snapshot, not mandatory persisted conversation history.
- Clarified document review findings before continuing Phase 2: OpenSpec task phases are authoritative over the older active issue phase sketch, and AIOps rule evaluation is stored in `diagnosis_run.self_evaluation.aiops_rule_evaluation`, not a separate table.
## Document Review Follow-up Before Phase 3 Gate
- Clarified AIOps SSE compatibility before Phase 5: keep SSE event name `message`, emit a JSON `SseMessage` with `type=metadata`, and preserve existing content message shape for report streaming.
- Clarified Feedback API before Phase 4: request `runId` is preferred, response always includes the bound `runId` and `fallbackToLatestRun`, and wrong-session run binding uses the existing failed feedback response path.
- Clarified run-list API before Phase 3 gate: `GET /api/chat/session/{sessionId}/runs` returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
## Phase 2 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
- Implemented unified Chat execution context carrying `sessionId` and `runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
- Switched valid Chat writes to create/update `chat_session` metadata and create one `diagnosis_run` per request.
- Switched Chat run completion, failure, self-evaluation, metrics, verifier support reads, gatekeeper validation, and evidence scoring to run-scoped data.
- Added `/api/chat` response `runId` and focused tests for valid run creation, invalid request no-run behavior, run-scoped trace consumers, and same-session multi-turn run creation.
- Phase 2 gate evidence is recorded in `phase-2-evidence.md`.
## Phase 3 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. The committed OpenSpec remained the execution source; a document review follow-up tightened DTO/SSE/listing contracts before the Phase 3 gate.
- Implemented latest-run trace resolution using `diagnosis_run.created_at DESC, id DESC`.
- Implemented exact trace lookup for `sessionId + runId` with run/session ownership validation.
- Changed trace details to read `agent_step` and `tool_invocation` by `run_id`, while retaining a legacy `diagnosis_session` fallback for historical compatibility.
- Added trace response fields for resolved `runId`, chat session metadata, run metadata, and per-row `runId`.
- Added lightweight `GET /api/chat/session/{sessionId}/runs` backed by `diagnosis_run` summaries.
- Added focused tests for latest trace, exact first trace, exact second trace, wrong-session rejection, missing session, read-only trace behavior, run listing, and legacy fallback.
- Phase 3 E2E used Maven startup with profile `mvp-demo`; evidence is recorded in `phase-3-evidence.md`.
## Phase 4 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were still unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, focused tests, API E2E, DB inspection, and logs.
- Added `FeedbackRequest.runId` and `FeedbackResponse.runId/fallbackToLatestRun`.
- Changed `FeedbackController` to pass `runId` through to `FeedbackService`.
- Changed `FeedbackService` so new feedback prefers exact `sessionId + runId`, validates ownership, falls back to latest run when `runId` is omitted, and writes new feedback to `diagnosis_run.feedback`.
- Preserved a legacy `DiagnosisSession` fallback only when no `diagnosis_run` exists, so old data can still receive feedback during the migration window.
- Added `CaseLibraryService.createFromRun`, using `diagnosis_run.run_id` as the new automatic `case_library.diagnosis_id`; `createFromSession` remains the legacy session-id path.
- Phase 4 evidence is recorded in `phase-4-evidence.md`.
- Document review follow-up: clarified that the legacy `DiagnosisSession` feedback path only applies when no `diagnosis_run` exists for the session. It returns no bound `runId` and is not the same as latest-run fallback.
## Phase 5 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools remain unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, dependency method inspection with `javap`, focused tests, E2E, DB inspection, and logs.
- Changed `/api/ai_ops` to allocate a `runId` before execution and emit a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` on SSE event name `message`.
- Changed `AiOpsService` to create one `diagnosis_run` with `agent_flow=AI_OPS` for each valid execution instead of writing new execution state to `diagnosis_session`.
- Changed AIOps execution context propagation to pass `sessionId/runId` through both `RunnableConfig.metadata` and `SessionContextHolder`.
- Changed AIOps final report, metrics, status, and `aiops_rule_evaluation` persistence to update the current `diagnosis_run`.
- Preserved `persistFinalReport(sessionId, finalReport, request)` as a historical `DiagnosisSession` compatibility path.
- Added focused tests for run-scoped final report evaluation, run-scoped metrics, distinct runs under the same AIOps session, and SSE metadata shape.
@@ -0,0 +1,145 @@
## Context
The current MVP persists diagnosis observability through `diagnosis_session`, `agent_step`, and `tool_invocation`. That model works for a single diagnosis per `sessionId`, but it conflates two lifecycles once a caller reuses the same `sessionId` for multi-turn conversation:
- conversation state: Redis `SessionContext.messageHistory` and session metadata;
- execution state: one diagnosis answer, trace, self-evaluation, and feedback target.
The E2E evidence in ISS-010 showed that Redis correctly preserved multi-turn context while MySQL mixed both rounds under the same `session_id`. This breaks trace replay, verifier/evaluation scoping, feedback targeting, and case-library provenance.
Constraints from existing work:
- `GET /api/diagnosis/{sessionId}/trace` is a read-only MVP/demo contract and must remain compatible.
- Current hooks and tools propagate `sessionId` through `RunnableConfig.metadata` and `SessionContextHolder`; `runId` must follow the same execution context boundary.
- Feedback currently writes `DiagnosisSession.feedback` and creates `case_library` from `DiagnosisSession.answer`.
- AIOps is a first-class traceable entry point and cannot be left permanently on the old mixed-run model.
## Goals / Non-Goals
**Goals:**
- Split conversation metadata from per-execution diagnosis state.
- Introduce `runId` as the official identifier for one replayable diagnosis execution.
- Preserve old `sessionId`-only callers by resolving the latest run where possible.
- Scope trace, evaluation, feedback, case creation, demo scripts, and Trace UI by run.
- Migrate historical data into compatibility runs without deleting the old table.
- Include Chat and AIOps in the same release-level change.
- Verify behavior with focused tests, E2E when needed, DB inspection, logs, and baseline drift checks.
**Non-Goals:**
- Do not persist full chat history in MySQL.
- Do not add `diagnosis_trace` or `trace_event`.
- Do not implement full run-list UI.
- Do not physically delete `diagnosis_session`.
- Do not attempt to reconstruct true historical round boundaries when only mixed `session_id` data exists.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Domain split | Add `chat_session` and `diagnosis_run` | Add `run_id` to `diagnosis_session` only | Separate tables keep conversation metadata and execution state from growing into one coupled table. |
| Trace detail storage | Reuse `agent_step` and `tool_invocation`, adding `run_id` | Add `diagnosis_trace` / `trace_event` | Existing detail tables already represent trace; isolation needs a run key, not a new event model. |
| API identity | `runId = "run-" + UUID` | Reuse short session id or DB id | Full UUID avoids collision and keeps external IDs independent from database internals. |
| Latest-run compatibility | `GET /api/diagnosis/{sessionId}/trace` resolves latest by `created_at DESC, id DESC` | Require `runId` immediately | Compatibility keeps existing demo/UI/scripts working while new clients migrate. `updated_at` is avoided because feedback/eval updates can reorder old runs. |
| Historical migration | Backfill one compatibility run per existing `diagnosis_session` | Try to split old mixed rows | Old rows do not contain reliable run boundaries. A compatibility run preserves auditability without inventing data. |
| Feedback fallback | Missing `runId` binds latest run and returns fallback metadata | Reject missing `runId` immediately | Short-term compatibility is needed for old clients; explicit fallback keeps ambiguity observable. |
| Case library provenance | New automatic cases store `diagnosis_id = run_id` | Add a new case-library column now | Existing column name can carry transitional provenance; docs and query logic must recognize old `session_id` and new `run_id`. |
| AIOps phase | Implement after Chat but before overall archive | Leave AIOps for a follow-up issue | AIOps is already a trace entry point; leaving it old-model would preserve the same bug on another endpoint. |
| Ownership integrity | Validate run/session ownership in application services; do not add database foreign keys in this change | Add foreign keys from `diagnosis_run`, `agent_step`, and `tool_invocation` | Existing historical/orphan compatibility data and rollback needs make additive, application-level validation safer for this release. |
## Data Model
```text
chat_session
id
session_id unique
status
message_pair_count
created_at
last_active_at
expires_at
diagnosis_run
id
run_id unique
session_id
query
status
agent_flow
answer
self_evaluation
feedback
total_duration_ms
total_token_count
step_count
tool_call_count
created_at
updated_at
agent_step
session_id
run_id
...
tool_invocation
session_id
run_id
...
```
`chat_session.expires_at` is nullable MySQL directory metadata and may be a best-effort Redis TTL snapshot when known. Redis TTL can expire `SessionContext.messageHistory`; persisted `diagnosis_run`, `agent_step`, and `tool_invocation` remain audit records.
Ownership between `chat_session`, `diagnosis_run`, `agent_step`, and `tool_invocation` is enforced by service-layer validation and indexed lookup in this change. The migration intentionally does not add database foreign keys so historical orphan trace detail rows and rollback paths remain compatible.
## API / Interface Impact
Interface level: L4.
- Database contract changes: new tables, new columns, backfill, indexes, and later non-null expectations for new writes.
- `/api/chat` response adds `runId`.
- `/api/ai_ops` SSE emits a compatible metadata message before report content. It keeps the existing SSE event name `message` and sends a JSON `SseMessage` with `type=metadata`; the metadata payload includes `sessionId` and `runId`. Report content continues to stream through the existing `type=content` message shape.
- Trace API accepts optional `runId`.
- Feedback request accepts preferred `runId`. For run-backed data, feedback response includes the actual bound `runId` and `fallbackToLatestRun`; wrong-session `runId`, missing session, and missing run use the existing failed feedback response path with HTTP 400 from `FeedbackController`. Historical `DiagnosisSession` fallback is retained only when no `diagnosis_run` exists for the session; that legacy path has no bound `runId` and is not treated as latest-run fallback.
- New run summary API: `GET /api/chat/session/{sessionId}/runs`, returned through the existing `ApiResponse<List<RunSummary>>` wrapper. A session with metadata but no runs returns an empty list; a missing session returns the existing not-found/error behavior.
Compatibility:
- Old `sessionId`-only trace and feedback calls bind to latest run when run-backed data exists.
- Historical feedback calls for sessions with no `diagnosis_run` may still bind to retained `diagnosis_session` data during the migration window.
- Old `diagnosis_session` is retained for rollback and historical comparison.
- New code must not keep writing new execution state into `diagnosis_session` after the write switch.
## Migration Plan
1. Add `chat_session` and `diagnosis_run`.
2. Add nullable `run_id` to `agent_step` and `tool_invocation`.
3. Backfill `diagnosis_run` from existing `diagnosis_session`.
4. Backfill old `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run.
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
6. Deploy repository and read-path compatibility.
7. Switch Chat write path to `chat_session + diagnosis_run`.
8. Switch Chat evaluation and verifier support reads to run-scoped data as part of the Chat write-path cutover.
9. Switch Trace run resolution and run listing.
10. Switch Feedback and case-library paths.
11. Switch AIOps write path, including `diagnosis_run.self_evaluation.aiops_rule_evaluation`.
12. Switch demo scripts, Trace UI, and MVP docs.
13. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
Rollback:
- Keep `diagnosis_session` intact during this change.
- Migrations are additive until constraints are tightened.
- If write-switch rollout fails, rollback code can read the retained old table while migrated compatibility rows remain harmless.
## Risks / Trade-offs
- [Risk] ThreadLocal/context propagation may miss `runId` in nested agent/tool calls. -> Mitigation: introduce a unified execution context carrying both IDs and test hook/tool recording.
- [Risk] Baseline metrics change because cross-round tool rows are no longer counted. -> Mitigation: run baseline diff and document expected drift.
- [Risk] Legacy `case_library.diagnosis_id` values are ambiguous. -> Mitigation: document transitional semantics and keep lookup logic aware of old `session_id` values.
- [Risk] AIOps SSE clients may not parse a new metadata event. -> Mitigation: add metadata in a compatible stream message and keep final report streaming behavior.
- [Risk] Historical mixed trace cannot be truly separated. -> Mitigation: call this out as compatibility data, not reconstructed truth.
## Open Questions
None blocking. Long-term tightening of missing feedback `runId` remains a follow-up decision after clients migrate.
@@ -0,0 +1,103 @@
# Phase 1 Evidence: Schema and Compatibility Foundation
## Scope Completed
- Added Flyway migration `V011__add_session_run_isolation.sql`.
- Added `chat_session` and `diagnosis_run` tables.
- Added nullable `run_id` columns and indexes to `agent_step` and `tool_invocation`.
- Backfilled compatibility runs from existing `diagnosis_session` rows.
- Backfilled existing step/tool rows where a matching `diagnosis_session.session_id` exists.
- Added `ChatSession`, `DiagnosisRun`, `ChatSessionRepository`, and `DiagnosisRunRepository`.
- Added run-scoped query/count methods to `AgentStepRepository` and `ToolInvocationRepository`.
- Added `DiagnosisRunRepositoryTest`.
## Verification
### Maven
```text
mvn -q "-Dtest=DiagnosisRunRepositoryTest" test
```
Result: passed.
```text
mvn -q "-Dtest=DiagnosisSessionRepositoryTest,AgentStepRepositoryTest,ToolInvocationRepositoryTest,DiagnosisRunRepositoryTest" test
```
Result: passed.
### Database Inspection
Checked with `scripts/query_mysql.py`.
```text
SELECT installed_rank, version, description, success
FROM flyway_schema_history
ORDER BY installed_rank DESC
LIMIT 5;
```
Relevant result:
```text
version = 011
description = add session run isolation
success = 1
```
```text
SELECT COUNT(*) AS chat_sessions FROM chat_session;
```
Result:
```text
chat_sessions = 83
```
```text
SELECT COUNT(*) AS diagnosis_runs,
SUM(CASE WHEN run_id IS NULL THEN 1 ELSE 0 END) AS null_run_ids
FROM diagnosis_run;
```
Result:
```text
diagnosis_runs = 83
null_run_ids = 0
```
```text
SELECT (SELECT COUNT(*) FROM agent_step WHERE run_id IS NULL) AS agent_step_null_run_id,
(SELECT COUNT(*) FROM tool_invocation WHERE run_id IS NULL) AS tool_invocation_null_run_id;
```
Result:
```text
agent_step_null_run_id = 0
tool_invocation_null_run_id = 1
```
The one remaining null `tool_invocation.run_id` is historical orphan data:
```text
id = 3
session_id = f9d2290d
tool_name = lookup_knowledge
created_at = 2026-06-26 14:17:35
```
It has no matching `diagnosis_session`, so V011 cannot safely infer a compatibility run. This matches the OpenSpec wording "when possible" for historical backfill.
### Logs
`logs/application.log` shows Hibernate using the new `run_id` fields in `agent_step` and `tool_invocation` inserts/selects during the repository verification.
## Notes
- Phase 1 intentionally does not switch Chat, Trace, Feedback, or AIOps write paths.
- New columns remain nullable until later phases switch new writes and verify coverage.
@@ -0,0 +1,129 @@
# Phase 2 Evidence: Chat Run Write Path
Date: 2026-07-10
## Scope
Phase 2 switched Chat writes from session-scoped execution state to run-scoped execution state:
- Chat creates/updates `chat_session` metadata.
- Each valid `/api/chat` execution creates one `diagnosis_run`.
- Chat execution context carries `sessionId + runId` through `RunnableConfig` and `SessionContextHolder`.
- `agent_step.run_id` and `tool_invocation.run_id` are written for Chat runs.
- Chat completion/failure/status/answer/self-evaluation/counts are written to `diagnosis_run`.
- verifier/gatekeeper/evaluation reads use run-scoped tool rows when `runId` is available.
- `/api/chat` response includes official `runId`.
## Static / Unit Verification
Commands:
```powershell
mvn -q clean test-compile
mvn -q "-Dtest=ChatControllerTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ExecutorGatekeeperServiceTest" test
openspec validate session-run-trace-isolation --strict
```
Result:
- `test-compile` passed.
- Focused Phase 2 tests passed.
- OpenSpec strict validation passed.
Focused coverage:
- valid Chat creates `chat_session` and `diagnosis_run`;
- invalid blank Chat request returns before `ChatService`, so no run is created;
- same `sessionId` across two Chat turns creates two distinct `runId` values;
- `ToolInvocationRecorder` copies `runId` from execution context;
- verifier trace summary reads by run;
- gatekeeper validates by run;
- evaluation writes rule evaluation to `diagnosis_run.self_evaluation`.
## E2E Verification
Startup command:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Process stdout/stderr:
- `target/e2e/phase2-mvn-20260710-185102.out.log`
- `target/e2e/phase2-mvn-20260710-185102.err.log`
Primary E2E session:
```text
sessionId = e2e-phase2-codex-20260710-1856
round 1 runId = run-24b6f04c-94a0-43cb-b94f-4f0141f9050d
round 2 runId = run-a7a2be77-697f-495a-a779-91afd1d8589c
```
HTTP evidence:
- `target/e2e/phase2-utf8-request-round1.json`
- `target/e2e/phase2-utf8-response-round1.json`
- `target/e2e/phase2-utf8-request-round2.json`
- `target/e2e/phase2-utf8-response-round2.json`
- both Chat responses returned `code=200`, `data.success=true`, the same `sessionId`, and distinct `runId` values.
Redis/session continuity evidence:
- `target/e2e/phase2-utf8-chat-session-response.json`
- response returned `messagePairCount=2`.
- logs show the second request entered with `会话历史消息对数: 1` and completed with `当前消息对数: 2`.
## Database Inspection
DB inspection used `scripts/query_mysql.py`.
Saved query outputs:
- `target/e2e/phase2-utf8-db-runs.txt`
- `target/e2e/phase2-utf8-db-chat-session.txt`
- `target/e2e/phase2-utf8-db-agent-steps.txt`
- `target/e2e/phase2-utf8-db-tool-invocations.txt`
- `target/e2e/phase2-utf8-db-missing-runid.txt`
Observed rows:
```text
diagnosis_run:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | SUCCESS | CHAT | step_count=8 | tool_call_count=12
run-a7a2be77-697f-495a-a779-91afd1d8589c | SUCCESS | CHAT | step_count=2 | tool_call_count=1
chat_session:
e2e-phase2-codex-20260710-1856 | ACTIVE | message_pair_count=2
agent_step grouped by run_id:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 8
run-a7a2be77-697f-495a-a779-91afd1d8589c | 2
tool_invocation grouped by run_id:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 12
run-a7a2be77-697f-495a-a779-91afd1d8589c | 1
missing run_id for this session:
agent_step = 0
tool_invocation = 0
```
## Log Review
Saved log excerpts:
- `target/e2e/phase2-utf8-application-log-excerpt.txt`
- `target/e2e/phase2-utf8-mvn-log-excerpt.txt`
Findings:
- Chat logs show second turn reused Redis history for the same session.
- `EvaluationService` wrote scoring results to both run ids.
- No E2E-specific application exception was observed for `e2e-phase2-codex-20260710-1856`.
- Earlier `/actuator/health` probes produced expected 500/no-resource noise because the actuator health endpoint is not exposed; the E2E readiness check used `/api/chat` instead.
## Notes
`codebase-retrieval` and LSP tools were not available in this environment. Call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
@@ -0,0 +1,136 @@
# Phase 3 Evidence: Trace Read Path and Run Listing
## Scope
Phase 3 implements run-scoped trace reads and lightweight run listing:
- `GET /api/diagnosis/{sessionId}/trace` resolves the latest run by `diagnosis_run.created_at DESC, id DESC`.
- `GET /api/diagnosis/{sessionId}/trace?runId=...` returns the exact run after validating run/session ownership.
- Trace responses include resolved run metadata and run-scoped step/tool rows.
- `GET /api/chat/session/{sessionId}/runs` returns lightweight run summaries without expanding trace detail rows.
## Verification Commands
- `mvn -q clean "-Dtest=DiagnosisTraceServiceTest" test`
- `mvn -q "-Dtest=DiagnosisTraceServiceTest,DiagnosisTraceEvaluatorTest,ChatControllerTest" test`
- `openspec validate session-run-trace-isolation --strict`
After the document review follow-up, OpenSpec strict validation was run again:
- `openspec validate session-run-trace-isolation --strict`
Result: passed.
## E2E Runtime
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured logs and artifacts:
- `target/e2e/phase3-mvn-20260710-191131.out.log`
- `target/e2e/phase3-mvn-20260710-191131.err.log`
- `target/e2e/phase3-request-round1.json`
- `target/e2e/phase3-response-round1.json`
- `target/e2e/phase3-request-round2.json`
- `target/e2e/phase3-response-round2.json`
- `target/e2e/phase3-trace-latest.json`
- `target/e2e/phase3-trace-first.json`
- `target/e2e/phase3-trace-second.json`
- `target/e2e/phase3-runs.json`
- `target/e2e/phase3-trace-wrong-session.json`
- `target/e2e/phase3-summary.json`
The E2E Maven process was stopped after evidence collection.
Note: `logs/application.log` was checked, but it did not contain the Phase 3 E2E session entries and its last write time was earlier than this E2E run. The Phase 3 runtime application logs were captured in the Maven stdout artifact above.
## E2E Summary
Session:
```text
sessionId = e2e-phase3-codex-20260710-1912
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
```
Observed behavior from `phase3-summary.json`:
```text
distinctRunIds = true
latestResolvedRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
firstTraceRunId = run-fdccbe21-e050-4f92-a741-a062aef59644
firstSteps = 9
firstTools = 13
secondTraceRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
secondSteps = 2
secondTools = 0
runListCount = 2
runListFirst = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
wrongSessionStatus = 400
```
Log evidence from `phase3-mvn-20260710-191131.out.log`:
- Round 1 request was received for `e2e-phase3-codex-20260710-1912`.
- Redis session was created for the first round and reused by the second round.
- Message pair count reached `2` after round 2.
- Latest trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
- Exact first trace request returned `run-fdccbe21-e050-4f92-a741-a062aef59644`.
- Exact second trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
- Wrong-session exact trace returned HTTP 400 with `runId does not belong to sessionId`.
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run` rows:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | SUCCESS | CHAT | step_count=2 | tool_call_count=0 | created_at=2026-07-10 19:15:40
run-fdccbe21-e050-4f92-a741-a062aef59644 | SUCCESS | CHAT | step_count=9 | tool_call_count=13 | created_at=2026-07-10 19:12:33
```
`agent_step` rows grouped by run:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | step_rows=2 | min_step=0 | max_step=1
run-fdccbe21-e050-4f92-a741-a062aef59644 | step_rows=9 | min_step=0 | max_step=4
```
`tool_invocation` rows grouped by run:
```text
run-fdccbe21-e050-4f92-a741-a062aef59644 | tool_rows=13
```
Missing run id checks for this session:
```text
agent_step missing run_id = 0
tool_invocation missing run_id = 0
```
`chat_session` metadata:
```text
session_id=e2e-phase3-codex-20260710-1912 | status=ACTIVE | message_pair_count=2
```
## Document Review Follow-up
Before closing Phase 3, the OpenSpec docs were tightened for upcoming phases:
- AIOps SSE metadata shape: keep SSE event name `message`, use JSON `SseMessage` with `type=metadata`, preserve existing content message shape.
- Feedback DTO contract: request `runId` is preferred; response includes bound `runId` and `fallbackToLatestRun`; wrong-session run binding fails instead of updating either run.
- Run-list API: returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
OpenSpec strict validation passed after these document changes.
## Conclusion
Phase 3 satisfies the run-scoped trace read and run-list contract. Same-session multi-turn E2E proves latest-run compatibility, exact-run replay, run-list ordering, run/session ownership rejection, and no missing `run_id` rows for new trace data.
@@ -0,0 +1,124 @@
# Phase 4 Evidence: Feedback and Case Library Run Binding
## Scope
Phase 4 implements run-scoped feedback and run-based automatic case creation:
- Feedback request accepts preferred `runId`.
- Feedback validates run/session ownership.
- Missing `runId` falls back to the latest run and returns `fallbackToLatestRun=true` plus the bound `runId`.
- New feedback persists to `diagnosis_run.feedback`.
- Useful feedback creates or reuses `case_library` from `diagnosis_run.query` and `diagnosis_run.answer`.
- New automatic `case_library.diagnosis_id` values store `run_id`; old `session_id` values remain supported through the legacy `DiagnosisSession` path.
## Focused Tests
Focused tests:
```text
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest" test
```
Result: passed.
Coverage:
- exact `sessionId + runId` feedback updates the specified run;
- missing `runId` binds to latest run and returns `fallbackToLatestRun=true`;
- wrong-session `runId` fails without saving feedback or case data;
- useful feedback creates a case from the run;
- case creation is idempotent by `diagnosis_id`;
- legacy `DiagnosisSession` case creation still stores `diagnosis_id=session_id`;
- `FeedbackController` passes `request.runId` to `FeedbackService`.
## E2E Feedback API Check
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured artifacts:
- `target/e2e/phase4-mvn-20260710-2016.out.log`
- `target/e2e/phase4-mvn-20260710-2016.err.log`
- `target/e2e/phase4-feedback-exact-run.json`
- `target/e2e/phase4-feedback-fallback-latest.json`
- `target/e2e/phase4-feedback-wrong-session.json`
- `target/e2e/phase4-summary.json`
The Maven process was stopped after evidence collection.
Inputs reused the Phase 3 E2E session:
```text
sessionId = e2e-phase3-codex-20260710-1912
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
```
API responses:
```text
exact run feedback:
status=200
runId=run-fdccbe21-e050-4f92-a741-a062aef59644
fallbackToLatestRun=false
caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
legacy fallback feedback:
status=200
runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94
fallbackToLatestRun=true
caseId=null
wrong-session feedback:
status=400
success=false
message=runId does not belong to sessionId
```
Log evidence from `phase4-mvn-20260710-2016.out.log`:
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-fdccbe21-e050-4f92-a741-a062aef59644, feedback=useful, fallbackToLatestRun=false, caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff`
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94, feedback=not_useful, fallbackToLatestRun=true, caseId=null`
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run.feedback`:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | feedback=not_useful | status=SUCCESS
run-fdccbe21-e050-4f92-a741-a062aef59644 | feedback=useful | status=SUCCESS
```
`case_library`:
```text
case_id=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
diagnosis_id=run-fdccbe21-e050-4f92-a741-a062aef59644
source_type=AUTO
title=支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
```
No automatic case was created for the fallback `not_useful` feedback on run2.
## Final Gate
The final Phase 4 gate was rerun after document review follow-up:
```text
mvn -q clean test-compile
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest,DiagnosisTraceServiceTest,ChatControllerTest" test
openspec validate session-run-trace-isolation --strict
git diff --check
```
Result: passed.
## Conclusion
Phase 4 satisfies run-scoped feedback, observable latest-run fallback, run-based useful case creation, and transitional old-data compatibility for `case_library.diagnosis_id`.
@@ -0,0 +1,157 @@
# Phase 5 Evidence: AIOps Run Isolation
## Scope
Phase 5 implements AIOps run isolation:
- `/api/ai_ops` allocates a `runId` before execution.
- The SSE stream keeps event name `message` and emits a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` before content.
- `AiOpsService` creates `diagnosis_run` rows with `agent_flow=AI_OPS`.
- AIOps Agent hooks and tool recording receive `sessionId + runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
- AIOps status, final report, metrics, and `self_evaluation.aiops_rule_evaluation` write to the current `diagnosis_run`.
- Historical `DiagnosisSession` final-report persistence remains available only through the legacy overload.
## Focused Tests
Focused tests:
```text
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest" test
```
Result: passed.
Coverage:
- AIOps final report updates `diagnosis_run.answer` and `diagnosis_run.self_evaluation`.
- Rule evaluation reads `tool_invocation` rows by `run_id`.
- AIOps run metrics count `agent_step` and `tool_invocation` rows by `run_id`.
- The same AIOps `sessionId` can create distinct `runId` values.
- SSE metadata uses `type=metadata` and carries `sessionId/runId`.
- Existing hook and tool-recorder tests cover `runId` propagation into `agent_step` and `tool_invocation`.
## E2E Runtime
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured artifacts:
- `target/e2e/phase5-mvn-20260710-204732.out.log`
- `target/e2e/phase5-mvn-20260710-204732.err.log`
- `target/e2e/phase5-aiops-request.json`
- `target/e2e/phase5-aiops-sse-response.txt`
- `target/e2e/phase5-mvn-20260710-205204.out.log`
- `target/e2e/phase5-mvn-20260710-205204.err.log`
- `target/e2e/phase5-aiops-request-2053.json`
- `target/e2e/phase5-aiops-sse-response-2053.txt`
- `target/e2e/phase5-aiops-trace-2053.json`
The Maven process was stopped after evidence collection.
First E2E attempt exposed an existing AIOps runtime integration bug:
```text
AI Ops 流程失败: mainAgent (ReactAgent) must be provided for supervisor agent
```
Diagnosis result:
- feedback loop: fixed `/api/ai_ops` request with `mvp-demo` profile;
- root cause: current `SupervisorAgent` dependency validates that `mainAgent(ReactAgent)` is set;
- fix: `AiOpsService.buildSupervisorAgent(...)` now sets Planner as `mainAgent` and Executor as sub-agent;
- regression coverage: `AiOpsServiceTest.buildSupervisorAgentSetsPlannerAsMainAgent`.
Successful E2E:
```text
sessionId = e2e-phase5-aiops-codex-20260710-2053
runId = run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
```
SSE response:
```text
event:message
data:{"type":"metadata","data":null,"sessionId":"e2e-phase5-aiops-codex-20260710-2053","runId":"run-84b8c02b-4d1e-4b24-ac88-222c88c8db26"}
...
event:message
data:{"type":"done","data":null,"sessionId":null,"runId":null}
```
Exact trace check:
```text
GET /api/diagnosis/e2e-phase5-aiops-codex-20260710-2053/trace?runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
```
Observed:
```text
code=200
runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
agentFlow=AI_OPS
summary.hasAiOpsRuleEvaluation=true
summary.persistedStepCount=1
summary.persistedToolCallCount=0
```
The E2E model did not call evidence tools. This is captured as rule-evaluation WARN rather than a run-isolation failure.
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run`:
```text
run_id=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
session_id=e2e-phase5-aiops-codex-20260710-2053
status=SUCCESS
agent_flow=AI_OPS
has_answer=1
has_aiops_eval=1
step_count=1
tool_call_count=0
```
`agent_step` grouped by run:
```text
run-84b8c02b-4d1e-4b24-ac88-222c88c8db26 | step_rows=1
```
`tool_invocation` grouped by run:
```text
(empty; the successful E2E did not invoke evidence tools)
```
Log evidence from `phase5-mvn-20260710-205204.out.log`:
- AIOps request was received with the expected `sessionId` and `runId`.
- `AiOpsService` started analysis and invoked the supervisor agent.
- AIOps orchestration completed and final report extraction ran.
- Exact trace was queried with the same `sessionId + runId`.
## Final Gate
Final Phase 5 gate:
```text
mvn -q clean test-compile
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest,DiagnosisTraceServiceTest" test
openspec validate session-run-trace-isolation --strict
git diff --check
```
Result: passed.
## Conclusion
Phase 5 satisfies AIOps run creation, SSE run metadata, run-scoped execution context propagation, run-scoped final report/evaluation/metrics writes, focused tests, and Maven E2E DB/log verification.
@@ -0,0 +1,253 @@
# Phase 6 Evidence: Demo, Trace UI, Documentation, and Verification
## Scope
Phase 6 completed the run-aware demo/documentation surface and final verification for `session-run-trace-isolation`.
Implemented:
- Demo scripts read Chat `runId`, query exact Trace with `?runId=...`, and submit feedback with `runId`.
- Trace UI accepts `?sessionId=...&runId=...` and calls the exact Trace API when `runId` is present.
- Chat UI remembers the latest run target and links to Trace Workbench with `sessionId + runId` when available.
- MVP table and architecture docs now describe `chat_session`, `diagnosis_run`, `agent_step.run_id`, `tool_invocation.run_id`, and transitional `case_library.diagnosis_id` semantics.
## Static Verification
Commands:
```powershell
node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js
$scripts = @(
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
Write-Host "Parsed $script"
}
openspec validate session-run-trace-isolation --strict
```
Result:
- JavaScript syntax: passed.
- PowerShell script parsing: passed.
- OpenSpec strict validation: passed.
## Focused Tests
Command:
```powershell
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
```
Result: passed.
## Final Same-Session Multi-Turn E2E
Startup command:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Startup log:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `target/e2e/phase6-mvn-20260710-211831.err.log`
Application readiness:
- `Started Main in 15.501 seconds`
- `ReadinessState changed to ACCEPTING_TRAFFIC`
E2E session:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
Artifacts:
- `target/e2e/phase6-chat1-20260710-2120.json`
- `target/e2e/phase6-chat2-20260710-2120.json`
- `target/e2e/phase6-trace-run1-20260710-2120.json`
- `target/e2e/phase6-trace-run2-20260710-2120.json`
- `target/e2e/phase6-trace-latest-20260710-2120.json`
- `target/e2e/phase6-e2e-summary-20260710-2120.json`
Observed:
```json
{
"sessionId": "e2e-phase6-chat-codex-20260710-2120",
"run1": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
"run2": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"chat1Success": true,
"chat2Success": true,
"trace1RunId": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
"trace2RunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"latestRunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"trace1Steps": 10,
"trace2Steps": 9,
"trace1Tools": 14,
"trace2Tools": 8
}
```
Interpretation:
- Two Chat requests reused the same `sessionId`.
- Each Chat request returned a distinct `runId`.
- Exact trace for run1 returned run1 only.
- Exact trace for run2 returned run2 only.
- Session-only Trace latest fallback returned run2.
## Database Inspection
Tool: `scripts/query_mysql.py`
`diagnosis_run`:
```text
run_id | session_id | status | agent_flow | step_count | tool_call_count | has_answer
-------------------------------------------------------------------------------------------------------------------------------------------------
run-e2a97696-4398-4abc-90e4-28f45c838f92 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 10 | 14 | 1
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 9 | 8 | 1
```
`agent_step` grouped by `run_id`:
```text
run_id | step_rows | min_step | max_step
--------------------------------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 9 | 0 | 5
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 10 | 0 | 6
```
`tool_invocation` grouped by `run_id`:
```text
run_id | tool_rows
----------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 8
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 14
```
Mixed row check:
```text
mixed_rows
----------
0
0
```
`chat_session` metadata:
```text
session_id | status | message_pair_count | last_active_at
---------------------------------------------------------------------------------------
e2e-phase6-chat-codex-20260710-2120 | ACTIVE | 2 | 2026-07-10 21:23:59
```
`GET /api/chat/session/e2e-phase6-chat-codex-20260710-2120` returned `messagePairCount=2`.
Interpretation:
- Run table has exactly two successful Chat runs for the E2E session.
- Step/tool counts match the exact Trace API responses.
- No `agent_step` or `tool_invocation` rows for this session have NULL or unexpected `run_id`.
- `chat_session` metadata confirms multi-turn context continuity at two message pairs.
## Log Inspection
Searched:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `logs/application.log`
- `logs/chat.log`
Patterns:
- `e2e-phase6-chat-codex-20260710-2120`
- `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
Result:
- Matching startup, Chat execution, run persistence, and trace lookup log lines were present in Maven output and `logs/application.log`.
## Baseline Drift
Focused baseline command:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
```
Broader baseline regression command from `mvp/eval/README.md`:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Result:
- Both commands passed.
- The baseline harness evaluates saved fixtures and does not depend on live DB/session tables.
- No baseline drift was observed.
## Final Gate Checks
Commands:
```powershell
rg -n "diagnosisSessionRepository\.save|new DiagnosisSession|DiagnosisSession\.builder|setAnswer\(|setSelfEvaluation\(|setFeedback\(|createFromSession|persistFinalReport\(sessionId, finalReport|save\(session\)" src\main\java\com\superbiz\agent -g "*.java"
rg -n "evaluate\(|evaluateRun\(|persistFinalReport\(|submitFeedback\(|createFromSession\(" src\main\java src\test\java -g "*.java"
git diff --check -- . ':!devflow/index.md'
openspec validate session-run-trace-isolation --strict
```
Result:
- New Chat write path calls `evaluationService.evaluateRun(...)` and writes `diagnosis_run`.
- New AIOps controller path calls `persistFinalReport(sessionId, runId, ...)` and writes `diagnosis_run`.
- Remaining `diagnosis_session` writes are legacy compatibility paths:
- `FeedbackService.submitLegacySessionFeedback(...)`
- `AiOpsService.persistLegacyFinalReport(...)`
- legacy `EvaluationService.evaluate(...)`
- `git diff --check`: passed after Markdown whitespace cleanup.
- OpenSpec strict validation: passed.
## Documentation Review Follow-up
After the final documentation review, the remaining demo helper docs were aligned with the run-aware contract:
- `mvp/demo/trace-inspection-checklist.md`
- `mvp/demo/payment-timeout-acceptance.md`
- `mvp/demo/interview-walkthrough.md`
- `mvp/tables/Agent步骤表-agent_step.md`
- `mvp/tables/README.md`
- `mvp/architecture/data-model.md`
- `mvp/architecture/session-trace-lifecycle.md`
The corrections remove session-only wording for Trace/Feedback and clarify that `run_id` is the execution isolation boundary while Trace API response order is the UI display contract.
Follow-up gate after these documentation fixes:
- `node --check src\main\resources\static\app.js`: passed.
- `node --check src\main\resources\static\trace.js`: passed.
- PowerShell demo script parsing: passed.
- `openspec validate session-run-trace-isolation --strict`: passed.
- `git diff --check -- . ':!devflow/index.md'`: passed.
## Notes
- The AGENTS-required `codebase-retrieval` and LSP tools were not available in this session. Fallback verification used OpenSpec context, `rg`, targeted file reads, focused tests, E2E, DB inspection, and log inspection.
@@ -0,0 +1,86 @@
# Change: Session / Run / Trace Isolation
## Problem
The current MVP uses the same `sessionId` for two different concepts:
- Redis `SessionContext` keeps multi-turn chat history for prompt context.
- MySQL `diagnosis_session`, `agent_step`, and `tool_invocation` persist diagnosis trace data for replay, verification, feedback, and evaluation.
End-to-end verification showed that two `/api/chat` calls with the same `sessionId` correctly reuse Redis context, but MySQL trace data is mixed under the same key:
- `diagnosis_session.query` is overwritten by the second round.
- `agent_step` and `tool_invocation` append rows from both rounds under the same `session_id`.
- Trace, verifier/evaluation, and feedback can read cross-round evidence.
This makes a trace no longer represent one replayable diagnosis run.
## Proposed Solution
Introduce a stable split between conversation state and execution state:
- `chat_session`: conversation metadata keyed by `session_id`.
- `diagnosis_run`: one execution/run keyed by `run_id`, belonging to a `session_id`.
- `agent_step` and `tool_invocation`: keep existing trace detail role, add `run_id` while retaining `session_id` for compatibility and coarse filtering.
`runId` becomes an official API field:
- `/api/chat` returns `sessionId + runId`.
- `/api/ai_ops` emits an SSE-compatible metadata message containing `sessionId` and `runId` before report content.
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
Trace remains an aggregate view of `diagnosis_run + agent_step + tool_invocation`; this change does not introduce a separate `diagnosis_trace` or `trace_event` table.
## Scope
- Add Flyway migrations and JPA entities/repositories for `chat_session` and `diagnosis_run`.
- Add nullable `run_id` to `agent_step` and `tool_invocation`, backfill historical data, then switch new writes to require run context.
- Move new Chat writes from `diagnosis_session` to `chat_session + diagnosis_run`.
- Update trace reads to resolve latest run or specified run.
- Add lightweight run list API: `GET /api/chat/session/{sessionId}/runs`.
- Update feedback and case-library creation to bind new data to `run_id`.
- Update AIOps to create and expose `runId` before this change is considered production complete.
- Update demo scripts and Trace UI with minimal `runId` support.
- Update relevant MVP table and architecture documentation.
- Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under `logs/`, database queries via `scripts/query_mysql.py`, and baseline drift checks.
## Non-Goals
- Do not add `diagnosis_trace` or `trace_event` in this change.
- Do not implement a full run-list UI.
- Do not remove the historical `diagnosis_session` table in this change.
- Do not change the Redis conversation window strategy.
- Do not persist full conversation history in MySQL; `chat_session` stores metadata only.
- Do not split historical mixed traces into true historical runs when the original run boundary is unavailable.
## Devflow Context Constraints
- `session-storage` established the current trace tables and decided that `sessionId` is propagated through `RunnableConfig.metadata` with `SessionContextHolder` as a fallback for tools.
- `confidence-feedback` established that useful feedback creates `case_library` from `DiagnosisSession.answer`, and that `feedback` does not change execution `status`.
- `mvp-demo-trace-acceptance` established `GET /api/diagnosis/{sessionId}/trace` as a read-only endpoint and demo scripts as part of the observable story.
- `aiops-traceable-diagnosis-entry` established `/api/ai_ops` as a traceable SSE entry point and made `sessionId` visible to callers.
- `data-model.md` and `session-trace-lifecycle.md` describe the current model as `diagnosis_session + agent_step + tool_invocation`, and list run id as a known follow-up.
- `devflow/glossary/CONTEXT.md` now defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks.
## Interface Impact
Level: L4 database/API contract migration with compatibility behavior.
- New API response field: `runId`.
- New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`.
- New API: `GET /api/chat/session/{sessionId}/runs`.
- Feedback request gains optional/preferred `runId`.
- Feedback response returns bound `runId` and `fallbackToLatestRun`.
- `/api/ai_ops` keeps SSE event name `message` and emits a `type=metadata` JSON message containing `sessionId` and `runId` before content.
- Database contract changes include new tables and new `run_id` columns.
- Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable.
## Risks
- Historical data has no true per-round boundary; backfill can only create compatibility runs from existing `diagnosis_session` rows.
- Context propagation through Agent hooks and tools is easy to break because it currently combines `RunnableConfig.metadata` and `SessionContextHolder`.
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
@@ -0,0 +1,29 @@
## MODIFIED Requirements
### Requirement: Diagnosis trace can be queried by session id
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
#### Scenario: Existing session latest trace is returned
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
#### Scenario: Existing session exact trace is returned
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system validates that `runId` belongs to `sessionId`
- **AND** it returns a success response containing only the trace data for that run
#### Scenario: Missing session returns not found
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
- **THEN** the system returns a 404 response using the existing session-not-found error contract
### Requirement: Trace aggregation is read-only
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
#### Scenario: Trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
#### Scenario: Exact trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
@@ -0,0 +1,175 @@
## ADDED Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
#### Scenario: Valid Chat execution creates session metadata and a run
- **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId`
- **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId`
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
- **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId`
#### Scenario: Invalid Chat request does not create a run
- **WHEN** a `/api/chat` request fails parameter validation before execution
- **THEN** the system SHALL NOT create a `diagnosis_run`
#### Scenario: Chat session stores metadata only
- **WHEN** a Chat request completes
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
- **AND** it SHALL NOT store full conversation message history
### Requirement: Chat responses SHALL expose run identity
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
#### Scenario: Chat response includes runId
- **WHEN** `/api/chat` returns a successful response
- **THEN** the response SHALL include `sessionId`
- **AND** the response SHALL include `runId` for the created diagnosis run
#### Scenario: Multi-turn Chat keeps one session and multiple runs
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
### Requirement: Trace details SHALL be scoped by run
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
#### Scenario: Agent steps are recorded with runId
- **WHEN** an Agent model step is persisted during a diagnosis run
- **THEN** the `agent_step` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Tool invocations are recorded with runId
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Run metrics count only current run rows
- **WHEN** a diagnosis run completes
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
### Requirement: Trace API SHALL support latest-run and exact-run queries
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
#### Scenario: Trace without runId resolves latest run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
- **AND** the response SHALL include the resolved `runId`
#### Scenario: Trace with runId returns exact run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
#### Scenario: Trace rejects run from another session
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
### Requirement: Session runs SHALL be listable without expanding trace details
The system SHALL provide a lightweight run-list API for a Chat Session.
#### Scenario: Run list returns summaries
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
#### Scenario: Run list handles session without runs
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
- **THEN** the system SHALL return a successful empty list
#### Scenario: Run list rejects missing session
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
- **THEN** the system SHALL use the existing not-found/error response behavior
### Requirement: Feedback SHALL bind to diagnosis runs
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
#### Scenario: Feedback with runId updates specified run
- **WHEN** a feedback request includes `sessionId` and `runId`
- **THEN** the system SHALL validate that the run belongs to the session
- **AND** it SHALL update feedback on that run
- **AND** the response SHALL include the actual bound `runId`
- **AND** the response SHALL include `fallbackToLatestRun=false`
#### Scenario: Feedback without runId falls back observably
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** at least one `diagnosis_run` exists for that session
- **THEN** the system SHALL bind feedback to the latest run for that session
- **AND** the response SHALL include `fallbackToLatestRun=true`
- **AND** the response SHALL include the actual bound `runId`
#### Scenario: Historical feedback without run-backed data remains compatible
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** no `diagnosis_run` exists for that session
- **AND** a historical `diagnosis_session` row exists for that session
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
- **AND** the response SHALL NOT claim latest-run fallback
- **AND** the response MAY omit `runId`
#### Scenario: Feedback rejects run from another session
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return a failed feedback response instead of updating either run
#### Scenario: Useful feedback creates case from run
- **WHEN** feedback for a run is `useful`
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
### Requirement: AIOps executions SHALL use run isolation
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
#### Scenario: AIOps creates run
- **WHEN** `/api/ai_ops` starts a valid execution
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
#### Scenario: AIOps SSE exposes runId
- **WHEN** `/api/ai_ops` streams response metadata to the caller
- **THEN** the stream SHALL send a compatible metadata message before report content
- **AND** the SSE event name SHALL remain `message`
- **AND** the message type SHALL be `metadata`
- **AND** the metadata payload SHALL expose the resolved `sessionId`
- **AND** the metadata payload SHALL expose the created `runId`
- **AND** report content SHALL continue to use the existing content message shape
### Requirement: Migration SHALL preserve historical trace access
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
#### Scenario: Historical session gets compatibility run
- **WHEN** migration runs on an existing `diagnosis_session` row
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
#### Scenario: Old table is retained
- **WHEN** the migration completes
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
### Requirement: Demo and Trace UI SHALL support runId
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
#### Scenario: Demo script queries exact trace
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
- **THEN** it SHALL include `runId` when querying the Trace API
#### Scenario: Trace UI honors URL runId
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
### Requirement: Run isolation SHALL be verified against baselines
The change SHALL verify both runtime behavior and evaluation baseline impact.
#### Scenario: Multi-turn E2E proves run isolation
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
- **THEN** database inspection SHALL show two `diagnosis_run` rows
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
- **AND** Redis session metadata SHALL still show multi-turn context continuity
#### Scenario: Baseline drift is checked
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
@@ -0,0 +1,59 @@
## 1. Schema and Compatibility Foundation
- [x] 1.1 Add Flyway migration for `chat_session` and `diagnosis_run` with indexes for unique `session_id`, unique `run_id`, latest-run lookup, and run summary listing.
- [x] 1.2 Add nullable `run_id` columns to `agent_step` and `tool_invocation` with indexes for run-scoped trace queries.
- [x] 1.3 Backfill one compatibility `diagnosis_run` for each existing `diagnosis_session` row.
- [x] 1.4 Backfill historical `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run for the same `session_id`.
- [x] 1.5 Add JPA entities and repositories for `ChatSession` and `DiagnosisRun`, including latest-run and run-id lookup methods.
- [x] 1.6 Add focused migration/repository verification that proves old data remains queryable and run lookup methods work.
- [x] 1.7 Phase 1 gate: run the smallest relevant test/build check, inspect DB migration behavior, update OpenSpec task status, archive phase evidence, and commit before starting Phase 2.
## 2. Chat Run Write Path
- [x] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
- [x] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
- [x] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
- [x] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
- [x] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
- [x] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
- [x] 2.7 Change `/api/chat` response DTO to include official `runId`.
- [x] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
- [x] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
## 3. Trace Read Path and Run Listing
- [x] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
- [x] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
- [x] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
- [x] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
- [x] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
- [x] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
- [x] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
## 4. Feedback and Case Library Run Binding
- [x] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
- [x] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
- [x] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
- [x] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
- [x] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
- [x] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
- [x] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
## 5. AIOps Run Isolation
- [x] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
- [x] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
- [x] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
- [x] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
- [x] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
- [x] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
## 6. Demo, Trace UI, Documentation, and Verification
- [x] 6.1 Update demo scripts to read `runId` from Chat/AIOps responses and pass `?runId=...` to Trace API.
- [x] 6.2 Update Trace UI to accept `?sessionId=...&runId=...` and query exact trace when `runId` is present.
- [x] 6.3 Update MVP table and architecture docs for `chat_session`, `diagnosis_run`, `run_id`, and transitional `case_library.diagnosis_id` semantics.
- [x] 6.4 Run final same-session multi-turn E2E using Maven startup if needed; collect DB evidence through `scripts/query_mysql.py` and inspect `logs/`.
- [x] 6.5 Run or explicitly evaluate the relevant baseline diff command and document whether drift is expected or a regression.
- [x] 6.6 Final gate: ensure all OpenSpec tasks are checked, no new writes depend on `diagnosis_session`, phase evidence is archived, final commit is created, and the change is ready for OpenSpec archive.