feat(trace): add session run isolation schema
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
committed: true
|
||||
change: session-run-trace-isolation
|
||||
validated: openspec validate session-run-trace-isolation --strict
|
||||
@@ -0,0 +1,135 @@
|
||||
# Decisions: session-run-trace-isolation
|
||||
|
||||
## sm-flow State
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Scale: complex
|
||||
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
|
||||
## Entry Summary
|
||||
|
||||
Problem: the same `sessionId` currently represents both multi-turn conversation context and one persisted diagnosis trace. Multi-turn Chat E2E proved that Redis context behaves correctly, but MySQL trace rows from different rounds are mixed under one `session_id`.
|
||||
|
||||
Expected result: split session metadata from per-run execution state, expose `runId` as the official run identifier, and make trace, feedback, evaluation, demo scripts, Trace UI, and AIOps read/write by run.
|
||||
|
||||
Known modules: Flyway/JPA entities/repositories, `ChatService`, `ChatController`, `AiOpsService`, `DiagnosisTraceService`, `EvaluationService`, `FeedbackService`, `CaseLibraryService`, `AgentLoggingHook`, `ToolInvocationRecorder`, `SessionContextHolder`, demo scripts, static Trace UI, MVP docs.
|
||||
|
||||
## Context Collection
|
||||
|
||||
`devflow/index.md`: relevant entries found.
|
||||
|
||||
Relevant historical decisions:
|
||||
|
||||
- `session-storage`: current trace persistence is `diagnosis_session + agent_step + tool_invocation`; `sessionId` propagation uses `RunnableConfig.metadata` with `SessionContextHolder` fallback for tools; AIOps records child agents only.
|
||||
- `confidence-feedback`: `FeedbackService` writes feedback, useful feedback creates `case_library`, and feedback must not change execution status.
|
||||
- `mvp-demo-trace-acceptance`: Trace API is read-only and demo artifacts/scripts are part of acceptance.
|
||||
- `aiops-traceable-diagnosis-entry`: `/api/ai_ops` is an SSE entry point that accepts optional alert payload and exposes `sessionId`.
|
||||
- `data-model.md`: current `case_library.diagnosis_id` maps to `diagnosis_session.session_id`; this must be treated as legacy data after the change.
|
||||
- `session-trace-lifecycle.md`: current docs already list run id as a future enhancement for multi-run sessions.
|
||||
|
||||
OpenSpec inputs that must be carried forward:
|
||||
|
||||
- Current specs mention `diagnosis_session` directly in trace, verifier, evidence, and demo requirements; new specs must either supersede or preserve compatibility for those requirements.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` must remain read-only.
|
||||
- Baseline/eval checks must expect changed counts only when the change is explained by run isolation.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | Question | Mode | Status |
|
||||
|---|---|---|---|
|
||||
| Q1 | Should this be an added field on existing `diagnosis_session`, or split tables? | user-interview | confirmed: split `chat_session` and `diagnosis_run`, keep existing trace detail tables |
|
||||
| Q2 | What is metadata in `chat_session`? | user-interview | confirmed: session directory fields only, not full message body |
|
||||
| Q3 | Is "message body" the full conversation history? | user-interview | confirmed: full conversation history stays in Redis `SessionContext.messageHistory` for now |
|
||||
| Q4 | Should `runId` be an official API field? | user-interview | confirmed: yes |
|
||||
| Q5 | How should trace work without `runId`? | user-interview | confirmed: default to latest run for compatibility |
|
||||
| Q6 | Should Feedback without `runId` fail or fall back? | user-interview | confirmed: short-term fallback to latest run, long-term may tighten |
|
||||
| Q7 | Should simple Chat Q&A create a run? | user-interview | confirmed: yes, every valid `/api/chat` execution creates a run |
|
||||
| Q8 | Should AIOps be included? | user-interview | confirmed: yes, same run isolation semantics |
|
||||
| Q9 | Should there be a separate trace table? | user-interview | confirmed: no, current `agent_step` and `tool_invocation` are enough for this phase |
|
||||
| Q10 | What should happen to old `diagnosis_session`? | user-interview | confirmed: keep it for history/rollback, new code stops writing it after migration |
|
||||
| Q11 | Does the issue need demo/Trace UI support? | user-interview | confirmed: yes, minimal `runId` support |
|
||||
| Q12 | What extra risks were found by document review? | evidence-driven | reported and patched into ISS-010 |
|
||||
|
||||
## Evidence-Driven Findings
|
||||
|
||||
- Code evidence: `SessionContext` contains `messageHistory` and `getMessagePairCount()`, supporting the decision that Redis holds hot conversation history while MySQL stores auditable per-run query/answer.
|
||||
- Code evidence: `CaseLibraryService.createFromSession` currently deduplicates by `session.getSessionId()` and maps answer/query from `DiagnosisSession`; this must change for new run-based data.
|
||||
- Documentation evidence: `mvp/architecture/data-model.md` states `case_library.diagnosis_id = diagnosis_session.session_id`; this becomes transitional legacy semantics.
|
||||
- Documentation evidence: existing trace OpenSpec requires `GET /api/diagnosis/{sessionId}/trace` to be read-only; run resolution must preserve that invariant.
|
||||
- E2E evidence from ISS-010: two Chat rounds with the same `sessionId` resulted in one overwritten `diagnosis_session` row and mixed step/tool rows.
|
||||
|
||||
## Confirmed Decisions
|
||||
|
||||
- `runId` format: `run-` + full UUID.
|
||||
- Latest run ordering: `diagnosis_run.created_at DESC, id DESC`, not `updated_at`.
|
||||
- `chat_session` stores metadata: `session_id`, `status`, `message_pair_count`, `created_at`, `last_active_at`, `expires_at`.
|
||||
- Per-run long-term audit stores `query` and `answer` in `diagnosis_run`.
|
||||
- `agent_step` and `tool_invocation` retain `session_id` and add `run_id`.
|
||||
- Missing feedback `runId` returns `fallbackToLatestRun=true` plus actual bound `runId`.
|
||||
- Historical mixed data is not split into multiple true runs.
|
||||
|
||||
## OpenSpec Backwrite Log
|
||||
|
||||
- Created `proposal.md` with problem, scope, non-goals, context constraints, interface impact, and risks.
|
||||
- ISS-010 patched to clarify phase boundary, migration order, AIOps same-change requirement, feedback fallback response, case-library transitional semantics, and Redis/MySQL TTL boundary.
|
||||
- Glossary patched to include `SessionContext.messageHistory` and its persistence boundary.
|
||||
- Created `design.md`, `specs/session-run-trace-isolation/spec.md`, `specs/mvp-demo-trace-acceptance/spec.md`, and `tasks.md`.
|
||||
- Architecture audit found that `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` also read tool rows by `sessionId`; Phase 2 tasks were updated to cover run-scoped reads.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Check | Result | Evidence |
|
||||
|---|---|---|
|
||||
| issue/context -> proposal | aligned | proposal carries E2E problem, split-table solution, compatibility, AIOps, feedback, demo/UI, and baseline scope |
|
||||
| proposal -> design | aligned | design records data model, API impact, migration plan, rollback, and key decisions |
|
||||
| design -> specs | aligned | specs cover session/run split, Chat, trace, feedback, AIOps, migration, demo/UI, E2E, and baseline behavior |
|
||||
| specs -> tasks | aligned | tasks implement schema, Chat write path, trace reads, feedback/case, AIOps, demo/UI/docs, and verification gates |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Input to output chain:
|
||||
|
||||
```text
|
||||
Chat/AIOps request
|
||||
-> ChatController / AIOps endpoint
|
||||
-> ChatService / AiOpsService
|
||||
-> execution context(sessionId, runId)
|
||||
-> AgentLoggingHook -> agent_step
|
||||
-> ToolInvocationRecorder -> tool_invocation
|
||||
-> verifier/gatekeeper/evaluation summary reads
|
||||
-> diagnosis_run answer/self_evaluation/status/counts
|
||||
-> Trace API / Feedback / CaseLibrary / Demo / Trace UI
|
||||
```
|
||||
|
||||
Audit conclusions:
|
||||
|
||||
- Data ownership is clearer with `chat_session` owning conversation metadata and `diagnosis_run` owning execution state; `agent_step` and `tool_invocation` remain trace details owned by one run.
|
||||
- The highest coupling risk is execution-context propagation because hooks and tools currently use both `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- The main read-path risk is missing one of the session-scoped consumers (`ToolTraceSummaryService`, `ExecutorGatekeeperService`, `EvaluationService`, trace, feedback, case creation).
|
||||
- Migration is additive and rollback-friendly until constraints are tightened; historical mixed data must be treated as compatibility data.
|
||||
- AIOps must complete before final archive because otherwise the system would still have one production entry point with mixed trace semantics.
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- Interface impact level: L4.
|
||||
- Required artifacts: proposal, design, specs, tasks.
|
||||
- Strict OpenSpec validation: passed with `openspec validate session-run-trace-isolation --strict`.
|
||||
- `.committed` marker: created after successful commit gate.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Reference migrations: `V005__create_session_storage.sql`, `V008__add_answer_to_diagnosis_session.sql`, `V010__add_relevance_level_to_tool_invocation.sql`.
|
||||
- Entity style: JPA entities use Lombok `@Data`, `@Builder`, `@NoArgsConstructor`, `@AllArgsConstructor`, `@PrePersist`, and `@PreUpdate` where timestamps need maintenance.
|
||||
- Repository style: Spring Data JPA repository interfaces with derived query methods returning `Optional<T>` or `List<T>`.
|
||||
- Test style: repository tests use `@DataJpaTest`, `@AutoConfigureTestDatabase(replace = NONE)`, Flyway enabled, and `ddl-auto=validate`.
|
||||
|
||||
## Phase 1 Apply Notes
|
||||
|
||||
- Implemented additive migration `V011__add_session_run_isolation.sql`.
|
||||
- Implemented `ChatSession` / `DiagnosisRun` entities and repositories.
|
||||
- Added nullable `runId` fields to `AgentStep` and `ToolInvocation`.
|
||||
- Added run-scoped repository methods for step/tool lookup and counts.
|
||||
- Added `DiagnosisRunRepositoryTest`.
|
||||
- Verification evidence is recorded in `phase-1-evidence.md`.
|
||||
- Historical DB inspection found one orphan `tool_invocation` row without a matching `diagnosis_session`; it remains `run_id = NULL` because no reliable compatibility run can be inferred.
|
||||
@@ -0,0 +1,139 @@
|
||||
## Context
|
||||
|
||||
The current MVP persists diagnosis observability through `diagnosis_session`, `agent_step`, and `tool_invocation`. That model works for a single diagnosis per `sessionId`, but it conflates two lifecycles once a caller reuses the same `sessionId` for multi-turn conversation:
|
||||
|
||||
- conversation state: Redis `SessionContext.messageHistory` and session metadata;
|
||||
- execution state: one diagnosis answer, trace, self-evaluation, and feedback target.
|
||||
|
||||
The E2E evidence in ISS-010 showed that Redis correctly preserved multi-turn context while MySQL mixed both rounds under the same `session_id`. This breaks trace replay, verifier/evaluation scoping, feedback targeting, and case-library provenance.
|
||||
|
||||
Constraints from existing work:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` is a read-only MVP/demo contract and must remain compatible.
|
||||
- Current hooks and tools propagate `sessionId` through `RunnableConfig.metadata` and `SessionContextHolder`; `runId` must follow the same execution context boundary.
|
||||
- Feedback currently writes `DiagnosisSession.feedback` and creates `case_library` from `DiagnosisSession.answer`.
|
||||
- AIOps is a first-class traceable entry point and cannot be left permanently on the old mixed-run model.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Split conversation metadata from per-execution diagnosis state.
|
||||
- Introduce `runId` as the official identifier for one replayable diagnosis execution.
|
||||
- Preserve old `sessionId`-only callers by resolving the latest run where possible.
|
||||
- Scope trace, evaluation, feedback, case creation, demo scripts, and Trace UI by run.
|
||||
- Migrate historical data into compatibility runs without deleting the old table.
|
||||
- Include Chat and AIOps in the same release-level change.
|
||||
- Verify behavior with focused tests, E2E when needed, DB inspection, logs, and baseline drift checks.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not persist full chat history in MySQL.
|
||||
- Do not add `diagnosis_trace` or `trace_event`.
|
||||
- Do not implement full run-list UI.
|
||||
- Do not physically delete `diagnosis_session`.
|
||||
- Do not attempt to reconstruct true historical round boundaries when only mixed `session_id` data exists.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Domain split | Add `chat_session` and `diagnosis_run` | Add `run_id` to `diagnosis_session` only | Separate tables keep conversation metadata and execution state from growing into one coupled table. |
|
||||
| Trace detail storage | Reuse `agent_step` and `tool_invocation`, adding `run_id` | Add `diagnosis_trace` / `trace_event` | Existing detail tables already represent trace; isolation needs a run key, not a new event model. |
|
||||
| API identity | `runId = "run-" + UUID` | Reuse short session id or DB id | Full UUID avoids collision and keeps external IDs independent from database internals. |
|
||||
| Latest-run compatibility | `GET /api/diagnosis/{sessionId}/trace` resolves latest by `created_at DESC, id DESC` | Require `runId` immediately | Compatibility keeps existing demo/UI/scripts working while new clients migrate. `updated_at` is avoided because feedback/eval updates can reorder old runs. |
|
||||
| Historical migration | Backfill one compatibility run per existing `diagnosis_session` | Try to split old mixed rows | Old rows do not contain reliable run boundaries. A compatibility run preserves auditability without inventing data. |
|
||||
| Feedback fallback | Missing `runId` binds latest run and returns fallback metadata | Reject missing `runId` immediately | Short-term compatibility is needed for old clients; explicit fallback keeps ambiguity observable. |
|
||||
| Case library provenance | New automatic cases store `diagnosis_id = run_id` | Add a new case-library column now | Existing column name can carry transitional provenance; docs and query logic must recognize old `session_id` and new `run_id`. |
|
||||
| AIOps phase | Implement after Chat but before overall archive | Leave AIOps for a follow-up issue | AIOps is already a trace entry point; leaving it old-model would preserve the same bug on another endpoint. |
|
||||
|
||||
## Data Model
|
||||
|
||||
```text
|
||||
chat_session
|
||||
id
|
||||
session_id unique
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
|
||||
diagnosis_run
|
||||
id
|
||||
run_id unique
|
||||
session_id
|
||||
query
|
||||
status
|
||||
agent_flow
|
||||
answer
|
||||
self_evaluation
|
||||
feedback
|
||||
total_duration_ms
|
||||
total_token_count
|
||||
step_count
|
||||
tool_call_count
|
||||
created_at
|
||||
updated_at
|
||||
|
||||
agent_step
|
||||
session_id
|
||||
run_id
|
||||
...
|
||||
|
||||
tool_invocation
|
||||
session_id
|
||||
run_id
|
||||
...
|
||||
```
|
||||
|
||||
`chat_session.expires_at` is MySQL directory metadata. Redis TTL can expire `SessionContext.messageHistory`; persisted `diagnosis_run`, `agent_step`, and `tool_invocation` remain audit records.
|
||||
|
||||
## API / Interface Impact
|
||||
|
||||
Interface level: L4.
|
||||
|
||||
- Database contract changes: new tables, new columns, backfill, indexes, and later non-null expectations for new writes.
|
||||
- `/api/chat` response adds `runId`.
|
||||
- `/api/ai_ops` SSE emits the resolved `runId`.
|
||||
- Trace API accepts optional `runId`.
|
||||
- Feedback request accepts preferred `runId` and returns fallback binding metadata when omitted.
|
||||
- New run summary API: `GET /api/chat/session/{sessionId}/runs`.
|
||||
|
||||
Compatibility:
|
||||
|
||||
- Old `sessionId`-only trace and feedback calls bind to latest run.
|
||||
- Old `diagnosis_session` is retained for rollback and historical comparison.
|
||||
- New code must not keep writing new execution state into `diagnosis_session` after the write switch.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add `chat_session` and `diagnosis_run`.
|
||||
2. Add nullable `run_id` to `agent_step` and `tool_invocation`.
|
||||
3. Backfill `diagnosis_run` from existing `diagnosis_session`.
|
||||
4. Backfill old `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run.
|
||||
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
|
||||
6. Deploy repository and read-path compatibility.
|
||||
7. Switch Chat write path to `chat_session + diagnosis_run`.
|
||||
8. Switch Trace, evaluation, feedback, case-library, demo, and UI paths.
|
||||
9. Switch AIOps write path.
|
||||
10. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
|
||||
|
||||
Rollback:
|
||||
|
||||
- Keep `diagnosis_session` intact during this change.
|
||||
- Migrations are additive until constraints are tightened.
|
||||
- If write-switch rollout fails, rollback code can read the retained old table while migrated compatibility rows remain harmless.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] ThreadLocal/context propagation may miss `runId` in nested agent/tool calls. -> Mitigation: introduce a unified execution context carrying both IDs and test hook/tool recording.
|
||||
- [Risk] Baseline metrics change because cross-round tool rows are no longer counted. -> Mitigation: run baseline diff and document expected drift.
|
||||
- [Risk] Legacy `case_library.diagnosis_id` values are ambiguous. -> Mitigation: document transitional semantics and keep lookup logic aware of old `session_id` values.
|
||||
- [Risk] AIOps SSE clients may not parse a new metadata event. -> Mitigation: add metadata in a compatible stream message and keep final report streaming behavior.
|
||||
- [Risk] Historical mixed trace cannot be truly separated. -> Mitigation: call this out as compatibility data, not reconstructed truth.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None blocking. Long-term tightening of missing feedback `runId` remains a follow-up decision after clients migrate.
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
# Phase 1 Evidence: Schema and Compatibility Foundation
|
||||
|
||||
## Scope Completed
|
||||
|
||||
- Added Flyway migration `V011__add_session_run_isolation.sql`.
|
||||
- Added `chat_session` and `diagnosis_run` tables.
|
||||
- Added nullable `run_id` columns and indexes to `agent_step` and `tool_invocation`.
|
||||
- Backfilled compatibility runs from existing `diagnosis_session` rows.
|
||||
- Backfilled existing step/tool rows where a matching `diagnosis_session.session_id` exists.
|
||||
- Added `ChatSession`, `DiagnosisRun`, `ChatSessionRepository`, and `DiagnosisRunRepository`.
|
||||
- Added run-scoped query/count methods to `AgentStepRepository` and `ToolInvocationRepository`.
|
||||
- Added `DiagnosisRunRepositoryTest`.
|
||||
|
||||
## Verification
|
||||
|
||||
### Maven
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=DiagnosisRunRepositoryTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=DiagnosisSessionRepositoryTest,AgentStepRepositoryTest,ToolInvocationRepositoryTest,DiagnosisRunRepositoryTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
### Database Inspection
|
||||
|
||||
Checked with `scripts/query_mysql.py`.
|
||||
|
||||
```text
|
||||
SELECT installed_rank, version, description, success
|
||||
FROM flyway_schema_history
|
||||
ORDER BY installed_rank DESC
|
||||
LIMIT 5;
|
||||
```
|
||||
|
||||
Relevant result:
|
||||
|
||||
```text
|
||||
version = 011
|
||||
description = add session run isolation
|
||||
success = 1
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT COUNT(*) AS chat_sessions FROM chat_session;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
chat_sessions = 83
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT COUNT(*) AS diagnosis_runs,
|
||||
SUM(CASE WHEN run_id IS NULL THEN 1 ELSE 0 END) AS null_run_ids
|
||||
FROM diagnosis_run;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
diagnosis_runs = 83
|
||||
null_run_ids = 0
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT (SELECT COUNT(*) FROM agent_step WHERE run_id IS NULL) AS agent_step_null_run_id,
|
||||
(SELECT COUNT(*) FROM tool_invocation WHERE run_id IS NULL) AS tool_invocation_null_run_id;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
agent_step_null_run_id = 0
|
||||
tool_invocation_null_run_id = 1
|
||||
```
|
||||
|
||||
The one remaining null `tool_invocation.run_id` is historical orphan data:
|
||||
|
||||
```text
|
||||
id = 3
|
||||
session_id = f9d2290d
|
||||
tool_name = lookup_knowledge
|
||||
created_at = 2026-06-26 14:17:35
|
||||
```
|
||||
|
||||
It has no matching `diagnosis_session`, so V011 cannot safely infer a compatibility run. This matches the OpenSpec wording "when possible" for historical backfill.
|
||||
|
||||
### Logs
|
||||
|
||||
`logs/application.log` shows Hibernate using the new `run_id` fields in `agent_step` and `tool_invocation` inserts/selects during the repository verification.
|
||||
|
||||
## Notes
|
||||
|
||||
- Phase 1 intentionally does not switch Chat, Trace, Feedback, or AIOps write paths.
|
||||
- New columns remain nullable until later phases switch new writes and verify coverage.
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
# Change: Session / Run / Trace Isolation
|
||||
|
||||
## Problem
|
||||
|
||||
The current MVP uses the same `sessionId` for two different concepts:
|
||||
|
||||
- Redis `SessionContext` keeps multi-turn chat history for prompt context.
|
||||
- MySQL `diagnosis_session`, `agent_step`, and `tool_invocation` persist diagnosis trace data for replay, verification, feedback, and evaluation.
|
||||
|
||||
End-to-end verification showed that two `/api/chat` calls with the same `sessionId` correctly reuse Redis context, but MySQL trace data is mixed under the same key:
|
||||
|
||||
- `diagnosis_session.query` is overwritten by the second round.
|
||||
- `agent_step` and `tool_invocation` append rows from both rounds under the same `session_id`.
|
||||
- Trace, verifier/evaluation, and feedback can read cross-round evidence.
|
||||
|
||||
This makes a trace no longer represent one replayable diagnosis run.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Introduce a stable split between conversation state and execution state:
|
||||
|
||||
- `chat_session`: conversation metadata keyed by `session_id`.
|
||||
- `diagnosis_run`: one execution/run keyed by `run_id`, belonging to a `session_id`.
|
||||
- `agent_step` and `tool_invocation`: keep existing trace detail role, add `run_id` while retaining `session_id` for compatibility and coarse filtering.
|
||||
|
||||
`runId` becomes an official API field:
|
||||
|
||||
- `/api/chat` returns `sessionId + runId`.
|
||||
- `/api/ai_ops` emits or returns `runId` in the SSE-compatible protocol.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
|
||||
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
|
||||
|
||||
Trace remains an aggregate view of `diagnosis_run + agent_step + tool_invocation`; this change does not introduce a separate `diagnosis_trace` or `trace_event` table.
|
||||
|
||||
## Scope
|
||||
|
||||
- Add Flyway migrations and JPA entities/repositories for `chat_session` and `diagnosis_run`.
|
||||
- Add nullable `run_id` to `agent_step` and `tool_invocation`, backfill historical data, then switch new writes to require run context.
|
||||
- Move new Chat writes from `diagnosis_session` to `chat_session + diagnosis_run`.
|
||||
- Update trace reads to resolve latest run or specified run.
|
||||
- Add lightweight run list API: `GET /api/chat/session/{sessionId}/runs`.
|
||||
- Update feedback and case-library creation to bind new data to `run_id`.
|
||||
- Update AIOps to create and expose `runId` before this change is considered production complete.
|
||||
- Update demo scripts and Trace UI with minimal `runId` support.
|
||||
- Update relevant MVP table and architecture documentation.
|
||||
- Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under `logs/`, database queries via `scripts/query_mysql.py`, and baseline drift checks.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Do not add `diagnosis_trace` or `trace_event` in this change.
|
||||
- Do not implement a full run-list UI.
|
||||
- Do not remove the historical `diagnosis_session` table in this change.
|
||||
- Do not change the Redis conversation window strategy.
|
||||
- Do not persist full conversation history in MySQL; `chat_session` stores metadata only.
|
||||
- Do not split historical mixed traces into true historical runs when the original run boundary is unavailable.
|
||||
|
||||
## Devflow Context Constraints
|
||||
|
||||
- `session-storage` established the current trace tables and decided that `sessionId` is propagated through `RunnableConfig.metadata` with `SessionContextHolder` as a fallback for tools.
|
||||
- `confidence-feedback` established that useful feedback creates `case_library` from `DiagnosisSession.answer`, and that `feedback` does not change execution `status`.
|
||||
- `mvp-demo-trace-acceptance` established `GET /api/diagnosis/{sessionId}/trace` as a read-only endpoint and demo scripts as part of the observable story.
|
||||
- `aiops-traceable-diagnosis-entry` established `/api/ai_ops` as a traceable SSE entry point and made `sessionId` visible to callers.
|
||||
- `data-model.md` and `session-trace-lifecycle.md` describe the current model as `diagnosis_session + agent_step + tool_invocation`, and list run id as a known follow-up.
|
||||
- `devflow/glossary/CONTEXT.md` now defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L4 database/API contract migration with compatibility behavior.
|
||||
|
||||
- New API response field: `runId`.
|
||||
- New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`.
|
||||
- New API: `GET /api/chat/session/{sessionId}/runs`.
|
||||
- Feedback request gains optional/preferred `runId`.
|
||||
- Database contract changes include new tables and new `run_id` columns.
|
||||
- Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable.
|
||||
|
||||
## Risks
|
||||
|
||||
- Historical data has no true per-round boundary; backfill can only create compatibility runs from existing `diagnosis_session` rows.
|
||||
- Context propagation through Agent hooks and tools is easy to break because it currently combines `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
|
||||
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
|
||||
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
|
||||
|
||||
@@ -0,0 +1,29 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
|
||||
|
||||
#### Scenario: Existing session latest trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
|
||||
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
|
||||
|
||||
#### Scenario: Existing session exact trace is returned
|
||||
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system validates that `runId` belongs to `sessionId`
|
||||
- **AND** it returns a success response containing only the trace data for that run
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
#### Scenario: Exact trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
|
||||
|
||||
+147
@@ -0,0 +1,147 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
|
||||
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
|
||||
|
||||
#### Scenario: Valid Chat execution creates session metadata and a run
|
||||
- **WHEN** a valid `/api/chat` request enters the Chat execution path with a `sessionId`
|
||||
- **THEN** the system SHALL ensure a `chat_session` row exists for that `sessionId`
|
||||
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
|
||||
- **AND** the `diagnosis_run.session_id` SHALL equal the request `sessionId`
|
||||
|
||||
#### Scenario: Invalid Chat request does not create a run
|
||||
- **WHEN** a `/api/chat` request fails parameter validation before execution
|
||||
- **THEN** the system SHALL NOT create a `diagnosis_run`
|
||||
|
||||
#### Scenario: Chat session stores metadata only
|
||||
- **WHEN** a Chat request completes
|
||||
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and expiration time
|
||||
- **AND** it SHALL NOT store full conversation message history
|
||||
|
||||
### Requirement: Chat responses SHALL expose run identity
|
||||
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
|
||||
|
||||
#### Scenario: Chat response includes runId
|
||||
- **WHEN** `/api/chat` returns a successful response
|
||||
- **THEN** the response SHALL include `sessionId`
|
||||
- **AND** the response SHALL include `runId` for the created diagnosis run
|
||||
|
||||
#### Scenario: Multi-turn Chat keeps one session and multiple runs
|
||||
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
|
||||
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
|
||||
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
|
||||
|
||||
### Requirement: Trace details SHALL be scoped by run
|
||||
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
|
||||
|
||||
#### Scenario: Agent steps are recorded with runId
|
||||
- **WHEN** an Agent model step is persisted during a diagnosis run
|
||||
- **THEN** the `agent_step` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Tool invocations are recorded with runId
|
||||
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
|
||||
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Run metrics count only current run rows
|
||||
- **WHEN** a diagnosis run completes
|
||||
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
|
||||
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
|
||||
|
||||
### Requirement: Trace API SHALL support latest-run and exact-run queries
|
||||
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
|
||||
|
||||
#### Scenario: Trace without runId resolves latest run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
|
||||
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
|
||||
- **AND** the response SHALL include the resolved `runId`
|
||||
|
||||
#### Scenario: Trace with runId returns exact run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
|
||||
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
|
||||
|
||||
#### Scenario: Trace rejects run from another session
|
||||
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
|
||||
|
||||
### Requirement: Session runs SHALL be listable without expanding trace details
|
||||
The system SHALL provide a lightweight run-list API for a Chat Session.
|
||||
|
||||
#### Scenario: Run list returns summaries
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
|
||||
- **THEN** the system SHALL return run summaries from `diagnosis_run`
|
||||
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
|
||||
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
|
||||
|
||||
### Requirement: Feedback SHALL bind to diagnosis runs
|
||||
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
|
||||
|
||||
#### Scenario: Feedback with runId updates specified run
|
||||
- **WHEN** a feedback request includes `sessionId` and `runId`
|
||||
- **THEN** the system SHALL validate that the run belongs to the session
|
||||
- **AND** it SHALL update feedback on that run
|
||||
|
||||
#### Scenario: Feedback without runId falls back observably
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **THEN** the system SHALL bind feedback to the latest run for that session
|
||||
- **AND** the response or log SHALL include `fallbackToLatestRun=true`
|
||||
- **AND** the response or log SHALL include the actual bound `runId`
|
||||
|
||||
#### Scenario: Useful feedback creates case from run
|
||||
- **WHEN** feedback for a run is `useful`
|
||||
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
|
||||
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
|
||||
|
||||
### Requirement: AIOps executions SHALL use run isolation
|
||||
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
|
||||
|
||||
#### Scenario: AIOps creates run
|
||||
- **WHEN** `/api/ai_ops` starts a valid execution
|
||||
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
|
||||
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
|
||||
|
||||
#### Scenario: AIOps SSE exposes runId
|
||||
- **WHEN** `/api/ai_ops` streams response metadata to the caller
|
||||
- **THEN** the stream SHALL expose the resolved `sessionId`
|
||||
- **AND** it SHALL expose the created `runId`
|
||||
|
||||
### Requirement: Migration SHALL preserve historical trace access
|
||||
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
|
||||
|
||||
#### Scenario: Historical session gets compatibility run
|
||||
- **WHEN** migration runs on an existing `diagnosis_session` row
|
||||
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
|
||||
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
|
||||
|
||||
#### Scenario: Old table is retained
|
||||
- **WHEN** the migration completes
|
||||
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
|
||||
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
|
||||
|
||||
### Requirement: Demo and Trace UI SHALL support runId
|
||||
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
|
||||
|
||||
#### Scenario: Demo script queries exact trace
|
||||
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
|
||||
- **THEN** it SHALL include `runId` when querying the Trace API
|
||||
|
||||
#### Scenario: Trace UI honors URL runId
|
||||
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
|
||||
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
|
||||
### Requirement: Run isolation SHALL be verified against baselines
|
||||
The change SHALL verify both runtime behavior and evaluation baseline impact.
|
||||
|
||||
#### Scenario: Multi-turn E2E proves run isolation
|
||||
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
|
||||
- **THEN** database inspection SHALL show two `diagnosis_run` rows
|
||||
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
|
||||
- **AND** Redis session metadata SHALL still show multi-turn context continuity
|
||||
|
||||
#### Scenario: Baseline drift is checked
|
||||
- **WHEN** verification is complete
|
||||
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
|
||||
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
## 1. Schema and Compatibility Foundation
|
||||
|
||||
- [x] 1.1 Add Flyway migration for `chat_session` and `diagnosis_run` with indexes for unique `session_id`, unique `run_id`, latest-run lookup, and run summary listing.
|
||||
- [x] 1.2 Add nullable `run_id` columns to `agent_step` and `tool_invocation` with indexes for run-scoped trace queries.
|
||||
- [x] 1.3 Backfill one compatibility `diagnosis_run` for each existing `diagnosis_session` row.
|
||||
- [x] 1.4 Backfill historical `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run for the same `session_id`.
|
||||
- [x] 1.5 Add JPA entities and repositories for `ChatSession` and `DiagnosisRun`, including latest-run and run-id lookup methods.
|
||||
- [x] 1.6 Add focused migration/repository verification that proves old data remains queryable and run lookup methods work.
|
||||
- [x] 1.7 Phase 1 gate: run the smallest relevant test/build check, inspect DB migration behavior, update OpenSpec task status, archive phase evidence, and commit before starting Phase 2.
|
||||
|
||||
## 2. Chat Run Write Path
|
||||
|
||||
- [ ] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
|
||||
- [ ] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
|
||||
- [ ] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
|
||||
- [ ] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
|
||||
- [ ] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
|
||||
- [ ] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
|
||||
- [ ] 2.7 Change `/api/chat` response DTO to include official `runId`.
|
||||
- [ ] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
|
||||
- [ ] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
|
||||
|
||||
## 3. Trace Read Path and Run Listing
|
||||
|
||||
- [ ] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
|
||||
- [ ] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
|
||||
- [ ] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
|
||||
- [ ] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
|
||||
- [ ] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
|
||||
- [ ] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
|
||||
- [ ] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
|
||||
|
||||
## 4. Feedback and Case Library Run Binding
|
||||
|
||||
- [ ] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
|
||||
- [ ] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
|
||||
- [ ] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
|
||||
- [ ] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
|
||||
- [ ] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
|
||||
- [ ] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
|
||||
- [ ] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
|
||||
|
||||
## 5. AIOps Run Isolation
|
||||
|
||||
- [ ] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
|
||||
- [ ] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
|
||||
- [ ] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
|
||||
- [ ] 5.4 Change AIOps final report, status, counts, and `aiops_rule_evaluation` writes to the current `diagnosis_run`.
|
||||
- [ ] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
|
||||
- [ ] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
|
||||
|
||||
## 6. Demo, Trace UI, Documentation, and Verification
|
||||
|
||||
- [ ] 6.1 Update demo scripts to read `runId` from Chat/AIOps responses and pass `?runId=...` to Trace API.
|
||||
- [ ] 6.2 Update Trace UI to accept `?sessionId=...&runId=...` and query exact trace when `runId` is present.
|
||||
- [ ] 6.3 Update MVP table and architecture docs for `chat_session`, `diagnosis_run`, `run_id`, and transitional `case_library.diagnosis_id` semantics.
|
||||
- [ ] 6.4 Run final same-session multi-turn E2E using Maven startup if needed; collect DB evidence through `scripts/query_mysql.py` and inspect `logs/`.
|
||||
- [ ] 6.5 Run or explicitly evaluate the relevant baseline diff command and document whether drift is expected or a regression.
|
||||
- [ ] 6.6 Final gate: ensure all OpenSpec tasks are checked, no new writes depend on `diagnosis_session`, phase evidence is archived, final commit is created, and the change is ready for OpenSpec archive.
|
||||
Reference in New Issue
Block a user