docs(openspec): archive session run isolation

This commit is contained in:
zhuyongxin
2026-07-10 22:52:39 +08:00
parent f9df94377b
commit 3578709896
21 changed files with 511 additions and 45 deletions
@@ -1,24 +1,33 @@
## Purpose
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same diagnosis run.
## Requirements
### Requirement: Diagnosis trace can be queried by session id
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
#### Scenario: Existing session trace is returned
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
#### Scenario: Existing session latest trace is returned
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
#### Scenario: Existing session exact trace is returned
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system validates that `runId` belongs to `sessionId`
- **AND** it returns a success response containing only the trace data for that run
#### Scenario: Missing session returns not found
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
- **THEN** the system returns a 404 response using the existing session-not-found error contract
### Requirement: Trace aggregation is read-only
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
#### Scenario: Trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
#### Scenario: Exact trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
### Requirement: MVP demo profile is available
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
@@ -28,11 +37,11 @@ The system SHALL provide an `mvp-demo` Spring profile that documents the demo ru
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
### Requirement: End-to-end MVP acceptance case is documented
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same `sessionId + runId`.
#### Scenario: Reviewer follows the acceptance case
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
- **THEN** they can run the application, submit a diagnosis question, query the exact trace endpoint, and submit feedback for the same run
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
@@ -51,10 +60,11 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
#### Scenario: Demo script sends the fixed diagnosis request
- **WHEN** the demo script is executed against a running local service
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
- **AND** it SHALL read the returned run id for exact trace and feedback calls
#### Scenario: Demo script captures review artifacts
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
- **THEN** it SHALL write chat, exact trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
@@ -62,8 +72,8 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same session id
- **AND** it SHALL submit useful feedback for that session
- **AND** it SHALL fetch the trace for the same `sessionId + runId`
- **AND** it SHALL submit useful feedback for that run
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
#### Scenario: interview demo check fails with actionable readiness output
@@ -77,7 +87,7 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and run-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
@@ -85,7 +95,7 @@ The MVP demo SHALL document which trace fields to inspect for evidence, verifier
### Requirement: MVP demo SHALL provide a browser trace workbench
The MVP demo SHALL provide a browser-accessible static page for inspecting one
diagnosis trace by session id using the existing read-only Trace API.
diagnosis trace by session id and optional run id using the existing read-only Trace API.
#### Scenario: Existing trace renders in the workbench
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
@@ -93,6 +103,11 @@ diagnosis trace by session id using the existing read-only Trace API.
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
invocations, verifier evaluation, and final answer when present
#### Scenario: Exact trace renders in the workbench
- **WHEN** a reviewer opens the Trace workbench with `?sessionId=...&runId=...`
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace?runId=...`
- **AND** it SHALL render only that run's trace data
#### Scenario: Trace workbench handles missing or failed traces
- **WHEN** the Trace API returns an error or the session id is empty
- **THEN** the page SHALL show a clear error or empty state without mutating any
@@ -0,0 +1,181 @@
# session-run-trace-isolation Specification
## Purpose
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
## Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
#### Scenario: Valid Chat execution creates session metadata and a run
- **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId`
- **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId`
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
- **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId`
#### Scenario: Invalid Chat request does not create a run
- **WHEN** a `/api/chat` request fails parameter validation before execution
- **THEN** the system SHALL NOT create a `diagnosis_run`
#### Scenario: Chat session stores metadata only
- **WHEN** a Chat request completes
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
- **AND** it SHALL NOT store full conversation message history
### Requirement: Chat responses SHALL expose run identity
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
#### Scenario: Chat response includes runId
- **WHEN** `/api/chat` returns a successful response
- **THEN** the response SHALL include `sessionId`
- **AND** the response SHALL include `runId` for the created diagnosis run
#### Scenario: Multi-turn Chat keeps one session and multiple runs
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
### Requirement: Trace details SHALL be scoped by run
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
#### Scenario: Agent steps are recorded with runId
- **WHEN** an Agent model step is persisted during a diagnosis run
- **THEN** the `agent_step` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Tool invocations are recorded with runId
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Run metrics count only current run rows
- **WHEN** a diagnosis run completes
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
### Requirement: Trace API SHALL support latest-run and exact-run queries
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
#### Scenario: Trace without runId resolves latest run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
- **AND** the response SHALL include the resolved `runId`
#### Scenario: Trace with runId returns exact run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
#### Scenario: Trace rejects run from another session
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
### Requirement: Session runs SHALL be listable without expanding trace details
The system SHALL provide a lightweight run-list API for a Chat Session.
#### Scenario: Run list returns summaries
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
#### Scenario: Run list handles session without runs
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
- **THEN** the system SHALL return a successful empty list
#### Scenario: Run list rejects missing session
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
- **THEN** the system SHALL use the existing not-found/error response behavior
### Requirement: Feedback SHALL bind to diagnosis runs
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
#### Scenario: Feedback with runId updates specified run
- **WHEN** a feedback request includes `sessionId` and `runId`
- **THEN** the system SHALL validate that the run belongs to the session
- **AND** it SHALL update feedback on that run
- **AND** the response SHALL include the actual bound `runId`
- **AND** the response SHALL include `fallbackToLatestRun=false`
#### Scenario: Feedback without runId falls back observably
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** at least one `diagnosis_run` exists for that session
- **THEN** the system SHALL bind feedback to the latest run for that session
- **AND** the response SHALL include `fallbackToLatestRun=true`
- **AND** the response SHALL include the actual bound `runId`
#### Scenario: Historical feedback without run-backed data remains compatible
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** no `diagnosis_run` exists for that session
- **AND** a historical `diagnosis_session` row exists for that session
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
- **AND** the response SHALL NOT claim latest-run fallback
- **AND** the response MAY omit `runId`
#### Scenario: Feedback rejects run from another session
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return a failed feedback response instead of updating either run
#### Scenario: Useful feedback creates case from run
- **WHEN** feedback for a run is `useful`
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
### Requirement: AIOps executions SHALL use run isolation
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
#### Scenario: AIOps creates run
- **WHEN** `/api/ai_ops` starts a valid execution
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
#### Scenario: AIOps SSE exposes runId
- **WHEN** `/api/ai_ops` streams response metadata to the caller
- **THEN** the stream SHALL send a compatible metadata message before report content
- **AND** the SSE event name SHALL remain `message`
- **AND** the message type SHALL be `metadata`
- **AND** the metadata payload SHALL expose the resolved `sessionId`
- **AND** the metadata payload SHALL expose the created `runId`
- **AND** report content SHALL continue to use the existing content message shape
### Requirement: Migration SHALL preserve historical trace access
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
#### Scenario: Historical session gets compatibility run
- **WHEN** migration runs on an existing `diagnosis_session` row
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
#### Scenario: Old table is retained
- **WHEN** the migration completes
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
### Requirement: Demo and Trace UI SHALL support runId
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
#### Scenario: Demo script queries exact trace
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
- **THEN** it SHALL include `runId` when querying the Trace API
#### Scenario: Trace UI honors URL runId
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
### Requirement: Run isolation SHALL be verified against baselines
The change SHALL verify both runtime behavior and evaluation baseline impact.
#### Scenario: Multi-turn E2E proves run isolation
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
- **THEN** database inspection SHALL show two `diagnosis_run` rows
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
- **AND** Redis session metadata SHALL still show multi-turn context continuity
#### Scenario: Baseline drift is checked
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression