feat(trace): isolate chat runs
This commit is contained in:
@@ -2,7 +2,7 @@
|
||||
|
||||
## sm-flow State
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Checkpoint: Apply / Phase 2
|
||||
- Scale: complex
|
||||
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
@@ -141,3 +141,13 @@ Audit conclusions:
|
||||
- Clarified AIOps SSE compatibility: emit a metadata message containing `sessionId` and `runId` before report content while preserving the existing content stream shape.
|
||||
- Clarified that run/session ownership is enforced by service-layer validation and indexed lookup in this change; database foreign keys are intentionally deferred to preserve compatibility with historical orphan detail rows and rollback.
|
||||
- Clarified that `chat_session.expires_at` is nullable directory metadata / best-effort TTL snapshot, not mandatory persisted conversation history.
|
||||
- Clarified document review findings before continuing Phase 2: OpenSpec task phases are authoritative over the older active issue phase sketch, and AIOps rule evaluation is stored in `diagnosis_run.self_evaluation.aiops_rule_evaluation`, not a separate table.
|
||||
|
||||
## Phase 2 Apply Notes
|
||||
|
||||
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
|
||||
- Implemented unified Chat execution context carrying `sessionId` and `runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- Switched valid Chat writes to create/update `chat_session` metadata and create one `diagnosis_run` per request.
|
||||
- Switched Chat run completion, failure, self-evaluation, metrics, verifier support reads, gatekeeper validation, and evidence scoring to run-scoped data.
|
||||
- Added `/api/chat` response `runId` and focused tests for valid run creation, invalid request no-run behavior, run-scoped trace consumers, and same-session multi-turn run creation.
|
||||
- Phase 2 gate evidence is recorded in `phase-2-evidence.md`.
|
||||
|
||||
@@ -118,9 +118,12 @@ Compatibility:
|
||||
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
|
||||
6. Deploy repository and read-path compatibility.
|
||||
7. Switch Chat write path to `chat_session + diagnosis_run`.
|
||||
8. Switch Trace, evaluation, feedback, case-library, demo, and UI paths.
|
||||
9. Switch AIOps write path.
|
||||
10. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
|
||||
8. Switch Chat evaluation and verifier support reads to run-scoped data as part of the Chat write-path cutover.
|
||||
9. Switch Trace run resolution and run listing.
|
||||
10. Switch Feedback and case-library paths.
|
||||
11. Switch AIOps write path, including `diagnosis_run.self_evaluation.aiops_rule_evaluation`.
|
||||
12. Switch demo scripts, Trace UI, and MVP docs.
|
||||
13. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
|
||||
|
||||
Rollback:
|
||||
|
||||
|
||||
@@ -0,0 +1,129 @@
|
||||
# Phase 2 Evidence: Chat Run Write Path
|
||||
|
||||
Date: 2026-07-10
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 2 switched Chat writes from session-scoped execution state to run-scoped execution state:
|
||||
|
||||
- Chat creates/updates `chat_session` metadata.
|
||||
- Each valid `/api/chat` execution creates one `diagnosis_run`.
|
||||
- Chat execution context carries `sessionId + runId` through `RunnableConfig` and `SessionContextHolder`.
|
||||
- `agent_step.run_id` and `tool_invocation.run_id` are written for Chat runs.
|
||||
- Chat completion/failure/status/answer/self-evaluation/counts are written to `diagnosis_run`.
|
||||
- verifier/gatekeeper/evaluation reads use run-scoped tool rows when `runId` is available.
|
||||
- `/api/chat` response includes official `runId`.
|
||||
|
||||
## Static / Unit Verification
|
||||
|
||||
Commands:
|
||||
|
||||
```powershell
|
||||
mvn -q clean test-compile
|
||||
mvn -q "-Dtest=ChatControllerTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ExecutorGatekeeperServiceTest" test
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- `test-compile` passed.
|
||||
- Focused Phase 2 tests passed.
|
||||
- OpenSpec strict validation passed.
|
||||
|
||||
Focused coverage:
|
||||
|
||||
- valid Chat creates `chat_session` and `diagnosis_run`;
|
||||
- invalid blank Chat request returns before `ChatService`, so no run is created;
|
||||
- same `sessionId` across two Chat turns creates two distinct `runId` values;
|
||||
- `ToolInvocationRecorder` copies `runId` from execution context;
|
||||
- verifier trace summary reads by run;
|
||||
- gatekeeper validates by run;
|
||||
- evaluation writes rule evaluation to `diagnosis_run.self_evaluation`.
|
||||
|
||||
## E2E Verification
|
||||
|
||||
Startup command:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Process stdout/stderr:
|
||||
|
||||
- `target/e2e/phase2-mvn-20260710-185102.out.log`
|
||||
- `target/e2e/phase2-mvn-20260710-185102.err.log`
|
||||
|
||||
Primary E2E session:
|
||||
|
||||
```text
|
||||
sessionId = e2e-phase2-codex-20260710-1856
|
||||
round 1 runId = run-24b6f04c-94a0-43cb-b94f-4f0141f9050d
|
||||
round 2 runId = run-a7a2be77-697f-495a-a779-91afd1d8589c
|
||||
```
|
||||
|
||||
HTTP evidence:
|
||||
|
||||
- `target/e2e/phase2-utf8-request-round1.json`
|
||||
- `target/e2e/phase2-utf8-response-round1.json`
|
||||
- `target/e2e/phase2-utf8-request-round2.json`
|
||||
- `target/e2e/phase2-utf8-response-round2.json`
|
||||
- both Chat responses returned `code=200`, `data.success=true`, the same `sessionId`, and distinct `runId` values.
|
||||
|
||||
Redis/session continuity evidence:
|
||||
|
||||
- `target/e2e/phase2-utf8-chat-session-response.json`
|
||||
- response returned `messagePairCount=2`.
|
||||
- logs show the second request entered with `会话历史消息对数: 1` and completed with `当前消息对数: 2`.
|
||||
|
||||
## Database Inspection
|
||||
|
||||
DB inspection used `scripts/query_mysql.py`.
|
||||
|
||||
Saved query outputs:
|
||||
|
||||
- `target/e2e/phase2-utf8-db-runs.txt`
|
||||
- `target/e2e/phase2-utf8-db-chat-session.txt`
|
||||
- `target/e2e/phase2-utf8-db-agent-steps.txt`
|
||||
- `target/e2e/phase2-utf8-db-tool-invocations.txt`
|
||||
- `target/e2e/phase2-utf8-db-missing-runid.txt`
|
||||
|
||||
Observed rows:
|
||||
|
||||
```text
|
||||
diagnosis_run:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | SUCCESS | CHAT | step_count=8 | tool_call_count=12
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | SUCCESS | CHAT | step_count=2 | tool_call_count=1
|
||||
|
||||
chat_session:
|
||||
e2e-phase2-codex-20260710-1856 | ACTIVE | message_pair_count=2
|
||||
|
||||
agent_step grouped by run_id:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 8
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | 2
|
||||
|
||||
tool_invocation grouped by run_id:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 12
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | 1
|
||||
|
||||
missing run_id for this session:
|
||||
agent_step = 0
|
||||
tool_invocation = 0
|
||||
```
|
||||
|
||||
## Log Review
|
||||
|
||||
Saved log excerpts:
|
||||
|
||||
- `target/e2e/phase2-utf8-application-log-excerpt.txt`
|
||||
- `target/e2e/phase2-utf8-mvn-log-excerpt.txt`
|
||||
|
||||
Findings:
|
||||
|
||||
- Chat logs show second turn reused Redis history for the same session.
|
||||
- `EvaluationService` wrote scoring results to both run ids.
|
||||
- No E2E-specific application exception was observed for `e2e-phase2-codex-20260710-1856`.
|
||||
- Earlier `/actuator/health` probes produced expected 500/no-resource noise because the actuator health endpoint is not exposed; the E2E readiness check used `/api/chat` instead.
|
||||
|
||||
## Notes
|
||||
|
||||
`codebase-retrieval` and LSP tools were not available in this environment. Call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
|
||||
@@ -26,7 +26,7 @@ Introduce a stable split between conversation state and execution state:
|
||||
`runId` becomes an official API field:
|
||||
|
||||
- `/api/chat` returns `sessionId + runId`.
|
||||
- `/api/ai_ops` emits or returns `runId` in the SSE-compatible protocol.
|
||||
- `/api/ai_ops` emits an SSE-compatible metadata message containing `sessionId` and `runId` before report content.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
|
||||
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
|
||||
@@ -82,4 +82,3 @@ Level: L4 database/API contract migration with compatibility behavior.
|
||||
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
|
||||
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
|
||||
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
|
||||
|
||||
|
||||
@@ -61,6 +61,7 @@ The system SHALL allow callers to query a diagnosis trace by `sessionId` alone f
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
|
||||
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
|
||||
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
|
||||
|
||||
#### Scenario: Trace rejects run from another session
|
||||
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
|
||||
@@ -101,6 +102,7 @@ The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops`
|
||||
- **WHEN** `/api/ai_ops` starts a valid execution
|
||||
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
|
||||
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
|
||||
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
|
||||
|
||||
#### Scenario: AIOps SSE exposes runId
|
||||
- **WHEN** `/api/ai_ops` streams response metadata to the caller
|
||||
|
||||
@@ -10,15 +10,15 @@
|
||||
|
||||
## 2. Chat Run Write Path
|
||||
|
||||
- [ ] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
|
||||
- [ ] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
|
||||
- [ ] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
|
||||
- [ ] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
|
||||
- [ ] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
|
||||
- [ ] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
|
||||
- [ ] 2.7 Change `/api/chat` response DTO to include official `runId`.
|
||||
- [ ] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
|
||||
- [ ] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
|
||||
- [x] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
|
||||
- [x] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
|
||||
- [x] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
|
||||
- [x] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
|
||||
- [x] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
|
||||
- [x] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
|
||||
- [x] 2.7 Change `/api/chat` response DTO to include official `runId`.
|
||||
- [x] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
|
||||
- [x] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
|
||||
|
||||
## 3. Trace Read Path and Run Listing
|
||||
|
||||
@@ -45,7 +45,7 @@
|
||||
- [ ] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
|
||||
- [ ] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
|
||||
- [ ] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
|
||||
- [ ] 5.4 Change AIOps final report, status, counts, and `aiops_rule_evaluation` writes to the current `diagnosis_run`.
|
||||
- [ ] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
|
||||
- [ ] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
|
||||
- [ ] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user