feat(trace): isolate chat runs

This commit is contained in:
zhuyongxin
2026-07-10 19:02:04 +08:00
parent 6fdbd34bab
commit 26d5529280
21 changed files with 610 additions and 134 deletions
@@ -2,7 +2,7 @@
## sm-flow State
- Checkpoint: Discover
- Checkpoint: Apply / Phase 2
- Scale: complex
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
- Change slug: `session-run-trace-isolation`
@@ -141,3 +141,13 @@ Audit conclusions:
- Clarified AIOps SSE compatibility: emit a metadata message containing `sessionId` and `runId` before report content while preserving the existing content stream shape.
- Clarified that run/session ownership is enforced by service-layer validation and indexed lookup in this change; database foreign keys are intentionally deferred to preserve compatibility with historical orphan detail rows and rollback.
- Clarified that `chat_session.expires_at` is nullable directory metadata / best-effort TTL snapshot, not mandatory persisted conversation history.
- Clarified document review findings before continuing Phase 2: OpenSpec task phases are authoritative over the older active issue phase sketch, and AIOps rule evaluation is stored in `diagnosis_run.self_evaluation.aiops_rule_evaluation`, not a separate table.
## Phase 2 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
- Implemented unified Chat execution context carrying `sessionId` and `runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
- Switched valid Chat writes to create/update `chat_session` metadata and create one `diagnosis_run` per request.
- Switched Chat run completion, failure, self-evaluation, metrics, verifier support reads, gatekeeper validation, and evidence scoring to run-scoped data.
- Added `/api/chat` response `runId` and focused tests for valid run creation, invalid request no-run behavior, run-scoped trace consumers, and same-session multi-turn run creation.
- Phase 2 gate evidence is recorded in `phase-2-evidence.md`.
@@ -118,9 +118,12 @@ Compatibility:
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
6. Deploy repository and read-path compatibility.
7. Switch Chat write path to `chat_session + diagnosis_run`.
8. Switch Trace, evaluation, feedback, case-library, demo, and UI paths.
9. Switch AIOps write path.
10. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
8. Switch Chat evaluation and verifier support reads to run-scoped data as part of the Chat write-path cutover.
9. Switch Trace run resolution and run listing.
10. Switch Feedback and case-library paths.
11. Switch AIOps write path, including `diagnosis_run.self_evaluation.aiops_rule_evaluation`.
12. Switch demo scripts, Trace UI, and MVP docs.
13. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
Rollback:
@@ -0,0 +1,129 @@
# Phase 2 Evidence: Chat Run Write Path
Date: 2026-07-10
## Scope
Phase 2 switched Chat writes from session-scoped execution state to run-scoped execution state:
- Chat creates/updates `chat_session` metadata.
- Each valid `/api/chat` execution creates one `diagnosis_run`.
- Chat execution context carries `sessionId + runId` through `RunnableConfig` and `SessionContextHolder`.
- `agent_step.run_id` and `tool_invocation.run_id` are written for Chat runs.
- Chat completion/failure/status/answer/self-evaluation/counts are written to `diagnosis_run`.
- verifier/gatekeeper/evaluation reads use run-scoped tool rows when `runId` is available.
- `/api/chat` response includes official `runId`.
## Static / Unit Verification
Commands:
```powershell
mvn -q clean test-compile
mvn -q "-Dtest=ChatControllerTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ExecutorGatekeeperServiceTest" test
openspec validate session-run-trace-isolation --strict
```
Result:
- `test-compile` passed.
- Focused Phase 2 tests passed.
- OpenSpec strict validation passed.
Focused coverage:
- valid Chat creates `chat_session` and `diagnosis_run`;
- invalid blank Chat request returns before `ChatService`, so no run is created;
- same `sessionId` across two Chat turns creates two distinct `runId` values;
- `ToolInvocationRecorder` copies `runId` from execution context;
- verifier trace summary reads by run;
- gatekeeper validates by run;
- evaluation writes rule evaluation to `diagnosis_run.self_evaluation`.
## E2E Verification
Startup command:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Process stdout/stderr:
- `target/e2e/phase2-mvn-20260710-185102.out.log`
- `target/e2e/phase2-mvn-20260710-185102.err.log`
Primary E2E session:
```text
sessionId = e2e-phase2-codex-20260710-1856
round 1 runId = run-24b6f04c-94a0-43cb-b94f-4f0141f9050d
round 2 runId = run-a7a2be77-697f-495a-a779-91afd1d8589c
```
HTTP evidence:
- `target/e2e/phase2-utf8-request-round1.json`
- `target/e2e/phase2-utf8-response-round1.json`
- `target/e2e/phase2-utf8-request-round2.json`
- `target/e2e/phase2-utf8-response-round2.json`
- both Chat responses returned `code=200`, `data.success=true`, the same `sessionId`, and distinct `runId` values.
Redis/session continuity evidence:
- `target/e2e/phase2-utf8-chat-session-response.json`
- response returned `messagePairCount=2`.
- logs show the second request entered with `会话历史消息对数: 1` and completed with `当前消息对数: 2`.
## Database Inspection
DB inspection used `scripts/query_mysql.py`.
Saved query outputs:
- `target/e2e/phase2-utf8-db-runs.txt`
- `target/e2e/phase2-utf8-db-chat-session.txt`
- `target/e2e/phase2-utf8-db-agent-steps.txt`
- `target/e2e/phase2-utf8-db-tool-invocations.txt`
- `target/e2e/phase2-utf8-db-missing-runid.txt`
Observed rows:
```text
diagnosis_run:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | SUCCESS | CHAT | step_count=8 | tool_call_count=12
run-a7a2be77-697f-495a-a779-91afd1d8589c | SUCCESS | CHAT | step_count=2 | tool_call_count=1
chat_session:
e2e-phase2-codex-20260710-1856 | ACTIVE | message_pair_count=2
agent_step grouped by run_id:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 8
run-a7a2be77-697f-495a-a779-91afd1d8589c | 2
tool_invocation grouped by run_id:
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 12
run-a7a2be77-697f-495a-a779-91afd1d8589c | 1
missing run_id for this session:
agent_step = 0
tool_invocation = 0
```
## Log Review
Saved log excerpts:
- `target/e2e/phase2-utf8-application-log-excerpt.txt`
- `target/e2e/phase2-utf8-mvn-log-excerpt.txt`
Findings:
- Chat logs show second turn reused Redis history for the same session.
- `EvaluationService` wrote scoring results to both run ids.
- No E2E-specific application exception was observed for `e2e-phase2-codex-20260710-1856`.
- Earlier `/actuator/health` probes produced expected 500/no-resource noise because the actuator health endpoint is not exposed; the E2E readiness check used `/api/chat` instead.
## Notes
`codebase-retrieval` and LSP tools were not available in this environment. Call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
@@ -26,7 +26,7 @@ Introduce a stable split between conversation state and execution state:
`runId` becomes an official API field:
- `/api/chat` returns `sessionId + runId`.
- `/api/ai_ops` emits or returns `runId` in the SSE-compatible protocol.
- `/api/ai_ops` emits an SSE-compatible metadata message containing `sessionId` and `runId` before report content.
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
@@ -82,4 +82,3 @@ Level: L4 database/API contract migration with compatibility behavior.
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
@@ -61,6 +61,7 @@ The system SHALL allow callers to query a diagnosis trace by `sessionId` alone f
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
#### Scenario: Trace rejects run from another session
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
@@ -101,6 +102,7 @@ The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops`
- **WHEN** `/api/ai_ops` starts a valid execution
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
#### Scenario: AIOps SSE exposes runId
- **WHEN** `/api/ai_ops` streams response metadata to the caller
@@ -10,15 +10,15 @@
## 2. Chat Run Write Path
- [ ] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
- [ ] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
- [ ] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
- [ ] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
- [ ] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
- [ ] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
- [ ] 2.7 Change `/api/chat` response DTO to include official `runId`.
- [ ] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
- [ ] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
- [x] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
- [x] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
- [x] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
- [x] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
- [x] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
- [x] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
- [x] 2.7 Change `/api/chat` response DTO to include official `runId`.
- [x] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
- [x] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
## 3. Trace Read Path and Run Listing
@@ -45,7 +45,7 @@
- [ ] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
- [ ] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
- [ ] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
- [ ] 5.4 Change AIOps final report, status, counts, and `aiops_rule_evaluation` writes to the current `diagnosis_run`.
- [ ] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
- [ ] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
- [ ] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.