feat(trace): add run-scoped trace reads

This commit is contained in:
zhuyongxin
2026-07-10 20:07:23 +08:00
parent 26d5529280
commit 027aed1eeb
11 changed files with 620 additions and 111 deletions
@@ -143,6 +143,12 @@ Audit conclusions:
- Clarified that `chat_session.expires_at` is nullable directory metadata / best-effort TTL snapshot, not mandatory persisted conversation history.
- Clarified document review findings before continuing Phase 2: OpenSpec task phases are authoritative over the older active issue phase sketch, and AIOps rule evaluation is stored in `diagnosis_run.self_evaluation.aiops_rule_evaluation`, not a separate table.
## Document Review Follow-up Before Phase 3 Gate
- Clarified AIOps SSE compatibility before Phase 5: keep SSE event name `message`, emit a JSON `SseMessage` with `type=metadata`, and preserve existing content message shape for report streaming.
- Clarified Feedback API before Phase 4: request `runId` is preferred, response always includes the bound `runId` and `fallbackToLatestRun`, and wrong-session run binding uses the existing failed feedback response path.
- Clarified run-list API before Phase 3 gate: `GET /api/chat/session/{sessionId}/runs` returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
## Phase 2 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
@@ -151,3 +157,14 @@ Audit conclusions:
- Switched Chat run completion, failure, self-evaluation, metrics, verifier support reads, gatekeeper validation, and evidence scoring to run-scoped data.
- Added `/api/chat` response `runId` and focused tests for valid run creation, invalid request no-run behavior, run-scoped trace consumers, and same-session multi-turn run creation.
- Phase 2 gate evidence is recorded in `phase-2-evidence.md`.
## Phase 3 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. The committed OpenSpec remained the execution source; a document review follow-up tightened DTO/SSE/listing contracts before the Phase 3 gate.
- Implemented latest-run trace resolution using `diagnosis_run.created_at DESC, id DESC`.
- Implemented exact trace lookup for `sessionId + runId` with run/session ownership validation.
- Changed trace details to read `agent_step` and `tool_invocation` by `run_id`, while retaining a legacy `diagnosis_session` fallback for historical compatibility.
- Added trace response fields for resolved `runId`, chat session metadata, run metadata, and per-row `runId`.
- Added lightweight `GET /api/chat/session/{sessionId}/runs` backed by `diagnosis_run` summaries.
- Added focused tests for latest trace, exact first trace, exact second trace, wrong-session rejection, missing session, read-only trace behavior, run listing, and legacy fallback.
- Phase 3 E2E used Maven startup with profile `mvp-demo`; evidence is recorded in `phase-3-evidence.md`.
@@ -98,10 +98,10 @@ Interface level: L4.
- Database contract changes: new tables, new columns, backfill, indexes, and later non-null expectations for new writes.
- `/api/chat` response adds `runId`.
- `/api/ai_ops` SSE emits a compatible metadata message before report content. The metadata payload includes `sessionId` and `runId`; report content continues to stream through the existing content message shape.
- `/api/ai_ops` SSE emits a compatible metadata message before report content. It keeps the existing SSE event name `message` and sends a JSON `SseMessage` with `type=metadata`; the metadata payload includes `sessionId` and `runId`. Report content continues to stream through the existing `type=content` message shape.
- Trace API accepts optional `runId`.
- Feedback request accepts preferred `runId` and returns fallback binding metadata when omitted.
- New run summary API: `GET /api/chat/session/{sessionId}/runs`.
- Feedback request accepts preferred `runId`. Feedback response includes the actual bound `runId` and `fallbackToLatestRun`; wrong-session `runId`, missing session, and missing run use the existing failed feedback response path with HTTP 400 from `FeedbackController`.
- New run summary API: `GET /api/chat/session/{sessionId}/runs`, returned through the existing `ApiResponse<List<RunSummary>>` wrapper. A session with metadata but no runs returns an empty list; a missing session returns the existing not-found/error behavior.
Compatibility:
@@ -0,0 +1,136 @@
# Phase 3 Evidence: Trace Read Path and Run Listing
## Scope
Phase 3 implements run-scoped trace reads and lightweight run listing:
- `GET /api/diagnosis/{sessionId}/trace` resolves the latest run by `diagnosis_run.created_at DESC, id DESC`.
- `GET /api/diagnosis/{sessionId}/trace?runId=...` returns the exact run after validating run/session ownership.
- Trace responses include resolved run metadata and run-scoped step/tool rows.
- `GET /api/chat/session/{sessionId}/runs` returns lightweight run summaries without expanding trace detail rows.
## Verification Commands
- `mvn -q clean "-Dtest=DiagnosisTraceServiceTest" test`
- `mvn -q "-Dtest=DiagnosisTraceServiceTest,DiagnosisTraceEvaluatorTest,ChatControllerTest" test`
- `openspec validate session-run-trace-isolation --strict`
After the document review follow-up, OpenSpec strict validation was run again:
- `openspec validate session-run-trace-isolation --strict`
Result: passed.
## E2E Runtime
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured logs and artifacts:
- `target/e2e/phase3-mvn-20260710-191131.out.log`
- `target/e2e/phase3-mvn-20260710-191131.err.log`
- `target/e2e/phase3-request-round1.json`
- `target/e2e/phase3-response-round1.json`
- `target/e2e/phase3-request-round2.json`
- `target/e2e/phase3-response-round2.json`
- `target/e2e/phase3-trace-latest.json`
- `target/e2e/phase3-trace-first.json`
- `target/e2e/phase3-trace-second.json`
- `target/e2e/phase3-runs.json`
- `target/e2e/phase3-trace-wrong-session.json`
- `target/e2e/phase3-summary.json`
The E2E Maven process was stopped after evidence collection.
Note: `logs/application.log` was checked, but it did not contain the Phase 3 E2E session entries and its last write time was earlier than this E2E run. The Phase 3 runtime application logs were captured in the Maven stdout artifact above.
## E2E Summary
Session:
```text
sessionId = e2e-phase3-codex-20260710-1912
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
```
Observed behavior from `phase3-summary.json`:
```text
distinctRunIds = true
latestResolvedRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
firstTraceRunId = run-fdccbe21-e050-4f92-a741-a062aef59644
firstSteps = 9
firstTools = 13
secondTraceRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
secondSteps = 2
secondTools = 0
runListCount = 2
runListFirst = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
wrongSessionStatus = 400
```
Log evidence from `phase3-mvn-20260710-191131.out.log`:
- Round 1 request was received for `e2e-phase3-codex-20260710-1912`.
- Redis session was created for the first round and reused by the second round.
- Message pair count reached `2` after round 2.
- Latest trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
- Exact first trace request returned `run-fdccbe21-e050-4f92-a741-a062aef59644`.
- Exact second trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
- Wrong-session exact trace returned HTTP 400 with `runId does not belong to sessionId`.
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run` rows:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | SUCCESS | CHAT | step_count=2 | tool_call_count=0 | created_at=2026-07-10 19:15:40
run-fdccbe21-e050-4f92-a741-a062aef59644 | SUCCESS | CHAT | step_count=9 | tool_call_count=13 | created_at=2026-07-10 19:12:33
```
`agent_step` rows grouped by run:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | step_rows=2 | min_step=0 | max_step=1
run-fdccbe21-e050-4f92-a741-a062aef59644 | step_rows=9 | min_step=0 | max_step=4
```
`tool_invocation` rows grouped by run:
```text
run-fdccbe21-e050-4f92-a741-a062aef59644 | tool_rows=13
```
Missing run id checks for this session:
```text
agent_step missing run_id = 0
tool_invocation missing run_id = 0
```
`chat_session` metadata:
```text
session_id=e2e-phase3-codex-20260710-1912 | status=ACTIVE | message_pair_count=2
```
## Document Review Follow-up
Before closing Phase 3, the OpenSpec docs were tightened for upcoming phases:
- AIOps SSE metadata shape: keep SSE event name `message`, use JSON `SseMessage` with `type=metadata`, preserve existing content message shape.
- Feedback DTO contract: request `runId` is preferred; response includes bound `runId` and `fallbackToLatestRun`; wrong-session run binding fails instead of updating either run.
- Run-list API: returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
OpenSpec strict validation passed after these document changes.
## Conclusion
Phase 3 satisfies the run-scoped trace read and run-list contract. Same-session multi-turn E2E proves latest-run compatibility, exact-run replay, run-list ordering, run/session ownership rejection, and no missing `run_id` rows for new trace data.
@@ -72,6 +72,8 @@ Level: L4 database/API contract migration with compatibility behavior.
- New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`.
- New API: `GET /api/chat/session/{sessionId}/runs`.
- Feedback request gains optional/preferred `runId`.
- Feedback response returns bound `runId` and `fallbackToLatestRun`.
- `/api/ai_ops` keeps SSE event name `message` and emits a `type=metadata` JSON message containing `sessionId` and `runId` before content.
- Database contract changes include new tables and new `run_id` columns.
- Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable.
@@ -72,10 +72,18 @@ The system SHALL provide a lightweight run-list API for a Chat Session.
#### Scenario: Run list returns summaries
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
- **THEN** the system SHALL return run summaries from `diagnosis_run`
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
#### Scenario: Run list handles session without runs
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
- **THEN** the system SHALL return a successful empty list
#### Scenario: Run list rejects missing session
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
- **THEN** the system SHALL use the existing not-found/error response behavior
### Requirement: Feedback SHALL bind to diagnosis runs
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
@@ -83,6 +91,8 @@ The system SHALL bind new feedback to a diagnosis run rather than an ambiguous m
- **WHEN** a feedback request includes `sessionId` and `runId`
- **THEN** the system SHALL validate that the run belongs to the session
- **AND** it SHALL update feedback on that run
- **AND** the response SHALL include the actual bound `runId`
- **AND** the response SHALL include `fallbackToLatestRun=false`
#### Scenario: Feedback without runId falls back observably
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
@@ -90,6 +100,10 @@ The system SHALL bind new feedback to a diagnosis run rather than an ambiguous m
- **AND** the response SHALL include `fallbackToLatestRun=true`
- **AND** the response SHALL include the actual bound `runId`
#### Scenario: Feedback rejects run from another session
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return a failed feedback response instead of updating either run
#### Scenario: Useful feedback creates case from run
- **WHEN** feedback for a run is `useful`
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
@@ -107,8 +121,11 @@ The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops`
#### Scenario: AIOps SSE exposes runId
- **WHEN** `/api/ai_ops` streams response metadata to the caller
- **THEN** the stream SHALL send a compatible metadata message before report content
- **AND** the SSE event name SHALL remain `message`
- **AND** the message type SHALL be `metadata`
- **AND** the metadata payload SHALL expose the resolved `sessionId`
- **AND** the metadata payload SHALL expose the created `runId`
- **AND** report content SHALL continue to use the existing content message shape
### Requirement: Migration SHALL preserve historical trace access
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
@@ -22,13 +22,13 @@
## 3. Trace Read Path and Run Listing
- [ ] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
- [ ] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
- [ ] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
- [ ] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
- [ ] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
- [ ] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
- [ ] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
- [x] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
- [x] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
- [x] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
- [x] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
- [x] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
- [x] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
- [x] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
## 4. Feedback and Case Library Run Binding