feat(trace): bind feedback to runs

This commit is contained in:
zhuyongxin
2026-07-10 20:29:56 +08:00
parent 027aed1eeb
commit d928a1968a
14 changed files with 527 additions and 23 deletions
@@ -2,7 +2,7 @@
## sm-flow State
- Checkpoint: Apply / Phase 2
- Checkpoint: Apply / Phase 5 ready
- Scale: complex
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
- Change slug: `session-run-trace-isolation`
@@ -168,3 +168,14 @@ Audit conclusions:
- Added lightweight `GET /api/chat/session/{sessionId}/runs` backed by `diagnosis_run` summaries.
- Added focused tests for latest trace, exact first trace, exact second trace, wrong-session rejection, missing session, read-only trace behavior, run listing, and legacy fallback.
- Phase 3 E2E used Maven startup with profile `mvp-demo`; evidence is recorded in `phase-3-evidence.md`.
## Phase 4 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were still unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, focused tests, API E2E, DB inspection, and logs.
- Added `FeedbackRequest.runId` and `FeedbackResponse.runId/fallbackToLatestRun`.
- Changed `FeedbackController` to pass `runId` through to `FeedbackService`.
- Changed `FeedbackService` so new feedback prefers exact `sessionId + runId`, validates ownership, falls back to latest run when `runId` is omitted, and writes new feedback to `diagnosis_run.feedback`.
- Preserved a legacy `DiagnosisSession` fallback only when no `diagnosis_run` exists, so old data can still receive feedback during the migration window.
- Added `CaseLibraryService.createFromRun`, using `diagnosis_run.run_id` as the new automatic `case_library.diagnosis_id`; `createFromSession` remains the legacy session-id path.
- Phase 4 evidence is recorded in `phase-4-evidence.md`.
- Document review follow-up: clarified that the legacy `DiagnosisSession` feedback path only applies when no `diagnosis_run` exists for the session. It returns no bound `runId` and is not the same as latest-run fallback.
@@ -100,12 +100,13 @@ Interface level: L4.
- `/api/chat` response adds `runId`.
- `/api/ai_ops` SSE emits a compatible metadata message before report content. It keeps the existing SSE event name `message` and sends a JSON `SseMessage` with `type=metadata`; the metadata payload includes `sessionId` and `runId`. Report content continues to stream through the existing `type=content` message shape.
- Trace API accepts optional `runId`.
- Feedback request accepts preferred `runId`. Feedback response includes the actual bound `runId` and `fallbackToLatestRun`; wrong-session `runId`, missing session, and missing run use the existing failed feedback response path with HTTP 400 from `FeedbackController`.
- Feedback request accepts preferred `runId`. For run-backed data, feedback response includes the actual bound `runId` and `fallbackToLatestRun`; wrong-session `runId`, missing session, and missing run use the existing failed feedback response path with HTTP 400 from `FeedbackController`. Historical `DiagnosisSession` fallback is retained only when no `diagnosis_run` exists for the session; that legacy path has no bound `runId` and is not treated as latest-run fallback.
- New run summary API: `GET /api/chat/session/{sessionId}/runs`, returned through the existing `ApiResponse<List<RunSummary>>` wrapper. A session with metadata but no runs returns an empty list; a missing session returns the existing not-found/error behavior.
Compatibility:
- Old `sessionId`-only trace and feedback calls bind to latest run.
- Old `sessionId`-only trace and feedback calls bind to latest run when run-backed data exists.
- Historical feedback calls for sessions with no `diagnosis_run` may still bind to retained `diagnosis_session` data during the migration window.
- Old `diagnosis_session` is retained for rollback and historical comparison.
- New code must not keep writing new execution state into `diagnosis_session` after the write switch.
@@ -0,0 +1,124 @@
# Phase 4 Evidence: Feedback and Case Library Run Binding
## Scope
Phase 4 implements run-scoped feedback and run-based automatic case creation:
- Feedback request accepts preferred `runId`.
- Feedback validates run/session ownership.
- Missing `runId` falls back to the latest run and returns `fallbackToLatestRun=true` plus the bound `runId`.
- New feedback persists to `diagnosis_run.feedback`.
- Useful feedback creates or reuses `case_library` from `diagnosis_run.query` and `diagnosis_run.answer`.
- New automatic `case_library.diagnosis_id` values store `run_id`; old `session_id` values remain supported through the legacy `DiagnosisSession` path.
## Focused Tests
Focused tests:
```text
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest" test
```
Result: passed.
Coverage:
- exact `sessionId + runId` feedback updates the specified run;
- missing `runId` binds to latest run and returns `fallbackToLatestRun=true`;
- wrong-session `runId` fails without saving feedback or case data;
- useful feedback creates a case from the run;
- case creation is idempotent by `diagnosis_id`;
- legacy `DiagnosisSession` case creation still stores `diagnosis_id=session_id`;
- `FeedbackController` passes `request.runId` to `FeedbackService`.
## E2E Feedback API Check
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured artifacts:
- `target/e2e/phase4-mvn-20260710-2016.out.log`
- `target/e2e/phase4-mvn-20260710-2016.err.log`
- `target/e2e/phase4-feedback-exact-run.json`
- `target/e2e/phase4-feedback-fallback-latest.json`
- `target/e2e/phase4-feedback-wrong-session.json`
- `target/e2e/phase4-summary.json`
The Maven process was stopped after evidence collection.
Inputs reused the Phase 3 E2E session:
```text
sessionId = e2e-phase3-codex-20260710-1912
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
```
API responses:
```text
exact run feedback:
status=200
runId=run-fdccbe21-e050-4f92-a741-a062aef59644
fallbackToLatestRun=false
caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
legacy fallback feedback:
status=200
runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94
fallbackToLatestRun=true
caseId=null
wrong-session feedback:
status=400
success=false
message=runId does not belong to sessionId
```
Log evidence from `phase4-mvn-20260710-2016.out.log`:
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-fdccbe21-e050-4f92-a741-a062aef59644, feedback=useful, fallbackToLatestRun=false, caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff`
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94, feedback=not_useful, fallbackToLatestRun=true, caseId=null`
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run.feedback`:
```text
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | feedback=not_useful | status=SUCCESS
run-fdccbe21-e050-4f92-a741-a062aef59644 | feedback=useful | status=SUCCESS
```
`case_library`:
```text
case_id=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
diagnosis_id=run-fdccbe21-e050-4f92-a741-a062aef59644
source_type=AUTO
title=支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
```
No automatic case was created for the fallback `not_useful` feedback on run2.
## Final Gate
The final Phase 4 gate was rerun after document review follow-up:
```text
mvn -q clean test-compile
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest,DiagnosisTraceServiceTest,ChatControllerTest" test
openspec validate session-run-trace-isolation --strict
git diff --check
```
Result: passed.
## Conclusion
Phase 4 satisfies run-scoped feedback, observable latest-run fallback, run-based useful case creation, and transitional old-data compatibility for `case_library.diagnosis_id`.
@@ -96,10 +96,19 @@ The system SHALL bind new feedback to a diagnosis run rather than an ambiguous m
#### Scenario: Feedback without runId falls back observably
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** at least one `diagnosis_run` exists for that session
- **THEN** the system SHALL bind feedback to the latest run for that session
- **AND** the response SHALL include `fallbackToLatestRun=true`
- **AND** the response SHALL include the actual bound `runId`
#### Scenario: Historical feedback without run-backed data remains compatible
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **AND** no `diagnosis_run` exists for that session
- **AND** a historical `diagnosis_session` row exists for that session
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
- **AND** the response SHALL NOT claim latest-run fallback
- **AND** the response MAY omit `runId`
#### Scenario: Feedback rejects run from another session
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return a failed feedback response instead of updating either run
@@ -32,13 +32,13 @@
## 4. Feedback and Case Library Run Binding
- [ ] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
- [ ] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
- [ ] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
- [ ] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
- [ ] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
- [ ] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
- [ ] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
- [x] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
- [x] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
- [x] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
- [x] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
- [x] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
- [x] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
- [x] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
## 5. AIOps Run Isolation