docs(openspec): archive session run isolation

This commit is contained in:
zhuyongxin
2026-07-10 22:52:39 +08:00
parent f9df94377b
commit 3578709896
21 changed files with 511 additions and 45 deletions
@@ -0,0 +1,253 @@
# Phase 6 Evidence: Demo, Trace UI, Documentation, and Verification
## Scope
Phase 6 completed the run-aware demo/documentation surface and final verification for `session-run-trace-isolation`.
Implemented:
- Demo scripts read Chat `runId`, query exact Trace with `?runId=...`, and submit feedback with `runId`.
- Trace UI accepts `?sessionId=...&runId=...` and calls the exact Trace API when `runId` is present.
- Chat UI remembers the latest run target and links to Trace Workbench with `sessionId + runId` when available.
- MVP table and architecture docs now describe `chat_session`, `diagnosis_run`, `agent_step.run_id`, `tool_invocation.run_id`, and transitional `case_library.diagnosis_id` semantics.
## Static Verification
Commands:
```powershell
node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js
$scripts = @(
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
Write-Host "Parsed $script"
}
openspec validate session-run-trace-isolation --strict
```
Result:
- JavaScript syntax: passed.
- PowerShell script parsing: passed.
- OpenSpec strict validation: passed.
## Focused Tests
Command:
```powershell
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
```
Result: passed.
## Final Same-Session Multi-Turn E2E
Startup command:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Startup log:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `target/e2e/phase6-mvn-20260710-211831.err.log`
Application readiness:
- `Started Main in 15.501 seconds`
- `ReadinessState changed to ACCEPTING_TRAFFIC`
E2E session:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
Artifacts:
- `target/e2e/phase6-chat1-20260710-2120.json`
- `target/e2e/phase6-chat2-20260710-2120.json`
- `target/e2e/phase6-trace-run1-20260710-2120.json`
- `target/e2e/phase6-trace-run2-20260710-2120.json`
- `target/e2e/phase6-trace-latest-20260710-2120.json`
- `target/e2e/phase6-e2e-summary-20260710-2120.json`
Observed:
```json
{
"sessionId": "e2e-phase6-chat-codex-20260710-2120",
"run1": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
"run2": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"chat1Success": true,
"chat2Success": true,
"trace1RunId": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
"trace2RunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"latestRunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
"trace1Steps": 10,
"trace2Steps": 9,
"trace1Tools": 14,
"trace2Tools": 8
}
```
Interpretation:
- Two Chat requests reused the same `sessionId`.
- Each Chat request returned a distinct `runId`.
- Exact trace for run1 returned run1 only.
- Exact trace for run2 returned run2 only.
- Session-only Trace latest fallback returned run2.
## Database Inspection
Tool: `scripts/query_mysql.py`
`diagnosis_run`:
```text
run_id | session_id | status | agent_flow | step_count | tool_call_count | has_answer
-------------------------------------------------------------------------------------------------------------------------------------------------
run-e2a97696-4398-4abc-90e4-28f45c838f92 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 10 | 14 | 1
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 9 | 8 | 1
```
`agent_step` grouped by `run_id`:
```text
run_id | step_rows | min_step | max_step
--------------------------------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 9 | 0 | 5
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 10 | 0 | 6
```
`tool_invocation` grouped by `run_id`:
```text
run_id | tool_rows
----------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 8
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 14
```
Mixed row check:
```text
mixed_rows
----------
0
0
```
`chat_session` metadata:
```text
session_id | status | message_pair_count | last_active_at
---------------------------------------------------------------------------------------
e2e-phase6-chat-codex-20260710-2120 | ACTIVE | 2 | 2026-07-10 21:23:59
```
`GET /api/chat/session/e2e-phase6-chat-codex-20260710-2120` returned `messagePairCount=2`.
Interpretation:
- Run table has exactly two successful Chat runs for the E2E session.
- Step/tool counts match the exact Trace API responses.
- No `agent_step` or `tool_invocation` rows for this session have NULL or unexpected `run_id`.
- `chat_session` metadata confirms multi-turn context continuity at two message pairs.
## Log Inspection
Searched:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `logs/application.log`
- `logs/chat.log`
Patterns:
- `e2e-phase6-chat-codex-20260710-2120`
- `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
Result:
- Matching startup, Chat execution, run persistence, and trace lookup log lines were present in Maven output and `logs/application.log`.
## Baseline Drift
Focused baseline command:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
```
Broader baseline regression command from `mvp/eval/README.md`:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Result:
- Both commands passed.
- The baseline harness evaluates saved fixtures and does not depend on live DB/session tables.
- No baseline drift was observed.
## Final Gate Checks
Commands:
```powershell
rg -n "diagnosisSessionRepository\.save|new DiagnosisSession|DiagnosisSession\.builder|setAnswer\(|setSelfEvaluation\(|setFeedback\(|createFromSession|persistFinalReport\(sessionId, finalReport|save\(session\)" src\main\java\com\superbiz\agent -g "*.java"
rg -n "evaluate\(|evaluateRun\(|persistFinalReport\(|submitFeedback\(|createFromSession\(" src\main\java src\test\java -g "*.java"
git diff --check -- . ':!devflow/index.md'
openspec validate session-run-trace-isolation --strict
```
Result:
- New Chat write path calls `evaluationService.evaluateRun(...)` and writes `diagnosis_run`.
- New AIOps controller path calls `persistFinalReport(sessionId, runId, ...)` and writes `diagnosis_run`.
- Remaining `diagnosis_session` writes are legacy compatibility paths:
- `FeedbackService.submitLegacySessionFeedback(...)`
- `AiOpsService.persistLegacyFinalReport(...)`
- legacy `EvaluationService.evaluate(...)`
- `git diff --check`: passed after Markdown whitespace cleanup.
- OpenSpec strict validation: passed.
## Documentation Review Follow-up
After the final documentation review, the remaining demo helper docs were aligned with the run-aware contract:
- `mvp/demo/trace-inspection-checklist.md`
- `mvp/demo/payment-timeout-acceptance.md`
- `mvp/demo/interview-walkthrough.md`
- `mvp/tables/Agent步骤表-agent_step.md`
- `mvp/tables/README.md`
- `mvp/architecture/data-model.md`
- `mvp/architecture/session-trace-lifecycle.md`
The corrections remove session-only wording for Trace/Feedback and clarify that `run_id` is the execution isolation boundary while Trace API response order is the UI display contract.
Follow-up gate after these documentation fixes:
- `node --check src\main\resources\static\app.js`: passed.
- `node --check src\main\resources\static\trace.js`: passed.
- PowerShell demo script parsing: passed.
- `openspec validate session-run-trace-isolation --strict`: passed.
- `git diff --check -- . ':!devflow/index.md'`: passed.
## Notes
- The AGENTS-required `codebase-retrieval` and LSP tools were not available in this session. Fallback verification used OpenSpec context, `rg`, targeted file reads, focused tests, E2E, DB inspection, and log inspection.