# Phase 6 Evidence: Demo, Trace UI, Documentation, and Verification ## Scope Phase 6 completed the run-aware demo/documentation surface and final verification for `session-run-trace-isolation`. Implemented: - Demo scripts read Chat `runId`, query exact Trace with `?runId=...`, and submit feedback with `runId`. - Trace UI accepts `?sessionId=...&runId=...` and calls the exact Trace API when `runId` is present. - Chat UI remembers the latest run target and links to Trace Workbench with `sessionId + runId` when available. - MVP table and architecture docs now describe `chat_session`, `diagnosis_run`, `agent_step.run_id`, `tool_invocation.run_id`, and transitional `case_library.diagnosis_id` semantics. ## Static Verification Commands: ```powershell node --check src\main\resources\static\app.js node --check src\main\resources\static\trace.js $scripts = @( 'mvp\demo\scripts\run-payment-timeout-demo.ps1', 'mvp\demo\scripts\run-interview-demo-check.ps1' ) foreach ($script in $scripts) { [scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null Write-Host "Parsed $script" } openspec validate session-run-trace-isolation --strict ``` Result: - JavaScript syntax: passed. - PowerShell script parsing: passed. - OpenSpec strict validation: passed. ## Focused Tests Command: ```powershell mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test ``` Result: passed. ## Final Same-Session Multi-Turn E2E Startup command: ```powershell mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo ``` Startup log: - `target/e2e/phase6-mvn-20260710-211831.out.log` - `target/e2e/phase6-mvn-20260710-211831.err.log` Application readiness: - `Started Main in 15.501 seconds` - `ReadinessState changed to ACCEPTING_TRAFFIC` E2E session: - `sessionId`: `e2e-phase6-chat-codex-20260710-2120` - `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92` - `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172` Artifacts: - `target/e2e/phase6-chat1-20260710-2120.json` - `target/e2e/phase6-chat2-20260710-2120.json` - `target/e2e/phase6-trace-run1-20260710-2120.json` - `target/e2e/phase6-trace-run2-20260710-2120.json` - `target/e2e/phase6-trace-latest-20260710-2120.json` - `target/e2e/phase6-e2e-summary-20260710-2120.json` Observed: ```json { "sessionId": "e2e-phase6-chat-codex-20260710-2120", "run1": "run-e2a97696-4398-4abc-90e4-28f45c838f92", "run2": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172", "chat1Success": true, "chat2Success": true, "trace1RunId": "run-e2a97696-4398-4abc-90e4-28f45c838f92", "trace2RunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172", "latestRunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172", "trace1Steps": 10, "trace2Steps": 9, "trace1Tools": 14, "trace2Tools": 8 } ``` Interpretation: - Two Chat requests reused the same `sessionId`. - Each Chat request returned a distinct `runId`. - Exact trace for run1 returned run1 only. - Exact trace for run2 returned run2 only. - Session-only Trace latest fallback returned run2. ## Database Inspection Tool: `scripts/query_mysql.py` `diagnosis_run`: ```text run_id | session_id | status | agent_flow | step_count | tool_call_count | has_answer ------------------------------------------------------------------------------------------------------------------------------------------------- run-e2a97696-4398-4abc-90e4-28f45c838f92 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 10 | 14 | 1 run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 9 | 8 | 1 ``` `agent_step` grouped by `run_id`: ```text run_id | step_rows | min_step | max_step -------------------------------------------------------------------------- run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 9 | 0 | 5 run-e2a97696-4398-4abc-90e4-28f45c838f92 | 10 | 0 | 6 ``` `tool_invocation` grouped by `run_id`: ```text run_id | tool_rows ---------------------------------------------------- run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 8 run-e2a97696-4398-4abc-90e4-28f45c838f92 | 14 ``` Mixed row check: ```text mixed_rows ---------- 0 0 ``` `chat_session` metadata: ```text session_id | status | message_pair_count | last_active_at --------------------------------------------------------------------------------------- e2e-phase6-chat-codex-20260710-2120 | ACTIVE | 2 | 2026-07-10 21:23:59 ``` `GET /api/chat/session/e2e-phase6-chat-codex-20260710-2120` returned `messagePairCount=2`. Interpretation: - Run table has exactly two successful Chat runs for the E2E session. - Step/tool counts match the exact Trace API responses. - No `agent_step` or `tool_invocation` rows for this session have NULL or unexpected `run_id`. - `chat_session` metadata confirms multi-turn context continuity at two message pairs. ## Log Inspection Searched: - `target/e2e/phase6-mvn-20260710-211831.out.log` - `logs/application.log` - `logs/chat.log` Patterns: - `e2e-phase6-chat-codex-20260710-2120` - `run-e2a97696-4398-4abc-90e4-28f45c838f92` - `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172` Result: - Matching startup, Chat execution, run persistence, and trace lookup log lines were present in Maven output and `logs/application.log`. ## Baseline Drift Focused baseline command: ```powershell mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test ``` Broader baseline regression command from `mvp/eval/README.md`: ```powershell mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test ``` Result: - Both commands passed. - The baseline harness evaluates saved fixtures and does not depend on live DB/session tables. - No baseline drift was observed. ## Final Gate Checks Commands: ```powershell rg -n "diagnosisSessionRepository\.save|new DiagnosisSession|DiagnosisSession\.builder|setAnswer\(|setSelfEvaluation\(|setFeedback\(|createFromSession|persistFinalReport\(sessionId, finalReport|save\(session\)" src\main\java\com\superbiz\agent -g "*.java" rg -n "evaluate\(|evaluateRun\(|persistFinalReport\(|submitFeedback\(|createFromSession\(" src\main\java src\test\java -g "*.java" git diff --check -- . ':!devflow/index.md' openspec validate session-run-trace-isolation --strict ``` Result: - New Chat write path calls `evaluationService.evaluateRun(...)` and writes `diagnosis_run`. - New AIOps controller path calls `persistFinalReport(sessionId, runId, ...)` and writes `diagnosis_run`. - Remaining `diagnosis_session` writes are legacy compatibility paths: - `FeedbackService.submitLegacySessionFeedback(...)` - `AiOpsService.persistLegacyFinalReport(...)` - legacy `EvaluationService.evaluate(...)` - `git diff --check`: passed after Markdown whitespace cleanup. - OpenSpec strict validation: passed. ## Documentation Review Follow-up After the final documentation review, the remaining demo helper docs were aligned with the run-aware contract: - `mvp/demo/trace-inspection-checklist.md` - `mvp/demo/payment-timeout-acceptance.md` - `mvp/demo/interview-walkthrough.md` - `mvp/tables/Agent步骤表-agent_step.md` - `mvp/tables/README.md` - `mvp/architecture/data-model.md` - `mvp/architecture/session-trace-lifecycle.md` The corrections remove session-only wording for Trace/Feedback and clarify that `run_id` is the execution isolation boundary while Trace API response order is the UI display contract. Follow-up gate after these documentation fixes: - `node --check src\main\resources\static\app.js`: passed. - `node --check src\main\resources\static\trace.js`: passed. - PowerShell demo script parsing: passed. - `openspec validate session-run-trace-isolation --strict`: passed. - `git diff --check -- . ':!devflow/index.md'`: passed. ## Notes - The AGENTS-required `codebase-retrieval` and LSP tools were not available in this session. Fallback verification used OpenSpec context, `rg`, targeted file reads, focused tests, E2E, DB inspection, and log inspection.