Files

8.1 KiB

Phase 6 Evidence: Demo, Trace UI, Documentation, and Verification

Scope

Phase 6 completed the run-aware demo/documentation surface and final verification for session-run-trace-isolation.

Implemented:

  • Demo scripts read Chat runId, query exact Trace with ?runId=..., and submit feedback with runId.
  • Trace UI accepts ?sessionId=...&runId=... and calls the exact Trace API when runId is present.
  • Chat UI remembers the latest run target and links to Trace Workbench with sessionId + runId when available.
  • MVP table and architecture docs now describe chat_session, diagnosis_run, agent_step.run_id, tool_invocation.run_id, and transitional case_library.diagnosis_id semantics.

Static Verification

Commands:

node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js

$scripts = @(
  'mvp\demo\scripts\run-payment-timeout-demo.ps1',
  'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
  [scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
  Write-Host "Parsed $script"
}

openspec validate session-run-trace-isolation --strict

Result:

  • JavaScript syntax: passed.
  • PowerShell script parsing: passed.
  • OpenSpec strict validation: passed.

Focused Tests

Command:

mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test

Result: passed.

Final Same-Session Multi-Turn E2E

Startup command:

mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo

Startup log:

  • target/e2e/phase6-mvn-20260710-211831.out.log
  • target/e2e/phase6-mvn-20260710-211831.err.log

Application readiness:

  • Started Main in 15.501 seconds
  • ReadinessState changed to ACCEPTING_TRAFFIC

E2E session:

  • sessionId: e2e-phase6-chat-codex-20260710-2120
  • run1: run-e2a97696-4398-4abc-90e4-28f45c838f92
  • run2: run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172

Artifacts:

  • target/e2e/phase6-chat1-20260710-2120.json
  • target/e2e/phase6-chat2-20260710-2120.json
  • target/e2e/phase6-trace-run1-20260710-2120.json
  • target/e2e/phase6-trace-run2-20260710-2120.json
  • target/e2e/phase6-trace-latest-20260710-2120.json
  • target/e2e/phase6-e2e-summary-20260710-2120.json

Observed:

{
  "sessionId": "e2e-phase6-chat-codex-20260710-2120",
  "run1": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
  "run2": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
  "chat1Success": true,
  "chat2Success": true,
  "trace1RunId": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
  "trace2RunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
  "latestRunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
  "trace1Steps": 10,
  "trace2Steps": 9,
  "trace1Tools": 14,
  "trace2Tools": 8
}

Interpretation:

  • Two Chat requests reused the same sessionId.
  • Each Chat request returned a distinct runId.
  • Exact trace for run1 returned run1 only.
  • Exact trace for run2 returned run2 only.
  • Session-only Trace latest fallback returned run2.

Database Inspection

Tool: scripts/query_mysql.py

diagnosis_run:

run_id                                   | session_id                          | status  | agent_flow | step_count | tool_call_count | has_answer
-------------------------------------------------------------------------------------------------------------------------------------------------
run-e2a97696-4398-4abc-90e4-28f45c838f92 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT       | 10         | 14              | 1
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT       | 9          | 8               | 1

agent_step grouped by run_id:

run_id                                   | step_rows | min_step | max_step
--------------------------------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 9         | 0        | 5
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 10        | 0        | 6

tool_invocation grouped by run_id:

run_id                                   | tool_rows
----------------------------------------------------
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 8
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 14

Mixed row check:

mixed_rows
----------
0
0

chat_session metadata:

session_id                          | status | message_pair_count | last_active_at
---------------------------------------------------------------------------------------
e2e-phase6-chat-codex-20260710-2120 | ACTIVE | 2                  | 2026-07-10 21:23:59

GET /api/chat/session/e2e-phase6-chat-codex-20260710-2120 returned messagePairCount=2.

Interpretation:

  • Run table has exactly two successful Chat runs for the E2E session.
  • Step/tool counts match the exact Trace API responses.
  • No agent_step or tool_invocation rows for this session have NULL or unexpected run_id.
  • chat_session metadata confirms multi-turn context continuity at two message pairs.

Log Inspection

Searched:

  • target/e2e/phase6-mvn-20260710-211831.out.log
  • logs/application.log
  • logs/chat.log

Patterns:

  • e2e-phase6-chat-codex-20260710-2120
  • run-e2a97696-4398-4abc-90e4-28f45c838f92
  • run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172

Result:

  • Matching startup, Chat execution, run persistence, and trace lookup log lines were present in Maven output and logs/application.log.

Baseline Drift

Focused baseline command:

mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test

Broader baseline regression command from mvp/eval/README.md:

mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test

Result:

  • Both commands passed.
  • The baseline harness evaluates saved fixtures and does not depend on live DB/session tables.
  • No baseline drift was observed.

Final Gate Checks

Commands:

rg -n "diagnosisSessionRepository\.save|new DiagnosisSession|DiagnosisSession\.builder|setAnswer\(|setSelfEvaluation\(|setFeedback\(|createFromSession|persistFinalReport\(sessionId, finalReport|save\(session\)" src\main\java\com\superbiz\agent -g "*.java"

rg -n "evaluate\(|evaluateRun\(|persistFinalReport\(|submitFeedback\(|createFromSession\(" src\main\java src\test\java -g "*.java"

git diff --check -- . ':!devflow/index.md'
openspec validate session-run-trace-isolation --strict

Result:

  • New Chat write path calls evaluationService.evaluateRun(...) and writes diagnosis_run.
  • New AIOps controller path calls persistFinalReport(sessionId, runId, ...) and writes diagnosis_run.
  • Remaining diagnosis_session writes are legacy compatibility paths:
    • FeedbackService.submitLegacySessionFeedback(...)
    • AiOpsService.persistLegacyFinalReport(...)
    • legacy EvaluationService.evaluate(...)
  • git diff --check: passed after Markdown whitespace cleanup.
  • OpenSpec strict validation: passed.

Documentation Review Follow-up

After the final documentation review, the remaining demo helper docs were aligned with the run-aware contract:

  • mvp/demo/trace-inspection-checklist.md
  • mvp/demo/payment-timeout-acceptance.md
  • mvp/demo/interview-walkthrough.md
  • mvp/tables/Agent步骤表-agent_step.md
  • mvp/tables/README.md
  • mvp/architecture/data-model.md
  • mvp/architecture/session-trace-lifecycle.md

The corrections remove session-only wording for Trace/Feedback and clarify that run_id is the execution isolation boundary while Trace API response order is the UI display contract.

Follow-up gate after these documentation fixes:

  • node --check src\main\resources\static\app.js: passed.
  • node --check src\main\resources\static\trace.js: passed.
  • PowerShell demo script parsing: passed.
  • openspec validate session-run-trace-isolation --strict: passed.
  • git diff --check -- . ':!devflow/index.md': passed.

Notes

  • The AGENTS-required codebase-retrieval and LSP tools were not available in this session. Fallback verification used OpenSpec context, rg, targeted file reads, focused tests, E2E, DB inspection, and log inspection.