130 lines
4.2 KiB
Markdown
130 lines
4.2 KiB
Markdown
# Phase 2 Evidence: Chat Run Write Path
|
|
|
|
Date: 2026-07-10
|
|
|
|
## Scope
|
|
|
|
Phase 2 switched Chat writes from session-scoped execution state to run-scoped execution state:
|
|
|
|
- Chat creates/updates `chat_session` metadata.
|
|
- Each valid `/api/chat` execution creates one `diagnosis_run`.
|
|
- Chat execution context carries `sessionId + runId` through `RunnableConfig` and `SessionContextHolder`.
|
|
- `agent_step.run_id` and `tool_invocation.run_id` are written for Chat runs.
|
|
- Chat completion/failure/status/answer/self-evaluation/counts are written to `diagnosis_run`.
|
|
- verifier/gatekeeper/evaluation reads use run-scoped tool rows when `runId` is available.
|
|
- `/api/chat` response includes official `runId`.
|
|
|
|
## Static / Unit Verification
|
|
|
|
Commands:
|
|
|
|
```powershell
|
|
mvn -q clean test-compile
|
|
mvn -q "-Dtest=ChatControllerTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ExecutorGatekeeperServiceTest" test
|
|
openspec validate session-run-trace-isolation --strict
|
|
```
|
|
|
|
Result:
|
|
|
|
- `test-compile` passed.
|
|
- Focused Phase 2 tests passed.
|
|
- OpenSpec strict validation passed.
|
|
|
|
Focused coverage:
|
|
|
|
- valid Chat creates `chat_session` and `diagnosis_run`;
|
|
- invalid blank Chat request returns before `ChatService`, so no run is created;
|
|
- same `sessionId` across two Chat turns creates two distinct `runId` values;
|
|
- `ToolInvocationRecorder` copies `runId` from execution context;
|
|
- verifier trace summary reads by run;
|
|
- gatekeeper validates by run;
|
|
- evaluation writes rule evaluation to `diagnosis_run.self_evaluation`.
|
|
|
|
## E2E Verification
|
|
|
|
Startup command:
|
|
|
|
```powershell
|
|
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
|
```
|
|
|
|
Process stdout/stderr:
|
|
|
|
- `target/e2e/phase2-mvn-20260710-185102.out.log`
|
|
- `target/e2e/phase2-mvn-20260710-185102.err.log`
|
|
|
|
Primary E2E session:
|
|
|
|
```text
|
|
sessionId = e2e-phase2-codex-20260710-1856
|
|
round 1 runId = run-24b6f04c-94a0-43cb-b94f-4f0141f9050d
|
|
round 2 runId = run-a7a2be77-697f-495a-a779-91afd1d8589c
|
|
```
|
|
|
|
HTTP evidence:
|
|
|
|
- `target/e2e/phase2-utf8-request-round1.json`
|
|
- `target/e2e/phase2-utf8-response-round1.json`
|
|
- `target/e2e/phase2-utf8-request-round2.json`
|
|
- `target/e2e/phase2-utf8-response-round2.json`
|
|
- both Chat responses returned `code=200`, `data.success=true`, the same `sessionId`, and distinct `runId` values.
|
|
|
|
Redis/session continuity evidence:
|
|
|
|
- `target/e2e/phase2-utf8-chat-session-response.json`
|
|
- response returned `messagePairCount=2`.
|
|
- logs show the second request entered with `会话历史消息对数: 1` and completed with `当前消息对数: 2`.
|
|
|
|
## Database Inspection
|
|
|
|
DB inspection used `scripts/query_mysql.py`.
|
|
|
|
Saved query outputs:
|
|
|
|
- `target/e2e/phase2-utf8-db-runs.txt`
|
|
- `target/e2e/phase2-utf8-db-chat-session.txt`
|
|
- `target/e2e/phase2-utf8-db-agent-steps.txt`
|
|
- `target/e2e/phase2-utf8-db-tool-invocations.txt`
|
|
- `target/e2e/phase2-utf8-db-missing-runid.txt`
|
|
|
|
Observed rows:
|
|
|
|
```text
|
|
diagnosis_run:
|
|
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | SUCCESS | CHAT | step_count=8 | tool_call_count=12
|
|
run-a7a2be77-697f-495a-a779-91afd1d8589c | SUCCESS | CHAT | step_count=2 | tool_call_count=1
|
|
|
|
chat_session:
|
|
e2e-phase2-codex-20260710-1856 | ACTIVE | message_pair_count=2
|
|
|
|
agent_step grouped by run_id:
|
|
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 8
|
|
run-a7a2be77-697f-495a-a779-91afd1d8589c | 2
|
|
|
|
tool_invocation grouped by run_id:
|
|
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 12
|
|
run-a7a2be77-697f-495a-a779-91afd1d8589c | 1
|
|
|
|
missing run_id for this session:
|
|
agent_step = 0
|
|
tool_invocation = 0
|
|
```
|
|
|
|
## Log Review
|
|
|
|
Saved log excerpts:
|
|
|
|
- `target/e2e/phase2-utf8-application-log-excerpt.txt`
|
|
- `target/e2e/phase2-utf8-mvn-log-excerpt.txt`
|
|
|
|
Findings:
|
|
|
|
- Chat logs show second turn reused Redis history for the same session.
|
|
- `EvaluationService` wrote scoring results to both run ids.
|
|
- No E2E-specific application exception was observed for `e2e-phase2-codex-20260710-1856`.
|
|
- Earlier `/actuator/health` probes produced expected 500/no-resource noise because the actuator health endpoint is not exposed; the E2E readiness check used `/api/chat` instead.
|
|
|
|
## Notes
|
|
|
|
`codebase-retrieval` and LSP tools were not available in this environment. Call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
|