feat(trace): add session run isolation schema

This commit is contained in:
zhuyongxin
2026-07-10 17:47:56 +08:00
parent 841437fa06
commit 52bf0302c6
21 changed files with 1718 additions and 1 deletions
+17 -1
View File
@@ -72,9 +72,10 @@
### SessionContext ### SessionContext
- 定义:会话上下文数据类,存储在 Redis 中的会话数据 - 定义:会话上下文数据类,存储在 Redis 中的会话数据
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL - 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer) - 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
- 使用场景:多轮对话上下文管理、工具调用历史追踪 - 使用场景:多轮对话上下文管理、工具调用历史追踪
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
### ToolCall ### ToolCall
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果 - 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
@@ -87,6 +88,21 @@
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall - 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
- 使用场景:分布式会话管理、Agent 状态维护 - 使用场景:分布式会话管理、Agent 状态维护
### Chat Session
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
### Diagnosis Run
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
### Diagnosis Trace
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
### Flyway ### Flyway
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行 - 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true - 配置:spring.flyway.enabled=true, baseline-on-migrate=true
+1
View File
@@ -18,6 +18,7 @@
|---|---|---|---|---| |---|---|---|---|---|
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) | | ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) | | ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 高 | 方案已确认,待 OpenSpec | [active/ISS-010-session-run-trace-isolation.md](active/ISS-010-session-run-trace-isolation.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) | | executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) | | rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
@@ -0,0 +1,559 @@
# ISS-010 同 session 多轮诊断 Trace 隔离
**状态**:方案已确认,待 OpenSpec
**严重程度**:高
**发现时间**:2026-07-10
**来源**:同一 `sessionId` 多轮 Chat E2E 验证
---
## 背景
当前 Chat 链路同时存在两类“会话”语义:
```text
Redis SessionContext
-> 保存同一 sessionId 的多轮对话历史
-> 用于下一轮模型上下文
MySQL diagnosis_session / agent_step / tool_invocation
-> 保存诊断 Trace
-> 用于 Trace API、Verifier、Evidence score、Feedback 和评测
```
多轮对话需要继续复用 `sessionId`,否则无法保留上下文。但一次诊断 Trace 应该是可独立回放、可独立评分、可独立反馈的执行单元。
当前实现只按 `sessionId` 关联 Trace,导致同一个 `sessionId` 下多轮诊断的 step/tool 记录混在一起。
---
## E2E 证据
本次使用 `mvp-demo` profile 通过 Maven 启动服务,并用同一个 `sessionId` 连续请求两轮 `/api/chat`:
```text
sessionId = e2e-multiturn-codex-20260710-1615
round 1 = 支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
round 2 = 基于上一轮结论,只列出目前最缺的三类证据,以及下一步应该优先查哪个系统。
```
验证结果:
- 第一轮成功,走多 Agent:`planner -> executor -> verifier -> composer`。
- 第二轮成功,日志显示进入请求时 `会话历史消息对数: 1`,说明 Redis 历史上下文被复用。
- `/api/chat/session/{sessionId}` 返回 `messagePairCount=2`。
- `diagnosis_session` 只有一行,`query` 被第二轮问题覆盖。
- `agent_step` 返回 14 行,包含第一轮多 Agent step 和第二轮 `intelligent_assistant` step。
- `tool_invocation` 返回 19 行,包含两轮工具调用。
- `self_evaluation.verifier_evaluation` 仍保留第一轮 Verifier 结果;第二轮简单问答没有新的 Verifier,但 rule evaluation 会基于同 session 全部工具调用重新计算。
关键入库形态:
```text
diagnosis_session
session_id = e2e-multiturn-codex-20260710-1615
query = round 2 question
status = SUCCESS
step_count = 14
tool_call_count = 19
agent_step
round 1: planner, executor..., verifier, composer
round 2: intelligent_assistant...
tool_invocation
round 1 tools + round 2 tools all under same session_id
```
本次验证产物保存在:
- `target/e2e/request-round1.json`
- `target/e2e/response-round1.json`
- `target/e2e/request-round2.json`
- `target/e2e/response-round2.json`
- `target/e2e/trace-after-round2.json`
---
## 核心问题
### P0:Trace 不是单次诊断的稳定回放
`GET /api/diagnosis/{sessionId}/trace` 会聚合同一 `sessionId` 下所有 `agent_step` 和 `tool_invocation`。
多轮之后,Trace 不再表示某一轮诊断,而是混合历史执行轨迹。
### P0:Verifier 和评分可能读取跨轮证据
Verifier、Gatekeeper、`ToolTraceSummaryService` 和 `EvaluationService` 当前主要按 `sessionId` 查询工具调用。
如果上一轮和当前轮证据混在一起,当前轮可能引用或评分到历史工具结果。
### P1:反馈语义不清晰
`feedback` 当前在 `diagnosis_session` 上按 `sessionId` 保存。
多轮之后,用户反馈的是哪一轮答案不再明确。`useful` 反馈沉淀到 `case_library` 时也可能关联到最新主表答案,而不是用户实际评价的那一轮。
### P1:`diagnosis_session` 字段被覆盖但子表追加
主表 `query/answer/status/self_evaluation/step_count/tool_call_count` 表示最新运行或混合统计,子表却保留多轮历史。
这会让 Trace summary、数据库统计和人工排查产生歧义。
---
## 已确认决策
### D1:`runId` 是正式 API 字段
`/api/chat` 和 `/api/ai_ops` 的响应或 SSE 消息需要暴露本次执行的 `runId`。
```text
sessionId = 多轮对话上下文 ID
runId = 本轮诊断执行 ID
```
新客户端应优先用 `runId` 查询 Trace 和提交 Feedback。旧客户端只传 `sessionId` 时,服务端兼容解析该 session 的最新 run。
### D2:拆分会话态和运行态
不再把 session、run、trace 全部塞进 `diagnosis_session` 一张主表。
新增两张主表:
```text
chat_session
-> 多轮对话上下文主表
diagnosis_run
-> 单次诊断执行主表
```
Trace 继续使用现有明细表表达:
```text
agent_step
tool_invocation
```
暂不新增单独的 `diagnosis_trace` 或 `trace_event` 主表。
### D3:`runId` 格式
使用 `run-` + UUID 全量字符串。
```text
run-550e8400-e29b-41d4-a716-446655440000
```
### D4:Trace API 兼容旧路径
```text
GET /api/diagnosis/{sessionId}/trace
-> 查该 session 最新 run
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
-> 查指定 run
```
指定 `runId` 时必须校验该 run 属于 path 中的 `sessionId`。
最新 run 建议按 `diagnosis_run.created_at DESC, id DESC` 解析,避免旧 run 因反馈或异步评分更新 `updated_at` 后被误认为最新。
### D5:Feedback 优先绑定 run
Feedback request 支持 `runId`。
- 有 `runId`:绑定指定 run。
- 无 `runId`:短期兼容绑定该 `sessionId` 最新 run,并显式标记 fallback。
- `case_library.diagnosis_id` 新数据保存 `run_id`。
兼容语义:历史 `case_library.diagnosis_id` 可能保存 `diagnosis_session.session_id`;本 change 之后自动沉淀的新数据保存 `diagnosis_run.run_id`。查询、幂等和文档需要在过渡期识别两种来源,避免把旧案例误判为无效数据。
### D6:所有 `/api/chat` 执行请求都创建 run
只要请求通过参数校验并进入 `ChatService.executeChatWithStrategy`,就创建新的 diagnosis run。
- 简单问答也创建 run。
- 复杂诊断也创建 run。
- 空问题等参数校验失败不创建 run。
### D7:AIOps 同步纳入 run 隔离
每次 `/api/ai_ops` 执行也创建新的 diagnosis run。AIOps 的 step、tool invocation 和 rule evaluation 都按 `runId` 隔离。
阶段说明:AIOps 可作为独立实现切片排在 Chat 之后,但必须在本 change 整体完成前落地;Chat-only 的中间状态只能作为过渡验证状态,不能作为生产完成状态归档。
### D8:`chat_session` 第一阶段只保存会话元数据
`chat_session` 是会话目录/索引表,不保存完整对话历史正文。
建议保存:
```text
session_id
status
message_pair_count
created_at
last_active_at
expires_at
```
完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`,用于下一轮 prompt 上下文。
每轮需要长期审计的用户问题和最终回答保存到 `diagnosis_run.query` / `diagnosis_run.answer`。
`chat_session.expires_at` 只表示 MySQL 会话目录的过期/清理元数据;Redis TTL 到期后,`SessionContext.messageHistory` 可能不再存在,但已经持久化的 `diagnosis_run`、`agent_step` 和 `tool_invocation` 仍作为审计记录保留。
如果未来需要长期保存完整聊天历史,再单独设计 `chat_message` 表,不在本阶段引入。
### D9:旧 `diagnosis_session` 表保留但新代码不再写入
新增 `diagnosis_run` 后,旧 `diagnosis_session` 不立即删除、不立即改造成 view、不直接重命名。
迁移策略:
1. 新增 `chat_session` / `diagnosis_run`。
2. 为 `agent_step` / `tool_invocation` 新增 nullable `run_id`。
3. 将旧 `diagnosis_session` 数据迁移/复制为 `diagnosis_run` 兼容记录。
4. 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
5. 增加必要索引和查询方法,先保持兼容读取。
6. 新代码切换为只写 `chat_session` 和 `diagnosis_run`,并为新 step/tool 写入 `run_id`。
7. 验证新旧数据 `run_id` 覆盖情况后,再将新写路径要求 `run_id` 非空,并补充索引/约束。
8. 旧 `diagnosis_session` 暂时保留,用于历史核对和回滚窗口。
9. 后续确认无依赖后,再单独归档或删除旧表。
### D10:提供轻量 run 列表 API
新增轻量查询接口,用于查看一个 Chat Session 下有哪些 Diagnosis Run。
```text
GET /api/chat/session/{sessionId}/runs
```
建议返回字段:
```text
runId
sessionId
query
status
agentFlow
answerPreview
stepCount
toolCallCount
createdAt
updatedAt
```
该接口只读 `diagnosis_run` 主表,不展开 `agent_step` / `tool_invocation` 大字段。
### D11:Feedback 缺少 `runId` 时短期兼容,长期收紧
Feedback 新协议优先要求 `runId`。
短期兼容策略:
- 有 `runId`:绑定指定 run。
- 无 `runId`:绑定该 `sessionId` 最新 run。
- 无 `runId` fallback 时,在响应或日志中明确标记 `fallbackToLatestRun=true`,并返回实际绑定的 `runId`。
长期收紧策略:
- 当前端、demo 脚本和外部调用方都完成 `runId` 传递后,再评估是否将缺少 `runId` 改为参数错误。
### D12:同步更新 demo 脚本和 Trace UI 的 `runId` 最小支持
本 issue 实施范围包含 demo 脚本和 Trace UI 的最小协议适配。
范围:
- Demo 脚本读取 `/api/chat` 或 `/api/ai_ops` 返回的 `runId`。
- Demo 脚本查询 Trace 时传 `?runId=...`。
- Trace UI 支持 URL 参数 `?sessionId=...&runId=...`。
- Trace UI 查询时如果有 `runId`,带上 `runId`。
- 不在本阶段实现完整 run 列表 UI。
---
## 目标语义
引入明确的 `sessionId` / `runId` 分层:
```text
sessionId = 多轮对话上下文
runId = 单次诊断执行 / 单次可回放 Trace
```
目标关系:
```text
chat_session(sessionId)
-> one conversation context
-> conversation metadata / TTL / last active state
Redis SessionContext(sessionId)
-> hot messageHistory cache
-> supports prompt context window
diagnosis_run(runId, sessionId)
-> one diagnosis run
agent_step(runId, sessionId)
-> steps of one run
tool_invocation(runId, sessionId)
-> tool calls of one run
```
---
## 建议方案
采用“拆分主表 + 复用现有 Trace 明细表”的方案:
1. 新增 `chat_session`。
2. 新增 `diagnosis_run`。
3. 逐步迁移当前 `diagnosis_session` 语义到 `diagnosis_run`。
4. `agent_step` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
5. `tool_invocation` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
6. Trace API 聚合 `diagnosis_run + agent_step + tool_invocation`。
建议核心字段:
```text
chat_session
id
session_id unique
status
message_pair_count
created_at
last_active_at
expires_at
diagnosis_run
id
run_id unique
session_id
query
status
agent_flow
answer
self_evaluation
feedback
total_duration_ms
total_token_count
step_count
tool_call_count
created_at
updated_at
agent_step(run_id, step_index)
tool_invocation(run_id, id)
```
理由:
- `chat_session` 只表达会话态,避免会话上下文和诊断结果混在一起。
- `diagnosis_run` 只表达一次执行,天然隔离每轮 Trace、评分和反馈。
- `agent_step` / `tool_invocation` 已足够表达 Trace 明细,暂不需要额外 trace 主表。
- 后续如果需要统一时间线,再增加 `trace_event`,不阻塞本次隔离。
---
## 分阶段计划
### Phase 0:协议基线和数据边界
目标:先把语义定死,避免实现中反复。
已确认基线:
1. `/api/chat` 是否返回 `runId`。
2. `GET /api/diagnosis/{sessionId}/trace` 默认查最新 run 还是要求显式传 `runId`。
3. Feedback 是否优先绑定 `runId`,只有旧请求缺失 `runId` 时才回退最新 run。
4. AIOps 是否和 Chat 同步接入 `runId`。
5. 简单问答是否也创建 diagnosis run。
建议默认:
- `/api/chat` 返回 `sessionId + runId`。
- `GET /api/diagnosis/{sessionId}/trace` 兼容查最新 run。
- `GET /api/diagnosis/{sessionId}/trace?runId=...` 查指定 run。
- Feedback 优先按 `runId` 绑定。
- Chat 简单问答也创建 run。
- AIOps 同步接入 run 隔离。
### Phase 1:Schema 迁移和历史数据兼容
目标:引入 `chat_session` / `diagnosis_run`,并保留旧数据可查询。
任务:
- Flyway 新增 `chat_session`。
- Flyway 新增 `diagnosis_run`。
- 为旧 `diagnosis_session` 生成兼容 `diagnosis_run` 记录。
- 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
- 增加 `find latest run by sessionId` 查询。
- 增加 `find by runId` 查询。
- 保留旧 `diagnosis_session` 一段时间,新代码不再写入。
验收:
- 旧 session 的 Trace 仍可查。
- 新索引存在。
- 不改变旧 `/api/chat` 必需字段。
- `agent_step` / `tool_invocation` 支持 nullable `run_id` 并完成旧数据回填。
- 新增 repository 查询可以按 `sessionId` 找最新 run、按 `runId` 找指定 run。
### Phase 2:Chat 写入切到 runId
目标:每轮 `/api/chat` 创建一个新的 run,step/tool 按 run 隔离。
任务:
- `ChatService` 每次执行生成新的 `runId`。
- `ChatController` 确保 `chat_session` 存在并更新会话态。
- `ChatService` 按 `runId` 创建 `diagnosis_run`。
- `AgentLoggingHook` 写入 `agent_step.run_id`。
- `ToolInvocationRecorder` 写入 `tool_invocation.run_id`。
- `SessionContextHolder` 或新的上下文 holder 同时携带 `sessionId + runId`。
- `backfillSessionMetrics` 按 `runId` 统计。
- `EvaluationService` 按 `runId` 读取工具调用。
验收:
- 同一 `sessionId` 连续两轮后,`diagnosis_run` 有两行不同 `run_id`。
- `/api/chat` 响应增加正式字段 `runId`。
- 两轮 `agent_step` / `tool_invocation` 分别按各自 `run_id` 查询。
- Redis `messagePairCount` 仍为 2,证明上下文不被破坏。
### Phase 3:Trace API 兼容和精确查询
目标:Trace API 可查最新 run,也可查指定 run,并能列出一个 session 下的 run。
任务:
- `GET /api/diagnosis/{sessionId}/trace` 从 `diagnosis_run` 默认解析最新 run。
- 增加 `runId` query 参数。
- Trace response 增加 `runId`。
- 新增 `GET /api/chat/session/{sessionId}/runs`。
- Demo 脚本读取响应中的 `runId` 并查询精确 Trace。
- Trace UI 支持 `sessionId + runId` 最小查询。
验收:
- 不传 `runId` 返回最新 run。
- 传第一轮 `runId` 只返回第一轮 step/tool。
- 传第二轮 `runId` 只返回第二轮 step/tool。
- run 列表 API 只返回轻量 run 摘要,不展开 trace 明细。
- Demo 脚本输出摘要包含 `runId`。
- Trace UI 可以通过 URL 参数打开指定 run。
### Phase 4:Feedback 和 CaseLibrary 绑定 run
目标:反馈明确评价哪一轮诊断。
任务:
- Feedback request 支持 `runId`。
- 旧请求只有 `sessionId` 时短期绑定最新 run,并显式标记 fallback。
- `case_library.diagnosis_id` 新数据保存 `run_id`。
- `CaseLibraryService` 以 run 为来源生成 case,并用 `run_id` 做新数据幂等键。
- 文档说明 `diagnosis_id` 的过渡语义:旧数据可能是 `session_id`,新数据是 `run_id`。
验收:
- 同 session 多轮后,对第一轮提交 feedback 不会覆盖第二轮。
- useful 生成 case 时能定位到对应 run 的 query/answer。
### Phase 5:AIOps 同步 run 隔离
目标:AIOps 使用同样的 run 语义,避免另一条入口继续混杂。
任务:
- `AiOpsService` 生成并返回/透出 `runId`。
- AIOps `agent_step` / `tool_invocation` / rule evaluation 按 `runId` 隔离。
- AIOps Trace 查询兼容 `sessionId + runId`。
验收:
- 同一 AIOps `sessionId` 重跑不会混合 step/tool。
- AIOps rule evaluation 只读取当前 run 工具调用。
---
## 暂不做
1. 暂不新增 `diagnosis_trace` 或 `trace_event` 主表。
2. 暂不做完整 run 列表 UI。
3. 暂不删除历史 Trace 数据。
4. 暂不改变 Redis 多轮上下文窗口策略。
5. 暂不立即物理删除旧 `diagnosis_session` 表。
---
## 风险
### 1. 兼容风险
现有脚本、Trace 页面和反馈接口可能只知道 `sessionId`。
缓解:保留 `sessionId` 默认查最新 run 的行为。
### 2. 异步上下文风险
工具调用和 Agent hook 依赖 ThreadLocal / RunnableConfig 传递上下文。
缓解:统一上下文对象,明确 `sessionId` 和 `runId` 必须同时传递。
### 3. 历史数据回填风险
旧数据没有真实 run 边界,只能按当前 `diagnosis_session` 生成一条兼容 `diagnosis_run`。
缓解:旧数据视为单 run,不尝试拆分历史混合数据。
### 4. 评分口径变化风险
按 `runId` 隔离后,工具调用数和 evidence score 可能下降,但语义更正确。
缓解:更新 eval fixture 和 baseline,记录这是预期行为变化。
---
## 已收敛问题
本 issue 当前已经收敛以下设计边界:
- `runId` 是正式 API 字段。
- `chat_session` 和 `diagnosis_run` 拆分为两张主表。
- Trace 明细继续由 `agent_step` / `tool_invocation` 承载。
- `chat_session` 只保存元数据,不保存完整对话历史。
- 旧 `diagnosis_session` 保留但新代码不再写入。
- 提供轻量 run 列表 API。
- Feedback 缺少 `runId` 时短期兼容、长期收紧。
- Demo 脚本和 Trace UI 做 `runId` 最小支持。
---
## 相关文件
- `mvp/architecture/session-trace-lifecycle.md`
- `mvp/architecture/data-model.md`
- `mvp/tables/诊断会话表-diagnosis_session.md`
- `mvp/tables/Agent步骤表-agent_step.md`
- `mvp/tables/工具调用表-tool_invocation.md`
- `mvp/tables/案例库表-case_library.md`
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/EvaluationService.java`
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`
- `src/main/java/com/superbiz/agent/service/CaseLibraryService.java`
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/util/SessionContextHolder.java`
- `src/main/resources/db/migration/V005__create_session_storage.sql`
@@ -0,0 +1,3 @@
committed: true
change: session-run-trace-isolation
validated: openspec validate session-run-trace-isolation --strict
@@ -0,0 +1,135 @@
# Decisions: session-run-trace-isolation
## sm-flow State
- Checkpoint: Discover
- Scale: complex
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
- Change slug: `session-run-trace-isolation`
## Entry Summary
Problem: the same `sessionId` currently represents both multi-turn conversation context and one persisted diagnosis trace. Multi-turn Chat E2E proved that Redis context behaves correctly, but MySQL trace rows from different rounds are mixed under one `session_id`.
Expected result: split session metadata from per-run execution state, expose `runId` as the official run identifier, and make trace, feedback, evaluation, demo scripts, Trace UI, and AIOps read/write by run.
Known modules: Flyway/JPA entities/repositories, `ChatService`, `ChatController`, `AiOpsService`, `DiagnosisTraceService`, `EvaluationService`, `FeedbackService`, `CaseLibraryService`, `AgentLoggingHook`, `ToolInvocationRecorder`, `SessionContextHolder`, demo scripts, static Trace UI, MVP docs.
## Context Collection
`devflow/index.md`: relevant entries found.
Relevant historical decisions:
- `session-storage`: current trace persistence is `diagnosis_session + agent_step + tool_invocation`; `sessionId` propagation uses `RunnableConfig.metadata` with `SessionContextHolder` fallback for tools; AIOps records child agents only.
- `confidence-feedback`: `FeedbackService` writes feedback, useful feedback creates `case_library`, and feedback must not change execution status.
- `mvp-demo-trace-acceptance`: Trace API is read-only and demo artifacts/scripts are part of acceptance.
- `aiops-traceable-diagnosis-entry`: `/api/ai_ops` is an SSE entry point that accepts optional alert payload and exposes `sessionId`.
- `data-model.md`: current `case_library.diagnosis_id` maps to `diagnosis_session.session_id`; this must be treated as legacy data after the change.
- `session-trace-lifecycle.md`: current docs already list run id as a future enhancement for multi-run sessions.
OpenSpec inputs that must be carried forward:
- Current specs mention `diagnosis_session` directly in trace, verifier, evidence, and demo requirements; new specs must either supersede or preserve compatibility for those requirements.
- `GET /api/diagnosis/{sessionId}/trace` must remain read-only.
- Baseline/eval checks must expect changed counts only when the change is explained by run isolation.
## Question Pool
| # | Question | Mode | Status |
|---|---|---|---|
| Q1 | Should this be an added field on existing `diagnosis_session`, or split tables? | user-interview | confirmed: split `chat_session` and `diagnosis_run`, keep existing trace detail tables |
| Q2 | What is metadata in `chat_session`? | user-interview | confirmed: session directory fields only, not full message body |
| Q3 | Is "message body" the full conversation history? | user-interview | confirmed: full conversation history stays in Redis `SessionContext.messageHistory` for now |
| Q4 | Should `runId` be an official API field? | user-interview | confirmed: yes |
| Q5 | How should trace work without `runId`? | user-interview | confirmed: default to latest run for compatibility |
| Q6 | Should Feedback without `runId` fail or fall back? | user-interview | confirmed: short-term fallback to latest run, long-term may tighten |
| Q7 | Should simple Chat Q&A create a run? | user-interview | confirmed: yes, every valid `/api/chat` execution creates a run |
| Q8 | Should AIOps be included? | user-interview | confirmed: yes, same run isolation semantics |
| Q9 | Should there be a separate trace table? | user-interview | confirmed: no, current `agent_step` and `tool_invocation` are enough for this phase |
| Q10 | What should happen to old `diagnosis_session`? | user-interview | confirmed: keep it for history/rollback, new code stops writing it after migration |
| Q11 | Does the issue need demo/Trace UI support? | user-interview | confirmed: yes, minimal `runId` support |
| Q12 | What extra risks were found by document review? | evidence-driven | reported and patched into ISS-010 |
## Evidence-Driven Findings
- Code evidence: `SessionContext` contains `messageHistory` and `getMessagePairCount()`, supporting the decision that Redis holds hot conversation history while MySQL stores auditable per-run query/answer.
- Code evidence: `CaseLibraryService.createFromSession` currently deduplicates by `session.getSessionId()` and maps answer/query from `DiagnosisSession`; this must change for new run-based data.
- Documentation evidence: `mvp/architecture/data-model.md` states `case_library.diagnosis_id = diagnosis_session.session_id`; this becomes transitional legacy semantics.
- Documentation evidence: existing trace OpenSpec requires `GET /api/diagnosis/{sessionId}/trace` to be read-only; run resolution must preserve that invariant.
- E2E evidence from ISS-010: two Chat rounds with the same `sessionId` resulted in one overwritten `diagnosis_session` row and mixed step/tool rows.
## Confirmed Decisions
- `runId` format: `run-` + full UUID.
- Latest run ordering: `diagnosis_run.created_at DESC, id DESC`, not `updated_at`.
- `chat_session` stores metadata: `session_id`, `status`, `message_pair_count`, `created_at`, `last_active_at`, `expires_at`.
- Per-run long-term audit stores `query` and `answer` in `diagnosis_run`.
- `agent_step` and `tool_invocation` retain `session_id` and add `run_id`.
- Missing feedback `runId` returns `fallbackToLatestRun=true` plus actual bound `runId`.
- Historical mixed data is not split into multiple true runs.
## OpenSpec Backwrite Log
- Created `proposal.md` with problem, scope, non-goals, context constraints, interface impact, and risks.
- ISS-010 patched to clarify phase boundary, migration order, AIOps same-change requirement, feedback fallback response, case-library transitional semantics, and Redis/MySQL TTL boundary.
- Glossary patched to include `SessionContext.messageHistory` and its persistence boundary.
- Created `design.md`, `specs/session-run-trace-isolation/spec.md`, `specs/mvp-demo-trace-acceptance/spec.md`, and `tasks.md`.
- Architecture audit found that `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` also read tool rows by `sessionId`; Phase 2 tasks were updated to cover run-scoped reads.
## Cross-Artifact Alignment
| Check | Result | Evidence |
|---|---|---|
| issue/context -> proposal | aligned | proposal carries E2E problem, split-table solution, compatibility, AIOps, feedback, demo/UI, and baseline scope |
| proposal -> design | aligned | design records data model, API impact, migration plan, rollback, and key decisions |
| design -> specs | aligned | specs cover session/run split, Chat, trace, feedback, AIOps, migration, demo/UI, E2E, and baseline behavior |
| specs -> tasks | aligned | tasks implement schema, Chat write path, trace reads, feedback/case, AIOps, demo/UI/docs, and verification gates |
## Architecture Audit
Input to output chain:
```text
Chat/AIOps request
-> ChatController / AIOps endpoint
-> ChatService / AiOpsService
-> execution context(sessionId, runId)
-> AgentLoggingHook -> agent_step
-> ToolInvocationRecorder -> tool_invocation
-> verifier/gatekeeper/evaluation summary reads
-> diagnosis_run answer/self_evaluation/status/counts
-> Trace API / Feedback / CaseLibrary / Demo / Trace UI
```
Audit conclusions:
- Data ownership is clearer with `chat_session` owning conversation metadata and `diagnosis_run` owning execution state; `agent_step` and `tool_invocation` remain trace details owned by one run.
- The highest coupling risk is execution-context propagation because hooks and tools currently use both `RunnableConfig.metadata` and `SessionContextHolder`.
- The main read-path risk is missing one of the session-scoped consumers (`ToolTraceSummaryService`, `ExecutorGatekeeperService`, `EvaluationService`, trace, feedback, case creation).
- Migration is additive and rollback-friendly until constraints are tightened; historical mixed data must be treated as compatibility data.
- AIOps must complete before final archive because otherwise the system would still have one production entry point with mixed trace semantics.
## Commit Gate Preflight
- Interface impact level: L4.
- Required artifacts: proposal, design, specs, tasks.
- Strict OpenSpec validation: passed with `openspec validate session-run-trace-isolation --strict`.
- `.committed` marker: created after successful commit gate.
## Pre-apply Research
- Reference migrations: `V005__create_session_storage.sql`, `V008__add_answer_to_diagnosis_session.sql`, `V010__add_relevance_level_to_tool_invocation.sql`.
- Entity style: JPA entities use Lombok `@Data`, `@Builder`, `@NoArgsConstructor`, `@AllArgsConstructor`, `@PrePersist`, and `@PreUpdate` where timestamps need maintenance.
- Repository style: Spring Data JPA repository interfaces with derived query methods returning `Optional<T>` or `List<T>`.
- Test style: repository tests use `@DataJpaTest`, `@AutoConfigureTestDatabase(replace = NONE)`, Flyway enabled, and `ddl-auto=validate`.
## Phase 1 Apply Notes
- Implemented additive migration `V011__add_session_run_isolation.sql`.
- Implemented `ChatSession` / `DiagnosisRun` entities and repositories.
- Added nullable `runId` fields to `AgentStep` and `ToolInvocation`.
- Added run-scoped repository methods for step/tool lookup and counts.
- Added `DiagnosisRunRepositoryTest`.
- Verification evidence is recorded in `phase-1-evidence.md`.
- Historical DB inspection found one orphan `tool_invocation` row without a matching `diagnosis_session`; it remains `run_id = NULL` because no reliable compatibility run can be inferred.
@@ -0,0 +1,139 @@
## Context
The current MVP persists diagnosis observability through `diagnosis_session`, `agent_step`, and `tool_invocation`. That model works for a single diagnosis per `sessionId`, but it conflates two lifecycles once a caller reuses the same `sessionId` for multi-turn conversation:
- conversation state: Redis `SessionContext.messageHistory` and session metadata;
- execution state: one diagnosis answer, trace, self-evaluation, and feedback target.
The E2E evidence in ISS-010 showed that Redis correctly preserved multi-turn context while MySQL mixed both rounds under the same `session_id`. This breaks trace replay, verifier/evaluation scoping, feedback targeting, and case-library provenance.
Constraints from existing work:
- `GET /api/diagnosis/{sessionId}/trace` is a read-only MVP/demo contract and must remain compatible.
- Current hooks and tools propagate `sessionId` through `RunnableConfig.metadata` and `SessionContextHolder`; `runId` must follow the same execution context boundary.
- Feedback currently writes `DiagnosisSession.feedback` and creates `case_library` from `DiagnosisSession.answer`.
- AIOps is a first-class traceable entry point and cannot be left permanently on the old mixed-run model.
## Goals / Non-Goals
**Goals:**
- Split conversation metadata from per-execution diagnosis state.
- Introduce `runId` as the official identifier for one replayable diagnosis execution.
- Preserve old `sessionId`-only callers by resolving the latest run where possible.
- Scope trace, evaluation, feedback, case creation, demo scripts, and Trace UI by run.
- Migrate historical data into compatibility runs without deleting the old table.
- Include Chat and AIOps in the same release-level change.
- Verify behavior with focused tests, E2E when needed, DB inspection, logs, and baseline drift checks.
**Non-Goals:**
- Do not persist full chat history in MySQL.
- Do not add `diagnosis_trace` or `trace_event`.
- Do not implement full run-list UI.
- Do not physically delete `diagnosis_session`.
- Do not attempt to reconstruct true historical round boundaries when only mixed `session_id` data exists.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Domain split | Add `chat_session` and `diagnosis_run` | Add `run_id` to `diagnosis_session` only | Separate tables keep conversation metadata and execution state from growing into one coupled table. |
| Trace detail storage | Reuse `agent_step` and `tool_invocation`, adding `run_id` | Add `diagnosis_trace` / `trace_event` | Existing detail tables already represent trace; isolation needs a run key, not a new event model. |
| API identity | `runId = "run-" + UUID` | Reuse short session id or DB id | Full UUID avoids collision and keeps external IDs independent from database internals. |
| Latest-run compatibility | `GET /api/diagnosis/{sessionId}/trace` resolves latest by `created_at DESC, id DESC` | Require `runId` immediately | Compatibility keeps existing demo/UI/scripts working while new clients migrate. `updated_at` is avoided because feedback/eval updates can reorder old runs. |
| Historical migration | Backfill one compatibility run per existing `diagnosis_session` | Try to split old mixed rows | Old rows do not contain reliable run boundaries. A compatibility run preserves auditability without inventing data. |
| Feedback fallback | Missing `runId` binds latest run and returns fallback metadata | Reject missing `runId` immediately | Short-term compatibility is needed for old clients; explicit fallback keeps ambiguity observable. |
| Case library provenance | New automatic cases store `diagnosis_id = run_id` | Add a new case-library column now | Existing column name can carry transitional provenance; docs and query logic must recognize old `session_id` and new `run_id`. |
| AIOps phase | Implement after Chat but before overall archive | Leave AIOps for a follow-up issue | AIOps is already a trace entry point; leaving it old-model would preserve the same bug on another endpoint. |
## Data Model
```text
chat_session
id
session_id unique
status
message_pair_count
created_at
last_active_at
expires_at
diagnosis_run
id
run_id unique
session_id
query
status
agent_flow
answer
self_evaluation
feedback
total_duration_ms
total_token_count
step_count
tool_call_count
created_at
updated_at
agent_step
session_id
run_id
...
tool_invocation
session_id
run_id
...
```
`chat_session.expires_at` is MySQL directory metadata. Redis TTL can expire `SessionContext.messageHistory`; persisted `diagnosis_run`, `agent_step`, and `tool_invocation` remain audit records.
## API / Interface Impact
Interface level: L4.
- Database contract changes: new tables, new columns, backfill, indexes, and later non-null expectations for new writes.
- `/api/chat` response adds `runId`.
- `/api/ai_ops` SSE emits the resolved `runId`.
- Trace API accepts optional `runId`.
- Feedback request accepts preferred `runId` and returns fallback binding metadata when omitted.
- New run summary API: `GET /api/chat/session/{sessionId}/runs`.
Compatibility:
- Old `sessionId`-only trace and feedback calls bind to latest run.
- Old `diagnosis_session` is retained for rollback and historical comparison.
- New code must not keep writing new execution state into `diagnosis_session` after the write switch.
## Migration Plan
1. Add `chat_session` and `diagnosis_run`.
2. Add nullable `run_id` to `agent_step` and `tool_invocation`.
3. Backfill `diagnosis_run` from existing `diagnosis_session`.
4. Backfill old `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run.
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
6. Deploy repository and read-path compatibility.
7. Switch Chat write path to `chat_session + diagnosis_run`.
8. Switch Trace, evaluation, feedback, case-library, demo, and UI paths.
9. Switch AIOps write path.
10. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
Rollback:
- Keep `diagnosis_session` intact during this change.
- Migrations are additive until constraints are tightened.
- If write-switch rollout fails, rollback code can read the retained old table while migrated compatibility rows remain harmless.
## Risks / Trade-offs
- [Risk] ThreadLocal/context propagation may miss `runId` in nested agent/tool calls. -> Mitigation: introduce a unified execution context carrying both IDs and test hook/tool recording.
- [Risk] Baseline metrics change because cross-round tool rows are no longer counted. -> Mitigation: run baseline diff and document expected drift.
- [Risk] Legacy `case_library.diagnosis_id` values are ambiguous. -> Mitigation: document transitional semantics and keep lookup logic aware of old `session_id` values.
- [Risk] AIOps SSE clients may not parse a new metadata event. -> Mitigation: add metadata in a compatible stream message and keep final report streaming behavior.
- [Risk] Historical mixed trace cannot be truly separated. -> Mitigation: call this out as compatibility data, not reconstructed truth.
## Open Questions
None blocking. Long-term tightening of missing feedback `runId` remains a follow-up decision after clients migrate.
@@ -0,0 +1,103 @@
# Phase 1 Evidence: Schema and Compatibility Foundation
## Scope Completed
- Added Flyway migration `V011__add_session_run_isolation.sql`.
- Added `chat_session` and `diagnosis_run` tables.
- Added nullable `run_id` columns and indexes to `agent_step` and `tool_invocation`.
- Backfilled compatibility runs from existing `diagnosis_session` rows.
- Backfilled existing step/tool rows where a matching `diagnosis_session.session_id` exists.
- Added `ChatSession`, `DiagnosisRun`, `ChatSessionRepository`, and `DiagnosisRunRepository`.
- Added run-scoped query/count methods to `AgentStepRepository` and `ToolInvocationRepository`.
- Added `DiagnosisRunRepositoryTest`.
## Verification
### Maven
```text
mvn -q "-Dtest=DiagnosisRunRepositoryTest" test
```
Result: passed.
```text
mvn -q "-Dtest=DiagnosisSessionRepositoryTest,AgentStepRepositoryTest,ToolInvocationRepositoryTest,DiagnosisRunRepositoryTest" test
```
Result: passed.
### Database Inspection
Checked with `scripts/query_mysql.py`.
```text
SELECT installed_rank, version, description, success
FROM flyway_schema_history
ORDER BY installed_rank DESC
LIMIT 5;
```
Relevant result:
```text
version = 011
description = add session run isolation
success = 1
```
```text
SELECT COUNT(*) AS chat_sessions FROM chat_session;
```
Result:
```text
chat_sessions = 83
```
```text
SELECT COUNT(*) AS diagnosis_runs,
SUM(CASE WHEN run_id IS NULL THEN 1 ELSE 0 END) AS null_run_ids
FROM diagnosis_run;
```
Result:
```text
diagnosis_runs = 83
null_run_ids = 0
```
```text
SELECT (SELECT COUNT(*) FROM agent_step WHERE run_id IS NULL) AS agent_step_null_run_id,
(SELECT COUNT(*) FROM tool_invocation WHERE run_id IS NULL) AS tool_invocation_null_run_id;
```
Result:
```text
agent_step_null_run_id = 0
tool_invocation_null_run_id = 1
```
The one remaining null `tool_invocation.run_id` is historical orphan data:
```text
id = 3
session_id = f9d2290d
tool_name = lookup_knowledge
created_at = 2026-06-26 14:17:35
```
It has no matching `diagnosis_session`, so V011 cannot safely infer a compatibility run. This matches the OpenSpec wording "when possible" for historical backfill.
### Logs
`logs/application.log` shows Hibernate using the new `run_id` fields in `agent_step` and `tool_invocation` inserts/selects during the repository verification.
## Notes
- Phase 1 intentionally does not switch Chat, Trace, Feedback, or AIOps write paths.
- New columns remain nullable until later phases switch new writes and verify coverage.
@@ -0,0 +1,85 @@
# Change: Session / Run / Trace Isolation
## Problem
The current MVP uses the same `sessionId` for two different concepts:
- Redis `SessionContext` keeps multi-turn chat history for prompt context.
- MySQL `diagnosis_session`, `agent_step`, and `tool_invocation` persist diagnosis trace data for replay, verification, feedback, and evaluation.
End-to-end verification showed that two `/api/chat` calls with the same `sessionId` correctly reuse Redis context, but MySQL trace data is mixed under the same key:
- `diagnosis_session.query` is overwritten by the second round.
- `agent_step` and `tool_invocation` append rows from both rounds under the same `session_id`.
- Trace, verifier/evaluation, and feedback can read cross-round evidence.
This makes a trace no longer represent one replayable diagnosis run.
## Proposed Solution
Introduce a stable split between conversation state and execution state:
- `chat_session`: conversation metadata keyed by `session_id`.
- `diagnosis_run`: one execution/run keyed by `run_id`, belonging to a `session_id`.
- `agent_step` and `tool_invocation`: keep existing trace detail role, add `run_id` while retaining `session_id` for compatibility and coarse filtering.
`runId` becomes an official API field:
- `/api/chat` returns `sessionId + runId`.
- `/api/ai_ops` emits or returns `runId` in the SSE-compatible protocol.
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
Trace remains an aggregate view of `diagnosis_run + agent_step + tool_invocation`; this change does not introduce a separate `diagnosis_trace` or `trace_event` table.
## Scope
- Add Flyway migrations and JPA entities/repositories for `chat_session` and `diagnosis_run`.
- Add nullable `run_id` to `agent_step` and `tool_invocation`, backfill historical data, then switch new writes to require run context.
- Move new Chat writes from `diagnosis_session` to `chat_session + diagnosis_run`.
- Update trace reads to resolve latest run or specified run.
- Add lightweight run list API: `GET /api/chat/session/{sessionId}/runs`.
- Update feedback and case-library creation to bind new data to `run_id`.
- Update AIOps to create and expose `runId` before this change is considered production complete.
- Update demo scripts and Trace UI with minimal `runId` support.
- Update relevant MVP table and architecture documentation.
- Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under `logs/`, database queries via `scripts/query_mysql.py`, and baseline drift checks.
## Non-Goals
- Do not add `diagnosis_trace` or `trace_event` in this change.
- Do not implement a full run-list UI.
- Do not remove the historical `diagnosis_session` table in this change.
- Do not change the Redis conversation window strategy.
- Do not persist full conversation history in MySQL; `chat_session` stores metadata only.
- Do not split historical mixed traces into true historical runs when the original run boundary is unavailable.
## Devflow Context Constraints
- `session-storage` established the current trace tables and decided that `sessionId` is propagated through `RunnableConfig.metadata` with `SessionContextHolder` as a fallback for tools.
- `confidence-feedback` established that useful feedback creates `case_library` from `DiagnosisSession.answer`, and that `feedback` does not change execution `status`.
- `mvp-demo-trace-acceptance` established `GET /api/diagnosis/{sessionId}/trace` as a read-only endpoint and demo scripts as part of the observable story.
- `aiops-traceable-diagnosis-entry` established `/api/ai_ops` as a traceable SSE entry point and made `sessionId` visible to callers.
- `data-model.md` and `session-trace-lifecycle.md` describe the current model as `diagnosis_session + agent_step + tool_invocation`, and list run id as a known follow-up.
- `devflow/glossary/CONTEXT.md` now defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks.
## Interface Impact
Level: L4 database/API contract migration with compatibility behavior.
- New API response field: `runId`.
- New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`.
- New API: `GET /api/chat/session/{sessionId}/runs`.
- Feedback request gains optional/preferred `runId`.
- Database contract changes include new tables and new `run_id` columns.
- Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable.
## Risks
- Historical data has no true per-round boundary; backfill can only create compatibility runs from existing `diagnosis_session` rows.
- Context propagation through Agent hooks and tools is easy to break because it currently combines `RunnableConfig.metadata` and `SessionContextHolder`.
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
@@ -0,0 +1,29 @@
## MODIFIED Requirements
### Requirement: Diagnosis trace can be queried by session id
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
#### Scenario: Existing session latest trace is returned
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
#### Scenario: Existing session exact trace is returned
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system validates that `runId` belongs to `sessionId`
- **AND** it returns a success response containing only the trace data for that run
#### Scenario: Missing session returns not found
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
- **THEN** the system returns a 404 response using the existing session-not-found error contract
### Requirement: Trace aggregation is read-only
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
#### Scenario: Trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
#### Scenario: Exact trace query does not change persisted state
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
@@ -0,0 +1,147 @@
## ADDED Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
#### Scenario: Valid Chat execution creates session metadata and a run
- **WHEN** a valid `/api/chat` request enters the Chat execution path with a `sessionId`
- **THEN** the system SHALL ensure a `chat_session` row exists for that `sessionId`
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
- **AND** the `diagnosis_run.session_id` SHALL equal the request `sessionId`
#### Scenario: Invalid Chat request does not create a run
- **WHEN** a `/api/chat` request fails parameter validation before execution
- **THEN** the system SHALL NOT create a `diagnosis_run`
#### Scenario: Chat session stores metadata only
- **WHEN** a Chat request completes
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and expiration time
- **AND** it SHALL NOT store full conversation message history
### Requirement: Chat responses SHALL expose run identity
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
#### Scenario: Chat response includes runId
- **WHEN** `/api/chat` returns a successful response
- **THEN** the response SHALL include `sessionId`
- **AND** the response SHALL include `runId` for the created diagnosis run
#### Scenario: Multi-turn Chat keeps one session and multiple runs
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
### Requirement: Trace details SHALL be scoped by run
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
#### Scenario: Agent steps are recorded with runId
- **WHEN** an Agent model step is persisted during a diagnosis run
- **THEN** the `agent_step` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Tool invocations are recorded with runId
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
- **AND** it SHALL retain the current `session_id`
#### Scenario: Run metrics count only current run rows
- **WHEN** a diagnosis run completes
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
### Requirement: Trace API SHALL support latest-run and exact-run queries
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
#### Scenario: Trace without runId resolves latest run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
- **AND** the response SHALL include the resolved `runId`
#### Scenario: Trace with runId returns exact run
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
#### Scenario: Trace rejects run from another session
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
### Requirement: Session runs SHALL be listable without expanding trace details
The system SHALL provide a lightweight run-list API for a Chat Session.
#### Scenario: Run list returns summaries
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
- **THEN** the system SHALL return run summaries from `diagnosis_run`
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
### Requirement: Feedback SHALL bind to diagnosis runs
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
#### Scenario: Feedback with runId updates specified run
- **WHEN** a feedback request includes `sessionId` and `runId`
- **THEN** the system SHALL validate that the run belongs to the session
- **AND** it SHALL update feedback on that run
#### Scenario: Feedback without runId falls back observably
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
- **THEN** the system SHALL bind feedback to the latest run for that session
- **AND** the response or log SHALL include `fallbackToLatestRun=true`
- **AND** the response or log SHALL include the actual bound `runId`
#### Scenario: Useful feedback creates case from run
- **WHEN** feedback for a run is `useful`
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
### Requirement: AIOps executions SHALL use run isolation
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
#### Scenario: AIOps creates run
- **WHEN** `/api/ai_ops` starts a valid execution
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
#### Scenario: AIOps SSE exposes runId
- **WHEN** `/api/ai_ops` streams response metadata to the caller
- **THEN** the stream SHALL expose the resolved `sessionId`
- **AND** it SHALL expose the created `runId`
### Requirement: Migration SHALL preserve historical trace access
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
#### Scenario: Historical session gets compatibility run
- **WHEN** migration runs on an existing `diagnosis_session` row
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
#### Scenario: Old table is retained
- **WHEN** the migration completes
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
### Requirement: Demo and Trace UI SHALL support runId
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
#### Scenario: Demo script queries exact trace
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
- **THEN** it SHALL include `runId` when querying the Trace API
#### Scenario: Trace UI honors URL runId
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
### Requirement: Run isolation SHALL be verified against baselines
The change SHALL verify both runtime behavior and evaluation baseline impact.
#### Scenario: Multi-turn E2E proves run isolation
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
- **THEN** database inspection SHALL show two `diagnosis_run` rows
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
- **AND** Redis session metadata SHALL still show multi-turn context continuity
#### Scenario: Baseline drift is checked
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
@@ -0,0 +1,59 @@
## 1. Schema and Compatibility Foundation
- [x] 1.1 Add Flyway migration for `chat_session` and `diagnosis_run` with indexes for unique `session_id`, unique `run_id`, latest-run lookup, and run summary listing.
- [x] 1.2 Add nullable `run_id` columns to `agent_step` and `tool_invocation` with indexes for run-scoped trace queries.
- [x] 1.3 Backfill one compatibility `diagnosis_run` for each existing `diagnosis_session` row.
- [x] 1.4 Backfill historical `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run for the same `session_id`.
- [x] 1.5 Add JPA entities and repositories for `ChatSession` and `DiagnosisRun`, including latest-run and run-id lookup methods.
- [x] 1.6 Add focused migration/repository verification that proves old data remains queryable and run lookup methods work.
- [x] 1.7 Phase 1 gate: run the smallest relevant test/build check, inspect DB migration behavior, update OpenSpec task status, archive phase evidence, and commit before starting Phase 2.
## 2. Chat Run Write Path
- [ ] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
- [ ] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
- [ ] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
- [ ] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
- [ ] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
- [ ] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
- [ ] 2.7 Change `/api/chat` response DTO to include official `runId`.
- [ ] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
- [ ] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
## 3. Trace Read Path and Run Listing
- [ ] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
- [ ] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
- [ ] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
- [ ] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
- [ ] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
- [ ] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
- [ ] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
## 4. Feedback and Case Library Run Binding
- [ ] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
- [ ] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
- [ ] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
- [ ] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
- [ ] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
- [ ] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
- [ ] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
## 5. AIOps Run Isolation
- [ ] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
- [ ] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
- [ ] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
- [ ] 5.4 Change AIOps final report, status, counts, and `aiops_rule_evaluation` writes to the current `diagnosis_run`.
- [ ] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
- [ ] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
## 6. Demo, Trace UI, Documentation, and Verification
- [ ] 6.1 Update demo scripts to read `runId` from Chat/AIOps responses and pass `?runId=...` to Trace API.
- [ ] 6.2 Update Trace UI to accept `?sessionId=...&runId=...` and query exact trace when `runId` is present.
- [ ] 6.3 Update MVP table and architecture docs for `chat_session`, `diagnosis_run`, `run_id`, and transitional `case_library.diagnosis_id` semantics.
- [ ] 6.4 Run final same-session multi-turn E2E using Maven startup if needed; collect DB evidence through `scripts/query_mysql.py` and inspect `logs/`.
- [ ] 6.5 Run or explicitly evaluate the relevant baseline diff command and document whether drift is expected or a regression.
- [ ] 6.6 Final gate: ensure all OpenSpec tasks are checked, no new writes depend on `diagnosis_session`, phase evidence is archived, final commit is created, and the change is ready for OpenSpec archive.
@@ -15,6 +15,7 @@ import java.time.LocalDateTime;
@Entity @Entity
@Table(name = "agent_step", indexes = { @Table(name = "agent_step", indexes = {
@Index(name = "idx_session_step", columnList = "session_id, step_index"), @Index(name = "idx_session_step", columnList = "session_id, step_index"),
@Index(name = "idx_agent_step_run_step", columnList = "run_id, step_index"),
@Index(name = "idx_agent_name", columnList = "agent_name") @Index(name = "idx_agent_name", columnList = "agent_name")
}) })
@Data @Data
@@ -30,6 +31,9 @@ public class AgentStep {
@Column(name = "session_id", nullable = false, length = 64) @Column(name = "session_id", nullable = false, length = 64)
private String sessionId; private String sessionId;
@Column(name = "run_id", length = 64)
private String runId;
@Column(name = "step_index", nullable = false) @Column(name = "step_index", nullable = false)
private Integer stepIndex; private Integer stepIndex;
@@ -0,0 +1,64 @@
package com.superbiz.agent.domain.entity;
import jakarta.persistence.*;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
import java.time.LocalDateTime;
/**
* Chat session metadata entity.
* Full message history remains in Redis SessionContext.
*/
@Entity
@Table(name = "chat_session", indexes = {
@Index(name = "idx_chat_session_last_active", columnList = "last_active_at"),
@Index(name = "idx_chat_session_status", columnList = "status"),
@Index(name = "idx_chat_session_expires_at", columnList = "expires_at")
})
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class ChatSession {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(name = "session_id", unique = true, nullable = false, length = 64)
private String sessionId;
@Column(name = "status", length = 16)
private String status = "ACTIVE";
@Column(name = "message_pair_count")
private Integer messagePairCount = 0;
@Column(name = "created_at", nullable = false, updatable = false)
private LocalDateTime createdAt;
@Column(name = "last_active_at")
private LocalDateTime lastActiveAt;
@Column(name = "expires_at")
private LocalDateTime expiresAt;
@PrePersist
protected void onCreate() {
LocalDateTime now = LocalDateTime.now();
createdAt = now;
if (lastActiveAt == null) {
lastActiveAt = now;
}
if (status == null) {
status = "ACTIVE";
}
if (messagePairCount == null) {
messagePairCount = 0;
}
}
}
@@ -0,0 +1,91 @@
package com.superbiz.agent.domain.entity;
import jakarta.persistence.*;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
import org.hibernate.annotations.JdbcTypeCode;
import org.hibernate.type.SqlTypes;
import java.time.LocalDateTime;
/**
* One diagnosis execution run.
*/
@Entity
@Table(name = "diagnosis_run", indexes = {
@Index(name = "idx_diagnosis_run_session_created", columnList = "session_id, created_at, id"),
@Index(name = "idx_diagnosis_run_session_run", columnList = "session_id, run_id"),
@Index(name = "idx_diagnosis_run_status", columnList = "status"),
@Index(name = "idx_diagnosis_run_agent_flow", columnList = "agent_flow")
})
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class DiagnosisRun {
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;
@Column(name = "run_id", unique = true, nullable = false, length = 64)
private String runId;
@Column(name = "session_id", nullable = false, length = 64)
private String sessionId;
@Column(name = "query", nullable = false, columnDefinition = "TEXT")
private String query;
@Column(name = "status", length = 16)
private String status = "PENDING";
@Column(name = "agent_flow", length = 32)
private String agentFlow;
@Column(name = "answer", columnDefinition = "LONGTEXT")
private String answer;
@JdbcTypeCode(SqlTypes.JSON)
@Column(name = "self_evaluation", columnDefinition = "JSON")
private String selfEvaluation;
@Column(name = "feedback", length = 16)
private String feedback;
@Column(name = "total_duration_ms")
private Integer totalDurationMs;
@Column(name = "total_token_count")
private Integer totalTokenCount;
@Column(name = "step_count")
private Integer stepCount;
@Column(name = "tool_call_count")
private Integer toolCallCount;
@Column(name = "created_at", nullable = false, updatable = false)
private LocalDateTime createdAt;
@Column(name = "updated_at")
private LocalDateTime updatedAt;
@PrePersist
protected void onCreate() {
LocalDateTime now = LocalDateTime.now();
createdAt = now;
updatedAt = now;
if (status == null) {
status = "PENDING";
}
}
@PreUpdate
protected void onUpdate() {
updatedAt = LocalDateTime.now();
}
}
@@ -17,6 +17,7 @@ import java.time.LocalDateTime;
@Entity @Entity
@Table(name = "tool_invocation", indexes = { @Table(name = "tool_invocation", indexes = {
@Index(name = "idx_session_id", columnList = "session_id"), @Index(name = "idx_session_id", columnList = "session_id"),
@Index(name = "idx_tool_invocation_run_id", columnList = "run_id, id"),
@Index(name = "idx_tool_name", columnList = "tool_name"), @Index(name = "idx_tool_name", columnList = "tool_name"),
@Index(name = "idx_retrieval_layer", columnList = "retrieval_layer") @Index(name = "idx_retrieval_layer", columnList = "retrieval_layer")
}) })
@@ -33,6 +34,9 @@ public class ToolInvocation {
@Column(name = "session_id", nullable = false, length = 64) @Column(name = "session_id", nullable = false, length = 64)
private String sessionId; private String sessionId;
@Column(name = "run_id", length = 64)
private String runId;
@Column(name = "step_id") @Column(name = "step_id")
private Long stepId; private Long stepId;
@@ -22,8 +22,18 @@ public interface AgentStepRepository extends JpaRepository<AgentStep, Long> {
*/ */
List<AgentStep> findBySessionId(String sessionId); List<AgentStep> findBySessionId(String sessionId);
/**
* 根据运行ID查询所有步骤(按步骤号排序)。
*/
List<AgentStep> findByRunIdOrderByStepIndex(String runId);
/** /**
* 统计某个会话的步骤数 * 统计某个会话的步骤数
*/ */
int countBySessionId(String sessionId); int countBySessionId(String sessionId);
/**
* 统计某个运行的步骤数。
*/
int countByRunId(String runId);
} }
@@ -0,0 +1,14 @@
package com.superbiz.agent.repository;
import com.superbiz.agent.domain.entity.ChatSession;
import org.springframework.data.jpa.repository.JpaRepository;
import org.springframework.stereotype.Repository;
import java.util.Optional;
@Repository
public interface ChatSessionRepository extends JpaRepository<ChatSession, Long> {
Optional<ChatSession> findBySessionId(String sessionId);
}
@@ -0,0 +1,21 @@
package com.superbiz.agent.repository;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import org.springframework.data.jpa.repository.JpaRepository;
import org.springframework.stereotype.Repository;
import java.util.List;
import java.util.Optional;
@Repository
public interface DiagnosisRunRepository extends JpaRepository<DiagnosisRun, Long> {
Optional<DiagnosisRun> findByRunId(String runId);
Optional<DiagnosisRun> findBySessionIdAndRunId(String sessionId, String runId);
Optional<DiagnosisRun> findFirstBySessionIdOrderByCreatedAtDescIdDesc(String sessionId);
List<DiagnosisRun> findBySessionIdOrderByCreatedAtDescIdDesc(String sessionId);
}
@@ -22,11 +22,21 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
*/ */
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId); List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
/**
* 根据运行ID按创建顺序查询所有工具调用
*/
List<ToolInvocation> findByRunIdOrderByIdAsc(String runId);
/** /**
* 根据会话ID统计真实工具调用次数 * 根据会话ID统计真实工具调用次数
*/ */
long countBySessionId(String sessionId); long countBySessionId(String sessionId);
/**
* 根据运行ID统计真实工具调用次数
*/
long countByRunId(String runId);
/** /**
* 根据工具名查询所有调用 * 根据工具名查询所有调用
*/ */
@@ -0,0 +1,111 @@
-- V011: split conversation metadata from diagnosis execution runs.
CREATE TABLE chat_session (
id BIGINT PRIMARY KEY AUTO_INCREMENT,
session_id VARCHAR(64) UNIQUE NOT NULL COMMENT 'Chat session id / conversation context id',
status VARCHAR(16) DEFAULT 'ACTIVE' COMMENT 'ACTIVE/EXPIRED/CLOSED',
message_pair_count INT DEFAULT 0 COMMENT 'Cached Redis message pair count snapshot',
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
last_active_at DATETIME DEFAULT CURRENT_TIMESTAMP,
expires_at DATETIME COMMENT 'Directory metadata only; Redis message history may expire independently',
INDEX idx_chat_session_last_active (last_active_at),
INDEX idx_chat_session_status (status),
INDEX idx_chat_session_expires_at (expires_at)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Chat session metadata table';
CREATE TABLE diagnosis_run (
id BIGINT PRIMARY KEY AUTO_INCREMENT,
run_id VARCHAR(64) UNIQUE NOT NULL COMMENT 'Diagnosis run id, format run-uuid',
session_id VARCHAR(64) NOT NULL COMMENT 'Owner chat_session.session_id',
query TEXT NOT NULL COMMENT 'User question or AIOps alert summary for this run',
status VARCHAR(16) DEFAULT 'PENDING' COMMENT 'PENDING/RUNNING/SUCCESS/FAILED',
agent_flow VARCHAR(32) COMMENT 'CHAT / AI_OPS',
answer LONGTEXT COMMENT 'Final answer/report for this run',
self_evaluation JSON COMMENT 'Run-scoped self evaluation payload',
feedback VARCHAR(16) COMMENT 'User feedback for this run',
total_duration_ms INT COMMENT 'Run duration in milliseconds',
total_token_count INT COMMENT 'Run token count',
step_count INT COMMENT 'Run agent step count',
tool_call_count INT COMMENT 'Run tool invocation count',
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
INDEX idx_diagnosis_run_session_created (session_id, created_at, id),
INDEX idx_diagnosis_run_session_run (session_id, run_id),
INDEX idx_diagnosis_run_status (status),
INDEX idx_diagnosis_run_agent_flow (agent_flow)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Diagnosis execution run table';
ALTER TABLE agent_step
ADD COLUMN run_id VARCHAR(64) NULL COMMENT '关联 diagnosis_run.run_id' AFTER session_id,
ADD INDEX idx_agent_step_run_step (run_id, step_index);
ALTER TABLE tool_invocation
ADD COLUMN run_id VARCHAR(64) NULL COMMENT '关联 diagnosis_run.run_id' AFTER session_id,
ADD INDEX idx_tool_invocation_run_id (run_id, id);
INSERT INTO chat_session (
session_id,
status,
message_pair_count,
created_at,
last_active_at,
expires_at
)
SELECT
ds.session_id,
'ACTIVE',
0,
ds.created_at,
COALESCE(ds.updated_at, ds.created_at),
NULL
FROM diagnosis_session ds;
INSERT INTO diagnosis_run (
run_id,
session_id,
query,
status,
agent_flow,
answer,
self_evaluation,
feedback,
total_duration_ms,
total_token_count,
step_count,
tool_call_count,
created_at,
updated_at
)
SELECT
CONCAT('run-', UUID()),
ds.session_id,
ds.query,
ds.status,
ds.agent_flow,
ds.answer,
ds.self_evaluation,
ds.feedback,
ds.total_duration_ms,
ds.total_token_count,
ds.step_count,
ds.tool_call_count,
ds.created_at,
ds.updated_at
FROM diagnosis_session ds;
UPDATE agent_step ast
JOIN diagnosis_run dr ON dr.session_id = ast.session_id
SET ast.run_id = dr.run_id
WHERE ast.run_id IS NULL;
UPDATE tool_invocation ti
JOIN diagnosis_run dr ON dr.session_id = ti.session_id
SET ti.run_id = dr.run_id
WHERE ti.run_id IS NULL;
@@ -0,0 +1,112 @@
package com.superbiz.agent.repository;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.ChatSession;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.domain.entity.ToolInvocation;
import org.junit.jupiter.api.Test;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.boot.test.autoconfigure.jdbc.AutoConfigureTestDatabase;
import org.springframework.boot.test.autoconfigure.orm.jpa.DataJpaTest;
import org.springframework.test.context.TestPropertySource;
import java.util.List;
import java.util.UUID;
import static org.junit.jupiter.api.Assertions.*;
@DataJpaTest
@AutoConfigureTestDatabase(replace = AutoConfigureTestDatabase.Replace.NONE)
@TestPropertySource(properties = {
"spring.flyway.enabled=true",
"spring.jpa.hibernate.ddl-auto=validate",
"spring.jpa.show-sql=true"
})
class DiagnosisRunRepositoryTest {
@Autowired
private ChatSessionRepository chatSessionRepository;
@Autowired
private DiagnosisRunRepository diagnosisRunRepository;
@Autowired
private AgentStepRepository agentStepRepository;
@Autowired
private ToolInvocationRepository toolInvocationRepository;
@Test
void saveAndFindLatestRunBySessionId() {
String sessionId = "test-session-" + UUID.randomUUID();
DiagnosisRun firstRun = saveRun(sessionId, "run-" + UUID.randomUUID(), "first question");
DiagnosisRun secondRun = saveRun(sessionId, "run-" + UUID.randomUUID(), "second question");
DiagnosisRun latest = diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId)
.orElseThrow();
assertEquals(secondRun.getRunId(), latest.getRunId());
assertEquals(firstRun.getRunId(), diagnosisRunRepository.findByRunId(firstRun.getRunId()).orElseThrow().getRunId());
assertTrue(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, secondRun.getRunId()).isPresent());
assertEquals(2, diagnosisRunRepository.findBySessionIdOrderByCreatedAtDescIdDesc(sessionId).size());
}
@Test
void stepAndToolCanBeQueriedByRunId() {
String sessionId = "test-session-" + UUID.randomUUID();
String runId = "run-" + UUID.randomUUID();
saveRun(sessionId, runId, "run scoped trace");
agentStepRepository.save(AgentStep.builder()
.sessionId(sessionId)
.runId(runId)
.stepIndex(1)
.agentName("executor")
.hasToolCall(true)
.build());
agentStepRepository.save(AgentStep.builder()
.sessionId(sessionId)
.runId(runId)
.stepIndex(0)
.agentName("planner")
.hasToolCall(false)
.build());
toolInvocationRepository.save(ToolInvocation.builder()
.sessionId(sessionId)
.runId(runId)
.toolName("lookup_knowledge")
.inputParams("{\"query\":\"payment timeout\"}")
.success(true)
.build());
List<AgentStep> steps = agentStepRepository.findByRunIdOrderByStepIndex(runId);
List<ToolInvocation> tools = toolInvocationRepository.findByRunIdOrderByIdAsc(runId);
assertEquals(2, steps.size());
assertEquals("planner", steps.get(0).getAgentName());
assertEquals(1, tools.size());
assertEquals(runId, tools.get(0).getRunId());
assertEquals(2, agentStepRepository.countByRunId(runId));
assertEquals(1, toolInvocationRepository.countByRunId(runId));
}
private DiagnosisRun saveRun(String sessionId, String runId, String query) {
chatSessionRepository.findBySessionId(sessionId)
.orElseGet(() -> chatSessionRepository.save(ChatSession.builder()
.sessionId(sessionId)
.status("ACTIVE")
.messagePairCount(0)
.build()));
return diagnosisRunRepository.save(DiagnosisRun.builder()
.sessionId(sessionId)
.runId(runId)
.query(query)
.status("SUCCESS")
.agentFlow("CHAT")
.answer("answer for " + query)
.build());
}
}