Files
SuperBizAgent-java/mvp/architecture/session-trace-lifecycle.md
T

211 lines
6.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 会话与 Trace 生命周期
**更新日期**:2026-07-08
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
## 1. 定位
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
```text
diagnosis_session
-> agent_step
-> tool_invocation
```
因此本文描述的是当前可运行链路:
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
- `agent_step` 保存每个 Agent 模型调用。
- `tool_invocation` 保存工具调用事实。
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
## 2. 生命周期总图
```mermaid
flowchart TD
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
Resolve --> Create["create or reset diagnosis_session"]
Create --> Running["status = RUNNING"]
Running --> Agent["Agent workflow"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Tool["Evidence tools"]
Tool --> Invocation["tool_invocation"]
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
Agent --> Final{"workflow result"}
Final -->|success| Success["status = SUCCESS, answer saved"]
Final -->|failed| Failed["status = FAILED"]
Success --> Evaluation["self_evaluation merge"]
Failed --> Evaluation
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
Success --> Feedback["POST /api/feedback"]
Feedback --> Case["useful -> case_library"]
```
## 3. sessionId 规则
| 链路 | sessionId 来源 |
|---|---|
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
| Trace | URL path 中的 `{sessionId}` |
| Feedback | request body 中的 `sessionId` |
设计含义:
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
## 4. 状态流转
```mermaid
stateDiagram-v2
[*] --> PENDING
PENDING --> RUNNING: start diagnosis
RUNNING --> SUCCESS: workflow completed
RUNNING --> FAILED: exception / empty state
SUCCESS --> SUCCESS: feedback submitted
FAILED --> FAILED: feedback submitted
```
字段边界:
| 字段 | 含义 |
|---|---|
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
| `answer` | Agent 最终返回给用户的报告或答复 |
| `self_evaluation` | 系统自评估 JSON |
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
## 5. agent_step 写入
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
```mermaid
sequenceDiagram
autonumber
participant Agent as ReactAgent
participant Hook as AgentLoggingHook
participant DB as agent_step
Agent->>Hook: before_model(messages, sessionId)
Hook->>DB: insert step_index / agent_name / model_input
Agent-->>Agent: model call
Agent->>Hook: after_model(messages, sessionId)
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
```
当前记录:
- `session_id`
- `step_index`
- `agent_name`
- `model_input`
- `model_output`
- `thought`
- `has_tool_call`
- `duration_ms`
- `token_count`
## 6. tool_invocation 写入
工具调用记录真实工具事实,不记录模型猜测。
关键字段:
```text
session_id
step_id
tool_name
input_params
output_preview
output_length
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
-> evidence_refs
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对日志、指标和知识库工具,`retrieval_details.evidence_refs` 会记录 Gatekeeper 可核验的最小证据引用:
```json
{
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "最小证据文本"
}
]
}
```
当工具明确没有返回匹配证据时,可以记录 `raw_path=$.no_evidence`。该路径只表示“本次工具查询未检索到匹配证据”,不表示问题被排除。
## 7. Trace API 聚合
```text
GET /api/diagnosis/{sessionId}/trace
```
聚合逻辑:
```text
diagnosis_session by sessionId
+ agent_step ordered by step_index
+ tool_invocation ordered by id
-> DiagnosisTraceResponse
```
Trace 视图回答的问题:
- 这次诊断是否成功?
- 哪些 Agent 参与了?
- 每一步模型输入输出是什么摘要?
- 调用了哪些工具?
- 工具返回了什么证据?
- Gatekeeper / Verifier / Composer / AIOps rule 是否通过?
- 用户是否反馈有用?
## 8. Chat 与 AIOps 差异
| 维度 | Chat | AIOps |
|---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 |
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
## 9. 清理与边界
当前会话持久化边界:
- MySQL Trace 记录是主要可回放来源。
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
- Redis 主会话存储是历史设计,不作为当前架构事实。
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
## 10. 后续增强
可考虑:
1. Trace API 增加更结构化的 `self_evaluation` 展示。
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
4. 为 Trace 增加导出能力,服务面试演示和回归分析。