Compare commits
9
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
30d3296043 | ||
|
|
3578709896 | ||
|
|
f9df94377b | ||
|
|
78c1477198 | ||
|
|
d928a1968a | ||
|
|
027aed1eeb | ||
|
|
26d5529280 | ||
|
|
6fdbd34bab | ||
|
|
52bf0302c6 |
@@ -72,9 +72,10 @@
|
||||
|
||||
### SessionContext
|
||||
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
|
||||
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
|
||||
- 使用场景:多轮对话上下文管理、工具调用历史追踪
|
||||
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
|
||||
|
||||
### ToolCall
|
||||
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
|
||||
@@ -87,6 +88,21 @@
|
||||
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
|
||||
- 使用场景:分布式会话管理、Agent 状态维护
|
||||
|
||||
### Chat Session
|
||||
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
|
||||
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
|
||||
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
|
||||
|
||||
### Diagnosis Run
|
||||
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
|
||||
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
|
||||
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
|
||||
|
||||
### Diagnosis Trace
|
||||
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
|
||||
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
||||
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
||||
|
||||
### Flyway
|
||||
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
||||
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
||||
|
||||
+31
-30
@@ -2,33 +2,34 @@
|
||||
|
||||
## 项目
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|
|
||||
| 2026-07-09 | interview-demo-quality-audit | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
# Acceptance
|
||||
|
||||
## 实现结果
|
||||
|
||||
- OpenSpec tasks: `42/42` complete。
|
||||
- Phase commits:
|
||||
- `52bf030 feat(trace): add session run isolation schema`
|
||||
- `6fdbd34 docs(openspec): tighten run isolation contract`
|
||||
- `26d5529 feat(trace): isolate chat runs`
|
||||
- `027aed1 feat(trace): add run-scoped trace reads`
|
||||
- `d928a19 feat(trace): bind feedback to runs`
|
||||
- `78c1477 feat(trace): isolate aiops runs`
|
||||
- `f9df943 feat(trace): finish run-aware demo verification`
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
node --check src\main\resources\static\app.js
|
||||
node --check src\main\resources\static\trace.js
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check -- . ':!devflow/index.md'
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
PowerShell demo 脚本解析:
|
||||
|
||||
```powershell
|
||||
$scripts = @(
|
||||
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||
)
|
||||
foreach ($script in $scripts) {
|
||||
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||
}
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Focused tests:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Baseline / regression:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:通过,无 baseline drift。
|
||||
|
||||
## E2E 验证
|
||||
|
||||
使用 Maven 启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
E2E 使用同一 `sessionId` 连续两轮 Chat:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
验证结果:
|
||||
|
||||
- Chat1 / Chat2 均成功。
|
||||
- run1 exact trace 返回 run1。
|
||||
- run2 exact trace 返回 run2。
|
||||
- session-only latest trace 返回 run2。
|
||||
- DB 中同一 session 有两条 `diagnosis_run`。
|
||||
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
|
||||
- `chat_session.message_pair_count = 2`。
|
||||
|
||||
## 日志验证
|
||||
|
||||
检查:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
|
||||
|
||||
## 剩余风险 / 后续事项
|
||||
|
||||
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
|
||||
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
|
||||
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
|
||||
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
|
||||
@@ -0,0 +1,41 @@
|
||||
# Session / Run / Trace Isolation
|
||||
|
||||
## 背景
|
||||
|
||||
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
|
||||
|
||||
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
|
||||
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
|
||||
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
|
||||
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
|
||||
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
|
||||
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
|
||||
|
||||
## 范围
|
||||
|
||||
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
|
||||
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
|
||||
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
|
||||
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
|
||||
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
|
||||
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
- 不实现完整 run-list UI。
|
||||
- 不删除旧 `diagnosis_session`。
|
||||
- 不改变 Redis 对话历史窗口策略。
|
||||
- 不把完整多轮正文历史持久化到 MySQL。
|
||||
- 不尝试把历史混合 trace 还原成真实多轮边界。
|
||||
|
||||
## 关联
|
||||
|
||||
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
- 分档: complex
|
||||
@@ -0,0 +1,48 @@
|
||||
# Decisions
|
||||
|
||||
## 核心决策
|
||||
|
||||
| 决策 | 选择 | 理由 |
|
||||
|---|---|---|
|
||||
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
|
||||
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
|
||||
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
|
||||
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
|
||||
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
|
||||
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
|
||||
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
|
||||
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
|
||||
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
|
||||
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
|
||||
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
|
||||
- `runId` 是正式 API 字段。
|
||||
- Trace 缺少 `runId` 时短期默认查 latest run。
|
||||
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
|
||||
- 每次有效 Chat/AIOps 都创建 run。
|
||||
- 不新增 `diagnosis_trace` / `trace_event` 主表。
|
||||
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
|
||||
- demo 脚本和 Trace UI 做最小 `runId` 支持。
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4。
|
||||
|
||||
- 新 API 响应字段:`runId`。
|
||||
- Trace API 新 query 参数:`runId`。
|
||||
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback request 新增 optional/preferred `runId`。
|
||||
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
|
||||
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
|
||||
- DB contract 新增两张表和两个 `run_id` 列。
|
||||
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
|
||||
|
||||
## 风险接受
|
||||
|
||||
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
|
||||
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
|
||||
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
|
||||
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
|
||||
@@ -0,0 +1,76 @@
|
||||
# Evidence
|
||||
|
||||
## 上下文证据
|
||||
|
||||
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
|
||||
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
|
||||
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
|
||||
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
|
||||
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
|
||||
|
||||
## 实现证据
|
||||
|
||||
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
|
||||
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
|
||||
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
|
||||
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
|
||||
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
|
||||
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
|
||||
|
||||
## E2E 证据
|
||||
|
||||
Maven 启动命令:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
日志:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
E2E session:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
结果:
|
||||
|
||||
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
|
||||
- 两轮返回不同 `runId`。
|
||||
- run1 exact trace 只返回 run1。
|
||||
- run2 exact trace 只返回 run2。
|
||||
- session-only Trace latest fallback 返回 run2。
|
||||
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
|
||||
|
||||
## DB 证据
|
||||
|
||||
通过 `scripts/query_mysql.py` 检查:
|
||||
|
||||
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
|
||||
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
|
||||
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
|
||||
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
|
||||
|
||||
## Baseline 证据
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
- 两组 baseline / regression 命令通过。
|
||||
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
|
||||
- 未观察到 baseline drift。
|
||||
|
||||
## 工具限制
|
||||
|
||||
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
|
||||
+7
-4
@@ -1,6 +1,6 @@
|
||||
# SuperBizAgent MVP 文档
|
||||
|
||||
**更新日期**:2026-07-09
|
||||
**更新日期**:2026-07-10
|
||||
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||
|
||||
@@ -30,7 +30,7 @@
|
||||
|
||||
## 当前系统一句话
|
||||
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
|
||||
## 文档结构
|
||||
|
||||
@@ -83,7 +83,8 @@ mvp/
|
||||
- `VectorSearchService` 是检索稳定门面。
|
||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
|
||||
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
||||
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||
|
||||
## 关键运行链路
|
||||
@@ -93,7 +94,8 @@ Chat
|
||||
-> ChatService
|
||||
-> Planner / Executor / Verifier
|
||||
-> evidence tools
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
-> chat_session / diagnosis_run
|
||||
-> agent_step.run_id / tool_invocation.run_id
|
||||
-> DiagnosisTraceService
|
||||
|
||||
AIOps
|
||||
@@ -102,6 +104,7 @@ AIOps
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / lookup_knowledge
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> diagnosis_run(agent_flow=AI_OPS)
|
||||
-> DiagnosisTraceService
|
||||
|
||||
RAG
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-10
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -29,7 +29,7 @@
|
||||
|
||||
## 当前架构一句话
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
|
||||
@@ -42,13 +42,15 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
|
||||
ChatService --> Session
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
@@ -57,7 +59,8 @@ flowchart TB
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> Session
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
@@ -103,7 +106,7 @@ sequenceDiagram
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant M as chat_composer
|
||||
participant S as diagnosis_session
|
||||
participant R as diagnosis_run
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
@@ -115,13 +118,13 @@ sequenceDiagram
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>S: 写入 verifier_evaluation
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>S: 保存 Composer 最终 answer
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
```
|
||||
|
||||
|
||||
@@ -65,7 +65,8 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
@@ -130,9 +131,10 @@ RAG Retrieval
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
@@ -163,21 +165,22 @@ sequenceDiagram
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 diagnosis_session.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace
|
||||
Trace->>DB: 聚合 session / step / tool
|
||||
Chat->>DB: 保存 diagnosis_run.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
|
||||
@@ -194,13 +197,14 @@ POST /api/chat
|
||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 diagnosis_session
|
||||
-> 保存 agent_step
|
||||
-> 保存 tool_invocation
|
||||
-> 合并 self_evaluation.verifier_evaluation
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
|
||||
@@ -255,7 +259,7 @@ POST /api/ai_ops
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 self_evaluation.aiops_rule_evaluation
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
```
|
||||
|
||||
@@ -317,23 +321,30 @@ RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以三张表为核心:
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> 一次诊断会话的主记录
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
@@ -343,7 +354,8 @@ tool_invocation
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
@@ -353,11 +365,12 @@ tool_invocation
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话状态和最终报告。
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
|
||||
+70
-194
@@ -1,31 +1,44 @@
|
||||
# 数据模型总览
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
|
||||
|
||||
核心数据分三组:
|
||||
|
||||
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
|
||||
- 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
|
||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||
- 反馈沉淀:`case_library`
|
||||
|
||||
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
|
||||
|
||||
## 2. 总体关系
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
diagnosis_session ||--o{ agent_step : has
|
||||
diagnosis_session ||--o{ tool_invocation : has
|
||||
diagnosis_session ||--o| case_library : creates_when_useful
|
||||
chat_session ||--o{ diagnosis_run : owns
|
||||
diagnosis_run ||--o{ agent_step : has
|
||||
diagnosis_run ||--o{ tool_invocation : has
|
||||
diagnosis_run ||--o| case_library : creates_when_useful
|
||||
api_document ||--o{ milvus_chunk : indexed_as
|
||||
knowledge_domain ||--o{ api_document : groups
|
||||
|
||||
diagnosis_session {
|
||||
chat_session {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar status
|
||||
int message_pair_count
|
||||
datetime last_active_at
|
||||
datetime expires_at
|
||||
}
|
||||
|
||||
diagnosis_run {
|
||||
bigint id
|
||||
varchar run_id
|
||||
varchar session_id
|
||||
text query
|
||||
varchar status
|
||||
varchar agent_flow
|
||||
@@ -37,42 +50,23 @@ erDiagram
|
||||
agent_step {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
int step_index
|
||||
varchar agent_name
|
||||
text model_input
|
||||
text model_output
|
||||
text thought
|
||||
boolean has_tool_call
|
||||
}
|
||||
|
||||
tool_invocation {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
bigint step_id
|
||||
varchar tool_name
|
||||
json input_params
|
||||
text output_preview
|
||||
varchar retrieval_layer
|
||||
json retrieval_details
|
||||
varchar relevance_level
|
||||
varchar dedup_reason
|
||||
}
|
||||
|
||||
api_document {
|
||||
bigint id
|
||||
varchar doc_id
|
||||
varchar file_name
|
||||
varchar file_path
|
||||
varchar status
|
||||
int chunk_count
|
||||
text metadata
|
||||
}
|
||||
|
||||
knowledge_domain {
|
||||
bigint id
|
||||
varchar domain_id
|
||||
varchar description
|
||||
text when_to_retrieve
|
||||
int document_count
|
||||
}
|
||||
|
||||
case_library {
|
||||
@@ -84,58 +78,52 @@ erDiagram
|
||||
text root_cause
|
||||
text solution
|
||||
}
|
||||
|
||||
milvus_chunk {
|
||||
varchar id
|
||||
text content
|
||||
json metadata
|
||||
vector vector
|
||||
}
|
||||
```
|
||||
|
||||
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||
|
||||
## 3. 诊断 Trace 模型
|
||||
## 3. 会话与运行模型
|
||||
|
||||
### diagnosis_session
|
||||
### chat_session
|
||||
|
||||
会话级主记录。
|
||||
|
||||
关键字段:
|
||||
`chat_session` 是会话目录表,保存 `sessionId` 的元数据:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
|
||||
| `query` | 用户原始问题或 AIOps 输入摘要 |
|
||||
| `status` | 执行状态 |
|
||||
| `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
|
||||
| `status` | 会话目录状态 |
|
||||
| `message_pair_count` | Redis 对话轮次数快照 |
|
||||
| `last_active_at` | 最近活跃时间 |
|
||||
| `expires_at` | 可为空的目录 TTL 元数据 |
|
||||
|
||||
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
|
||||
### diagnosis_run
|
||||
|
||||
`diagnosis_run` 是一次可回放诊断执行的主记录:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `run_id` | 运行 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | 所属 `chat_session.session_id` |
|
||||
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | 本次执行状态 |
|
||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||
| `answer` | 最终答复或告警报告 |
|
||||
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 用户反馈 |
|
||||
| `answer` | 本次运行最终答复或告警报告 |
|
||||
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 本次运行的用户反馈 |
|
||||
|
||||
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||
|
||||
## 4. Trace 明细模型
|
||||
|
||||
### agent_step
|
||||
|
||||
记录模型调用步骤。
|
||||
|
||||
用途:
|
||||
|
||||
- 回放 Agent 推理过程。
|
||||
- 查看 Planner / Executor / Verifier 的输入输出摘要。
|
||||
- 统计 step count、duration、token count。
|
||||
`agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
### tool_invocation
|
||||
|
||||
记录工具调用事实。
|
||||
|
||||
用途:
|
||||
|
||||
- 给 Trace API 展示证据。
|
||||
- 给 Gatekeeper 提供 `retrieval_details.evidence_refs` 引用验真源。
|
||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
||||
- 给 `EvaluationService` 计算 evidence score。
|
||||
- 给 RAG eval 和人工排查提供检索细节。
|
||||
|
||||
`retrieval_details.evidence_refs` 是当前 Chat 证据链路的关键字段:
|
||||
`tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -149,83 +137,28 @@ erDiagram
|
||||
}
|
||||
```
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `evidence_status` | 工具证据状态,例如 `supported`、`no_evidence`、`deduped`、`failed` |
|
||||
| `evidence_refs[].raw_path` | Executor 可引用的稳定路径,例如 `$.logs[0]`、`$.alerts[0]`、`$.evidence_blocks[0]`、`$.no_evidence` |
|
||||
| `evidence_refs[].text` | 系统抽取的最小证据文本,Gatekeeper 用它核对 `evidence_excerpt` |
|
||||
|
||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||
|
||||
## 4. 知识库模型
|
||||
|
||||
### api_document
|
||||
|
||||
MySQL 中的文档元数据表。
|
||||
|
||||
职责:
|
||||
|
||||
- 管理上传文件。
|
||||
- 保存 file hash,用于去重。
|
||||
- 记录索引状态和 chunk 数量。
|
||||
- 保存 frontmatter JSON。
|
||||
|
||||
### knowledge_domain
|
||||
|
||||
领域级元数据。
|
||||
|
||||
职责:
|
||||
|
||||
- 按 category 聚合文档。
|
||||
- 存储领域描述。
|
||||
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
|
||||
|
||||
### Milvus/Zilliz metadata
|
||||
|
||||
向量 collection 中每个 chunk 的 metadata 主要包括:
|
||||
|
||||
```text
|
||||
docId
|
||||
_source
|
||||
chunkIndex
|
||||
totalChunks
|
||||
title
|
||||
breadcrumb
|
||||
category
|
||||
```
|
||||
|
||||
这些字段支撑:
|
||||
|
||||
- category filter。
|
||||
- source 展示。
|
||||
- breadcrumb 上下文。
|
||||
- docId 删除和重建索引。
|
||||
- evidence block 构造。
|
||||
|
||||
## 5. 反馈沉淀模型
|
||||
|
||||
### case_library
|
||||
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||
|
||||
当前自动映射:
|
||||
|
||||
| 字段 | 来源 |
|
||||
|---|---|
|
||||
| `case_id` | UUID |
|
||||
| `diagnosis_id` | `diagnosis_session.session_id` |
|
||||
| `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
|
||||
| `source_type` | `AUTO` |
|
||||
| `fault_category` | 当前默认 `GENERAL` |
|
||||
| `title` | session query 前 100 字符 |
|
||||
| `root_cause` | session answer |
|
||||
| `solution` | session answer |
|
||||
| `title` | run query 前 100 字符 |
|
||||
| `root_cause` | run answer |
|
||||
| `solution` | run answer |
|
||||
| `created_by` | `system` |
|
||||
|
||||
## 6. self_evaluation 结构
|
||||
|
||||
`diagnosis_session.self_evaluation` 是 JSON 容器:
|
||||
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -235,78 +168,21 @@ category
|
||||
}
|
||||
```
|
||||
|
||||
边界:
|
||||
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||
|
||||
- `rule_evaluation` 评估证据收集充分度。
|
||||
- `verifier_evaluation` 评估 Chat 结构化 claims 是否能由已验真证据推出,并保存 Gatekeeper、Verifier、Composer 的审计数据。
|
||||
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
|
||||
|
||||
当前 `verifier_evaluation` 关键结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
```
|
||||
|
||||
必要审计字段:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` |
|
||||
| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info |
|
||||
| `gatekeeper_result` | 引用真实性校验结果,包括 rule set version、checked bindings、failed rules、warnings、errors |
|
||||
| `composer_output` | Composer 最终表达及解析状态 |
|
||||
| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 |
|
||||
|
||||
## 7. 数据写入时序
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant API as API
|
||||
participant Svc as ChatService/AiOpsService
|
||||
participant Session as diagnosis_session
|
||||
participant Agent as Agent
|
||||
participant Step as agent_step
|
||||
participant Tool as tool_invocation
|
||||
participant Eval as self_evaluation
|
||||
participant Feedback as case_library
|
||||
|
||||
API->>Svc: request
|
||||
Svc->>Session: create/update RUNNING
|
||||
Agent->>Step: before/after model
|
||||
Agent->>Tool: tool call record
|
||||
Svc->>Session: SUCCESS/FAILED + answer
|
||||
Svc->>Eval: merge evaluation
|
||||
API->>Svc: feedback useful
|
||||
Svc->>Feedback: create case
|
||||
```
|
||||
|
||||
## 8. 当前边界和后续
|
||||
## 7. 当前边界和后续
|
||||
|
||||
当前边界:
|
||||
|
||||
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
|
||||
- `tool_invocation.step_id` 可为空。
|
||||
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
|
||||
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
|
||||
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||
- `diagnosis_session` 只作为历史兼容和回滚表保留。
|
||||
|
||||
后续可增强:
|
||||
|
||||
1. 增加 run id,支持同 session 多次独立诊断。
|
||||
2. 强化 `tool_invocation.step_id` 关联。
|
||||
3. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
1. 强化 `tool_invocation.step_id` 关联。
|
||||
2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||
3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
|
||||
@@ -391,7 +391,7 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
|
||||
|
||||
## 7. Trace Persistence
|
||||
|
||||
`diagnosis_session.self_evaluation.verifier_evaluation` 持久化:
|
||||
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||||
|
||||
```json
|
||||
{
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 反馈与自评估架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||
|
||||
@@ -8,8 +8,8 @@
|
||||
|
||||
反馈架构包含两条闭环:
|
||||
|
||||
1. 系统自评估:基于工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
|
||||
1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
|
||||
|
||||
当前重要边界:
|
||||
|
||||
@@ -21,7 +21,7 @@
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
|
||||
Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
|
||||
|
||||
subgraph SelfEval["Self evaluation"]
|
||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||
@@ -40,24 +40,24 @@ flowchart TD
|
||||
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||
VerifierEval --> Merge
|
||||
AiOpsEval --> Merge
|
||||
Merge --> SelfJson["diagnosis_session.self_evaluation"]
|
||||
Merge --> SelfJson["diagnosis_run.self_evaluation"]
|
||||
|
||||
subgraph UserFeedback["User feedback"]
|
||||
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||
API --> FeedbackService["FeedbackService"]
|
||||
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
|
||||
FeedbackService --> FeedbackField["diagnosis_run.feedback"]
|
||||
FeedbackService --> Useful{"feedback == useful?"}
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
|
||||
CaseService --> Case["case_library"]
|
||||
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||
end
|
||||
|
||||
Session --> UI
|
||||
Run --> UI
|
||||
```
|
||||
|
||||
## 3. self_evaluation JSON
|
||||
|
||||
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
|
||||
`SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
|
||||
|
||||
当前结构:
|
||||
|
||||
@@ -149,7 +149,7 @@ flowchart LR
|
||||
Composer --> ComposerOutput["composer_output"]
|
||||
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||
ComposerOutput --> Merge
|
||||
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
|
||||
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
|
||||
```
|
||||
|
||||
Verifier 输出:
|
||||
@@ -198,6 +198,7 @@ POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"runId": "run-xxx",
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
@@ -209,6 +210,8 @@ Content-Type: application/json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"runId": "run-xxx",
|
||||
"fallbackToLatestRun": false,
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
@@ -217,10 +220,16 @@ Content-Type: application/json
|
||||
|
||||
| feedback | 行为 |
|
||||
|---|---|
|
||||
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
|
||||
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
|
||||
| `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
|
||||
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
兼容行为:
|
||||
|
||||
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
|
||||
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
|
||||
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
|
||||
|
||||
## 8. 案例沉淀
|
||||
|
||||
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||
@@ -230,7 +239,7 @@ Content-Type: application/json
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| `caseId` | UUID |
|
||||
| `diagnosisId` | `DiagnosisSession.sessionId` |
|
||||
| `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
|
||||
| `sourceType` | `AUTO` |
|
||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||
| `title` | `query` 前 100 字符 |
|
||||
@@ -241,7 +250,7 @@ Content-Type: application/json
|
||||
幂等性:
|
||||
|
||||
```text
|
||||
case_library.diagnosisId == sessionId
|
||||
case_library.diagnosisId == runId
|
||||
-> existing case: return existing
|
||||
-> missing case: create new
|
||||
```
|
||||
@@ -260,9 +269,9 @@ Trace API 会展示:
|
||||
|
||||
| 视角 | 数据来源 |
|
||||
|---|---|
|
||||
| 执行是否成功 | `diagnosis_session.status` |
|
||||
| 执行是否成功 | `diagnosis_run.status` |
|
||||
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||
| 用户是否认可 | `diagnosis_session.feedback` |
|
||||
| 用户是否认可 | `diagnosis_run.feedback` |
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
|
||||
@@ -36,7 +36,7 @@ flowchart TB
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Session["diagnosis_session"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
@@ -49,7 +49,7 @@ flowchart TB
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Session --> TraceAPI["DiagnosisTraceService"]
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -110,10 +110,10 @@ sequenceDiagram
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId)
|
||||
H->>DB: 写入 model_input / step_index / agent_name
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId)
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
@@ -126,6 +126,8 @@ sequenceDiagram
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
@@ -210,7 +212,7 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.verifier_evaluation
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
@@ -229,7 +231,7 @@ AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 1. 一句话
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
|
||||
|
||||
## 2. 一张图
|
||||
|
||||
@@ -34,14 +34,15 @@ flowchart TB
|
||||
AiOpsFlow --> Trace
|
||||
Tools --> Trace
|
||||
|
||||
Trace --> Session["diagnosis_session"]
|
||||
Trace --> ChatSession["chat_session"]
|
||||
Trace --> Run["diagnosis_run"]
|
||||
Trace --> Step["agent_step"]
|
||||
Trace --> Invocation["tool_invocation"]
|
||||
|
||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||
Verifier --> SelfEval["self_evaluation"]
|
||||
|
||||
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -61,15 +62,15 @@ Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具
|
||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||
|
||||
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
|
||||
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
||||
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
|
||||
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
||||
```
|
||||
|
||||
## 4. 五个亮点
|
||||
|
||||
| 亮点 | 怎么讲 |
|
||||
|---|---|
|
||||
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
|
||||
| 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
|
||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||
|
||||
@@ -1,69 +1,68 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
```
|
||||
|
||||
因此本文描述的是当前可运行链路:
|
||||
|
||||
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
|
||||
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
|
||||
- `agent_step` 保存每个 Agent 模型调用。
|
||||
- `tool_invocation` 保存工具调用事实。
|
||||
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||
|
||||
## 2. 生命周期总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Create["create or reset diagnosis_session"]
|
||||
Create --> Running["status = RUNNING"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["status = FAILED"]
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
Success --> Evaluation["self_evaluation merge"]
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Success --> Feedback["POST /api/feedback"]
|
||||
Feedback --> Case["useful -> case_library"]
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
|
||||
## 3. sessionId 规则
|
||||
## 3. ID 规则
|
||||
|
||||
| 链路 | sessionId 来源 |
|
||||
|---|---|
|
||||
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
|
||||
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
|
||||
| Trace | URL path 中的 `{sessionId}` |
|
||||
| Feedback | request body 中的 `sessionId` |
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
|
||||
设计含义:
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
|
||||
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
|
||||
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||
|
||||
## 4. 状态流转
|
||||
## 4. 运行状态流转
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
@@ -77,14 +76,14 @@ stateDiagram-v2
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 含义 |
|
||||
|---|---|
|
||||
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
|
||||
| `answer` | Agent 最终返回给用户的报告或答复 |
|
||||
| `self_evaluation` | 系统自评估 JSON |
|
||||
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
@@ -93,93 +92,49 @@ stateDiagram-v2
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as ReactAgent
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId)
|
||||
Hook->>DB: insert step_index / agent_name / model_input
|
||||
Agent-->>Agent: model call
|
||||
Agent->>Hook: after_model(messages, sessionId)
|
||||
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
当前记录:
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录真实工具事实,不记录模型猜测。
|
||||
|
||||
关键字段:
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
session_id
|
||||
step_id
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
output_length
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对日志、指标和知识库工具,`retrieval_details.evidence_refs` 会记录 Gatekeeper 可核验的最小证据引用:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "最小证据文本"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
当工具明确没有返回匹配证据时,可以记录 `raw_path=$.no_evidence`。该路径只表示“本次工具查询未检索到匹配证据”,不表示问题被排除。
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_session by sessionId
|
||||
+ agent_step ordered by step_index
|
||||
+ tool_invocation ordered by id
|
||||
diagnosis_run by sessionId + runId
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
Trace 视图回答的问题:
|
||||
|
||||
- 这次诊断是否成功?
|
||||
- 哪些 Agent 参与了?
|
||||
- 每一步模型输入输出是什么摘要?
|
||||
- 调用了哪些工具?
|
||||
- 工具返回了什么证据?
|
||||
- Gatekeeper / Verifier / Composer / AIOps rule 是否通过?
|
||||
- 用户是否反馈有用?
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
@@ -189,22 +144,17 @@ Trace 视图回答的问题:
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
当前会话持久化边界:
|
||||
|
||||
- MySQL Trace 记录是主要可回放来源。
|
||||
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
|
||||
- Redis 主会话存储是历史设计,不作为当前架构事实。
|
||||
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
可考虑:
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
|
||||
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
|
||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||
|
||||
+25
-8
@@ -67,10 +67,23 @@ Invoke-RestMethod `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
@@ -78,13 +91,15 @@ Invoke-RestMethod `
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
@@ -96,6 +111,7 @@ Invoke-RestMethod `
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
@@ -109,7 +125,8 @@ Invoke-RestMethod `
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- 后续 Trace 中 `data.session.feedback = useful`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
@@ -135,19 +152,19 @@ Invoke-WebRequest `
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace:
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
@@ -155,7 +172,7 @@ Invoke-RestMethod `
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> 证据工具
|
||||
@@ -168,7 +185,7 @@ Chat 主线:
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## 1. 目标
|
||||
|
||||
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
|
||||
## 2. 输入
|
||||
|
||||
@@ -25,14 +25,14 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
|
||||
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
|
||||
2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
|
||||
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
|
||||
## 4. 已知边界
|
||||
|
||||
|
||||
@@ -52,12 +52,12 @@ Invoke-RestMethod `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
然后查询同一 session:
|
||||
然后保留响应里的 `runId`,查询同一 run 的 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
|
||||
```
|
||||
|
||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## 为什么不用普通 Chatbot?
|
||||
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。每次运行都能通过同一个 `sessionId` 回放。
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
|
||||
|
||||
## 为什么 RAG 要做成显式工具?
|
||||
|
||||
|
||||
@@ -13,7 +13,7 @@
|
||||
关键主张不是“模型回答了一次”,而是:
|
||||
|
||||
```text
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
|
||||
```
|
||||
|
||||
## 2. Demo 流程
|
||||
@@ -23,7 +23,7 @@
|
||||
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||
5. 指出证据工具和 verifier evaluation。
|
||||
6. 提交 feedback,并展示它挂在同一个 session 上。
|
||||
6. 提交 feedback,并展示它挂在当前 run 上。
|
||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||
|
||||
## 3. 命令
|
||||
@@ -59,7 +59,7 @@ mvp/demo/output/chat-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
|
||||
```
|
||||
|
||||
### 4.2 证据 Trace
|
||||
@@ -112,7 +112,7 @@ mvp/demo/output/feedback-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
feedback 会挂在同一个 diagnosis session 上。
|
||||
feedback 会挂在当前 diagnosis run 上。
|
||||
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||
```
|
||||
|
||||
|
||||
@@ -12,14 +12,16 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. Chat 返回成功答复,且 session id 与请求一致。
|
||||
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
|
||||
2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||
4. 可以使用同一个 session id 提交反馈。
|
||||
5. 后续 Trace 查询能看到已持久化的 feedback 值。
|
||||
4. 可以使用同一个 session id 和本次 run id 提交反馈。
|
||||
5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
|
||||
|
||||
## 4. 需要检查的 Trace 字段
|
||||
|
||||
- `data.runId`
|
||||
- `data.run.runId`
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
|
||||
@@ -78,9 +78,14 @@ $chat = Invoke-RestMethod @chatRequest
|
||||
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
|
||||
$traceRequest = @{
|
||||
Method = "Get"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
}
|
||||
$trace = Invoke-RestMethod @traceRequest
|
||||
|
||||
@@ -89,6 +94,7 @@ $trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
@@ -136,6 +142,7 @@ $summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
baseUrl = $BaseUrl
|
||||
chatSuccess = $chat.data.success
|
||||
verdict = $verdict
|
||||
|
||||
@@ -26,15 +26,22 @@ $chat = Invoke-RestMethod `
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
|
||||
}
|
||||
Write-Host "RunId: $runId"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
|
||||
@@ -53,7 +53,7 @@ mvp/demo/output/interview-demo-summary.json
|
||||
|
||||
```text
|
||||
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
|
||||
固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
|
||||
```
|
||||
|
||||
## 3. 展示用户答案
|
||||
@@ -68,6 +68,7 @@ mvp/demo/output/chat-response.json
|
||||
|
||||
```text
|
||||
data.sessionId
|
||||
data.runId
|
||||
data.answer
|
||||
```
|
||||
|
||||
@@ -76,7 +77,7 @@ data.answer
|
||||
```text
|
||||
这是用户看到的答案。
|
||||
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||
接下来我用同一个 sessionId 查 trace。
|
||||
接下来我用同一个 sessionId 加 runId 查 trace。
|
||||
```
|
||||
|
||||
## 4. 展示 Trace
|
||||
@@ -182,7 +183,7 @@ caseId
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
用户反馈 useful 会写回同一个 diagnosis_session。
|
||||
用户反馈 useful 会写回当前 diagnosis_run。
|
||||
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||
|
||||
这里 status 和 feedback 是分开的:
|
||||
|
||||
@@ -6,14 +6,15 @@
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
@@ -48,10 +49,10 @@
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same session
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP Issues 索引
|
||||
|
||||
**更新日期**:2026-07-09
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## 目录约定
|
||||
@@ -59,6 +59,7 @@
|
||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 已实施 | [archived/ISS-007-verifier-evidence-summary-fidelity.md](archived/ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||
| ISS-008 | Executor 窄范围查询越界 | 已修复 | [archived/ISS-008-executor-narrow-scope-overreach.md](archived/ISS-008-executor-narrow-scope-overreach.md) |
|
||||
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 已修复 | [archived/ISS-009-negative-observation-no-evidence-reference.md](archived/ISS-009-negative-observation-no-evidence-reference.md) |
|
||||
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 已归档 | [archived/ISS-010-session-run-trace-isolation.md](archived/ISS-010-session-run-trace-isolation.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 已归档 | [archived/diagnosis-eval-baseline-diff.md](archived/diagnosis-eval-baseline-diff.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 已归档 | [archived/expand-diagnosis-eval-fixtures.md](archived/expand-diagnosis-eval-fixtures.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 已归档 | [archived/mvp-demo-interview-runbook.md](archived/mvp-demo-interview-runbook.md) |
|
||||
|
||||
@@ -0,0 +1,586 @@
|
||||
# ISS-010 同 session 多轮诊断 Trace 隔离
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-10
|
||||
**来源**:同一 `sessionId` 多轮 Chat E2E 验证
|
||||
|
||||
**归档日期**:2026-07-10
|
||||
**OpenSpec**:`openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
**实现提交**:`52bf030`、`26d5529`、`027aed1`、`d928a19`、`78c1477`、`f9df943`
|
||||
**归档提交**:`3578709`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
归档结论:当前 MVP 已将“会话态”和“运行态”拆开。`sessionId` 表示多轮会话目录和 Redis 上下文;`runId` 表示一次可回放诊断执行。Trace、Feedback、Evaluation、AIOps 和案例沉淀的新路径都按 `runId` 隔离。
|
||||
|
||||
原问题中 Chat 链路同时存在两类“会话”语义:
|
||||
|
||||
```text
|
||||
Redis SessionContext
|
||||
-> 保存同一 sessionId 的多轮对话历史
|
||||
-> 用于下一轮模型上下文
|
||||
|
||||
MySQL diagnosis_session / agent_step / tool_invocation
|
||||
-> 保存诊断 Trace
|
||||
-> 用于 Trace API、Verifier、Evidence score、Feedback 和评测
|
||||
```
|
||||
|
||||
多轮对话需要继续复用 `sessionId`,否则无法保留上下文。但一次诊断 Trace 应该是可独立回放、可独立评分、可独立反馈的执行单元。
|
||||
|
||||
当前实现只按 `sessionId` 关联 Trace,导致同一个 `sessionId` 下多轮诊断的 step/tool 记录混在一起。
|
||||
|
||||
当前实现已改为:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
```
|
||||
|
||||
旧 `diagnosis_session` 保留为历史兼容和回滚表,新 Chat/AIOps 执行不再写入新的运行态。
|
||||
|
||||
## 归档结果
|
||||
|
||||
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-10-session-run-trace-isolation`。
|
||||
- 主规格已同步到 `openspec/specs/session-run-trace-isolation/spec.md`。
|
||||
- `mvp/architecture/` 和 `mvp/tables/` 已更新为 `chat_session -> diagnosis_run -> agent_step/tool_invocation(run_id)` 模型。
|
||||
- Demo 脚本和 Trace UI 已支持 `sessionId + runId` 精确 Trace 和 Feedback。
|
||||
- Maven E2E、`scripts/query_mysql.py` DB 检查、`logs/` 日志检查和 baseline drift 检查均已通过;未观察到 baseline drift。
|
||||
- `devflow/projects/2026-07-10-session-run-trace-isolation/` 已保存 brief、evidence、decisions、acceptance。
|
||||
|
||||
---
|
||||
|
||||
## E2E 证据
|
||||
|
||||
本次使用 `mvp-demo` profile 通过 Maven 启动服务,并用同一个 `sessionId` 连续请求两轮 `/api/chat`:
|
||||
|
||||
```text
|
||||
sessionId = e2e-multiturn-codex-20260710-1615
|
||||
round 1 = 支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
|
||||
round 2 = 基于上一轮结论,只列出目前最缺的三类证据,以及下一步应该优先查哪个系统。
|
||||
```
|
||||
|
||||
验证结果:
|
||||
|
||||
- 第一轮成功,走多 Agent:`planner -> executor -> verifier -> composer`。
|
||||
- 第二轮成功,日志显示进入请求时 `会话历史消息对数: 1`,说明 Redis 历史上下文被复用。
|
||||
- `/api/chat/session/{sessionId}` 返回 `messagePairCount=2`。
|
||||
- `diagnosis_session` 只有一行,`query` 被第二轮问题覆盖。
|
||||
- `agent_step` 返回 14 行,包含第一轮多 Agent step 和第二轮 `intelligent_assistant` step。
|
||||
- `tool_invocation` 返回 19 行,包含两轮工具调用。
|
||||
- `self_evaluation.verifier_evaluation` 仍保留第一轮 Verifier 结果;第二轮简单问答没有新的 Verifier,但 rule evaluation 会基于同 session 全部工具调用重新计算。
|
||||
|
||||
关键入库形态:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
session_id = e2e-multiturn-codex-20260710-1615
|
||||
query = round 2 question
|
||||
status = SUCCESS
|
||||
step_count = 14
|
||||
tool_call_count = 19
|
||||
|
||||
agent_step
|
||||
round 1: planner, executor..., verifier, composer
|
||||
round 2: intelligent_assistant...
|
||||
|
||||
tool_invocation
|
||||
round 1 tools + round 2 tools all under same session_id
|
||||
```
|
||||
|
||||
本次验证产物保存在:
|
||||
|
||||
- `target/e2e/request-round1.json`
|
||||
- `target/e2e/response-round1.json`
|
||||
- `target/e2e/request-round2.json`
|
||||
- `target/e2e/response-round2.json`
|
||||
- `target/e2e/trace-after-round2.json`
|
||||
|
||||
---
|
||||
|
||||
## 核心问题
|
||||
|
||||
### P0:Trace 不是单次诊断的稳定回放
|
||||
|
||||
`GET /api/diagnosis/{sessionId}/trace` 会聚合同一 `sessionId` 下所有 `agent_step` 和 `tool_invocation`。
|
||||
|
||||
多轮之后,Trace 不再表示某一轮诊断,而是混合历史执行轨迹。
|
||||
|
||||
### P0:Verifier 和评分可能读取跨轮证据
|
||||
|
||||
Verifier、Gatekeeper、`ToolTraceSummaryService` 和 `EvaluationService` 当前主要按 `sessionId` 查询工具调用。
|
||||
|
||||
如果上一轮和当前轮证据混在一起,当前轮可能引用或评分到历史工具结果。
|
||||
|
||||
### P1:反馈语义不清晰
|
||||
|
||||
`feedback` 当前在 `diagnosis_session` 上按 `sessionId` 保存。
|
||||
|
||||
多轮之后,用户反馈的是哪一轮答案不再明确。`useful` 反馈沉淀到 `case_library` 时也可能关联到最新主表答案,而不是用户实际评价的那一轮。
|
||||
|
||||
### P1:`diagnosis_session` 字段被覆盖但子表追加
|
||||
|
||||
主表 `query/answer/status/self_evaluation/step_count/tool_call_count` 表示最新运行或混合统计,子表却保留多轮历史。
|
||||
|
||||
这会让 Trace summary、数据库统计和人工排查产生歧义。
|
||||
|
||||
---
|
||||
|
||||
## 已确认决策
|
||||
|
||||
### D1:`runId` 是正式 API 字段
|
||||
|
||||
`/api/chat` 和 `/api/ai_ops` 的响应或 SSE 消息需要暴露本次执行的 `runId`。
|
||||
|
||||
```text
|
||||
sessionId = 多轮对话上下文 ID
|
||||
runId = 本轮诊断执行 ID
|
||||
```
|
||||
|
||||
新客户端应优先用 `runId` 查询 Trace 和提交 Feedback。旧客户端只传 `sessionId` 时,服务端兼容解析该 session 的最新 run。
|
||||
|
||||
### D2:拆分会话态和运行态
|
||||
|
||||
不再把 session、run、trace 全部塞进 `diagnosis_session` 一张主表。
|
||||
|
||||
新增两张主表:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
-> 多轮对话上下文主表
|
||||
|
||||
diagnosis_run
|
||||
-> 单次诊断执行主表
|
||||
```
|
||||
|
||||
Trace 继续使用现有明细表表达:
|
||||
|
||||
```text
|
||||
agent_step
|
||||
tool_invocation
|
||||
```
|
||||
|
||||
暂不新增单独的 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
|
||||
### D3:`runId` 格式
|
||||
|
||||
使用 `run-` + UUID 全量字符串。
|
||||
|
||||
```text
|
||||
run-550e8400-e29b-41d4-a716-446655440000
|
||||
```
|
||||
|
||||
### D4:Trace API 兼容旧路径
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
-> 查该 session 最新 run
|
||||
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
|
||||
-> 查指定 run
|
||||
```
|
||||
|
||||
指定 `runId` 时必须校验该 run 属于 path 中的 `sessionId`。
|
||||
|
||||
最新 run 建议按 `diagnosis_run.created_at DESC, id DESC` 解析,避免旧 run 因反馈或异步评分更新 `updated_at` 后被误认为最新。
|
||||
|
||||
### D5:Feedback 优先绑定 run
|
||||
|
||||
Feedback request 支持 `runId`。
|
||||
|
||||
- 有 `runId`:绑定指定 run。
|
||||
- 无 `runId`:短期兼容绑定该 `sessionId` 最新 run,并显式标记 fallback。
|
||||
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||
|
||||
兼容语义:历史 `case_library.diagnosis_id` 可能保存 `diagnosis_session.session_id`;本 change 之后自动沉淀的新数据保存 `diagnosis_run.run_id`。查询、幂等和文档需要在过渡期识别两种来源,避免把旧案例误判为无效数据。
|
||||
|
||||
### D6:所有 `/api/chat` 执行请求都创建 run
|
||||
|
||||
只要请求通过参数校验并进入 `ChatService.executeChatWithStrategy`,就创建新的 diagnosis run。
|
||||
|
||||
- 简单问答也创建 run。
|
||||
- 复杂诊断也创建 run。
|
||||
- 空问题等参数校验失败不创建 run。
|
||||
|
||||
### D7:AIOps 同步纳入 run 隔离
|
||||
|
||||
每次 `/api/ai_ops` 执行也创建新的 diagnosis run。AIOps 的 step、tool invocation 和 rule evaluation 都按 `runId` 隔离。
|
||||
|
||||
阶段说明:AIOps 可作为独立实现切片排在 Chat 之后,但必须在本 change 整体完成前落地;Chat-only 的中间状态只能作为过渡验证状态,不能作为生产完成状态归档。
|
||||
|
||||
### D8:`chat_session` 第一阶段只保存会话元数据
|
||||
|
||||
`chat_session` 是会话目录/索引表,不保存完整对话历史正文。
|
||||
|
||||
建议保存:
|
||||
|
||||
```text
|
||||
session_id
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
```
|
||||
|
||||
完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`,用于下一轮 prompt 上下文。
|
||||
|
||||
每轮需要长期审计的用户问题和最终回答保存到 `diagnosis_run.query` / `diagnosis_run.answer`。
|
||||
|
||||
`chat_session.expires_at` 只表示 MySQL 会话目录的过期/清理元数据;Redis TTL 到期后,`SessionContext.messageHistory` 可能不再存在,但已经持久化的 `diagnosis_run`、`agent_step` 和 `tool_invocation` 仍作为审计记录保留。
|
||||
|
||||
如果未来需要长期保存完整聊天历史,再单独设计 `chat_message` 表,不在本阶段引入。
|
||||
|
||||
### D9:旧 `diagnosis_session` 表保留但新代码不再写入
|
||||
|
||||
新增 `diagnosis_run` 后,旧 `diagnosis_session` 不立即删除、不立即改造成 view、不直接重命名。
|
||||
|
||||
迁移策略:
|
||||
|
||||
1. 新增 `chat_session` / `diagnosis_run`。
|
||||
2. 为 `agent_step` / `tool_invocation` 新增 nullable `run_id`。
|
||||
3. 将旧 `diagnosis_session` 数据迁移/复制为 `diagnosis_run` 兼容记录。
|
||||
4. 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||
5. 增加必要索引和查询方法,先保持兼容读取。
|
||||
6. 新代码切换为只写 `chat_session` 和 `diagnosis_run`,并为新 step/tool 写入 `run_id`。
|
||||
7. 验证新旧数据 `run_id` 覆盖情况后,再将新写路径要求 `run_id` 非空,并补充索引/约束。
|
||||
8. 旧 `diagnosis_session` 暂时保留,用于历史核对和回滚窗口。
|
||||
9. 后续确认无依赖后,再单独归档或删除旧表。
|
||||
|
||||
### D10:提供轻量 run 列表 API
|
||||
|
||||
新增轻量查询接口,用于查看一个 Chat Session 下有哪些 Diagnosis Run。
|
||||
|
||||
```text
|
||||
GET /api/chat/session/{sessionId}/runs
|
||||
```
|
||||
|
||||
建议返回字段:
|
||||
|
||||
```text
|
||||
runId
|
||||
sessionId
|
||||
query
|
||||
status
|
||||
agentFlow
|
||||
answerPreview
|
||||
stepCount
|
||||
toolCallCount
|
||||
createdAt
|
||||
updatedAt
|
||||
```
|
||||
|
||||
该接口只读 `diagnosis_run` 主表,不展开 `agent_step` / `tool_invocation` 大字段。
|
||||
|
||||
### D11:Feedback 缺少 `runId` 时短期兼容,长期收紧
|
||||
|
||||
Feedback 新协议优先要求 `runId`。
|
||||
|
||||
短期兼容策略:
|
||||
|
||||
- 有 `runId`:绑定指定 run。
|
||||
- 无 `runId`:绑定该 `sessionId` 最新 run。
|
||||
- 无 `runId` fallback 时,在响应或日志中明确标记 `fallbackToLatestRun=true`,并返回实际绑定的 `runId`。
|
||||
|
||||
长期收紧策略:
|
||||
|
||||
- 当前端、demo 脚本和外部调用方都完成 `runId` 传递后,再评估是否将缺少 `runId` 改为参数错误。
|
||||
|
||||
### D12:同步更新 demo 脚本和 Trace UI 的 `runId` 最小支持
|
||||
|
||||
本 issue 实施范围包含 demo 脚本和 Trace UI 的最小协议适配。
|
||||
|
||||
范围:
|
||||
|
||||
- Demo 脚本读取 `/api/chat` 或 `/api/ai_ops` 返回的 `runId`。
|
||||
- Demo 脚本查询 Trace 时传 `?runId=...`。
|
||||
- Trace UI 支持 URL 参数 `?sessionId=...&runId=...`。
|
||||
- Trace UI 查询时如果有 `runId`,带上 `runId`。
|
||||
- 不在本阶段实现完整 run 列表 UI。
|
||||
|
||||
---
|
||||
|
||||
## 目标语义
|
||||
|
||||
引入明确的 `sessionId` / `runId` 分层:
|
||||
|
||||
```text
|
||||
sessionId = 多轮对话上下文
|
||||
runId = 单次诊断执行 / 单次可回放 Trace
|
||||
```
|
||||
|
||||
目标关系:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> one conversation context
|
||||
-> conversation metadata / TTL / last active state
|
||||
|
||||
Redis SessionContext(sessionId)
|
||||
-> hot messageHistory cache
|
||||
-> supports prompt context window
|
||||
|
||||
diagnosis_run(runId, sessionId)
|
||||
-> one diagnosis run
|
||||
|
||||
agent_step(runId, sessionId)
|
||||
-> steps of one run
|
||||
|
||||
tool_invocation(runId, sessionId)
|
||||
-> tool calls of one run
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 建议方案
|
||||
|
||||
采用“拆分主表 + 复用现有 Trace 明细表”的方案:
|
||||
|
||||
1. 新增 `chat_session`。
|
||||
2. 新增 `diagnosis_run`。
|
||||
3. 逐步迁移当前 `diagnosis_session` 语义到 `diagnosis_run`。
|
||||
4. `agent_step` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||
5. `tool_invocation` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||
6. Trace API 聚合 `diagnosis_run + agent_step + tool_invocation`。
|
||||
|
||||
建议核心字段:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
id
|
||||
session_id unique
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
|
||||
diagnosis_run
|
||||
id
|
||||
run_id unique
|
||||
session_id
|
||||
query
|
||||
status
|
||||
agent_flow
|
||||
answer
|
||||
self_evaluation
|
||||
feedback
|
||||
total_duration_ms
|
||||
total_token_count
|
||||
step_count
|
||||
tool_call_count
|
||||
created_at
|
||||
updated_at
|
||||
|
||||
agent_step(run_id, step_index)
|
||||
tool_invocation(run_id, id)
|
||||
```
|
||||
|
||||
理由:
|
||||
|
||||
- `chat_session` 只表达会话态,避免会话上下文和诊断结果混在一起。
|
||||
- `diagnosis_run` 只表达一次执行,天然隔离每轮 Trace、评分和反馈。
|
||||
- `agent_step` / `tool_invocation` 已足够表达 Trace 明细,暂不需要额外 trace 主表。
|
||||
- 后续如果需要统一时间线,再增加 `trace_event`,不阻塞本次隔离。
|
||||
|
||||
---
|
||||
|
||||
## 分阶段计划
|
||||
|
||||
### Phase 0:协议基线和数据边界
|
||||
|
||||
目标:先把语义定死,避免实现中反复。
|
||||
|
||||
已确认基线:
|
||||
|
||||
1. `/api/chat` 是否返回 `runId`。
|
||||
2. `GET /api/diagnosis/{sessionId}/trace` 默认查最新 run 还是要求显式传 `runId`。
|
||||
3. Feedback 是否优先绑定 `runId`,只有旧请求缺失 `runId` 时才回退最新 run。
|
||||
4. AIOps 是否和 Chat 同步接入 `runId`。
|
||||
5. 简单问答是否也创建 diagnosis run。
|
||||
|
||||
建议默认:
|
||||
|
||||
- `/api/chat` 返回 `sessionId + runId`。
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 兼容查最新 run。
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=...` 查指定 run。
|
||||
- Feedback 优先按 `runId` 绑定。
|
||||
- Chat 简单问答也创建 run。
|
||||
- AIOps 同步接入 run 隔离。
|
||||
|
||||
### Phase 1:Schema 迁移和历史数据兼容
|
||||
|
||||
目标:引入 `chat_session` / `diagnosis_run`,并保留旧数据可查询。
|
||||
|
||||
任务:
|
||||
|
||||
- Flyway 新增 `chat_session`。
|
||||
- Flyway 新增 `diagnosis_run`。
|
||||
- 为旧 `diagnosis_session` 生成兼容 `diagnosis_run` 记录。
|
||||
- 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||
- 增加 `find latest run by sessionId` 查询。
|
||||
- 增加 `find by runId` 查询。
|
||||
- 保留旧 `diagnosis_session` 一段时间,新代码不再写入。
|
||||
|
||||
验收:
|
||||
|
||||
- 旧 session 的 Trace 仍可查。
|
||||
- 新索引存在。
|
||||
- 不改变旧 `/api/chat` 必需字段。
|
||||
- `agent_step` / `tool_invocation` 支持 nullable `run_id` 并完成旧数据回填。
|
||||
- 新增 repository 查询可以按 `sessionId` 找最新 run、按 `runId` 找指定 run。
|
||||
|
||||
### Phase 2:Chat 写入切到 runId
|
||||
|
||||
目标:每轮 `/api/chat` 创建一个新的 run,step/tool 按 run 隔离。
|
||||
|
||||
任务:
|
||||
|
||||
- `ChatService` 每次执行生成新的 `runId`。
|
||||
- `ChatController` 确保 `chat_session` 存在并更新会话态。
|
||||
- `ChatService` 按 `runId` 创建 `diagnosis_run`。
|
||||
- `AgentLoggingHook` 写入 `agent_step.run_id`。
|
||||
- `ToolInvocationRecorder` 写入 `tool_invocation.run_id`。
|
||||
- `SessionContextHolder` 或新的上下文 holder 同时携带 `sessionId + runId`。
|
||||
- `backfillSessionMetrics` 按 `runId` 统计。
|
||||
- `EvaluationService` 按 `runId` 读取工具调用。
|
||||
|
||||
验收:
|
||||
|
||||
- 同一 `sessionId` 连续两轮后,`diagnosis_run` 有两行不同 `run_id`。
|
||||
- `/api/chat` 响应增加正式字段 `runId`。
|
||||
- 两轮 `agent_step` / `tool_invocation` 分别按各自 `run_id` 查询。
|
||||
- Redis `messagePairCount` 仍为 2,证明上下文不被破坏。
|
||||
|
||||
### Phase 3:Trace API 兼容和精确查询
|
||||
|
||||
目标:Trace API 可查最新 run,也可查指定 run,并能列出一个 session 下的 run。
|
||||
|
||||
任务:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 从 `diagnosis_run` 默认解析最新 run。
|
||||
- 增加 `runId` query 参数。
|
||||
- Trace response 增加 `runId`。
|
||||
- 新增 `GET /api/chat/session/{sessionId}/runs`。
|
||||
|
||||
验收:
|
||||
|
||||
- 不传 `runId` 返回最新 run。
|
||||
- 传第一轮 `runId` 只返回第一轮 step/tool。
|
||||
- 传第二轮 `runId` 只返回第二轮 step/tool。
|
||||
- run 列表 API 只返回轻量 run 摘要,不展开 trace 明细。
|
||||
- Demo 脚本和 Trace UI 的 `runId` 最小适配按 OpenSpec tasks 放到 Phase 6,避免 Phase 3 同时混入前端/脚本范围。
|
||||
|
||||
### Phase 4:Feedback 和 CaseLibrary 绑定 run
|
||||
|
||||
目标:反馈明确评价哪一轮诊断。
|
||||
|
||||
任务:
|
||||
|
||||
- Feedback request 支持 `runId`。
|
||||
- 旧请求只有 `sessionId` 时短期绑定最新 run,并显式标记 fallback。
|
||||
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||
- `CaseLibraryService` 以 run 为来源生成 case,并用 `run_id` 做新数据幂等键。
|
||||
- 文档说明 `diagnosis_id` 的过渡语义:旧数据可能是 `session_id`,新数据是 `run_id`。
|
||||
|
||||
验收:
|
||||
|
||||
- 同 session 多轮后,对第一轮提交 feedback 不会覆盖第二轮。
|
||||
- useful 生成 case 时能定位到对应 run 的 query/answer。
|
||||
|
||||
### Phase 5:AIOps 同步 run 隔离
|
||||
|
||||
目标:AIOps 使用同样的 run 语义,避免另一条入口继续混杂。
|
||||
|
||||
任务:
|
||||
|
||||
- `AiOpsService` 生成并返回/透出 `runId`。
|
||||
- AIOps `agent_step` / `tool_invocation` 按 `runId` 隔离。
|
||||
- AIOps rule evaluation 写入当前 run 的 `diagnosis_run.self_evaluation.aiops_rule_evaluation`。
|
||||
- AIOps Trace 查询兼容 `sessionId + runId`。
|
||||
|
||||
验收:
|
||||
|
||||
- 同一 AIOps `sessionId` 重跑不会混合 step/tool。
|
||||
- AIOps rule evaluation 只读取当前 run 工具调用。
|
||||
|
||||
---
|
||||
|
||||
## 暂不做
|
||||
|
||||
1. 暂不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
2. 暂不做完整 run 列表 UI。
|
||||
3. 暂不删除历史 Trace 数据。
|
||||
4. 暂不改变 Redis 多轮上下文窗口策略。
|
||||
5. 暂不立即物理删除旧 `diagnosis_session` 表。
|
||||
|
||||
---
|
||||
|
||||
## 风险
|
||||
|
||||
### 1. 兼容风险
|
||||
|
||||
现有脚本、Trace 页面和反馈接口可能只知道 `sessionId`。
|
||||
|
||||
缓解:保留 `sessionId` 默认查最新 run 的行为。
|
||||
|
||||
### 2. 异步上下文风险
|
||||
|
||||
工具调用和 Agent hook 依赖 ThreadLocal / RunnableConfig 传递上下文。
|
||||
|
||||
缓解:统一上下文对象,明确 `sessionId` 和 `runId` 必须同时传递。
|
||||
|
||||
### 3. 历史数据回填风险
|
||||
|
||||
旧数据没有真实 run 边界,只能按当前 `diagnosis_session` 生成一条兼容 `diagnosis_run`。
|
||||
|
||||
缓解:旧数据视为单 run,不尝试拆分历史混合数据。
|
||||
|
||||
### 4. 评分口径变化风险
|
||||
|
||||
按 `runId` 隔离后,工具调用数和 evidence score 可能下降,但语义更正确。
|
||||
|
||||
缓解:更新 eval fixture 和 baseline,记录这是预期行为变化。
|
||||
|
||||
---
|
||||
|
||||
## 已收敛问题
|
||||
|
||||
本 issue 当前已经收敛以下设计边界:
|
||||
|
||||
- `runId` 是正式 API 字段。
|
||||
- `chat_session` 和 `diagnosis_run` 拆分为两张主表。
|
||||
- Trace 明细继续由 `agent_step` / `tool_invocation` 承载。
|
||||
- `chat_session` 只保存元数据,不保存完整对话历史。
|
||||
- 旧 `diagnosis_session` 保留但新代码不再写入。
|
||||
- 提供轻量 run 列表 API。
|
||||
- Feedback 缺少 `runId` 时短期兼容、长期收紧。
|
||||
- Demo 脚本和 Trace UI 做 `runId` 最小支持。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/architecture/session-trace-lifecycle.md`
|
||||
- `mvp/architecture/data-model.md`
|
||||
- `mvp/tables/聊天会话表-chat_session.md`
|
||||
- `mvp/tables/诊断运行表-diagnosis_run.md`
|
||||
- `mvp/tables/诊断会话表-diagnosis_session.md`
|
||||
- `mvp/tables/Agent步骤表-agent_step.md`
|
||||
- `mvp/tables/工具调用表-tool_invocation.md`
|
||||
- `mvp/tables/案例库表-case_library.md`
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/EvaluationService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/CaseLibraryService.java`
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/util/SessionContextHolder.java`
|
||||
- `src/main/resources/db/migration/V005__create_session_storage.sql`
|
||||
@@ -5,14 +5,15 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_step` 记录一次诊断过程中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。页面展示执行链路时应优先按 `step_index` 排序。
|
||||
`agent_step` 记录一次诊断运行中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。`run_id` 是执行隔离边界;Trace 页面展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 关联 `diagnosis_session.session_id` |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
||||
| `model_input` | TEXT | 否 | 模型输入摘要;`V006` 已从 JSON 改为 TEXT |
|
||||
@@ -27,15 +28,18 @@
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_session_step` | `session_id, step_index` | Trace 页面按会话和步骤顺序查询 |
|
||||
| `idx_session_step` | `session_id, step_index` | 历史兼容和粗粒度排查 |
|
||||
| `idx_agent_step_run_step` | `run_id, step_index` | 按运行筛选步骤并辅助顺序查询 |
|
||||
| `idx_agent_name` | `agent_name` | 按 Agent 类型筛选 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `agent_step.session_id` 逻辑关联 `diagnosis_session.session_id`。
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 前端展示步骤时应按 `step_index` 排序,而不是按 `created_at` 或数据库返回顺序。
|
||||
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
|
||||
- 新 Trace、Verifier 和评测读路径应按 `run_id` 取数,避免同一 `sessionId` 多轮诊断混入。
|
||||
- Verifier 应在 Executor 循环完成后出现;如果 `step_index` 中 Verifier 提前,通常意味着编排或记录顺序有问题。
|
||||
|
||||
+14
-8
@@ -1,6 +1,6 @@
|
||||
# MVP 数据表索引
|
||||
|
||||
**更新日期**:2026-07-09
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前表文档入口
|
||||
|
||||
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
||||
@@ -9,26 +9,32 @@
|
||||
|
||||
| 表 | 用途 | 文档 |
|
||||
|---|---|---|
|
||||
| `diagnosis_session` | 会话级主记录,保存 query、状态、最终答案和自评估 | [诊断会话表-diagnosis_session.md](诊断会话表-diagnosis_session.md) |
|
||||
| `agent_step` | Agent 步骤记录,按 `step_index` 回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||
| `tool_invocation` | 工具调用记录,支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||
| `chat_session` | 会话目录元数据,保存同一个 `sessionId` 的多轮会话状态快照 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
|
||||
| `diagnosis_run` | 运行级主记录,保存一次 Chat/AIOps 诊断的 query、状态、答案、自评估和反馈 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
|
||||
| `agent_step` | Agent 步骤记录,按 `run_id` 隔离回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||
| `tool_invocation` | 工具调用记录,按 `run_id` 支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
||||
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
||||
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
||||
| `diagnosis_session` | 历史兼容和回滚表,新执行写入不再依赖它 | [诊断会话表-diagnosis_session.md](诊断会话表-diagnosis_session.md) |
|
||||
|
||||
## 已归档表
|
||||
|
||||
| 表 | 归档原因 | 文档 |
|
||||
|---|---|---|
|
||||
| `diagnosis_record` | 已由 `V007` 删除,被 `diagnosis_session + agent_step + tool_invocation` 替代 | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
||||
| `diagnosis_record` | 已由 `V007` 删除,历史上被 `diagnosis_session + agent_step + tool_invocation` 替代;当前新模型是 `chat_session + diagnosis_run + trace detail` | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
||||
|
||||
## 核心关系
|
||||
|
||||
```text
|
||||
chat_session.session_id
|
||||
-> diagnosis_run.session_id
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
||||
|
||||
diagnosis_session.session_id
|
||||
-> agent_step.session_id
|
||||
-> tool_invocation.session_id
|
||||
-> case_library.diagnosis_id
|
||||
-> historical compatibility / rollback only
|
||||
|
||||
api_document.doc_id
|
||||
-> vector chunk metadata.docId / doc_id
|
||||
|
||||
@@ -5,14 +5,15 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`tool_invocation` 记录 Agent 显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
|
||||
`tool_invocation` 记录 Agent 在一次诊断运行中显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 关联 `diagnosis_session.session_id` |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
||||
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
||||
| `input_params` | JSON | 是 | 工具入参 |
|
||||
@@ -34,13 +35,15 @@
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_session_id` | `session_id` | 按会话查询工具调用 |
|
||||
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
|
||||
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
||||
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
||||
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `tool_invocation.session_id` 逻辑关联 `diagnosis_session.session_id`。
|
||||
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `tool_invocation.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前不强制。
|
||||
|
||||
## 关键 JSON
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`case_library` 保存高质量诊断案例,用于后续相似案例推荐和知识沉淀。当前自动沉淀路径来自 `useful` 用户反馈:系统把 `diagnosis_session` 中的 query 和 answer 映射为案例内容。
|
||||
`case_library` 保存高质量诊断案例,用于后续相似案例推荐和知识沉淀。当前自动沉淀路径来自 `useful` 用户反馈:新数据把 `diagnosis_run` 中的 query 和 answer 映射为案例内容。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -13,7 +13,7 @@
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `case_id` | VARCHAR(64) | 是 | 案例唯一 ID |
|
||||
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断会话;当前自动生成时存 `diagnosis_session.session_id` |
|
||||
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断来源;新自动生成时存 `diagnosis_run.run_id`,历史数据可能是 `diagnosis_session.session_id` |
|
||||
| `source_type` | VARCHAR(16) | 否 | 来源类型:`AUTO` 或 `MANUAL` |
|
||||
| `fault_category` | VARCHAR(32) | 否 | 故障类别,实体侧使用 `FaultCategory` |
|
||||
| `fault_source` | VARCHAR(128) | 否 | 故障源,例如服务、系统或省份 |
|
||||
@@ -35,16 +35,17 @@
|
||||
| `idx_error_code` | `error_code` | 按错误码精确匹配 |
|
||||
| `idx_fault_source` | `fault_source` | 按故障源筛选 |
|
||||
| `idx_fault_target` | `fault_target(100)` | 按故障目标筛选 |
|
||||
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源会话 |
|
||||
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源运行或历史会话 |
|
||||
| `idx_reference_count` | `reference_count` | 推荐排序 |
|
||||
| `idx_created_at` | `created_at` | 时间排序 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `case_library.diagnosis_id` 当前逻辑关联 `diagnosis_session.session_id`,不是旧的 `diagnosis_record`。
|
||||
- `case_library.diagnosis_id` 是过渡字段:新自动案例逻辑关联 `diagnosis_run.run_id`,历史自动案例可能仍是 `diagnosis_session.session_id`。
|
||||
- 人工录入案例可以不填写 `diagnosis_id`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 旧文档里提到的 `diagnosis_record` 已被 `V007` 删除,不再是当前主模型。
|
||||
- 查询新自动案例时优先按 `run_id` 追溯;遇到旧值时再按历史 `session_id` 解释。
|
||||
- 当前自动沉淀仍比较粗:`root_cause` 和 `solution` 都可能来自完整 answer。后续可从结构化结论中拆分根因、证据和修复建议。
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
# 聊天会话表:chat_session
|
||||
|
||||
**状态**:当前会话目录表
|
||||
**来源**:`V011__add_session_run_isolation.sql`、`ChatSession`
|
||||
|
||||
## 定位
|
||||
|
||||
`chat_session` 保存多轮 Chat 会话的元数据,用于把同一个 `sessionId` 下的多次诊断运行组织在一起。它不保存完整对话历史;正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 会话目录 ID,外部 API 仍通过它定位会话 |
|
||||
| `status` | VARCHAR(16) | 否 | `ACTIVE`、`EXPIRED`、`CLOSED` |
|
||||
| `message_pair_count` | INT | 否 | Redis 会话中问答轮次数的快照 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `last_active_at` | DATETIME | 否 | 最近活跃时间 |
|
||||
| `expires_at` | DATETIME | 否 | 目录元数据,可为空;Redis 消息历史可独立过期 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `session_id` unique | `session_id` | 会话目录唯一约束 |
|
||||
| `idx_chat_session_last_active` | `last_active_at` | 最近会话列表和排查 |
|
||||
| `idx_chat_session_status` | `status` | 按状态筛选 |
|
||||
| `idx_chat_session_expires_at` | `expires_at` | 过期目录排查 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `diagnosis_run.session_id` 逻辑关联 `chat_session.session_id`。
|
||||
- 当前不强制数据库外键,服务层校验 run/session ownership。
|
||||
- 一个 `chat_session` 可以拥有多个 `diagnosis_run`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- `chat_session` 是会话元数据,不是诊断执行记录。
|
||||
- 不要把 query、answer、self_evaluation、feedback 写入该表;这些属于 `diagnosis_run`。
|
||||
@@ -1,18 +1,18 @@
|
||||
# 诊断会话表:diagnosis_session
|
||||
|
||||
**状态**:当前主表
|
||||
**状态**:历史兼容和回滚表
|
||||
**来源**:`V005__create_session_storage.sql`、`V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession`
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_session` 是一次 Chat 或 AIOps 诊断的会话级主记录,负责保存用户问题、执行状态、最终答案、总体统计和自评估结果。
|
||||
`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,新 Chat/AIOps 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 会话唯一 ID,Trace API 和反馈接口使用它 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 旧版会话唯一 ID,也是兼容 run 回填来源 |
|
||||
| `query` | TEXT | 是 | 用户原始问题或 AIOps 输入摘要 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||
@@ -37,11 +37,12 @@
|
||||
|
||||
## 关系
|
||||
|
||||
- `agent_step.session_id` 逻辑关联 `diagnosis_session.session_id`。
|
||||
- `tool_invocation.session_id` 逻辑关联 `diagnosis_session.session_id`。
|
||||
- `case_library.diagnosis_id` 在自动生成案例时保存 `diagnosis_session.session_id`。
|
||||
- 迁移时每条 `diagnosis_session` 会生成一条兼容 `diagnosis_run`。
|
||||
- 历史 `agent_step.run_id` 和 `tool_invocation.run_id` 会尽量从兼容 `diagnosis_run` 回填。
|
||||
- 历史自动案例可能仍使用 `case_library.diagnosis_id = diagnosis_session.session_id`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 当前没有数据库外键,Trace 聚合依赖 `session_id`。
|
||||
- `self_evaluation` 是扩展容器,里面可能包含 `rule_evaluation`、`verifier_evaluation`、`aiops_rule_evaluation`。
|
||||
- 新执行不应再把 query、answer、self_evaluation、feedback、统计计数写入本表。
|
||||
- 新 Trace 聚合优先读取 `diagnosis_run + agent_step.run_id + tool_invocation.run_id`。
|
||||
- 历史 fallback 仅在没有 run-backed 数据时读取本表。
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
# 诊断运行表:diagnosis_run
|
||||
|
||||
**状态**:当前诊断运行主表
|
||||
**来源**:`V011__add_session_run_isolation.sql`、`DiagnosisRun`
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都应优先按 `run_id` 绑定。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
|
||||
| `query` | TEXT | 是 | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||
| `answer` | LONGTEXT | 否 | 本次运行的最终答复或告警报告 |
|
||||
| `self_evaluation` | JSON | 否 | 本次运行的 rule、verifier、aiops 自评估容器 |
|
||||
| `feedback` | VARCHAR(16) | 否 | 本次运行的用户反馈 |
|
||||
| `total_duration_ms` | INT | 否 | 本次运行总耗时 |
|
||||
| `total_token_count` | INT | 否 | 本次运行 Token 消耗 |
|
||||
| `step_count` | INT | 否 | 本次运行的 Agent 步骤数 |
|
||||
| `tool_call_count` | INT | 否 | 本次运行的工具调用数 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `run_id` unique | `run_id` | 运行唯一约束 |
|
||||
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下最新运行解析和运行列表 |
|
||||
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
|
||||
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
|
||||
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||
|
||||
## 关系
|
||||
|
||||
- `diagnosis_run.session_id` 逻辑关联 `chat_session.session_id`。
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- 新的自动案例沉淀使用 `case_library.diagnosis_id = diagnosis_run.run_id`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 未带 `runId` 时只为兼容解析 latest run;新 demo 和新客户端应传 `runId`。
|
||||
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
|
||||
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
|
||||
@@ -0,0 +1 @@
|
||||
archive-ready
|
||||
@@ -0,0 +1,3 @@
|
||||
committed: true
|
||||
change: session-run-trace-isolation
|
||||
validated: openspec validate session-run-trace-isolation --strict
|
||||
@@ -0,0 +1,191 @@
|
||||
# Decisions: session-run-trace-isolation
|
||||
|
||||
## sm-flow State
|
||||
|
||||
- Checkpoint: Apply / Phase 6 ready
|
||||
- Scale: complex
|
||||
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
|
||||
## Entry Summary
|
||||
|
||||
Problem: the same `sessionId` currently represents both multi-turn conversation context and one persisted diagnosis trace. Multi-turn Chat E2E proved that Redis context behaves correctly, but MySQL trace rows from different rounds are mixed under one `session_id`.
|
||||
|
||||
Expected result: split session metadata from per-run execution state, expose `runId` as the official run identifier, and make trace, feedback, evaluation, demo scripts, Trace UI, and AIOps read/write by run.
|
||||
|
||||
Known modules: Flyway/JPA entities/repositories, `ChatService`, `ChatController`, `AiOpsService`, `DiagnosisTraceService`, `EvaluationService`, `FeedbackService`, `CaseLibraryService`, `AgentLoggingHook`, `ToolInvocationRecorder`, `SessionContextHolder`, demo scripts, static Trace UI, MVP docs.
|
||||
|
||||
## Context Collection
|
||||
|
||||
`devflow/index.md`: relevant entries found.
|
||||
|
||||
Relevant historical decisions:
|
||||
|
||||
- `session-storage`: current trace persistence is `diagnosis_session + agent_step + tool_invocation`; `sessionId` propagation uses `RunnableConfig.metadata` with `SessionContextHolder` fallback for tools; AIOps records child agents only.
|
||||
- `confidence-feedback`: `FeedbackService` writes feedback, useful feedback creates `case_library`, and feedback must not change execution status.
|
||||
- `mvp-demo-trace-acceptance`: Trace API is read-only and demo artifacts/scripts are part of acceptance.
|
||||
- `aiops-traceable-diagnosis-entry`: `/api/ai_ops` is an SSE entry point that accepts optional alert payload and exposes `sessionId`.
|
||||
- `data-model.md`: current `case_library.diagnosis_id` maps to `diagnosis_session.session_id`; this must be treated as legacy data after the change.
|
||||
- `session-trace-lifecycle.md`: current docs already list run id as a future enhancement for multi-run sessions.
|
||||
|
||||
OpenSpec inputs that must be carried forward:
|
||||
|
||||
- Current specs mention `diagnosis_session` directly in trace, verifier, evidence, and demo requirements; new specs must either supersede or preserve compatibility for those requirements.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` must remain read-only.
|
||||
- Baseline/eval checks must expect changed counts only when the change is explained by run isolation.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | Question | Mode | Status |
|
||||
|---|---|---|---|
|
||||
| Q1 | Should this be an added field on existing `diagnosis_session`, or split tables? | user-interview | confirmed: split `chat_session` and `diagnosis_run`, keep existing trace detail tables |
|
||||
| Q2 | What is metadata in `chat_session`? | user-interview | confirmed: session directory fields only, not full message body |
|
||||
| Q3 | Is "message body" the full conversation history? | user-interview | confirmed: full conversation history stays in Redis `SessionContext.messageHistory` for now |
|
||||
| Q4 | Should `runId` be an official API field? | user-interview | confirmed: yes |
|
||||
| Q5 | How should trace work without `runId`? | user-interview | confirmed: default to latest run for compatibility |
|
||||
| Q6 | Should Feedback without `runId` fail or fall back? | user-interview | confirmed: short-term fallback to latest run, long-term may tighten |
|
||||
| Q7 | Should simple Chat Q&A create a run? | user-interview | confirmed: yes, every valid `/api/chat` execution creates a run |
|
||||
| Q8 | Should AIOps be included? | user-interview | confirmed: yes, same run isolation semantics |
|
||||
| Q9 | Should there be a separate trace table? | user-interview | confirmed: no, current `agent_step` and `tool_invocation` are enough for this phase |
|
||||
| Q10 | What should happen to old `diagnosis_session`? | user-interview | confirmed: keep it for history/rollback, new code stops writing it after migration |
|
||||
| Q11 | Does the issue need demo/Trace UI support? | user-interview | confirmed: yes, minimal `runId` support |
|
||||
| Q12 | What extra risks were found by document review? | evidence-driven | reported and patched into ISS-010 |
|
||||
|
||||
## Evidence-Driven Findings
|
||||
|
||||
- Code evidence: `SessionContext` contains `messageHistory` and `getMessagePairCount()`, supporting the decision that Redis holds hot conversation history while MySQL stores auditable per-run query/answer.
|
||||
- Code evidence: `CaseLibraryService.createFromSession` currently deduplicates by `session.getSessionId()` and maps answer/query from `DiagnosisSession`; this must change for new run-based data.
|
||||
- Documentation evidence: `mvp/architecture/data-model.md` states `case_library.diagnosis_id = diagnosis_session.session_id`; this becomes transitional legacy semantics.
|
||||
- Documentation evidence: existing trace OpenSpec requires `GET /api/diagnosis/{sessionId}/trace` to be read-only; run resolution must preserve that invariant.
|
||||
- E2E evidence from ISS-010: two Chat rounds with the same `sessionId` resulted in one overwritten `diagnosis_session` row and mixed step/tool rows.
|
||||
|
||||
## Confirmed Decisions
|
||||
|
||||
- `runId` format: `run-` + full UUID.
|
||||
- Latest run ordering: `diagnosis_run.created_at DESC, id DESC`, not `updated_at`.
|
||||
- `chat_session` stores metadata: `session_id`, `status`, `message_pair_count`, `created_at`, `last_active_at`, `expires_at`.
|
||||
- Per-run long-term audit stores `query` and `answer` in `diagnosis_run`.
|
||||
- `agent_step` and `tool_invocation` retain `session_id` and add `run_id`.
|
||||
- Missing feedback `runId` returns `fallbackToLatestRun=true` plus actual bound `runId`.
|
||||
- Historical mixed data is not split into multiple true runs.
|
||||
|
||||
## OpenSpec Backwrite Log
|
||||
|
||||
- Created `proposal.md` with problem, scope, non-goals, context constraints, interface impact, and risks.
|
||||
- ISS-010 patched to clarify phase boundary, migration order, AIOps same-change requirement, feedback fallback response, case-library transitional semantics, and Redis/MySQL TTL boundary.
|
||||
- Glossary patched to include `SessionContext.messageHistory` and its persistence boundary.
|
||||
- Created `design.md`, `specs/session-run-trace-isolation/spec.md`, `specs/mvp-demo-trace-acceptance/spec.md`, and `tasks.md`.
|
||||
- Architecture audit found that `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` also read tool rows by `sessionId`; Phase 2 tasks were updated to cover run-scoped reads.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Check | Result | Evidence |
|
||||
|---|---|---|
|
||||
| issue/context -> proposal | aligned | proposal carries E2E problem, split-table solution, compatibility, AIOps, feedback, demo/UI, and baseline scope |
|
||||
| proposal -> design | aligned | design records data model, API impact, migration plan, rollback, and key decisions |
|
||||
| design -> specs | aligned | specs cover session/run split, Chat, trace, feedback, AIOps, migration, demo/UI, E2E, and baseline behavior |
|
||||
| specs -> tasks | aligned | tasks implement schema, Chat write path, trace reads, feedback/case, AIOps, demo/UI/docs, and verification gates |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Input to output chain:
|
||||
|
||||
```text
|
||||
Chat/AIOps request
|
||||
-> ChatController / AIOps endpoint
|
||||
-> ChatService / AiOpsService
|
||||
-> execution context(sessionId, runId)
|
||||
-> AgentLoggingHook -> agent_step
|
||||
-> ToolInvocationRecorder -> tool_invocation
|
||||
-> verifier/gatekeeper/evaluation summary reads
|
||||
-> diagnosis_run answer/self_evaluation/status/counts
|
||||
-> Trace API / Feedback / CaseLibrary / Demo / Trace UI
|
||||
```
|
||||
|
||||
Audit conclusions:
|
||||
|
||||
- Data ownership is clearer with `chat_session` owning conversation metadata and `diagnosis_run` owning execution state; `agent_step` and `tool_invocation` remain trace details owned by one run.
|
||||
- The highest coupling risk is execution-context propagation because hooks and tools currently use both `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- The main read-path risk is missing one of the session-scoped consumers (`ToolTraceSummaryService`, `ExecutorGatekeeperService`, `EvaluationService`, trace, feedback, case creation).
|
||||
- Migration is additive and rollback-friendly until constraints are tightened; historical mixed data must be treated as compatibility data.
|
||||
- AIOps must complete before final archive because otherwise the system would still have one production entry point with mixed trace semantics.
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- Interface impact level: L4.
|
||||
- Required artifacts: proposal, design, specs, tasks.
|
||||
- Strict OpenSpec validation: passed with `openspec validate session-run-trace-isolation --strict`.
|
||||
- `.committed` marker: created after successful commit gate.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Reference migrations: `V005__create_session_storage.sql`, `V008__add_answer_to_diagnosis_session.sql`, `V010__add_relevance_level_to_tool_invocation.sql`.
|
||||
- Entity style: JPA entities use Lombok `@Data`, `@Builder`, `@NoArgsConstructor`, `@AllArgsConstructor`, `@PrePersist`, and `@PreUpdate` where timestamps need maintenance.
|
||||
- Repository style: Spring Data JPA repository interfaces with derived query methods returning `Optional<T>` or `List<T>`.
|
||||
- Test style: repository tests use `@DataJpaTest`, `@AutoConfigureTestDatabase(replace = NONE)`, Flyway enabled, and `ddl-auto=validate`.
|
||||
|
||||
## Phase 1 Apply Notes
|
||||
|
||||
- Implemented additive migration `V011__add_session_run_isolation.sql`.
|
||||
- Implemented `ChatSession` / `DiagnosisRun` entities and repositories.
|
||||
- Added nullable `runId` fields to `AgentStep` and `ToolInvocation`.
|
||||
- Added run-scoped repository methods for step/tool lookup and counts.
|
||||
- Added `DiagnosisRunRepositoryTest`.
|
||||
- Verification evidence is recorded in `phase-1-evidence.md`.
|
||||
- Historical DB inspection found one orphan `tool_invocation` row without a matching `diagnosis_session`; it remains `run_id = NULL` because no reliable compatibility run can be inferred.
|
||||
|
||||
## Document Review Follow-up
|
||||
|
||||
- Clarified that Chat creates runs for the effective resolved `sessionId`, including requests where the server generates a session id.
|
||||
- Tightened legacy feedback fallback so `fallbackToLatestRun=true` and the actual bound `runId` are response fields, not log-only evidence.
|
||||
- Clarified AIOps SSE compatibility: emit a metadata message containing `sessionId` and `runId` before report content while preserving the existing content stream shape.
|
||||
- Clarified that run/session ownership is enforced by service-layer validation and indexed lookup in this change; database foreign keys are intentionally deferred to preserve compatibility with historical orphan detail rows and rollback.
|
||||
- Clarified that `chat_session.expires_at` is nullable directory metadata / best-effort TTL snapshot, not mandatory persisted conversation history.
|
||||
- Clarified document review findings before continuing Phase 2: OpenSpec task phases are authoritative over the older active issue phase sketch, and AIOps rule evaluation is stored in `diagnosis_run.self_evaluation.aiops_rule_evaluation`, not a separate table.
|
||||
|
||||
## Document Review Follow-up Before Phase 3 Gate
|
||||
|
||||
- Clarified AIOps SSE compatibility before Phase 5: keep SSE event name `message`, emit a JSON `SseMessage` with `type=metadata`, and preserve existing content message shape for report streaming.
|
||||
- Clarified Feedback API before Phase 4: request `runId` is preferred, response always includes the bound `runId` and `fallbackToLatestRun`, and wrong-session run binding uses the existing failed feedback response path.
|
||||
- Clarified run-list API before Phase 3 gate: `GET /api/chat/session/{sessionId}/runs` returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
|
||||
|
||||
## Phase 2 Apply Notes
|
||||
|
||||
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
|
||||
- Implemented unified Chat execution context carrying `sessionId` and `runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- Switched valid Chat writes to create/update `chat_session` metadata and create one `diagnosis_run` per request.
|
||||
- Switched Chat run completion, failure, self-evaluation, metrics, verifier support reads, gatekeeper validation, and evidence scoring to run-scoped data.
|
||||
- Added `/api/chat` response `runId` and focused tests for valid run creation, invalid request no-run behavior, run-scoped trace consumers, and same-session multi-turn run creation.
|
||||
- Phase 2 gate evidence is recorded in `phase-2-evidence.md`.
|
||||
|
||||
## Phase 3 Apply Notes
|
||||
|
||||
- Capability source: `openspec-apply-change` + sm-flow apply protocol. The committed OpenSpec remained the execution source; a document review follow-up tightened DTO/SSE/listing contracts before the Phase 3 gate.
|
||||
- Implemented latest-run trace resolution using `diagnosis_run.created_at DESC, id DESC`.
|
||||
- Implemented exact trace lookup for `sessionId + runId` with run/session ownership validation.
|
||||
- Changed trace details to read `agent_step` and `tool_invocation` by `run_id`, while retaining a legacy `diagnosis_session` fallback for historical compatibility.
|
||||
- Added trace response fields for resolved `runId`, chat session metadata, run metadata, and per-row `runId`.
|
||||
- Added lightweight `GET /api/chat/session/{sessionId}/runs` backed by `diagnosis_run` summaries.
|
||||
- Added focused tests for latest trace, exact first trace, exact second trace, wrong-session rejection, missing session, read-only trace behavior, run listing, and legacy fallback.
|
||||
- Phase 3 E2E used Maven startup with profile `mvp-demo`; evidence is recorded in `phase-3-evidence.md`.
|
||||
|
||||
## Phase 4 Apply Notes
|
||||
|
||||
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were still unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, focused tests, API E2E, DB inspection, and logs.
|
||||
- Added `FeedbackRequest.runId` and `FeedbackResponse.runId/fallbackToLatestRun`.
|
||||
- Changed `FeedbackController` to pass `runId` through to `FeedbackService`.
|
||||
- Changed `FeedbackService` so new feedback prefers exact `sessionId + runId`, validates ownership, falls back to latest run when `runId` is omitted, and writes new feedback to `diagnosis_run.feedback`.
|
||||
- Preserved a legacy `DiagnosisSession` fallback only when no `diagnosis_run` exists, so old data can still receive feedback during the migration window.
|
||||
- Added `CaseLibraryService.createFromRun`, using `diagnosis_run.run_id` as the new automatic `case_library.diagnosis_id`; `createFromSession` remains the legacy session-id path.
|
||||
- Phase 4 evidence is recorded in `phase-4-evidence.md`.
|
||||
- Document review follow-up: clarified that the legacy `DiagnosisSession` feedback path only applies when no `diagnosis_run` exists for the session. It returns no bound `runId` and is not the same as latest-run fallback.
|
||||
|
||||
## Phase 5 Apply Notes
|
||||
|
||||
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools remain unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, dependency method inspection with `javap`, focused tests, E2E, DB inspection, and logs.
|
||||
- Changed `/api/ai_ops` to allocate a `runId` before execution and emit a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` on SSE event name `message`.
|
||||
- Changed `AiOpsService` to create one `diagnosis_run` with `agent_flow=AI_OPS` for each valid execution instead of writing new execution state to `diagnosis_session`.
|
||||
- Changed AIOps execution context propagation to pass `sessionId/runId` through both `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- Changed AIOps final report, metrics, status, and `aiops_rule_evaluation` persistence to update the current `diagnosis_run`.
|
||||
- Preserved `persistFinalReport(sessionId, finalReport, request)` as a historical `DiagnosisSession` compatibility path.
|
||||
- Added focused tests for run-scoped final report evaluation, run-scoped metrics, distinct runs under the same AIOps session, and SSE metadata shape.
|
||||
@@ -0,0 +1,145 @@
|
||||
## Context
|
||||
|
||||
The current MVP persists diagnosis observability through `diagnosis_session`, `agent_step`, and `tool_invocation`. That model works for a single diagnosis per `sessionId`, but it conflates two lifecycles once a caller reuses the same `sessionId` for multi-turn conversation:
|
||||
|
||||
- conversation state: Redis `SessionContext.messageHistory` and session metadata;
|
||||
- execution state: one diagnosis answer, trace, self-evaluation, and feedback target.
|
||||
|
||||
The E2E evidence in ISS-010 showed that Redis correctly preserved multi-turn context while MySQL mixed both rounds under the same `session_id`. This breaks trace replay, verifier/evaluation scoping, feedback targeting, and case-library provenance.
|
||||
|
||||
Constraints from existing work:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` is a read-only MVP/demo contract and must remain compatible.
|
||||
- Current hooks and tools propagate `sessionId` through `RunnableConfig.metadata` and `SessionContextHolder`; `runId` must follow the same execution context boundary.
|
||||
- Feedback currently writes `DiagnosisSession.feedback` and creates `case_library` from `DiagnosisSession.answer`.
|
||||
- AIOps is a first-class traceable entry point and cannot be left permanently on the old mixed-run model.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Split conversation metadata from per-execution diagnosis state.
|
||||
- Introduce `runId` as the official identifier for one replayable diagnosis execution.
|
||||
- Preserve old `sessionId`-only callers by resolving the latest run where possible.
|
||||
- Scope trace, evaluation, feedback, case creation, demo scripts, and Trace UI by run.
|
||||
- Migrate historical data into compatibility runs without deleting the old table.
|
||||
- Include Chat and AIOps in the same release-level change.
|
||||
- Verify behavior with focused tests, E2E when needed, DB inspection, logs, and baseline drift checks.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not persist full chat history in MySQL.
|
||||
- Do not add `diagnosis_trace` or `trace_event`.
|
||||
- Do not implement full run-list UI.
|
||||
- Do not physically delete `diagnosis_session`.
|
||||
- Do not attempt to reconstruct true historical round boundaries when only mixed `session_id` data exists.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Domain split | Add `chat_session` and `diagnosis_run` | Add `run_id` to `diagnosis_session` only | Separate tables keep conversation metadata and execution state from growing into one coupled table. |
|
||||
| Trace detail storage | Reuse `agent_step` and `tool_invocation`, adding `run_id` | Add `diagnosis_trace` / `trace_event` | Existing detail tables already represent trace; isolation needs a run key, not a new event model. |
|
||||
| API identity | `runId = "run-" + UUID` | Reuse short session id or DB id | Full UUID avoids collision and keeps external IDs independent from database internals. |
|
||||
| Latest-run compatibility | `GET /api/diagnosis/{sessionId}/trace` resolves latest by `created_at DESC, id DESC` | Require `runId` immediately | Compatibility keeps existing demo/UI/scripts working while new clients migrate. `updated_at` is avoided because feedback/eval updates can reorder old runs. |
|
||||
| Historical migration | Backfill one compatibility run per existing `diagnosis_session` | Try to split old mixed rows | Old rows do not contain reliable run boundaries. A compatibility run preserves auditability without inventing data. |
|
||||
| Feedback fallback | Missing `runId` binds latest run and returns fallback metadata | Reject missing `runId` immediately | Short-term compatibility is needed for old clients; explicit fallback keeps ambiguity observable. |
|
||||
| Case library provenance | New automatic cases store `diagnosis_id = run_id` | Add a new case-library column now | Existing column name can carry transitional provenance; docs and query logic must recognize old `session_id` and new `run_id`. |
|
||||
| AIOps phase | Implement after Chat but before overall archive | Leave AIOps for a follow-up issue | AIOps is already a trace entry point; leaving it old-model would preserve the same bug on another endpoint. |
|
||||
| Ownership integrity | Validate run/session ownership in application services; do not add database foreign keys in this change | Add foreign keys from `diagnosis_run`, `agent_step`, and `tool_invocation` | Existing historical/orphan compatibility data and rollback needs make additive, application-level validation safer for this release. |
|
||||
|
||||
## Data Model
|
||||
|
||||
```text
|
||||
chat_session
|
||||
id
|
||||
session_id unique
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
|
||||
diagnosis_run
|
||||
id
|
||||
run_id unique
|
||||
session_id
|
||||
query
|
||||
status
|
||||
agent_flow
|
||||
answer
|
||||
self_evaluation
|
||||
feedback
|
||||
total_duration_ms
|
||||
total_token_count
|
||||
step_count
|
||||
tool_call_count
|
||||
created_at
|
||||
updated_at
|
||||
|
||||
agent_step
|
||||
session_id
|
||||
run_id
|
||||
...
|
||||
|
||||
tool_invocation
|
||||
session_id
|
||||
run_id
|
||||
...
|
||||
```
|
||||
|
||||
`chat_session.expires_at` is nullable MySQL directory metadata and may be a best-effort Redis TTL snapshot when known. Redis TTL can expire `SessionContext.messageHistory`; persisted `diagnosis_run`, `agent_step`, and `tool_invocation` remain audit records.
|
||||
|
||||
Ownership between `chat_session`, `diagnosis_run`, `agent_step`, and `tool_invocation` is enforced by service-layer validation and indexed lookup in this change. The migration intentionally does not add database foreign keys so historical orphan trace detail rows and rollback paths remain compatible.
|
||||
|
||||
## API / Interface Impact
|
||||
|
||||
Interface level: L4.
|
||||
|
||||
- Database contract changes: new tables, new columns, backfill, indexes, and later non-null expectations for new writes.
|
||||
- `/api/chat` response adds `runId`.
|
||||
- `/api/ai_ops` SSE emits a compatible metadata message before report content. It keeps the existing SSE event name `message` and sends a JSON `SseMessage` with `type=metadata`; the metadata payload includes `sessionId` and `runId`. Report content continues to stream through the existing `type=content` message shape.
|
||||
- Trace API accepts optional `runId`.
|
||||
- Feedback request accepts preferred `runId`. For run-backed data, feedback response includes the actual bound `runId` and `fallbackToLatestRun`; wrong-session `runId`, missing session, and missing run use the existing failed feedback response path with HTTP 400 from `FeedbackController`. Historical `DiagnosisSession` fallback is retained only when no `diagnosis_run` exists for the session; that legacy path has no bound `runId` and is not treated as latest-run fallback.
|
||||
- New run summary API: `GET /api/chat/session/{sessionId}/runs`, returned through the existing `ApiResponse<List<RunSummary>>` wrapper. A session with metadata but no runs returns an empty list; a missing session returns the existing not-found/error behavior.
|
||||
|
||||
Compatibility:
|
||||
|
||||
- Old `sessionId`-only trace and feedback calls bind to latest run when run-backed data exists.
|
||||
- Historical feedback calls for sessions with no `diagnosis_run` may still bind to retained `diagnosis_session` data during the migration window.
|
||||
- Old `diagnosis_session` is retained for rollback and historical comparison.
|
||||
- New code must not keep writing new execution state into `diagnosis_session` after the write switch.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add `chat_session` and `diagnosis_run`.
|
||||
2. Add nullable `run_id` to `agent_step` and `tool_invocation`.
|
||||
3. Backfill `diagnosis_run` from existing `diagnosis_session`.
|
||||
4. Backfill old `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run.
|
||||
5. Add indexes for `session_id`, `run_id`, latest-run lookup, and trace-detail lookup.
|
||||
6. Deploy repository and read-path compatibility.
|
||||
7. Switch Chat write path to `chat_session + diagnosis_run`.
|
||||
8. Switch Chat evaluation and verifier support reads to run-scoped data as part of the Chat write-path cutover.
|
||||
9. Switch Trace run resolution and run listing.
|
||||
10. Switch Feedback and case-library paths.
|
||||
11. Switch AIOps write path, including `diagnosis_run.self_evaluation.aiops_rule_evaluation`.
|
||||
12. Switch demo scripts, Trace UI, and MVP docs.
|
||||
13. Verify no new rows are missing `run_id`; only then tighten application-level and, if safe, database-level non-null assumptions for new data.
|
||||
|
||||
Rollback:
|
||||
|
||||
- Keep `diagnosis_session` intact during this change.
|
||||
- Migrations are additive until constraints are tightened.
|
||||
- If write-switch rollout fails, rollback code can read the retained old table while migrated compatibility rows remain harmless.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] ThreadLocal/context propagation may miss `runId` in nested agent/tool calls. -> Mitigation: introduce a unified execution context carrying both IDs and test hook/tool recording.
|
||||
- [Risk] Baseline metrics change because cross-round tool rows are no longer counted. -> Mitigation: run baseline diff and document expected drift.
|
||||
- [Risk] Legacy `case_library.diagnosis_id` values are ambiguous. -> Mitigation: document transitional semantics and keep lookup logic aware of old `session_id` values.
|
||||
- [Risk] AIOps SSE clients may not parse a new metadata event. -> Mitigation: add metadata in a compatible stream message and keep final report streaming behavior.
|
||||
- [Risk] Historical mixed trace cannot be truly separated. -> Mitigation: call this out as compatibility data, not reconstructed truth.
|
||||
|
||||
## Open Questions
|
||||
|
||||
None blocking. Long-term tightening of missing feedback `runId` remains a follow-up decision after clients migrate.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Phase 1 Evidence: Schema and Compatibility Foundation
|
||||
|
||||
## Scope Completed
|
||||
|
||||
- Added Flyway migration `V011__add_session_run_isolation.sql`.
|
||||
- Added `chat_session` and `diagnosis_run` tables.
|
||||
- Added nullable `run_id` columns and indexes to `agent_step` and `tool_invocation`.
|
||||
- Backfilled compatibility runs from existing `diagnosis_session` rows.
|
||||
- Backfilled existing step/tool rows where a matching `diagnosis_session.session_id` exists.
|
||||
- Added `ChatSession`, `DiagnosisRun`, `ChatSessionRepository`, and `DiagnosisRunRepository`.
|
||||
- Added run-scoped query/count methods to `AgentStepRepository` and `ToolInvocationRepository`.
|
||||
- Added `DiagnosisRunRepositoryTest`.
|
||||
|
||||
## Verification
|
||||
|
||||
### Maven
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=DiagnosisRunRepositoryTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=DiagnosisSessionRepositoryTest,AgentStepRepositoryTest,ToolInvocationRepositoryTest,DiagnosisRunRepositoryTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
### Database Inspection
|
||||
|
||||
Checked with `scripts/query_mysql.py`.
|
||||
|
||||
```text
|
||||
SELECT installed_rank, version, description, success
|
||||
FROM flyway_schema_history
|
||||
ORDER BY installed_rank DESC
|
||||
LIMIT 5;
|
||||
```
|
||||
|
||||
Relevant result:
|
||||
|
||||
```text
|
||||
version = 011
|
||||
description = add session run isolation
|
||||
success = 1
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT COUNT(*) AS chat_sessions FROM chat_session;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
chat_sessions = 83
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT COUNT(*) AS diagnosis_runs,
|
||||
SUM(CASE WHEN run_id IS NULL THEN 1 ELSE 0 END) AS null_run_ids
|
||||
FROM diagnosis_run;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
diagnosis_runs = 83
|
||||
null_run_ids = 0
|
||||
```
|
||||
|
||||
```text
|
||||
SELECT (SELECT COUNT(*) FROM agent_step WHERE run_id IS NULL) AS agent_step_null_run_id,
|
||||
(SELECT COUNT(*) FROM tool_invocation WHERE run_id IS NULL) AS tool_invocation_null_run_id;
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
```text
|
||||
agent_step_null_run_id = 0
|
||||
tool_invocation_null_run_id = 1
|
||||
```
|
||||
|
||||
The one remaining null `tool_invocation.run_id` is historical orphan data:
|
||||
|
||||
```text
|
||||
id = 3
|
||||
session_id = f9d2290d
|
||||
tool_name = lookup_knowledge
|
||||
created_at = 2026-06-26 14:17:35
|
||||
```
|
||||
|
||||
It has no matching `diagnosis_session`, so V011 cannot safely infer a compatibility run. This matches the OpenSpec wording "when possible" for historical backfill.
|
||||
|
||||
### Logs
|
||||
|
||||
`logs/application.log` shows Hibernate using the new `run_id` fields in `agent_step` and `tool_invocation` inserts/selects during the repository verification.
|
||||
|
||||
## Notes
|
||||
|
||||
- Phase 1 intentionally does not switch Chat, Trace, Feedback, or AIOps write paths.
|
||||
- New columns remain nullable until later phases switch new writes and verify coverage.
|
||||
|
||||
@@ -0,0 +1,129 @@
|
||||
# Phase 2 Evidence: Chat Run Write Path
|
||||
|
||||
Date: 2026-07-10
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 2 switched Chat writes from session-scoped execution state to run-scoped execution state:
|
||||
|
||||
- Chat creates/updates `chat_session` metadata.
|
||||
- Each valid `/api/chat` execution creates one `diagnosis_run`.
|
||||
- Chat execution context carries `sessionId + runId` through `RunnableConfig` and `SessionContextHolder`.
|
||||
- `agent_step.run_id` and `tool_invocation.run_id` are written for Chat runs.
|
||||
- Chat completion/failure/status/answer/self-evaluation/counts are written to `diagnosis_run`.
|
||||
- verifier/gatekeeper/evaluation reads use run-scoped tool rows when `runId` is available.
|
||||
- `/api/chat` response includes official `runId`.
|
||||
|
||||
## Static / Unit Verification
|
||||
|
||||
Commands:
|
||||
|
||||
```powershell
|
||||
mvn -q clean test-compile
|
||||
mvn -q "-Dtest=ChatControllerTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ExecutorGatekeeperServiceTest" test
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- `test-compile` passed.
|
||||
- Focused Phase 2 tests passed.
|
||||
- OpenSpec strict validation passed.
|
||||
|
||||
Focused coverage:
|
||||
|
||||
- valid Chat creates `chat_session` and `diagnosis_run`;
|
||||
- invalid blank Chat request returns before `ChatService`, so no run is created;
|
||||
- same `sessionId` across two Chat turns creates two distinct `runId` values;
|
||||
- `ToolInvocationRecorder` copies `runId` from execution context;
|
||||
- verifier trace summary reads by run;
|
||||
- gatekeeper validates by run;
|
||||
- evaluation writes rule evaluation to `diagnosis_run.self_evaluation`.
|
||||
|
||||
## E2E Verification
|
||||
|
||||
Startup command:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Process stdout/stderr:
|
||||
|
||||
- `target/e2e/phase2-mvn-20260710-185102.out.log`
|
||||
- `target/e2e/phase2-mvn-20260710-185102.err.log`
|
||||
|
||||
Primary E2E session:
|
||||
|
||||
```text
|
||||
sessionId = e2e-phase2-codex-20260710-1856
|
||||
round 1 runId = run-24b6f04c-94a0-43cb-b94f-4f0141f9050d
|
||||
round 2 runId = run-a7a2be77-697f-495a-a779-91afd1d8589c
|
||||
```
|
||||
|
||||
HTTP evidence:
|
||||
|
||||
- `target/e2e/phase2-utf8-request-round1.json`
|
||||
- `target/e2e/phase2-utf8-response-round1.json`
|
||||
- `target/e2e/phase2-utf8-request-round2.json`
|
||||
- `target/e2e/phase2-utf8-response-round2.json`
|
||||
- both Chat responses returned `code=200`, `data.success=true`, the same `sessionId`, and distinct `runId` values.
|
||||
|
||||
Redis/session continuity evidence:
|
||||
|
||||
- `target/e2e/phase2-utf8-chat-session-response.json`
|
||||
- response returned `messagePairCount=2`.
|
||||
- logs show the second request entered with `会话历史消息对数: 1` and completed with `当前消息对数: 2`.
|
||||
|
||||
## Database Inspection
|
||||
|
||||
DB inspection used `scripts/query_mysql.py`.
|
||||
|
||||
Saved query outputs:
|
||||
|
||||
- `target/e2e/phase2-utf8-db-runs.txt`
|
||||
- `target/e2e/phase2-utf8-db-chat-session.txt`
|
||||
- `target/e2e/phase2-utf8-db-agent-steps.txt`
|
||||
- `target/e2e/phase2-utf8-db-tool-invocations.txt`
|
||||
- `target/e2e/phase2-utf8-db-missing-runid.txt`
|
||||
|
||||
Observed rows:
|
||||
|
||||
```text
|
||||
diagnosis_run:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | SUCCESS | CHAT | step_count=8 | tool_call_count=12
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | SUCCESS | CHAT | step_count=2 | tool_call_count=1
|
||||
|
||||
chat_session:
|
||||
e2e-phase2-codex-20260710-1856 | ACTIVE | message_pair_count=2
|
||||
|
||||
agent_step grouped by run_id:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 8
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | 2
|
||||
|
||||
tool_invocation grouped by run_id:
|
||||
run-24b6f04c-94a0-43cb-b94f-4f0141f9050d | 12
|
||||
run-a7a2be77-697f-495a-a779-91afd1d8589c | 1
|
||||
|
||||
missing run_id for this session:
|
||||
agent_step = 0
|
||||
tool_invocation = 0
|
||||
```
|
||||
|
||||
## Log Review
|
||||
|
||||
Saved log excerpts:
|
||||
|
||||
- `target/e2e/phase2-utf8-application-log-excerpt.txt`
|
||||
- `target/e2e/phase2-utf8-mvn-log-excerpt.txt`
|
||||
|
||||
Findings:
|
||||
|
||||
- Chat logs show second turn reused Redis history for the same session.
|
||||
- `EvaluationService` wrote scoring results to both run ids.
|
||||
- No E2E-specific application exception was observed for `e2e-phase2-codex-20260710-1856`.
|
||||
- Earlier `/actuator/health` probes produced expected 500/no-resource noise because the actuator health endpoint is not exposed; the E2E readiness check used `/api/chat` instead.
|
||||
|
||||
## Notes
|
||||
|
||||
`codebase-retrieval` and LSP tools were not available in this environment. Call-chain confirmation used OpenSpec context, `rg`, targeted file reads, compilation, focused tests, E2E, DB inspection, and logs.
|
||||
@@ -0,0 +1,136 @@
|
||||
# Phase 3 Evidence: Trace Read Path and Run Listing
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 3 implements run-scoped trace reads and lightweight run listing:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` resolves the latest run by `diagnosis_run.created_at DESC, id DESC`.
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=...` returns the exact run after validating run/session ownership.
|
||||
- Trace responses include resolved run metadata and run-scoped step/tool rows.
|
||||
- `GET /api/chat/session/{sessionId}/runs` returns lightweight run summaries without expanding trace detail rows.
|
||||
|
||||
## Verification Commands
|
||||
|
||||
- `mvn -q clean "-Dtest=DiagnosisTraceServiceTest" test`
|
||||
- `mvn -q "-Dtest=DiagnosisTraceServiceTest,DiagnosisTraceEvaluatorTest,ChatControllerTest" test`
|
||||
- `openspec validate session-run-trace-isolation --strict`
|
||||
|
||||
After the document review follow-up, OpenSpec strict validation was run again:
|
||||
|
||||
- `openspec validate session-run-trace-isolation --strict`
|
||||
|
||||
Result: passed.
|
||||
|
||||
## E2E Runtime
|
||||
|
||||
Maven startup:
|
||||
|
||||
```text
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Captured logs and artifacts:
|
||||
|
||||
- `target/e2e/phase3-mvn-20260710-191131.out.log`
|
||||
- `target/e2e/phase3-mvn-20260710-191131.err.log`
|
||||
- `target/e2e/phase3-request-round1.json`
|
||||
- `target/e2e/phase3-response-round1.json`
|
||||
- `target/e2e/phase3-request-round2.json`
|
||||
- `target/e2e/phase3-response-round2.json`
|
||||
- `target/e2e/phase3-trace-latest.json`
|
||||
- `target/e2e/phase3-trace-first.json`
|
||||
- `target/e2e/phase3-trace-second.json`
|
||||
- `target/e2e/phase3-runs.json`
|
||||
- `target/e2e/phase3-trace-wrong-session.json`
|
||||
- `target/e2e/phase3-summary.json`
|
||||
|
||||
The E2E Maven process was stopped after evidence collection.
|
||||
|
||||
Note: `logs/application.log` was checked, but it did not contain the Phase 3 E2E session entries and its last write time was earlier than this E2E run. The Phase 3 runtime application logs were captured in the Maven stdout artifact above.
|
||||
|
||||
## E2E Summary
|
||||
|
||||
Session:
|
||||
|
||||
```text
|
||||
sessionId = e2e-phase3-codex-20260710-1912
|
||||
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
|
||||
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
```
|
||||
|
||||
Observed behavior from `phase3-summary.json`:
|
||||
|
||||
```text
|
||||
distinctRunIds = true
|
||||
latestResolvedRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
firstTraceRunId = run-fdccbe21-e050-4f92-a741-a062aef59644
|
||||
firstSteps = 9
|
||||
firstTools = 13
|
||||
secondTraceRunId = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
secondSteps = 2
|
||||
secondTools = 0
|
||||
runListCount = 2
|
||||
runListFirst = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
wrongSessionStatus = 400
|
||||
```
|
||||
|
||||
Log evidence from `phase3-mvn-20260710-191131.out.log`:
|
||||
|
||||
- Round 1 request was received for `e2e-phase3-codex-20260710-1912`.
|
||||
- Redis session was created for the first round and reused by the second round.
|
||||
- Message pair count reached `2` after round 2.
|
||||
- Latest trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
|
||||
- Exact first trace request returned `run-fdccbe21-e050-4f92-a741-a062aef59644`.
|
||||
- Exact second trace request returned `run-a9b883ab-9cab-4eec-accf-38129bd2eb94`.
|
||||
- Wrong-session exact trace returned HTTP 400 with `runId does not belong to sessionId`.
|
||||
|
||||
## Database Inspection
|
||||
|
||||
Queried through `scripts/query_mysql.py`.
|
||||
|
||||
`diagnosis_run` rows:
|
||||
|
||||
```text
|
||||
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | SUCCESS | CHAT | step_count=2 | tool_call_count=0 | created_at=2026-07-10 19:15:40
|
||||
run-fdccbe21-e050-4f92-a741-a062aef59644 | SUCCESS | CHAT | step_count=9 | tool_call_count=13 | created_at=2026-07-10 19:12:33
|
||||
```
|
||||
|
||||
`agent_step` rows grouped by run:
|
||||
|
||||
```text
|
||||
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | step_rows=2 | min_step=0 | max_step=1
|
||||
run-fdccbe21-e050-4f92-a741-a062aef59644 | step_rows=9 | min_step=0 | max_step=4
|
||||
```
|
||||
|
||||
`tool_invocation` rows grouped by run:
|
||||
|
||||
```text
|
||||
run-fdccbe21-e050-4f92-a741-a062aef59644 | tool_rows=13
|
||||
```
|
||||
|
||||
Missing run id checks for this session:
|
||||
|
||||
```text
|
||||
agent_step missing run_id = 0
|
||||
tool_invocation missing run_id = 0
|
||||
```
|
||||
|
||||
`chat_session` metadata:
|
||||
|
||||
```text
|
||||
session_id=e2e-phase3-codex-20260710-1912 | status=ACTIVE | message_pair_count=2
|
||||
```
|
||||
|
||||
## Document Review Follow-up
|
||||
|
||||
Before closing Phase 3, the OpenSpec docs were tightened for upcoming phases:
|
||||
|
||||
- AIOps SSE metadata shape: keep SSE event name `message`, use JSON `SseMessage` with `type=metadata`, preserve existing content message shape.
|
||||
- Feedback DTO contract: request `runId` is preferred; response includes bound `runId` and `fallbackToLatestRun`; wrong-session run binding fails instead of updating either run.
|
||||
- Run-list API: returns `ApiResponse<List<RunSummary>>`, returns an empty list for an existing session with no runs, and uses existing missing-session error behavior when no session/run data exists.
|
||||
|
||||
OpenSpec strict validation passed after these document changes.
|
||||
|
||||
## Conclusion
|
||||
|
||||
Phase 3 satisfies the run-scoped trace read and run-list contract. Same-session multi-turn E2E proves latest-run compatibility, exact-run replay, run-list ordering, run/session ownership rejection, and no missing `run_id` rows for new trace data.
|
||||
@@ -0,0 +1,124 @@
|
||||
# Phase 4 Evidence: Feedback and Case Library Run Binding
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 4 implements run-scoped feedback and run-based automatic case creation:
|
||||
|
||||
- Feedback request accepts preferred `runId`.
|
||||
- Feedback validates run/session ownership.
|
||||
- Missing `runId` falls back to the latest run and returns `fallbackToLatestRun=true` plus the bound `runId`.
|
||||
- New feedback persists to `diagnosis_run.feedback`.
|
||||
- Useful feedback creates or reuses `case_library` from `diagnosis_run.query` and `diagnosis_run.answer`.
|
||||
- New automatic `case_library.diagnosis_id` values store `run_id`; old `session_id` values remain supported through the legacy `DiagnosisSession` path.
|
||||
|
||||
## Focused Tests
|
||||
|
||||
Focused tests:
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
Coverage:
|
||||
|
||||
- exact `sessionId + runId` feedback updates the specified run;
|
||||
- missing `runId` binds to latest run and returns `fallbackToLatestRun=true`;
|
||||
- wrong-session `runId` fails without saving feedback or case data;
|
||||
- useful feedback creates a case from the run;
|
||||
- case creation is idempotent by `diagnosis_id`;
|
||||
- legacy `DiagnosisSession` case creation still stores `diagnosis_id=session_id`;
|
||||
- `FeedbackController` passes `request.runId` to `FeedbackService`.
|
||||
|
||||
## E2E Feedback API Check
|
||||
|
||||
Maven startup:
|
||||
|
||||
```text
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Captured artifacts:
|
||||
|
||||
- `target/e2e/phase4-mvn-20260710-2016.out.log`
|
||||
- `target/e2e/phase4-mvn-20260710-2016.err.log`
|
||||
- `target/e2e/phase4-feedback-exact-run.json`
|
||||
- `target/e2e/phase4-feedback-fallback-latest.json`
|
||||
- `target/e2e/phase4-feedback-wrong-session.json`
|
||||
- `target/e2e/phase4-summary.json`
|
||||
|
||||
The Maven process was stopped after evidence collection.
|
||||
|
||||
Inputs reused the Phase 3 E2E session:
|
||||
|
||||
```text
|
||||
sessionId = e2e-phase3-codex-20260710-1912
|
||||
run1 = run-fdccbe21-e050-4f92-a741-a062aef59644
|
||||
run2 = run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
```
|
||||
|
||||
API responses:
|
||||
|
||||
```text
|
||||
exact run feedback:
|
||||
status=200
|
||||
runId=run-fdccbe21-e050-4f92-a741-a062aef59644
|
||||
fallbackToLatestRun=false
|
||||
caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
|
||||
|
||||
legacy fallback feedback:
|
||||
status=200
|
||||
runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94
|
||||
fallbackToLatestRun=true
|
||||
caseId=null
|
||||
|
||||
wrong-session feedback:
|
||||
status=400
|
||||
success=false
|
||||
message=runId does not belong to sessionId
|
||||
```
|
||||
|
||||
Log evidence from `phase4-mvn-20260710-2016.out.log`:
|
||||
|
||||
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-fdccbe21-e050-4f92-a741-a062aef59644, feedback=useful, fallbackToLatestRun=false, caseId=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff`
|
||||
- `反馈已记录: sessionId=e2e-phase3-codex-20260710-1912, runId=run-a9b883ab-9cab-4eec-accf-38129bd2eb94, feedback=not_useful, fallbackToLatestRun=true, caseId=null`
|
||||
|
||||
## Database Inspection
|
||||
|
||||
Queried through `scripts/query_mysql.py`.
|
||||
|
||||
`diagnosis_run.feedback`:
|
||||
|
||||
```text
|
||||
run-a9b883ab-9cab-4eec-accf-38129bd2eb94 | feedback=not_useful | status=SUCCESS
|
||||
run-fdccbe21-e050-4f92-a741-a062aef59644 | feedback=useful | status=SUCCESS
|
||||
```
|
||||
|
||||
`case_library`:
|
||||
|
||||
```text
|
||||
case_id=428b0cf8-d3b4-4f8a-a82e-4e815999d3ff
|
||||
diagnosis_id=run-fdccbe21-e050-4f92-a741-a062aef59644
|
||||
source_type=AUTO
|
||||
title=支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
|
||||
```
|
||||
|
||||
No automatic case was created for the fallback `not_useful` feedback on run2.
|
||||
|
||||
## Final Gate
|
||||
|
||||
The final Phase 4 gate was rerun after document review follow-up:
|
||||
|
||||
```text
|
||||
mvn -q clean test-compile
|
||||
mvn -q "-Dtest=FeedbackServiceTest,CaseLibraryServiceTest,FeedbackControllerTest,DiagnosisTraceServiceTest,ChatControllerTest" test
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
## Conclusion
|
||||
|
||||
Phase 4 satisfies run-scoped feedback, observable latest-run fallback, run-based useful case creation, and transitional old-data compatibility for `case_library.diagnosis_id`.
|
||||
@@ -0,0 +1,157 @@
|
||||
# Phase 5 Evidence: AIOps Run Isolation
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 5 implements AIOps run isolation:
|
||||
|
||||
- `/api/ai_ops` allocates a `runId` before execution.
|
||||
- The SSE stream keeps event name `message` and emits a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` before content.
|
||||
- `AiOpsService` creates `diagnosis_run` rows with `agent_flow=AI_OPS`.
|
||||
- AIOps Agent hooks and tool recording receive `sessionId + runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- AIOps status, final report, metrics, and `self_evaluation.aiops_rule_evaluation` write to the current `diagnosis_run`.
|
||||
- Historical `DiagnosisSession` final-report persistence remains available only through the legacy overload.
|
||||
|
||||
## Focused Tests
|
||||
|
||||
Focused tests:
|
||||
|
||||
```text
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
Coverage:
|
||||
|
||||
- AIOps final report updates `diagnosis_run.answer` and `diagnosis_run.self_evaluation`.
|
||||
- Rule evaluation reads `tool_invocation` rows by `run_id`.
|
||||
- AIOps run metrics count `agent_step` and `tool_invocation` rows by `run_id`.
|
||||
- The same AIOps `sessionId` can create distinct `runId` values.
|
||||
- SSE metadata uses `type=metadata` and carries `sessionId/runId`.
|
||||
- Existing hook and tool-recorder tests cover `runId` propagation into `agent_step` and `tool_invocation`.
|
||||
|
||||
## E2E Runtime
|
||||
|
||||
Maven startup:
|
||||
|
||||
```text
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Captured artifacts:
|
||||
|
||||
- `target/e2e/phase5-mvn-20260710-204732.out.log`
|
||||
- `target/e2e/phase5-mvn-20260710-204732.err.log`
|
||||
- `target/e2e/phase5-aiops-request.json`
|
||||
- `target/e2e/phase5-aiops-sse-response.txt`
|
||||
- `target/e2e/phase5-mvn-20260710-205204.out.log`
|
||||
- `target/e2e/phase5-mvn-20260710-205204.err.log`
|
||||
- `target/e2e/phase5-aiops-request-2053.json`
|
||||
- `target/e2e/phase5-aiops-sse-response-2053.txt`
|
||||
- `target/e2e/phase5-aiops-trace-2053.json`
|
||||
|
||||
The Maven process was stopped after evidence collection.
|
||||
|
||||
First E2E attempt exposed an existing AIOps runtime integration bug:
|
||||
|
||||
```text
|
||||
AI Ops 流程失败: mainAgent (ReactAgent) must be provided for supervisor agent
|
||||
```
|
||||
|
||||
Diagnosis result:
|
||||
|
||||
- feedback loop: fixed `/api/ai_ops` request with `mvp-demo` profile;
|
||||
- root cause: current `SupervisorAgent` dependency validates that `mainAgent(ReactAgent)` is set;
|
||||
- fix: `AiOpsService.buildSupervisorAgent(...)` now sets Planner as `mainAgent` and Executor as sub-agent;
|
||||
- regression coverage: `AiOpsServiceTest.buildSupervisorAgentSetsPlannerAsMainAgent`.
|
||||
|
||||
Successful E2E:
|
||||
|
||||
```text
|
||||
sessionId = e2e-phase5-aiops-codex-20260710-2053
|
||||
runId = run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
|
||||
```
|
||||
|
||||
SSE response:
|
||||
|
||||
```text
|
||||
event:message
|
||||
data:{"type":"metadata","data":null,"sessionId":"e2e-phase5-aiops-codex-20260710-2053","runId":"run-84b8c02b-4d1e-4b24-ac88-222c88c8db26"}
|
||||
|
||||
...
|
||||
|
||||
event:message
|
||||
data:{"type":"done","data":null,"sessionId":null,"runId":null}
|
||||
```
|
||||
|
||||
Exact trace check:
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/e2e-phase5-aiops-codex-20260710-2053/trace?runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
|
||||
```
|
||||
|
||||
Observed:
|
||||
|
||||
```text
|
||||
code=200
|
||||
runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
|
||||
agentFlow=AI_OPS
|
||||
summary.hasAiOpsRuleEvaluation=true
|
||||
summary.persistedStepCount=1
|
||||
summary.persistedToolCallCount=0
|
||||
```
|
||||
|
||||
The E2E model did not call evidence tools. This is captured as rule-evaluation WARN rather than a run-isolation failure.
|
||||
|
||||
## Database Inspection
|
||||
|
||||
Queried through `scripts/query_mysql.py`.
|
||||
|
||||
`diagnosis_run`:
|
||||
|
||||
```text
|
||||
run_id=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
|
||||
session_id=e2e-phase5-aiops-codex-20260710-2053
|
||||
status=SUCCESS
|
||||
agent_flow=AI_OPS
|
||||
has_answer=1
|
||||
has_aiops_eval=1
|
||||
step_count=1
|
||||
tool_call_count=0
|
||||
```
|
||||
|
||||
`agent_step` grouped by run:
|
||||
|
||||
```text
|
||||
run-84b8c02b-4d1e-4b24-ac88-222c88c8db26 | step_rows=1
|
||||
```
|
||||
|
||||
`tool_invocation` grouped by run:
|
||||
|
||||
```text
|
||||
(empty; the successful E2E did not invoke evidence tools)
|
||||
```
|
||||
|
||||
Log evidence from `phase5-mvn-20260710-205204.out.log`:
|
||||
|
||||
- AIOps request was received with the expected `sessionId` and `runId`.
|
||||
- `AiOpsService` started analysis and invoked the supervisor agent.
|
||||
- AIOps orchestration completed and final report extraction ran.
|
||||
- Exact trace was queried with the same `sessionId + runId`.
|
||||
|
||||
## Final Gate
|
||||
|
||||
Final Phase 5 gate:
|
||||
|
||||
```text
|
||||
mvn -q clean test-compile
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest,DiagnosisTraceServiceTest" test
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
## Conclusion
|
||||
|
||||
Phase 5 satisfies AIOps run creation, SSE run metadata, run-scoped execution context propagation, run-scoped final report/evaluation/metrics writes, focused tests, and Maven E2E DB/log verification.
|
||||
@@ -0,0 +1,253 @@
|
||||
# Phase 6 Evidence: Demo, Trace UI, Documentation, and Verification
|
||||
|
||||
## Scope
|
||||
|
||||
Phase 6 completed the run-aware demo/documentation surface and final verification for `session-run-trace-isolation`.
|
||||
|
||||
Implemented:
|
||||
|
||||
- Demo scripts read Chat `runId`, query exact Trace with `?runId=...`, and submit feedback with `runId`.
|
||||
- Trace UI accepts `?sessionId=...&runId=...` and calls the exact Trace API when `runId` is present.
|
||||
- Chat UI remembers the latest run target and links to Trace Workbench with `sessionId + runId` when available.
|
||||
- MVP table and architecture docs now describe `chat_session`, `diagnosis_run`, `agent_step.run_id`, `tool_invocation.run_id`, and transitional `case_library.diagnosis_id` semantics.
|
||||
|
||||
## Static Verification
|
||||
|
||||
Commands:
|
||||
|
||||
```powershell
|
||||
node --check src\main\resources\static\app.js
|
||||
node --check src\main\resources\static\trace.js
|
||||
|
||||
$scripts = @(
|
||||
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||
)
|
||||
foreach ($script in $scripts) {
|
||||
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||
Write-Host "Parsed $script"
|
||||
}
|
||||
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- JavaScript syntax: passed.
|
||||
- PowerShell script parsing: passed.
|
||||
- OpenSpec strict validation: passed.
|
||||
|
||||
## Focused Tests
|
||||
|
||||
Command:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||
```
|
||||
|
||||
Result: passed.
|
||||
|
||||
## Final Same-Session Multi-Turn E2E
|
||||
|
||||
Startup command:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
Startup log:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||
|
||||
Application readiness:
|
||||
|
||||
- `Started Main in 15.501 seconds`
|
||||
- `ReadinessState changed to ACCEPTING_TRAFFIC`
|
||||
|
||||
E2E session:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
Artifacts:
|
||||
|
||||
- `target/e2e/phase6-chat1-20260710-2120.json`
|
||||
- `target/e2e/phase6-chat2-20260710-2120.json`
|
||||
- `target/e2e/phase6-trace-run1-20260710-2120.json`
|
||||
- `target/e2e/phase6-trace-run2-20260710-2120.json`
|
||||
- `target/e2e/phase6-trace-latest-20260710-2120.json`
|
||||
- `target/e2e/phase6-e2e-summary-20260710-2120.json`
|
||||
|
||||
Observed:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "e2e-phase6-chat-codex-20260710-2120",
|
||||
"run1": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
|
||||
"run2": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
|
||||
"chat1Success": true,
|
||||
"chat2Success": true,
|
||||
"trace1RunId": "run-e2a97696-4398-4abc-90e4-28f45c838f92",
|
||||
"trace2RunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
|
||||
"latestRunId": "run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172",
|
||||
"trace1Steps": 10,
|
||||
"trace2Steps": 9,
|
||||
"trace1Tools": 14,
|
||||
"trace2Tools": 8
|
||||
}
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Two Chat requests reused the same `sessionId`.
|
||||
- Each Chat request returned a distinct `runId`.
|
||||
- Exact trace for run1 returned run1 only.
|
||||
- Exact trace for run2 returned run2 only.
|
||||
- Session-only Trace latest fallback returned run2.
|
||||
|
||||
## Database Inspection
|
||||
|
||||
Tool: `scripts/query_mysql.py`
|
||||
|
||||
`diagnosis_run`:
|
||||
|
||||
```text
|
||||
run_id | session_id | status | agent_flow | step_count | tool_call_count | has_answer
|
||||
-------------------------------------------------------------------------------------------------------------------------------------------------
|
||||
run-e2a97696-4398-4abc-90e4-28f45c838f92 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 10 | 14 | 1
|
||||
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | e2e-phase6-chat-codex-20260710-2120 | SUCCESS | CHAT | 9 | 8 | 1
|
||||
```
|
||||
|
||||
`agent_step` grouped by `run_id`:
|
||||
|
||||
```text
|
||||
run_id | step_rows | min_step | max_step
|
||||
--------------------------------------------------------------------------
|
||||
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 9 | 0 | 5
|
||||
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 10 | 0 | 6
|
||||
```
|
||||
|
||||
`tool_invocation` grouped by `run_id`:
|
||||
|
||||
```text
|
||||
run_id | tool_rows
|
||||
----------------------------------------------------
|
||||
run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172 | 8
|
||||
run-e2a97696-4398-4abc-90e4-28f45c838f92 | 14
|
||||
```
|
||||
|
||||
Mixed row check:
|
||||
|
||||
```text
|
||||
mixed_rows
|
||||
----------
|
||||
0
|
||||
0
|
||||
```
|
||||
|
||||
`chat_session` metadata:
|
||||
|
||||
```text
|
||||
session_id | status | message_pair_count | last_active_at
|
||||
---------------------------------------------------------------------------------------
|
||||
e2e-phase6-chat-codex-20260710-2120 | ACTIVE | 2 | 2026-07-10 21:23:59
|
||||
```
|
||||
|
||||
`GET /api/chat/session/e2e-phase6-chat-codex-20260710-2120` returned `messagePairCount=2`.
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Run table has exactly two successful Chat runs for the E2E session.
|
||||
- Step/tool counts match the exact Trace API responses.
|
||||
- No `agent_step` or `tool_invocation` rows for this session have NULL or unexpected `run_id`.
|
||||
- `chat_session` metadata confirms multi-turn context continuity at two message pairs.
|
||||
|
||||
## Log Inspection
|
||||
|
||||
Searched:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
Patterns:
|
||||
|
||||
- `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
Result:
|
||||
|
||||
- Matching startup, Chat execution, run persistence, and trace lookup log lines were present in Maven output and `logs/application.log`.
|
||||
|
||||
## Baseline Drift
|
||||
|
||||
Focused baseline command:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
```
|
||||
|
||||
Broader baseline regression command from `mvp/eval/README.md`:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- Both commands passed.
|
||||
- The baseline harness evaluates saved fixtures and does not depend on live DB/session tables.
|
||||
- No baseline drift was observed.
|
||||
|
||||
## Final Gate Checks
|
||||
|
||||
Commands:
|
||||
|
||||
```powershell
|
||||
rg -n "diagnosisSessionRepository\.save|new DiagnosisSession|DiagnosisSession\.builder|setAnswer\(|setSelfEvaluation\(|setFeedback\(|createFromSession|persistFinalReport\(sessionId, finalReport|save\(session\)" src\main\java\com\superbiz\agent -g "*.java"
|
||||
|
||||
rg -n "evaluate\(|evaluateRun\(|persistFinalReport\(|submitFeedback\(|createFromSession\(" src\main\java src\test\java -g "*.java"
|
||||
|
||||
git diff --check -- . ':!devflow/index.md'
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
```
|
||||
|
||||
Result:
|
||||
|
||||
- New Chat write path calls `evaluationService.evaluateRun(...)` and writes `diagnosis_run`.
|
||||
- New AIOps controller path calls `persistFinalReport(sessionId, runId, ...)` and writes `diagnosis_run`.
|
||||
- Remaining `diagnosis_session` writes are legacy compatibility paths:
|
||||
- `FeedbackService.submitLegacySessionFeedback(...)`
|
||||
- `AiOpsService.persistLegacyFinalReport(...)`
|
||||
- legacy `EvaluationService.evaluate(...)`
|
||||
- `git diff --check`: passed after Markdown whitespace cleanup.
|
||||
- OpenSpec strict validation: passed.
|
||||
|
||||
## Documentation Review Follow-up
|
||||
|
||||
After the final documentation review, the remaining demo helper docs were aligned with the run-aware contract:
|
||||
|
||||
- `mvp/demo/trace-inspection-checklist.md`
|
||||
- `mvp/demo/payment-timeout-acceptance.md`
|
||||
- `mvp/demo/interview-walkthrough.md`
|
||||
- `mvp/tables/Agent步骤表-agent_step.md`
|
||||
- `mvp/tables/README.md`
|
||||
- `mvp/architecture/data-model.md`
|
||||
- `mvp/architecture/session-trace-lifecycle.md`
|
||||
|
||||
The corrections remove session-only wording for Trace/Feedback and clarify that `run_id` is the execution isolation boundary while Trace API response order is the UI display contract.
|
||||
|
||||
Follow-up gate after these documentation fixes:
|
||||
|
||||
- `node --check src\main\resources\static\app.js`: passed.
|
||||
- `node --check src\main\resources\static\trace.js`: passed.
|
||||
- PowerShell demo script parsing: passed.
|
||||
- `openspec validate session-run-trace-isolation --strict`: passed.
|
||||
- `git diff --check -- . ':!devflow/index.md'`: passed.
|
||||
|
||||
## Notes
|
||||
|
||||
- The AGENTS-required `codebase-retrieval` and LSP tools were not available in this session. Fallback verification used OpenSpec context, `rg`, targeted file reads, focused tests, E2E, DB inspection, and log inspection.
|
||||
@@ -0,0 +1,86 @@
|
||||
# Change: Session / Run / Trace Isolation
|
||||
|
||||
## Problem
|
||||
|
||||
The current MVP uses the same `sessionId` for two different concepts:
|
||||
|
||||
- Redis `SessionContext` keeps multi-turn chat history for prompt context.
|
||||
- MySQL `diagnosis_session`, `agent_step`, and `tool_invocation` persist diagnosis trace data for replay, verification, feedback, and evaluation.
|
||||
|
||||
End-to-end verification showed that two `/api/chat` calls with the same `sessionId` correctly reuse Redis context, but MySQL trace data is mixed under the same key:
|
||||
|
||||
- `diagnosis_session.query` is overwritten by the second round.
|
||||
- `agent_step` and `tool_invocation` append rows from both rounds under the same `session_id`.
|
||||
- Trace, verifier/evaluation, and feedback can read cross-round evidence.
|
||||
|
||||
This makes a trace no longer represent one replayable diagnosis run.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Introduce a stable split between conversation state and execution state:
|
||||
|
||||
- `chat_session`: conversation metadata keyed by `session_id`.
|
||||
- `diagnosis_run`: one execution/run keyed by `run_id`, belonging to a `session_id`.
|
||||
- `agent_step` and `tool_invocation`: keep existing trace detail role, add `run_id` while retaining `session_id` for compatibility and coarse filtering.
|
||||
|
||||
`runId` becomes an official API field:
|
||||
|
||||
- `/api/chat` returns `sessionId + runId`.
|
||||
- `/api/ai_ops` emits an SSE-compatible metadata message containing `sessionId` and `runId` before report content.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` defaults to the latest run for compatibility.
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=run-...` returns the specified run after validating it belongs to the path `sessionId`.
|
||||
- Feedback prefers `runId`; missing `runId` temporarily falls back to the latest run and returns both `fallbackToLatestRun=true` and the actual bound `runId`.
|
||||
|
||||
Trace remains an aggregate view of `diagnosis_run + agent_step + tool_invocation`; this change does not introduce a separate `diagnosis_trace` or `trace_event` table.
|
||||
|
||||
## Scope
|
||||
|
||||
- Add Flyway migrations and JPA entities/repositories for `chat_session` and `diagnosis_run`.
|
||||
- Add nullable `run_id` to `agent_step` and `tool_invocation`, backfill historical data, then switch new writes to require run context.
|
||||
- Move new Chat writes from `diagnosis_session` to `chat_session + diagnosis_run`.
|
||||
- Update trace reads to resolve latest run or specified run.
|
||||
- Add lightweight run list API: `GET /api/chat/session/{sessionId}/runs`.
|
||||
- Update feedback and case-library creation to bind new data to `run_id`.
|
||||
- Update AIOps to create and expose `runId` before this change is considered production complete.
|
||||
- Update demo scripts and Trace UI with minimal `runId` support.
|
||||
- Update relevant MVP table and architecture documentation.
|
||||
- Verify with focused tests, an E2E multi-turn run using Maven when needed, logs under `logs/`, database queries via `scripts/query_mysql.py`, and baseline drift checks.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Do not add `diagnosis_trace` or `trace_event` in this change.
|
||||
- Do not implement a full run-list UI.
|
||||
- Do not remove the historical `diagnosis_session` table in this change.
|
||||
- Do not change the Redis conversation window strategy.
|
||||
- Do not persist full conversation history in MySQL; `chat_session` stores metadata only.
|
||||
- Do not split historical mixed traces into true historical runs when the original run boundary is unavailable.
|
||||
|
||||
## Devflow Context Constraints
|
||||
|
||||
- `session-storage` established the current trace tables and decided that `sessionId` is propagated through `RunnableConfig.metadata` with `SessionContextHolder` as a fallback for tools.
|
||||
- `confidence-feedback` established that useful feedback creates `case_library` from `DiagnosisSession.answer`, and that `feedback` does not change execution `status`.
|
||||
- `mvp-demo-trace-acceptance` established `GET /api/diagnosis/{sessionId}/trace` as a read-only endpoint and demo scripts as part of the observable story.
|
||||
- `aiops-traceable-diagnosis-entry` established `/api/ai_ops` as a traceable SSE entry point and made `sessionId` visible to callers.
|
||||
- `data-model.md` and `session-trace-lifecycle.md` describe the current model as `diagnosis_session + agent_step + tool_invocation`, and list run id as a known follow-up.
|
||||
- `devflow/glossary/CONTEXT.md` now defines Chat Session, Diagnosis Run, and Diagnosis Trace. These terms must be used consistently in design/specs/tasks.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L4 database/API contract migration with compatibility behavior.
|
||||
|
||||
- New API response field: `runId`.
|
||||
- New query parameter: `GET /api/diagnosis/{sessionId}/trace?runId=...`.
|
||||
- New API: `GET /api/chat/session/{sessionId}/runs`.
|
||||
- Feedback request gains optional/preferred `runId`.
|
||||
- Feedback response returns bound `runId` and `fallbackToLatestRun`.
|
||||
- `/api/ai_ops` keeps SSE event name `message` and emits a `type=metadata` JSON message containing `sessionId` and `runId` before content.
|
||||
- Database contract changes include new tables and new `run_id` columns.
|
||||
- Old callers that only pass `sessionId` remain compatible by binding to latest run, but this fallback must be observable.
|
||||
|
||||
## Risks
|
||||
|
||||
- Historical data has no true per-round boundary; backfill can only create compatibility runs from existing `diagnosis_session` rows.
|
||||
- Context propagation through Agent hooks and tools is easy to break because it currently combines `RunnableConfig.metadata` and `SessionContextHolder`.
|
||||
- Evidence score, verifier inputs, and baseline metrics may change after run isolation because cross-round tool rows are no longer counted.
|
||||
- Chat-only intermediate completion would leave AIOps as the remaining mixed-trace entry point; AIOps must be completed before overall archive.
|
||||
- `case_library.diagnosis_id` becomes transitional: old data may contain `session_id`, new data contains `run_id`.
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
|
||||
|
||||
#### Scenario: Existing session latest trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
|
||||
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
|
||||
|
||||
#### Scenario: Existing session exact trace is returned
|
||||
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system validates that `runId` belongs to `sessionId`
|
||||
- **AND** it returns a success response containing only the trace data for that run
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
#### Scenario: Exact trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
|
||||
|
||||
+175
@@ -0,0 +1,175 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
|
||||
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
|
||||
|
||||
#### Scenario: Valid Chat execution creates session metadata and a run
|
||||
- **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId`
|
||||
- **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId`
|
||||
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
|
||||
- **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId`
|
||||
|
||||
#### Scenario: Invalid Chat request does not create a run
|
||||
- **WHEN** a `/api/chat` request fails parameter validation before execution
|
||||
- **THEN** the system SHALL NOT create a `diagnosis_run`
|
||||
|
||||
#### Scenario: Chat session stores metadata only
|
||||
- **WHEN** a Chat request completes
|
||||
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
|
||||
- **AND** it SHALL NOT store full conversation message history
|
||||
|
||||
### Requirement: Chat responses SHALL expose run identity
|
||||
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
|
||||
|
||||
#### Scenario: Chat response includes runId
|
||||
- **WHEN** `/api/chat` returns a successful response
|
||||
- **THEN** the response SHALL include `sessionId`
|
||||
- **AND** the response SHALL include `runId` for the created diagnosis run
|
||||
|
||||
#### Scenario: Multi-turn Chat keeps one session and multiple runs
|
||||
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
|
||||
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
|
||||
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
|
||||
|
||||
### Requirement: Trace details SHALL be scoped by run
|
||||
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
|
||||
|
||||
#### Scenario: Agent steps are recorded with runId
|
||||
- **WHEN** an Agent model step is persisted during a diagnosis run
|
||||
- **THEN** the `agent_step` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Tool invocations are recorded with runId
|
||||
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
|
||||
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Run metrics count only current run rows
|
||||
- **WHEN** a diagnosis run completes
|
||||
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
|
||||
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
|
||||
|
||||
### Requirement: Trace API SHALL support latest-run and exact-run queries
|
||||
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
|
||||
|
||||
#### Scenario: Trace without runId resolves latest run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
|
||||
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
|
||||
- **AND** the response SHALL include the resolved `runId`
|
||||
|
||||
#### Scenario: Trace with runId returns exact run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
|
||||
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
|
||||
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
|
||||
|
||||
#### Scenario: Trace rejects run from another session
|
||||
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
|
||||
|
||||
### Requirement: Session runs SHALL be listable without expanding trace details
|
||||
The system SHALL provide a lightweight run-list API for a Chat Session.
|
||||
|
||||
#### Scenario: Run list returns summaries
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
|
||||
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
|
||||
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
|
||||
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
|
||||
|
||||
#### Scenario: Run list handles session without runs
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
|
||||
- **THEN** the system SHALL return a successful empty list
|
||||
|
||||
#### Scenario: Run list rejects missing session
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
|
||||
- **THEN** the system SHALL use the existing not-found/error response behavior
|
||||
|
||||
### Requirement: Feedback SHALL bind to diagnosis runs
|
||||
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
|
||||
|
||||
#### Scenario: Feedback with runId updates specified run
|
||||
- **WHEN** a feedback request includes `sessionId` and `runId`
|
||||
- **THEN** the system SHALL validate that the run belongs to the session
|
||||
- **AND** it SHALL update feedback on that run
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=false`
|
||||
|
||||
#### Scenario: Feedback without runId falls back observably
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** at least one `diagnosis_run` exists for that session
|
||||
- **THEN** the system SHALL bind feedback to the latest run for that session
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=true`
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
|
||||
#### Scenario: Historical feedback without run-backed data remains compatible
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** no `diagnosis_run` exists for that session
|
||||
- **AND** a historical `diagnosis_session` row exists for that session
|
||||
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
|
||||
- **AND** the response SHALL NOT claim latest-run fallback
|
||||
- **AND** the response MAY omit `runId`
|
||||
|
||||
#### Scenario: Feedback rejects run from another session
|
||||
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return a failed feedback response instead of updating either run
|
||||
|
||||
#### Scenario: Useful feedback creates case from run
|
||||
- **WHEN** feedback for a run is `useful`
|
||||
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
|
||||
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
|
||||
|
||||
### Requirement: AIOps executions SHALL use run isolation
|
||||
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
|
||||
|
||||
#### Scenario: AIOps creates run
|
||||
- **WHEN** `/api/ai_ops` starts a valid execution
|
||||
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
|
||||
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
|
||||
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
|
||||
|
||||
#### Scenario: AIOps SSE exposes runId
|
||||
- **WHEN** `/api/ai_ops` streams response metadata to the caller
|
||||
- **THEN** the stream SHALL send a compatible metadata message before report content
|
||||
- **AND** the SSE event name SHALL remain `message`
|
||||
- **AND** the message type SHALL be `metadata`
|
||||
- **AND** the metadata payload SHALL expose the resolved `sessionId`
|
||||
- **AND** the metadata payload SHALL expose the created `runId`
|
||||
- **AND** report content SHALL continue to use the existing content message shape
|
||||
|
||||
### Requirement: Migration SHALL preserve historical trace access
|
||||
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
|
||||
|
||||
#### Scenario: Historical session gets compatibility run
|
||||
- **WHEN** migration runs on an existing `diagnosis_session` row
|
||||
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
|
||||
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
|
||||
|
||||
#### Scenario: Old table is retained
|
||||
- **WHEN** the migration completes
|
||||
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
|
||||
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
|
||||
|
||||
### Requirement: Demo and Trace UI SHALL support runId
|
||||
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
|
||||
|
||||
#### Scenario: Demo script queries exact trace
|
||||
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
|
||||
- **THEN** it SHALL include `runId` when querying the Trace API
|
||||
|
||||
#### Scenario: Trace UI honors URL runId
|
||||
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
|
||||
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
|
||||
### Requirement: Run isolation SHALL be verified against baselines
|
||||
The change SHALL verify both runtime behavior and evaluation baseline impact.
|
||||
|
||||
#### Scenario: Multi-turn E2E proves run isolation
|
||||
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
|
||||
- **THEN** database inspection SHALL show two `diagnosis_run` rows
|
||||
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
|
||||
- **AND** Redis session metadata SHALL still show multi-turn context continuity
|
||||
|
||||
#### Scenario: Baseline drift is checked
|
||||
- **WHEN** verification is complete
|
||||
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
|
||||
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
|
||||
@@ -0,0 +1,59 @@
|
||||
## 1. Schema and Compatibility Foundation
|
||||
|
||||
- [x] 1.1 Add Flyway migration for `chat_session` and `diagnosis_run` with indexes for unique `session_id`, unique `run_id`, latest-run lookup, and run summary listing.
|
||||
- [x] 1.2 Add nullable `run_id` columns to `agent_step` and `tool_invocation` with indexes for run-scoped trace queries.
|
||||
- [x] 1.3 Backfill one compatibility `diagnosis_run` for each existing `diagnosis_session` row.
|
||||
- [x] 1.4 Backfill historical `agent_step.run_id` and `tool_invocation.run_id` from the compatibility run for the same `session_id`.
|
||||
- [x] 1.5 Add JPA entities and repositories for `ChatSession` and `DiagnosisRun`, including latest-run and run-id lookup methods.
|
||||
- [x] 1.6 Add focused migration/repository verification that proves old data remains queryable and run lookup methods work.
|
||||
- [x] 1.7 Phase 1 gate: run the smallest relevant test/build check, inspect DB migration behavior, update OpenSpec task status, archive phase evidence, and commit before starting Phase 2.
|
||||
|
||||
## 2. Chat Run Write Path
|
||||
|
||||
- [x] 2.1 Add a unified execution context that carries both `sessionId` and `runId` through Chat service, Agent hooks, and tool recording.
|
||||
- [x] 2.2 Change valid `/api/chat` executions to create or update `chat_session` metadata and create one new `diagnosis_run`.
|
||||
- [x] 2.3 Change `AgentLoggingHook` to write `agent_step.run_id` for Chat runs while retaining `session_id`.
|
||||
- [x] 2.4 Change `ToolInvocationRecorder` and evidence tools to write `tool_invocation.run_id` for Chat runs while retaining `session_id`.
|
||||
- [x] 2.5 Change Chat completion, failure, answer, self-evaluation, duration, token, step, and tool count writes from `diagnosis_session` to the current `diagnosis_run`.
|
||||
- [x] 2.6 Change `ToolTraceSummaryService`, `ExecutorGatekeeperService`, and `EvaluationService` Chat reads from session-scoped tool rows to run-scoped tool rows.
|
||||
- [x] 2.7 Change `/api/chat` response DTO to include official `runId`.
|
||||
- [x] 2.8 Add focused tests for valid Chat run creation, invalid request no-run behavior, run-scoped counts, run-scoped verifier/gatekeeper/evaluation reads, and multi-turn context preservation.
|
||||
- [x] 2.9 Phase 2 gate: run focused tests plus a same-session two-round Chat E2E when needed, inspect DB with `scripts/query_mysql.py`, review `logs/`, update task status, archive phase evidence, and commit before starting Phase 3.
|
||||
|
||||
## 3. Trace Read Path and Run Listing
|
||||
|
||||
- [x] 3.1 Change `DiagnosisTraceService` to resolve latest run by `diagnosis_run.created_at DESC, id DESC` when `runId` is omitted.
|
||||
- [x] 3.2 Add exact trace lookup for `sessionId + runId`, including validation that the run belongs to the session.
|
||||
- [x] 3.3 Change trace response DTOs to include resolved `runId` and run summary fields.
|
||||
- [x] 3.4 Add `GET /api/chat/session/{sessionId}/runs` returning lightweight run summaries without expanding trace detail rows.
|
||||
- [x] 3.5 Preserve read-only trace behavior for latest-run and exact-run requests.
|
||||
- [x] 3.6 Add tests for latest run, exact first run, exact second run, wrong-session run rejection, missing session, and read-only behavior.
|
||||
- [x] 3.7 Phase 3 gate: run focused trace tests and same-session E2E trace checks, update task status, archive phase evidence, and commit before starting Phase 4.
|
||||
|
||||
## 4. Feedback and Case Library Run Binding
|
||||
|
||||
- [x] 4.1 Change feedback request handling to prefer `runId` and validate run/session ownership.
|
||||
- [x] 4.2 Implement legacy feedback fallback to latest run with observable `fallbackToLatestRun=true` and actual bound `runId`.
|
||||
- [x] 4.3 Change feedback persistence to update `diagnosis_run.feedback` for new data.
|
||||
- [x] 4.4 Change `CaseLibraryService` to create automatic cases from `diagnosis_run.query` and `diagnosis_run.answer`.
|
||||
- [x] 4.5 Preserve transitional case-library semantics where old `diagnosis_id` values may be `session_id` and new automatic values are `run_id`.
|
||||
- [x] 4.6 Add tests for run-specific feedback, legacy fallback, useful case creation, idempotency, and old-data compatibility.
|
||||
- [x] 4.7 Phase 4 gate: run focused feedback/case tests, inspect DB with `scripts/query_mysql.py`, update task status, archive phase evidence, and commit before starting Phase 5.
|
||||
|
||||
## 5. AIOps Run Isolation
|
||||
|
||||
- [x] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
|
||||
- [x] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
|
||||
- [x] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
|
||||
- [x] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
|
||||
- [x] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
|
||||
- [x] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
|
||||
|
||||
## 6. Demo, Trace UI, Documentation, and Verification
|
||||
|
||||
- [x] 6.1 Update demo scripts to read `runId` from Chat/AIOps responses and pass `?runId=...` to Trace API.
|
||||
- [x] 6.2 Update Trace UI to accept `?sessionId=...&runId=...` and query exact trace when `runId` is present.
|
||||
- [x] 6.3 Update MVP table and architecture docs for `chat_session`, `diagnosis_run`, `run_id`, and transitional `case_library.diagnosis_id` semantics.
|
||||
- [x] 6.4 Run final same-session multi-turn E2E using Maven startup if needed; collect DB evidence through `scripts/query_mysql.py` and inspect `logs/`.
|
||||
- [x] 6.5 Run or explicitly evaluate the relevant baseline diff command and document whether drift is expected or a regression.
|
||||
- [x] 6.6 Final gate: ensure all OpenSpec tasks are checked, no new writes depend on `diagnosis_session`, phase evidence is archived, final commit is created, and the change is ready for OpenSpec archive.
|
||||
@@ -1,24 +1,33 @@
|
||||
## Purpose
|
||||
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same diagnosis run.
|
||||
## Requirements
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
|
||||
|
||||
#### Scenario: Existing session trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||
#### Scenario: Existing session latest trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
|
||||
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
|
||||
|
||||
#### Scenario: Existing session exact trace is returned
|
||||
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system validates that `runId` belongs to `sessionId`
|
||||
- **AND** it returns a success response containing only the trace data for that run
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
#### Scenario: Exact trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
|
||||
|
||||
### Requirement: MVP demo profile is available
|
||||
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||
@@ -28,11 +37,11 @@ The system SHALL provide an `mvp-demo` Spring profile that documents the demo ru
|
||||
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||
|
||||
### Requirement: End-to-end MVP acceptance case is documented
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same `sessionId + runId`.
|
||||
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the exact trace endpoint, and submit feedback for the same run
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
@@ -51,10 +60,11 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
- **AND** it SHALL read the returned run id for exact trace and feedback calls
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
- **THEN** it SHALL write chat, exact trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL be reproducible for interviews
|
||||
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
|
||||
@@ -62,8 +72,8 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
|
||||
#### Scenario: interview demo check script records an evidence bundle
|
||||
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
|
||||
- **THEN** the script SHALL submit a fixed Chat diagnosis request
|
||||
- **AND** it SHALL fetch the trace for the same session id
|
||||
- **AND** it SHALL submit useful feedback for that session
|
||||
- **AND** it SHALL fetch the trace for the same `sessionId + runId`
|
||||
- **AND** it SHALL submit useful feedback for that run
|
||||
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
|
||||
|
||||
#### Scenario: interview demo check fails with actionable readiness output
|
||||
@@ -77,7 +87,7 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
|
||||
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and run-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
@@ -85,7 +95,7 @@ The MVP demo SHALL document which trace fields to inspect for evidence, verifier
|
||||
|
||||
### Requirement: MVP demo SHALL provide a browser trace workbench
|
||||
The MVP demo SHALL provide a browser-accessible static page for inspecting one
|
||||
diagnosis trace by session id using the existing read-only Trace API.
|
||||
diagnosis trace by session id and optional run id using the existing read-only Trace API.
|
||||
|
||||
#### Scenario: Existing trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
|
||||
@@ -93,6 +103,11 @@ diagnosis trace by session id using the existing read-only Trace API.
|
||||
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
|
||||
invocations, verifier evaluation, and final answer when present
|
||||
|
||||
#### Scenario: Exact trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with `?sessionId=...&runId=...`
|
||||
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
- **AND** it SHALL render only that run's trace data
|
||||
|
||||
#### Scenario: Trace workbench handles missing or failed traces
|
||||
- **WHEN** the Trace API returns an error or the session id is empty
|
||||
- **THEN** the page SHALL show a clear error or empty state without mutating any
|
||||
|
||||
@@ -0,0 +1,181 @@
|
||||
# session-run-trace-isolation Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
|
||||
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
|
||||
|
||||
#### Scenario: Valid Chat execution creates session metadata and a run
|
||||
- **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId`
|
||||
- **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId`
|
||||
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
|
||||
- **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId`
|
||||
|
||||
#### Scenario: Invalid Chat request does not create a run
|
||||
- **WHEN** a `/api/chat` request fails parameter validation before execution
|
||||
- **THEN** the system SHALL NOT create a `diagnosis_run`
|
||||
|
||||
#### Scenario: Chat session stores metadata only
|
||||
- **WHEN** a Chat request completes
|
||||
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
|
||||
- **AND** it SHALL NOT store full conversation message history
|
||||
|
||||
### Requirement: Chat responses SHALL expose run identity
|
||||
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
|
||||
|
||||
#### Scenario: Chat response includes runId
|
||||
- **WHEN** `/api/chat` returns a successful response
|
||||
- **THEN** the response SHALL include `sessionId`
|
||||
- **AND** the response SHALL include `runId` for the created diagnosis run
|
||||
|
||||
#### Scenario: Multi-turn Chat keeps one session and multiple runs
|
||||
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
|
||||
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
|
||||
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
|
||||
|
||||
### Requirement: Trace details SHALL be scoped by run
|
||||
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
|
||||
|
||||
#### Scenario: Agent steps are recorded with runId
|
||||
- **WHEN** an Agent model step is persisted during a diagnosis run
|
||||
- **THEN** the `agent_step` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Tool invocations are recorded with runId
|
||||
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
|
||||
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Run metrics count only current run rows
|
||||
- **WHEN** a diagnosis run completes
|
||||
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
|
||||
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
|
||||
|
||||
### Requirement: Trace API SHALL support latest-run and exact-run queries
|
||||
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
|
||||
|
||||
#### Scenario: Trace without runId resolves latest run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
|
||||
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
|
||||
- **AND** the response SHALL include the resolved `runId`
|
||||
|
||||
#### Scenario: Trace with runId returns exact run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
|
||||
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
|
||||
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
|
||||
|
||||
#### Scenario: Trace rejects run from another session
|
||||
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
|
||||
|
||||
### Requirement: Session runs SHALL be listable without expanding trace details
|
||||
The system SHALL provide a lightweight run-list API for a Chat Session.
|
||||
|
||||
#### Scenario: Run list returns summaries
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
|
||||
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
|
||||
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
|
||||
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
|
||||
|
||||
#### Scenario: Run list handles session without runs
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
|
||||
- **THEN** the system SHALL return a successful empty list
|
||||
|
||||
#### Scenario: Run list rejects missing session
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
|
||||
- **THEN** the system SHALL use the existing not-found/error response behavior
|
||||
|
||||
### Requirement: Feedback SHALL bind to diagnosis runs
|
||||
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
|
||||
|
||||
#### Scenario: Feedback with runId updates specified run
|
||||
- **WHEN** a feedback request includes `sessionId` and `runId`
|
||||
- **THEN** the system SHALL validate that the run belongs to the session
|
||||
- **AND** it SHALL update feedback on that run
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=false`
|
||||
|
||||
#### Scenario: Feedback without runId falls back observably
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** at least one `diagnosis_run` exists for that session
|
||||
- **THEN** the system SHALL bind feedback to the latest run for that session
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=true`
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
|
||||
#### Scenario: Historical feedback without run-backed data remains compatible
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** no `diagnosis_run` exists for that session
|
||||
- **AND** a historical `diagnosis_session` row exists for that session
|
||||
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
|
||||
- **AND** the response SHALL NOT claim latest-run fallback
|
||||
- **AND** the response MAY omit `runId`
|
||||
|
||||
#### Scenario: Feedback rejects run from another session
|
||||
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return a failed feedback response instead of updating either run
|
||||
|
||||
#### Scenario: Useful feedback creates case from run
|
||||
- **WHEN** feedback for a run is `useful`
|
||||
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
|
||||
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
|
||||
|
||||
### Requirement: AIOps executions SHALL use run isolation
|
||||
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
|
||||
|
||||
#### Scenario: AIOps creates run
|
||||
- **WHEN** `/api/ai_ops` starts a valid execution
|
||||
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
|
||||
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
|
||||
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
|
||||
|
||||
#### Scenario: AIOps SSE exposes runId
|
||||
- **WHEN** `/api/ai_ops` streams response metadata to the caller
|
||||
- **THEN** the stream SHALL send a compatible metadata message before report content
|
||||
- **AND** the SSE event name SHALL remain `message`
|
||||
- **AND** the message type SHALL be `metadata`
|
||||
- **AND** the metadata payload SHALL expose the resolved `sessionId`
|
||||
- **AND** the metadata payload SHALL expose the created `runId`
|
||||
- **AND** report content SHALL continue to use the existing content message shape
|
||||
|
||||
### Requirement: Migration SHALL preserve historical trace access
|
||||
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
|
||||
|
||||
#### Scenario: Historical session gets compatibility run
|
||||
- **WHEN** migration runs on an existing `diagnosis_session` row
|
||||
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
|
||||
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
|
||||
|
||||
#### Scenario: Old table is retained
|
||||
- **WHEN** the migration completes
|
||||
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
|
||||
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
|
||||
|
||||
### Requirement: Demo and Trace UI SHALL support runId
|
||||
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
|
||||
|
||||
#### Scenario: Demo script queries exact trace
|
||||
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
|
||||
- **THEN** it SHALL include `runId` when querying the Trace API
|
||||
|
||||
#### Scenario: Trace UI honors URL runId
|
||||
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
|
||||
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
|
||||
### Requirement: Run isolation SHALL be verified against baselines
|
||||
The change SHALL verify both runtime behavior and evaluation baseline impact.
|
||||
|
||||
#### Scenario: Multi-turn E2E proves run isolation
|
||||
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
|
||||
- **THEN** database inspection SHALL show two `diagnosis_run` rows
|
||||
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
|
||||
- **AND** Redis session metadata SHALL still show multi-turn context continuity
|
||||
|
||||
#### Scenario: Baseline drift is checked
|
||||
- **WHEN** verification is complete
|
||||
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
|
||||
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
|
||||
@@ -5,8 +5,10 @@ import lombok.Getter;
|
||||
import lombok.Setter;
|
||||
import com.superbiz.agent.domain.model.SessionContext;
|
||||
import com.superbiz.agent.dto.AIOpsRequest;
|
||||
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||
import com.superbiz.agent.service.AiOpsService;
|
||||
import com.superbiz.agent.service.ChatService;
|
||||
import com.superbiz.agent.service.DiagnosisTraceService;
|
||||
import com.superbiz.agent.service.session.SessionManager;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -43,6 +45,9 @@ public class ChatController {
|
||||
@Autowired
|
||||
private ChatService chatService;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisTraceService diagnosisTraceService;
|
||||
|
||||
@Autowired
|
||||
private SessionManager sessionManager;
|
||||
|
||||
@@ -96,10 +101,11 @@ public class ChatController {
|
||||
// 更新会话历史
|
||||
session.addChatMessagePair(request.getQuestion(), fullAnswer, MAX_WINDOW_SIZE);
|
||||
sessionManager.updateSession(session);
|
||||
chatService.syncChatSessionMetadata(session.getSessionId(), session.getMessagePairCount());
|
||||
logger.info("已更新会话历史 - SessionId: {}, 当前消息对数: {}",
|
||||
session.getSessionId(), session.getMessagePairCount());
|
||||
|
||||
return ResponseEntity.ok(ApiResponse.success(ChatResponse.success(fullAnswer, result.sessionId())));
|
||||
return ResponseEntity.ok(ApiResponse.success(ChatResponse.success(fullAnswer, result.sessionId(), result.runId())));
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("对话失败", e);
|
||||
@@ -216,20 +222,21 @@ public class ChatController {
|
||||
public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) {
|
||||
SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢)
|
||||
String sessionId = aiOpsService.resolveSessionId(request);
|
||||
String runId = aiOpsService.newRunId();
|
||||
|
||||
executor.execute(() -> {
|
||||
try {
|
||||
logger.info("收到 AI 智能运维请求 - SessionId: {}, 启动多 Agent 协作流程", sessionId);
|
||||
logger.info("收到 AI 智能运维请求 - SessionId: {}, RunId: {}, 启动多 Agent 协作流程", sessionId, runId);
|
||||
|
||||
ChatModel chatModel = chatService.getChatModel();
|
||||
|
||||
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
|
||||
|
||||
emitter.send(SseEmitter.event().name("message").data(SseMessage.session(sessionId), MediaType.APPLICATION_JSON));
|
||||
emitter.send(SseEmitter.event().name("message").data(SseMessage.metadata(sessionId, runId), MediaType.APPLICATION_JSON));
|
||||
emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n")));
|
||||
|
||||
// 调用 AiOpsService 执行分析流程
|
||||
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId);
|
||||
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId, runId);
|
||||
|
||||
if (overAllStateOptional.isEmpty()) {
|
||||
emitter.send(SseEmitter.event().name("message")
|
||||
@@ -248,7 +255,7 @@ public class ChatController {
|
||||
if (finalReportOptional.isPresent()) {
|
||||
String finalReportText = finalReportOptional.get();
|
||||
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
|
||||
aiOpsService.persistFinalReport(sessionId, finalReportText, request);
|
||||
aiOpsService.persistFinalReport(sessionId, runId, finalReportText, request);
|
||||
|
||||
// 发送分隔线
|
||||
emitter.send(SseEmitter.event().name("message")
|
||||
@@ -324,6 +331,12 @@ public class ChatController {
|
||||
}
|
||||
}
|
||||
|
||||
@GetMapping("/chat/session/{sessionId}/runs")
|
||||
public ResponseEntity<ApiResponse<List<DiagnosisTraceResponse.RunSummary>>> listSessionRuns(
|
||||
@PathVariable String sessionId) {
|
||||
return ResponseEntity.ok(ApiResponse.success(diagnosisTraceService.listRunSummaries(sessionId)));
|
||||
}
|
||||
|
||||
// ==================== 辅助方法 ====================
|
||||
|
||||
private SessionContext getOrCreateSession(String sessionId) {
|
||||
@@ -413,12 +426,14 @@ public class ChatController {
|
||||
private String answer;
|
||||
private String errorMessage;
|
||||
private String sessionId;
|
||||
private String runId;
|
||||
|
||||
public static ChatResponse success(String answer, String sessionId) {
|
||||
public static ChatResponse success(String answer, String sessionId, String runId) {
|
||||
ChatResponse response = new ChatResponse();
|
||||
response.setSuccess(true);
|
||||
response.setAnswer(answer);
|
||||
response.setSessionId(sessionId);
|
||||
response.setRunId(runId);
|
||||
return response;
|
||||
}
|
||||
|
||||
@@ -437,8 +452,10 @@ public class ChatController {
|
||||
@Setter
|
||||
@Getter
|
||||
public static class SseMessage {
|
||||
private String type; // content: 内容块, error: 错误, done: 完成
|
||||
private String type; // metadata: 元数据, content: 内容块, error: 错误, done: 完成
|
||||
private String data;
|
||||
private String sessionId;
|
||||
private String runId;
|
||||
|
||||
public static SseMessage content(String data) {
|
||||
SseMessage message = new SseMessage();
|
||||
@@ -454,6 +471,14 @@ public class ChatController {
|
||||
return message;
|
||||
}
|
||||
|
||||
public static SseMessage metadata(String sessionId, String runId) {
|
||||
SseMessage message = new SseMessage();
|
||||
message.setType("metadata");
|
||||
message.setSessionId(sessionId);
|
||||
message.setRunId(runId);
|
||||
return message;
|
||||
}
|
||||
|
||||
public static SseMessage error(String errorMessage) {
|
||||
SseMessage message = new SseMessage();
|
||||
message.setType("error");
|
||||
|
||||
@@ -7,6 +7,7 @@ import lombok.RequiredArgsConstructor;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.web.bind.annotation.GetMapping;
|
||||
import org.springframework.web.bind.annotation.PathVariable;
|
||||
import org.springframework.web.bind.annotation.RequestParam;
|
||||
import org.springframework.web.bind.annotation.RequestMapping;
|
||||
import org.springframework.web.bind.annotation.RestController;
|
||||
|
||||
@@ -18,7 +19,10 @@ public class DiagnosisTraceController {
|
||||
private final DiagnosisTraceService diagnosisTraceService;
|
||||
|
||||
@GetMapping("/{sessionId}/trace")
|
||||
public ResponseEntity<Result<DiagnosisTraceResponse>> getTrace(@PathVariable String sessionId) {
|
||||
return ResponseEntity.ok(Result.success(diagnosisTraceService.getTrace(sessionId)));
|
||||
public ResponseEntity<Result<DiagnosisTraceResponse>> getTrace(
|
||||
@PathVariable String sessionId,
|
||||
@RequestParam(required = false) String runId
|
||||
) {
|
||||
return ResponseEntity.ok(Result.success(diagnosisTraceService.getTrace(sessionId, runId)));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -16,7 +16,11 @@ public class FeedbackController {
|
||||
|
||||
@PostMapping("/feedback")
|
||||
public ResponseEntity<FeedbackResponse> submitFeedback(@RequestBody FeedbackRequest request) {
|
||||
FeedbackResponse response = feedbackService.submitFeedback(request.getSessionId(), request.getFeedback());
|
||||
FeedbackResponse response = feedbackService.submitFeedback(
|
||||
request.getSessionId(),
|
||||
request.getRunId(),
|
||||
request.getFeedback()
|
||||
);
|
||||
if (!response.isSuccess()) {
|
||||
return ResponseEntity.badRequest().body(response);
|
||||
}
|
||||
|
||||
@@ -15,6 +15,7 @@ import java.time.LocalDateTime;
|
||||
@Entity
|
||||
@Table(name = "agent_step", indexes = {
|
||||
@Index(name = "idx_session_step", columnList = "session_id, step_index"),
|
||||
@Index(name = "idx_agent_step_run_step", columnList = "run_id, step_index"),
|
||||
@Index(name = "idx_agent_name", columnList = "agent_name")
|
||||
})
|
||||
@Data
|
||||
@@ -30,6 +31,9 @@ public class AgentStep {
|
||||
@Column(name = "session_id", nullable = false, length = 64)
|
||||
private String sessionId;
|
||||
|
||||
@Column(name = "run_id", length = 64)
|
||||
private String runId;
|
||||
|
||||
@Column(name = "step_index", nullable = false)
|
||||
private Integer stepIndex;
|
||||
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
package com.superbiz.agent.domain.entity;
|
||||
|
||||
import jakarta.persistence.*;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
|
||||
/**
|
||||
* Chat session metadata entity.
|
||||
* Full message history remains in Redis SessionContext.
|
||||
*/
|
||||
@Entity
|
||||
@Table(name = "chat_session", indexes = {
|
||||
@Index(name = "idx_chat_session_last_active", columnList = "last_active_at"),
|
||||
@Index(name = "idx_chat_session_status", columnList = "status"),
|
||||
@Index(name = "idx_chat_session_expires_at", columnList = "expires_at")
|
||||
})
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class ChatSession {
|
||||
|
||||
@Id
|
||||
@GeneratedValue(strategy = GenerationType.IDENTITY)
|
||||
private Long id;
|
||||
|
||||
@Column(name = "session_id", unique = true, nullable = false, length = 64)
|
||||
private String sessionId;
|
||||
|
||||
@Column(name = "status", length = 16)
|
||||
private String status = "ACTIVE";
|
||||
|
||||
@Column(name = "message_pair_count")
|
||||
private Integer messagePairCount = 0;
|
||||
|
||||
@Column(name = "created_at", nullable = false, updatable = false)
|
||||
private LocalDateTime createdAt;
|
||||
|
||||
@Column(name = "last_active_at")
|
||||
private LocalDateTime lastActiveAt;
|
||||
|
||||
@Column(name = "expires_at")
|
||||
private LocalDateTime expiresAt;
|
||||
|
||||
@PrePersist
|
||||
protected void onCreate() {
|
||||
LocalDateTime now = LocalDateTime.now();
|
||||
createdAt = now;
|
||||
if (lastActiveAt == null) {
|
||||
lastActiveAt = now;
|
||||
}
|
||||
if (status == null) {
|
||||
status = "ACTIVE";
|
||||
}
|
||||
if (messagePairCount == null) {
|
||||
messagePairCount = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
package com.superbiz.agent.domain.entity;
|
||||
|
||||
import jakarta.persistence.*;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
import org.hibernate.annotations.JdbcTypeCode;
|
||||
import org.hibernate.type.SqlTypes;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
|
||||
/**
|
||||
* One diagnosis execution run.
|
||||
*/
|
||||
@Entity
|
||||
@Table(name = "diagnosis_run", indexes = {
|
||||
@Index(name = "idx_diagnosis_run_session_created", columnList = "session_id, created_at, id"),
|
||||
@Index(name = "idx_diagnosis_run_session_run", columnList = "session_id, run_id"),
|
||||
@Index(name = "idx_diagnosis_run_status", columnList = "status"),
|
||||
@Index(name = "idx_diagnosis_run_agent_flow", columnList = "agent_flow")
|
||||
})
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class DiagnosisRun {
|
||||
|
||||
@Id
|
||||
@GeneratedValue(strategy = GenerationType.IDENTITY)
|
||||
private Long id;
|
||||
|
||||
@Column(name = "run_id", unique = true, nullable = false, length = 64)
|
||||
private String runId;
|
||||
|
||||
@Column(name = "session_id", nullable = false, length = 64)
|
||||
private String sessionId;
|
||||
|
||||
@Column(name = "query", nullable = false, columnDefinition = "TEXT")
|
||||
private String query;
|
||||
|
||||
@Column(name = "status", length = 16)
|
||||
private String status = "PENDING";
|
||||
|
||||
@Column(name = "agent_flow", length = 32)
|
||||
private String agentFlow;
|
||||
|
||||
@Column(name = "answer", columnDefinition = "LONGTEXT")
|
||||
private String answer;
|
||||
|
||||
@JdbcTypeCode(SqlTypes.JSON)
|
||||
@Column(name = "self_evaluation", columnDefinition = "JSON")
|
||||
private String selfEvaluation;
|
||||
|
||||
@Column(name = "feedback", length = 16)
|
||||
private String feedback;
|
||||
|
||||
@Column(name = "total_duration_ms")
|
||||
private Integer totalDurationMs;
|
||||
|
||||
@Column(name = "total_token_count")
|
||||
private Integer totalTokenCount;
|
||||
|
||||
@Column(name = "step_count")
|
||||
private Integer stepCount;
|
||||
|
||||
@Column(name = "tool_call_count")
|
||||
private Integer toolCallCount;
|
||||
|
||||
@Column(name = "created_at", nullable = false, updatable = false)
|
||||
private LocalDateTime createdAt;
|
||||
|
||||
@Column(name = "updated_at")
|
||||
private LocalDateTime updatedAt;
|
||||
|
||||
@PrePersist
|
||||
protected void onCreate() {
|
||||
LocalDateTime now = LocalDateTime.now();
|
||||
createdAt = now;
|
||||
updatedAt = now;
|
||||
if (status == null) {
|
||||
status = "PENDING";
|
||||
}
|
||||
}
|
||||
|
||||
@PreUpdate
|
||||
protected void onUpdate() {
|
||||
updatedAt = LocalDateTime.now();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,6 +17,7 @@ import java.time.LocalDateTime;
|
||||
@Entity
|
||||
@Table(name = "tool_invocation", indexes = {
|
||||
@Index(name = "idx_session_id", columnList = "session_id"),
|
||||
@Index(name = "idx_tool_invocation_run_id", columnList = "run_id, id"),
|
||||
@Index(name = "idx_tool_name", columnList = "tool_name"),
|
||||
@Index(name = "idx_retrieval_layer", columnList = "retrieval_layer")
|
||||
})
|
||||
@@ -33,6 +34,9 @@ public class ToolInvocation {
|
||||
@Column(name = "session_id", nullable = false, length = 64)
|
||||
private String sessionId;
|
||||
|
||||
@Column(name = "run_id", length = 64)
|
||||
private String runId;
|
||||
|
||||
@Column(name = "step_id")
|
||||
private Long stepId;
|
||||
|
||||
|
||||
@@ -15,11 +15,28 @@ import java.util.Map;
|
||||
@AllArgsConstructor
|
||||
public class DiagnosisTraceResponse {
|
||||
|
||||
private String runId;
|
||||
private ChatSessionTrace chatSession;
|
||||
private SessionTrace session;
|
||||
private RunTrace run;
|
||||
private List<AgentStepTrace> steps;
|
||||
private List<ToolInvocationTrace> toolInvocations;
|
||||
private TraceSummary summary;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class ChatSessionTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private String status;
|
||||
private Integer messagePairCount;
|
||||
private LocalDateTime createdAt;
|
||||
private LocalDateTime lastActiveAt;
|
||||
private LocalDateTime expiresAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@@ -42,6 +59,29 @@ public class DiagnosisTraceResponse {
|
||||
private LocalDateTime updatedAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class RunTrace {
|
||||
private Long id;
|
||||
private String runId;
|
||||
private String sessionId;
|
||||
private String query;
|
||||
private String status;
|
||||
private String agentFlow;
|
||||
private Integer totalDurationMs;
|
||||
private Integer totalTokenCount;
|
||||
private Integer stepCount;
|
||||
private Integer toolCallCount;
|
||||
private String answer;
|
||||
private String selfEvaluationRaw;
|
||||
private Map<String, Object> selfEvaluation;
|
||||
private String feedback;
|
||||
private LocalDateTime createdAt;
|
||||
private LocalDateTime updatedAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@@ -49,6 +89,7 @@ public class DiagnosisTraceResponse {
|
||||
public static class AgentStepTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private String runId;
|
||||
private Integer stepIndex;
|
||||
private String agentName;
|
||||
private String modelInput;
|
||||
@@ -67,6 +108,7 @@ public class DiagnosisTraceResponse {
|
||||
public static class ToolInvocationTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private String runId;
|
||||
private Long stepId;
|
||||
private String toolName;
|
||||
private String inputParamsRaw;
|
||||
@@ -92,6 +134,7 @@ public class DiagnosisTraceResponse {
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class TraceSummary {
|
||||
private String resolvedRunId;
|
||||
private int persistedStepCount;
|
||||
private int returnedStepCount;
|
||||
private int persistedToolCallCount;
|
||||
@@ -100,4 +143,21 @@ public class DiagnosisTraceResponse {
|
||||
private boolean hasAiOpsRuleEvaluation;
|
||||
private boolean hasFeedback;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class RunSummary {
|
||||
private String runId;
|
||||
private String sessionId;
|
||||
private String query;
|
||||
private String status;
|
||||
private String agentFlow;
|
||||
private String answerPreview;
|
||||
private Integer stepCount;
|
||||
private Integer toolCallCount;
|
||||
private LocalDateTime createdAt;
|
||||
private LocalDateTime updatedAt;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -7,5 +7,6 @@ import lombok.Setter;
|
||||
@Setter
|
||||
public class FeedbackRequest {
|
||||
private String sessionId;
|
||||
private String runId;
|
||||
private String feedback;
|
||||
}
|
||||
|
||||
@@ -9,4 +9,6 @@ public class FeedbackResponse {
|
||||
private boolean success;
|
||||
private String message;
|
||||
private String caseId;
|
||||
private String runId;
|
||||
private boolean fallbackToLatestRun;
|
||||
}
|
||||
|
||||
@@ -47,11 +47,13 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
String sessionId = resolveSessionId(config);
|
||||
String runId = resolveRunId(config);
|
||||
String traceScopeId = traceScopeId(sessionId, runId);
|
||||
boolean hasSession = sessionId != null;
|
||||
|
||||
int stepIndex = 0;
|
||||
if (hasSession) {
|
||||
stepIndex = stepCounters.merge(sessionId, 0, (oldValue, ignored) -> oldValue + 1);
|
||||
if (traceScopeId != null) {
|
||||
stepIndex = stepCounters.merge(traceScopeId, 0, (oldValue, ignored) -> oldValue + 1);
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
@@ -73,12 +75,13 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
try {
|
||||
AgentStep step = AgentStep.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.stepIndex(stepIndex)
|
||||
.agentName(agentName)
|
||||
.modelInput(buildModelInputSummary(previousMessages))
|
||||
.build();
|
||||
AgentStep saved = agentStepRepository.save(step);
|
||||
pendingSteps.put(sessionId + "_" + stepIndex, Map.of(
|
||||
pendingSteps.put(stepKey(traceScopeId, stepIndex), Map.of(
|
||||
"stepId", saved.getId(),
|
||||
"startTime", System.currentTimeMillis()
|
||||
));
|
||||
@@ -93,7 +96,9 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
@Override
|
||||
public AgentCommand afterModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
String sessionId = resolveSessionId(config);
|
||||
int stepIndex = sessionId == null ? 0 : stepCounters.getOrDefault(sessionId, 0);
|
||||
String runId = resolveRunId(config);
|
||||
String traceScopeId = traceScopeId(sessionId, runId);
|
||||
int stepIndex = traceScopeId == null ? 0 : stepCounters.getOrDefault(traceScopeId, 0);
|
||||
|
||||
log.info("========================================");
|
||||
log.info("*** [AgentTrace] agent={}, phase=after_model, stepIndex={}", agentName, stepIndex);
|
||||
@@ -122,7 +127,7 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
log.info("========================================");
|
||||
|
||||
if (sessionId != null) {
|
||||
String stepKey = sessionId + "_" + stepIndex;
|
||||
String stepKey = stepKey(traceScopeId, stepIndex);
|
||||
Map<String, Object> pending = pendingSteps.remove(stepKey);
|
||||
if (pending != null) {
|
||||
try {
|
||||
@@ -162,6 +167,23 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
}
|
||||
|
||||
private String resolveRunId(RunnableConfig config) {
|
||||
return config.metadata("runId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getRunId);
|
||||
}
|
||||
|
||||
private String traceScopeId(String sessionId, String runId) {
|
||||
if (runId != null && !runId.isBlank()) {
|
||||
return runId;
|
||||
}
|
||||
return sessionId;
|
||||
}
|
||||
|
||||
private String stepKey(String traceScopeId, int stepIndex) {
|
||||
return traceScopeId + "_" + stepIndex;
|
||||
}
|
||||
|
||||
private AssistantMessage findLastAssistant(List<Message> previousMessages) {
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage assistantMessage) {
|
||||
|
||||
@@ -59,13 +59,17 @@ public class VerifierInputHook extends MessagesModelHook {
|
||||
String sessionId = config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
String runId = config.metadata("runId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getRunId);
|
||||
String executorFinalAnswer = VerifierContextHolder.getExecutorFinalAnswer();
|
||||
if (executorFinalAnswer == null || executorFinalAnswer.isBlank()) {
|
||||
executorFinalAnswer = extractLastAssistantText(previousMessages);
|
||||
}
|
||||
|
||||
List<Map<String, Object>> toolTraceSummary =
|
||||
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
|
||||
List<Map<String, Object>> toolTraceSummary = runId == null || runId.isBlank()
|
||||
? toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer)
|
||||
: toolTraceSummaryService.buildVerifierTraceSummaryForRun(runId, executorFinalAnswer);
|
||||
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
|
||||
|
||||
ExecutorOutputParseResult parseResult = parseExecutorOutput(executorFinalAnswer);
|
||||
@@ -76,7 +80,7 @@ public class VerifierInputHook extends MessagesModelHook {
|
||||
VerifierContextHolder.setExecutorStructuredOutput(parseResult.structuredOutput());
|
||||
VerifierContextHolder.setExecutorOutputParseStatus(parseResult.status());
|
||||
|
||||
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
|
||||
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, runId, parseResult);
|
||||
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
|
||||
|
||||
Map<String, Object> verifierInput = new LinkedHashMap<>();
|
||||
@@ -96,11 +100,14 @@ public class VerifierInputHook extends MessagesModelHook {
|
||||
}
|
||||
}
|
||||
|
||||
private Map<String, Object> runGatekeeper(String sessionId, ExecutorOutputParseResult parseResult) {
|
||||
private Map<String, Object> runGatekeeper(String sessionId, String runId, ExecutorOutputParseResult parseResult) {
|
||||
if (executorGatekeeperService == null) {
|
||||
return passGatekeeperResult();
|
||||
}
|
||||
try {
|
||||
if (runId != null && !runId.isBlank()) {
|
||||
return executorGatekeeperService.validateRun(runId, parseResult.structuredOutput(), parseResult.status());
|
||||
}
|
||||
return executorGatekeeperService.validate(sessionId, parseResult.structuredOutput(), parseResult.status());
|
||||
} catch (Exception e) {
|
||||
log.error("Gatekeeper validation failed unexpectedly", e);
|
||||
|
||||
@@ -22,8 +22,18 @@ public interface AgentStepRepository extends JpaRepository<AgentStep, Long> {
|
||||
*/
|
||||
List<AgentStep> findBySessionId(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据运行ID查询所有步骤(按步骤号排序)。
|
||||
*/
|
||||
List<AgentStep> findByRunIdOrderByStepIndex(String runId);
|
||||
|
||||
/**
|
||||
* 统计某个会话的步骤数
|
||||
*/
|
||||
int countBySessionId(String sessionId);
|
||||
|
||||
/**
|
||||
* 统计某个运行的步骤数。
|
||||
*/
|
||||
int countByRunId(String runId);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
package com.superbiz.agent.repository;
|
||||
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import org.springframework.data.jpa.repository.JpaRepository;
|
||||
import org.springframework.stereotype.Repository;
|
||||
|
||||
import java.util.Optional;
|
||||
|
||||
@Repository
|
||||
public interface ChatSessionRepository extends JpaRepository<ChatSession, Long> {
|
||||
|
||||
Optional<ChatSession> findBySessionId(String sessionId);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
package com.superbiz.agent.repository;
|
||||
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import org.springframework.data.jpa.repository.JpaRepository;
|
||||
import org.springframework.stereotype.Repository;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Optional;
|
||||
|
||||
@Repository
|
||||
public interface DiagnosisRunRepository extends JpaRepository<DiagnosisRun, Long> {
|
||||
|
||||
Optional<DiagnosisRun> findByRunId(String runId);
|
||||
|
||||
Optional<DiagnosisRun> findBySessionIdAndRunId(String sessionId, String runId);
|
||||
|
||||
Optional<DiagnosisRun> findFirstBySessionIdOrderByCreatedAtDescIdDesc(String sessionId);
|
||||
|
||||
List<DiagnosisRun> findBySessionIdOrderByCreatedAtDescIdDesc(String sessionId);
|
||||
}
|
||||
|
||||
@@ -22,11 +22,21 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
|
||||
*/
|
||||
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据运行ID按创建顺序查询所有工具调用
|
||||
*/
|
||||
List<ToolInvocation> findByRunIdOrderByIdAsc(String runId);
|
||||
|
||||
/**
|
||||
* 根据会话ID统计真实工具调用次数
|
||||
*/
|
||||
long countBySessionId(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据运行ID统计真实工具调用次数
|
||||
*/
|
||||
long countByRunId(String runId);
|
||||
|
||||
/**
|
||||
* 根据工具名查询所有调用
|
||||
*/
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import org.springframework.ai.chat.model.ChatModel;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||
@@ -13,11 +14,13 @@ import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.dto.AIOpsRequest;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
@@ -65,6 +68,9 @@ public class AiOpsService {
|
||||
@Autowired
|
||||
private DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisRunRepository diagnosisRunRepository;
|
||||
|
||||
@Autowired
|
||||
private AgentStepRepository agentStepRepository;
|
||||
|
||||
@@ -89,46 +95,48 @@ public class AiOpsService {
|
||||
* @throws GraphRunnerException Agent 图执行失败时抛出
|
||||
*/
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null), newRunId());
|
||||
}
|
||||
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
AIOpsRequest request, String sessionId) throws GraphRunnerException {
|
||||
logger.info("Starting AI Ops multi-agent analysis");
|
||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId, newRunId());
|
||||
}
|
||||
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
AIOpsRequest request, String sessionId, String runId) throws GraphRunnerException {
|
||||
logger.info("Starting AI Ops multi-agent analysis");
|
||||
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
||||
String resolvedRunId = isBlank(runId) ? newRunId() : runId.trim();
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
||||
diagnosisSessionRepository.save(session);
|
||||
DiagnosisRun run = startDiagnosisRun(resolvedSessionId, resolvedRunId, request);
|
||||
|
||||
// 让工具调用、Hook 和知识库检索能够拿到当前诊断会话 ID。
|
||||
SessionContextHolder.setSessionId(resolvedSessionId);
|
||||
// 让工具调用、Hook 和知识库检索能够拿到当前诊断会话和运行 ID。
|
||||
SessionContextHolder.setContext(resolvedSessionId, resolvedRunId);
|
||||
|
||||
try {
|
||||
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
|
||||
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
|
||||
|
||||
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
|
||||
.name("ai_ops_supervisor")
|
||||
.description("Coordinates Planner and Executor agents")
|
||||
.model(chatModel)
|
||||
.systemPrompt(promptProperties.getSupervisor())
|
||||
.subAgents(List.of(plannerAgent, executorAgent))
|
||||
.build();
|
||||
SupervisorAgent supervisorAgent = buildSupervisorAgent(chatModel, plannerAgent, executorAgent);
|
||||
|
||||
String taskPrompt = buildTaskPrompt(request);
|
||||
RunnableConfig config = RunnableConfig.builder()
|
||||
.addMetadata("sessionId", resolvedSessionId)
|
||||
.addMetadata("runId", resolvedRunId)
|
||||
.build();
|
||||
|
||||
logger.info("Invoking AI Ops supervisor agent");
|
||||
|
||||
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
|
||||
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt, config);
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
run.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
||||
run.setTotalDurationMs((int) duration);
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
|
||||
if (stateOptional.isPresent()) {
|
||||
OverAllState state = stateOptional.get();
|
||||
@@ -139,8 +147,10 @@ public class AiOpsService {
|
||||
|
||||
return stateOptional;
|
||||
} catch (Exception e) {
|
||||
session.setStatus("FAILED");
|
||||
diagnosisSessionRepository.save(session);
|
||||
run.setStatus("FAILED");
|
||||
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
throw e;
|
||||
} finally {
|
||||
SessionContextHolder.clear();
|
||||
@@ -177,11 +187,34 @@ public class AiOpsService {
|
||||
return UUID.randomUUID().toString();
|
||||
}
|
||||
|
||||
public String newRunId() {
|
||||
return "run-" + UUID.randomUUID();
|
||||
}
|
||||
|
||||
public void persistFinalReport(String sessionId, String finalReport) {
|
||||
persistFinalReport(sessionId, finalReport, null);
|
||||
}
|
||||
|
||||
public void persistFinalReport(String sessionId, String finalReport, AIOpsRequest request) {
|
||||
persistLegacyFinalReport(sessionId, finalReport, request);
|
||||
}
|
||||
|
||||
public void persistFinalReport(String sessionId, String runId, String finalReport, AIOpsRequest request) {
|
||||
if (isBlank(sessionId) || isBlank(runId) || isBlank(finalReport)) {
|
||||
return;
|
||||
}
|
||||
diagnosisRunRepository.findBySessionIdAndRunId(sessionId.trim(), runId.trim()).ifPresent(run -> {
|
||||
run.setAnswer(finalReport);
|
||||
List<com.superbiz.agent.domain.entity.ToolInvocation> invocations =
|
||||
toolInvocationRepository.findByRunIdOrderByIdAsc(run.getRunId());
|
||||
Map<String, Object> evaluation = aiOpsRuleEvaluationService.evaluate(request, finalReport, invocations);
|
||||
run.setSelfEvaluation(selfEvaluationMergeService.mergeAiOpsRuleEvaluation(
|
||||
run.getSelfEvaluation(), evaluation));
|
||||
diagnosisRunRepository.save(run);
|
||||
});
|
||||
}
|
||||
|
||||
private void persistLegacyFinalReport(String sessionId, String finalReport, AIOpsRequest request) {
|
||||
if (isBlank(sessionId) || isBlank(finalReport)) {
|
||||
return;
|
||||
}
|
||||
@@ -261,21 +294,15 @@ public class AiOpsService {
|
||||
return prompt.toString();
|
||||
}
|
||||
|
||||
private DiagnosisSession startDiagnosisSession(String sessionId, AIOpsRequest request) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElseGet(() -> DiagnosisSession.builder()
|
||||
private DiagnosisRun startDiagnosisRun(String sessionId, String runId, AIOpsRequest request) {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.runId(runId)
|
||||
.sessionId(sessionId)
|
||||
.query(buildQuerySummary(request))
|
||||
.status("RUNNING")
|
||||
.agentFlow("AI_OPS")
|
||||
.build());
|
||||
session.setQuery(buildQuerySummary(request));
|
||||
session.setStatus("RUNNING");
|
||||
session.setAgentFlow("AI_OPS");
|
||||
session.setAnswer(null);
|
||||
session.setTotalDurationMs(null);
|
||||
session.setTotalTokenCount(null);
|
||||
session.setStepCount(null);
|
||||
session.setToolCallCount(null);
|
||||
return session;
|
||||
.build();
|
||||
return diagnosisRunRepository.save(run);
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -338,12 +365,23 @@ public class AiOpsService {
|
||||
};
|
||||
}
|
||||
|
||||
SupervisorAgent buildSupervisorAgent(ChatModel chatModel, ReactAgent plannerAgent, ReactAgent executorAgent) {
|
||||
return SupervisorAgent.builder()
|
||||
.name("ai_ops_supervisor")
|
||||
.description("Coordinates Planner and Executor agents")
|
||||
.model(chatModel)
|
||||
.mainAgent(plannerAgent)
|
||||
.systemPrompt(promptProperties.getSupervisor())
|
||||
.subAgents(List.of(executorAgent))
|
||||
.build();
|
||||
}
|
||||
|
||||
/**
|
||||
* 从 agent_step 和 tool_invocation 回填 diagnosis_session 的汇总指标。
|
||||
* 从 agent_step 和 tool_invocation 回填 diagnosis_run 的汇总指标。
|
||||
*/
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
private void backfillRunMetrics(DiagnosisRun run) {
|
||||
try {
|
||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||
List<AgentStep> steps = agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId());
|
||||
|
||||
int totalTokens = 0;
|
||||
int stepCount = 0;
|
||||
@@ -351,12 +389,12 @@ public class AiOpsService {
|
||||
stepCount++;
|
||||
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
||||
}
|
||||
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
|
||||
session.setTotalTokenCount(totalTokens);
|
||||
session.setStepCount(stepCount);
|
||||
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
long toolCallCount = toolInvocationRepository.countByRunId(run.getRunId());
|
||||
run.setTotalTokenCount(totalTokens);
|
||||
run.setStepCount(stepCount);
|
||||
run.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
} catch (Exception e) {
|
||||
logger.warn("Failed to backfill AI Ops session metrics, sessionId={}", session.getSessionId(), e);
|
||||
logger.warn("Failed to backfill AI Ops run metrics, runId={}", run.getRunId(), e);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.CaseLibrary;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.enums.FaultCategory;
|
||||
import com.superbiz.agent.domain.enums.SourceType;
|
||||
@@ -20,22 +21,31 @@ public class CaseLibraryService {
|
||||
@Autowired
|
||||
private CaseLibraryRepository caseLibraryRepository;
|
||||
|
||||
public CaseLibrary createFromSession(DiagnosisSession session) {
|
||||
return caseLibraryRepository.findByDiagnosisId(session.getSessionId())
|
||||
.orElseGet(() -> {
|
||||
String content = session.getAnswer();
|
||||
if (content == null || content.isBlank()) {
|
||||
content = session.getQuery() + "\n(自动提取失败,请人工补充)";
|
||||
public CaseLibrary createFromRun(DiagnosisRun run) {
|
||||
return createFromSource(run.getRunId(), run.getQuery(), run.getAnswer(), "runId=" + run.getRunId());
|
||||
}
|
||||
|
||||
String title = session.getQuery();
|
||||
public CaseLibrary createFromSession(DiagnosisSession session) {
|
||||
return createFromSource(session.getSessionId(), session.getQuery(), session.getAnswer(),
|
||||
"sessionId=" + session.getSessionId());
|
||||
}
|
||||
|
||||
private CaseLibrary createFromSource(String diagnosisId, String query, String answer, String logContext) {
|
||||
return caseLibraryRepository.findByDiagnosisId(diagnosisId)
|
||||
.orElseGet(() -> {
|
||||
String content = answer;
|
||||
if (content == null || content.isBlank()) {
|
||||
content = query + "\n(自动提取失败,请人工补充)";
|
||||
}
|
||||
|
||||
String title = query;
|
||||
if (title.length() > 100) {
|
||||
title = title.substring(0, 100);
|
||||
}
|
||||
|
||||
CaseLibrary caseLibrary = CaseLibrary.builder()
|
||||
.caseId(UUID.randomUUID().toString())
|
||||
.diagnosisId(session.getSessionId())
|
||||
.diagnosisId(diagnosisId)
|
||||
.sourceType(SourceType.AUTO)
|
||||
.faultCategory(FaultCategory.GENERAL)
|
||||
.title(title)
|
||||
@@ -46,7 +56,7 @@ public class CaseLibraryService {
|
||||
.build();
|
||||
|
||||
CaseLibrary saved = caseLibraryRepository.save(caseLibrary);
|
||||
logger.info("案例已沉淀: caseId={}, sessionId={}", saved.getCaseId(), session.getSessionId());
|
||||
logger.info("案例已沉淀: caseId={}, {}", saved.getCaseId(), logContext);
|
||||
return saved;
|
||||
});
|
||||
}
|
||||
|
||||
@@ -14,14 +14,16 @@ import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
||||
import com.superbiz.agent.hook.TokenUsageHolder;
|
||||
import com.superbiz.agent.hook.VerifierInputHook;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ChatSessionRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||
@@ -43,6 +45,7 @@ import org.springframework.stereotype.Service;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
@@ -62,8 +65,8 @@ public class ChatService {
|
||||
private static final String DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论,建议人工介入。";
|
||||
private static final String CHAT_PROMPT_AUDIT_VERSION = "chat-prompts-v1";
|
||||
|
||||
/** 封装 answer + 后端生成的 sessionId,用于 feedback 关联 */
|
||||
public record ChatResult(String answer, String sessionId) {}
|
||||
/** 封装 answer + 后端生成的 sessionId/runId,用于 feedback 关联 */
|
||||
public record ChatResult(String answer, String sessionId, String runId) {}
|
||||
|
||||
@Autowired
|
||||
private InternalDocsTools internalDocsTools;
|
||||
@@ -74,7 +77,7 @@ public class ChatService {
|
||||
@Autowired
|
||||
private QueryMetricsTools queryMetricsTools;
|
||||
|
||||
@Autowired(required = false) // Mock 模式下才注册,所以设置为 optional,真实环境通过mcp配置注入
|
||||
@Autowired(required = false) // Mock 模式下才注册,所以设置为 optional,真实环境通过 mcp 配置注入
|
||||
private QueryLogsTools queryLogsTools;
|
||||
|
||||
@Autowired(required = false)
|
||||
@@ -87,7 +90,10 @@ public class ChatService {
|
||||
private LookupKnowledgeTool lookupKnowledgeTool;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
private ChatSessionRepository chatSessionRepository;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisRunRepository diagnosisRunRepository;
|
||||
|
||||
@Autowired
|
||||
private AgentStepRepository agentStepRepository;
|
||||
@@ -131,7 +137,7 @@ public class ChatService {
|
||||
|
||||
@PostConstruct
|
||||
public void init() {
|
||||
// 加载 Prompt
|
||||
// 鍔犺浇 Prompt
|
||||
try {
|
||||
chatPlannerPrompt = new String(
|
||||
new ClassPathResource("prompts/chat-planner-prompt.md").getInputStream().readAllBytes(),
|
||||
@@ -186,7 +192,7 @@ public class ChatService {
|
||||
String role = msg.get("role");
|
||||
String content = msg.get("content");
|
||||
|
||||
// 🔧 过滤时间查询相关的历史消息,避免 LLM 复用旧的时间信息
|
||||
// 过滤时间查询相关的历史消息,避免 LLM 复用旧的时间信息
|
||||
if ("user".equals(role) && isTimeQuery(content)) {
|
||||
continue; // 跳过时间查询问题
|
||||
}
|
||||
@@ -229,7 +235,7 @@ public class ChatService {
|
||||
return false;
|
||||
}
|
||||
// 匹配日期时间格式:2026年5月31日、15:57、下午3点 等
|
||||
return content.matches(".*(\\d{4}年\\d{1,2}月\\d{1,2}日|\\d{1,2}:\\d{2}|[上下午]+\\d{1,2}[点时]).*");
|
||||
return content.matches(".*(\\d{4}.*\\d{1,2}.*\\d{1,2}.*|\\d{1,2}:\\d{2}|[上下]午\\s*\\d{1,2}[点时]).*");
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -295,7 +301,7 @@ public class ChatService {
|
||||
* 执行 ReactAgent 对话(非流式)
|
||||
* @param agent ReactAgent 实例
|
||||
* @param question 用户问题
|
||||
* @return ChatResult(answer + sessionId)
|
||||
* @return ChatResult(answer + sessionId + runId)
|
||||
*/
|
||||
public ChatResult executeChat(ReactAgent agent, String question) throws GraphRunnerException {
|
||||
return executeChat(agent, question, null);
|
||||
@@ -303,22 +309,24 @@ public class ChatService {
|
||||
|
||||
public ChatResult executeChat(ReactAgent agent, String question, String requestedSessionId) throws GraphRunnerException {
|
||||
logger.info("========================================");
|
||||
logger.info("📝 用户问题: {}", question);
|
||||
logger.info("用户问题: {}", question);
|
||||
|
||||
String sessionId = resolveSessionId(requestedSessionId);
|
||||
String runId = newRunId();
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
// 创建或更新诊断会话
|
||||
DiagnosisSession session = startDiagnosisSession(sessionId, question);
|
||||
diagnosisSessionRepository.save(session);
|
||||
// 创建或更新 Chat Session 元数据,并创建本次诊断 run
|
||||
ensureChatSession(sessionId, null);
|
||||
DiagnosisRun run = startDiagnosisRun(sessionId, runId, question);
|
||||
|
||||
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
||||
SessionContextHolder.setSessionId(sessionId);
|
||||
// 设置 ThreadLocal 上下文,供工具和 Hook 读取 sessionId/runId
|
||||
SessionContextHolder.setContext(sessionId, runId);
|
||||
|
||||
try {
|
||||
// 通过 RunnableConfig 将 sessionId 传入 Hook(线程安全,异步也兼容)
|
||||
// 通过 RunnableConfig 将 sessionId/runId 传入 Hook
|
||||
var config = RunnableConfig.builder()
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();
|
||||
|
||||
var response = agent.call(question, config);
|
||||
@@ -326,23 +334,27 @@ public class ChatService {
|
||||
|
||||
String answer = response.getText();
|
||||
|
||||
// 更新诊断会话
|
||||
session.setStatus("SUCCESS");
|
||||
session.setAnswer(answer);
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
// 更新诊断 run
|
||||
run.setStatus("SUCCESS");
|
||||
run.setAnswer(answer);
|
||||
run.setTotalDurationMs((int) duration);
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
|
||||
evaluationService.evaluate(sessionId, answer);
|
||||
evaluationService.evaluateRun(runId, answer);
|
||||
|
||||
logger.info("⏱️ 总耗时: {} ms", duration);
|
||||
logger.info("📏 输出长度: {} 字符", answer.length());
|
||||
logger.info("总耗时: {} ms", duration);
|
||||
logger.info("输出长度: {} 字符", answer.length());
|
||||
logger.info("========================================");
|
||||
|
||||
return new ChatResult(answer, sessionId);
|
||||
return new ChatResult(answer, sessionId, runId);
|
||||
} catch (Exception e) {
|
||||
session.setStatus("FAILED");
|
||||
diagnosisSessionRepository.save(session);
|
||||
String errorAnswer = "Execution failed: " + e.getMessage();
|
||||
run.setStatus("FAILED");
|
||||
run.setAnswer(errorAnswer);
|
||||
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
throw e;
|
||||
} finally {
|
||||
retrievedDocTracker.clearSession(sessionId);
|
||||
@@ -356,7 +368,7 @@ public class ChatService {
|
||||
* @param toolCallbacks 工具回调
|
||||
* @param question 用户问题
|
||||
* @param history 历史消息
|
||||
* @return ChatResult(answer + sessionId)
|
||||
* @return ChatResult(answer + sessionId + runId)
|
||||
*/
|
||||
public ChatResult executeChatWithStrategy(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
String question, List<Map<String, String>> history) throws GraphRunnerException {
|
||||
@@ -367,10 +379,10 @@ public class ChatService {
|
||||
String question, List<Map<String, String>> history,
|
||||
String requestedSessionId) throws GraphRunnerException {
|
||||
if (QuestionComplexity.isComplex(question)) {
|
||||
logger.info("📊 问题判定为复杂,使用多 Agent(Planner + Executor)执行");
|
||||
logger.info("问题判定为复杂,使用多 Agent(Planner + Executor)执行");
|
||||
return executeChatComplex(chatModel, toolCallbacks, question, history, requestedSessionId);
|
||||
} else {
|
||||
logger.info("📊 问题判定为简单,使用单 Agent 执行");
|
||||
logger.info("问题判定为简单,使用单 Agent 执行");
|
||||
String systemPrompt = buildSystemPrompt(history);
|
||||
ReactAgent agent = createReactAgent(chatModel, systemPrompt);
|
||||
return executeChat(agent, question, requestedSessionId);
|
||||
@@ -389,12 +401,13 @@ public class ChatService {
|
||||
String question, List<Map<String, String>> history,
|
||||
String requestedSessionId) throws GraphRunnerException {
|
||||
String sessionId = resolveSessionId(requestedSessionId);
|
||||
String runId = newRunId();
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
DiagnosisSession session = startDiagnosisSession(sessionId, question);
|
||||
diagnosisSessionRepository.save(session);
|
||||
ensureChatSession(sessionId, history == null ? null : history.size() / 2);
|
||||
DiagnosisRun run = startDiagnosisRun(sessionId, runId, question);
|
||||
|
||||
SessionContextHolder.setSessionId(sessionId);
|
||||
SessionContextHolder.setContext(sessionId, runId);
|
||||
VerifierContextHolder.setOriginalQuery(question);
|
||||
VerifierContextHolder.setRetryContext(null);
|
||||
VerifierContextHolder.setExecutorFinalAnswer(null);
|
||||
@@ -405,6 +418,7 @@ public class ChatService {
|
||||
String answer = null;
|
||||
RunnableConfig config = RunnableConfig.builder()
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();
|
||||
|
||||
for (int round = 1; round <= 2; round++) {
|
||||
@@ -427,7 +441,7 @@ public class ChatService {
|
||||
finalDecision = buildVerifierFallbackDecision(round, "workflow 未返回有效状态");
|
||||
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
|
||||
answer = renderResult.answer();
|
||||
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
|
||||
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
|
||||
break;
|
||||
}
|
||||
|
||||
@@ -449,21 +463,21 @@ public class ChatService {
|
||||
finalDecision = buildVerifierFallbackDecision(round, "verifier_output 缺失或无法解析");
|
||||
ComposerRenderResult renderResult = buildFixedFallbackAnswer(question, finalDecision);
|
||||
answer = renderResult.answer();
|
||||
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
|
||||
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
|
||||
break;
|
||||
}
|
||||
|
||||
if ("PASS".equals(finalDecision.verdict())) {
|
||||
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
|
||||
answer = renderResult.answer();
|
||||
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
|
||||
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
|
||||
break;
|
||||
}
|
||||
|
||||
if ("REJECT".equals(finalDecision.verdict())) {
|
||||
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
|
||||
answer = renderResult.answer();
|
||||
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
|
||||
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
|
||||
break;
|
||||
}
|
||||
|
||||
@@ -473,12 +487,12 @@ public class ChatService {
|
||||
if (!shouldRetry) {
|
||||
ComposerRenderResult renderResult = composeFinalAnswer(chatModel, question, finalDecision, config);
|
||||
answer = renderResult.answer();
|
||||
persistVerifierEvaluation(session, finalDecision, round, renderResult.audit());
|
||||
persistVerifierEvaluation(run, finalDecision, round, renderResult.audit());
|
||||
break;
|
||||
}
|
||||
|
||||
retryContext = buildRetryContext(finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
persistVerifierEvaluation(run, finalDecision, round);
|
||||
}
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
@@ -487,24 +501,28 @@ public class ChatService {
|
||||
answer = "抱歉,多 Agent 分析未能生成有效结论。";
|
||||
}
|
||||
|
||||
session.setStatus("SUCCESS");
|
||||
session.setAnswer(answer);
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
run.setStatus("SUCCESS");
|
||||
run.setAnswer(answer);
|
||||
run.setTotalDurationMs((int) duration);
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
|
||||
evaluationService.evaluate(sessionId, answer);
|
||||
evaluationService.evaluateRun(runId, answer);
|
||||
|
||||
logger.info("⏱️ 多 Agent 总耗时: {} ms", duration);
|
||||
logger.info("📏 输出长度: {} 字符", answer.length());
|
||||
logger.info("多 Agent 总耗时: {} ms", duration);
|
||||
logger.info("输出长度: {} 字符", answer.length());
|
||||
|
||||
return new ChatResult(answer, sessionId);
|
||||
return new ChatResult(answer, sessionId, runId);
|
||||
|
||||
} catch (Exception e) {
|
||||
session.setStatus("FAILED");
|
||||
diagnosisSessionRepository.save(session);
|
||||
String errorAnswer = "Execution failed: " + e.getMessage();
|
||||
run.setStatus("FAILED");
|
||||
run.setAnswer(errorAnswer);
|
||||
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
|
||||
backfillRunMetrics(run);
|
||||
diagnosisRunRepository.save(run);
|
||||
logger.error("多 Agent 执行失败", e);
|
||||
return new ChatResult("执行失败: " + e.getMessage(), sessionId);
|
||||
return new ChatResult(errorAnswer, sessionId, runId);
|
||||
} finally {
|
||||
retrievedDocTracker.clearSession(sessionId);
|
||||
SessionContextHolder.clear();
|
||||
@@ -516,7 +534,7 @@ public class ChatService {
|
||||
String retryContext) {
|
||||
StringBuilder prompt = new StringBuilder(chatPlannerPrompt);
|
||||
|
||||
// 注入 knowledge map
|
||||
// 娉ㄥ叆 knowledge map
|
||||
String knowledgeMap = knowledgeDomainService.buildKnowledgeMap();
|
||||
if (!knowledgeMap.isBlank()) {
|
||||
prompt.append("\n\n## 可用知识库\n\n").append(knowledgeMap);
|
||||
@@ -530,7 +548,7 @@ public class ChatService {
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
prompt.append("\n\n--- 鏈疆琛ヨ瘉鎹害鏉?---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_planner")
|
||||
@@ -580,7 +598,7 @@ public class ChatService {
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
prompt.append("\n\n--- 鏈疆琛ヨ瘉鎹害鏉?---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_executor")
|
||||
@@ -620,21 +638,40 @@ public class ChatService {
|
||||
return UUID.randomUUID().toString().substring(0, 8);
|
||||
}
|
||||
|
||||
private DiagnosisSession startDiagnosisSession(String sessionId, String question) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElseGet(() -> DiagnosisSession.builder()
|
||||
private String newRunId() {
|
||||
return "run-" + UUID.randomUUID();
|
||||
}
|
||||
|
||||
public void syncChatSessionMetadata(String sessionId, Integer messagePairCount) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return;
|
||||
}
|
||||
ensureChatSession(sessionId, messagePairCount);
|
||||
}
|
||||
|
||||
private ChatSession ensureChatSession(String sessionId, Integer messagePairCount) {
|
||||
ChatSession chatSession = chatSessionRepository.findBySessionId(sessionId)
|
||||
.orElseGet(() -> ChatSession.builder()
|
||||
.sessionId(sessionId)
|
||||
.agentFlow("CHAT")
|
||||
.status("ACTIVE")
|
||||
.build());
|
||||
session.setQuery(question);
|
||||
session.setStatus("RUNNING");
|
||||
session.setAgentFlow("CHAT");
|
||||
session.setAnswer(null);
|
||||
session.setTotalDurationMs(null);
|
||||
session.setTotalTokenCount(null);
|
||||
session.setStepCount(null);
|
||||
session.setToolCallCount(null);
|
||||
return session;
|
||||
chatSession.setStatus("ACTIVE");
|
||||
chatSession.setLastActiveAt(LocalDateTime.now());
|
||||
if (messagePairCount != null) {
|
||||
chatSession.setMessagePairCount(messagePairCount);
|
||||
}
|
||||
return chatSessionRepository.save(chatSession);
|
||||
}
|
||||
|
||||
private DiagnosisRun startDiagnosisRun(String sessionId, String runId, String question) {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.runId(runId)
|
||||
.sessionId(sessionId)
|
||||
.query(question)
|
||||
.status("RUNNING")
|
||||
.agentFlow("CHAT")
|
||||
.build();
|
||||
return diagnosisRunRepository.save(run);
|
||||
}
|
||||
|
||||
private String buildWorkflowInput(String question, String retryContext) {
|
||||
@@ -867,11 +904,11 @@ public class ChatService {
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round) {
|
||||
persistVerifierEvaluation(session, decision, round, null);
|
||||
private void persistVerifierEvaluation(DiagnosisRun run, VerifierDecision decision, int round) {
|
||||
persistVerifierEvaluation(run, decision, round, null);
|
||||
}
|
||||
|
||||
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round,
|
||||
private void persistVerifierEvaluation(DiagnosisRun run, VerifierDecision decision, int round,
|
||||
Map<String, Object> composerOutput) {
|
||||
if (decision == null) {
|
||||
return;
|
||||
@@ -899,9 +936,9 @@ public class ChatService {
|
||||
verifierEvaluation.put("composer_output", composerOutput);
|
||||
}
|
||||
|
||||
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(session.getSelfEvaluation(), verifierEvaluation);
|
||||
session.setSelfEvaluation(merged);
|
||||
diagnosisSessionRepository.save(session);
|
||||
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(run.getSelfEvaluation(), verifierEvaluation);
|
||||
run.setSelfEvaluation(merged);
|
||||
diagnosisRunRepository.save(run);
|
||||
}
|
||||
|
||||
private Map<String, Object> promptAuditSnapshot() {
|
||||
@@ -1244,7 +1281,10 @@ public class ChatService {
|
||||
|
||||
private List<String> buildNextStepSuggestionsFromTrace() {
|
||||
List<String> suggestions = new ArrayList<>();
|
||||
List<Map<String, Object>> toolSummary = toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null);
|
||||
String runId = SessionContextHolder.getRunId();
|
||||
List<Map<String, Object>> toolSummary = runId == null || runId.isBlank()
|
||||
? toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null)
|
||||
: toolTraceSummaryService.buildVerifierTraceSummaryForRun(runId, null);
|
||||
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
|
||||
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
|
||||
|
||||
@@ -1309,11 +1349,11 @@ public class ChatService {
|
||||
private record ComposerRenderResult(String answer, Map<String, Object> audit) {
|
||||
}
|
||||
|
||||
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_run */
|
||||
private void backfillRunMetrics(DiagnosisRun run) {
|
||||
try {
|
||||
List<com.superbiz.agent.domain.entity.AgentStep> steps =
|
||||
agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||
agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId());
|
||||
|
||||
int totalTokens = 0;
|
||||
int stepCount = 0;
|
||||
@@ -1321,12 +1361,12 @@ public class ChatService {
|
||||
stepCount++;
|
||||
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
||||
}
|
||||
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
|
||||
session.setTotalTokenCount(totalTokens);
|
||||
session.setStepCount(stepCount);
|
||||
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
long toolCallCount = toolInvocationRepository.countByRunId(run.getRunId());
|
||||
run.setTotalTokenCount(totalTokens);
|
||||
run.setStepCount(stepCount);
|
||||
run.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
} catch (Exception e) {
|
||||
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
||||
logger.warn("Failed to backfill run metrics: runId={}", run.getRunId(), e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,11 +3,15 @@ package com.superbiz.agent.service;
|
||||
import com.fasterxml.jackson.core.type.TypeReference;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||
import com.superbiz.agent.exception.SessionNotFoundException;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.ChatSessionRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import lombok.RequiredArgsConstructor;
|
||||
@@ -25,18 +29,76 @@ public class DiagnosisTraceService {
|
||||
};
|
||||
|
||||
private final DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
private final ChatSessionRepository chatSessionRepository;
|
||||
private final DiagnosisRunRepository diagnosisRunRepository;
|
||||
private final AgentStepRepository agentStepRepository;
|
||||
private final ToolInvocationRepository toolInvocationRepository;
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
public DiagnosisTraceResponse getTrace(String sessionId) {
|
||||
return getTrace(sessionId, null);
|
||||
}
|
||||
|
||||
public DiagnosisTraceResponse getTrace(String sessionId, String runId) {
|
||||
if (runId != null && !runId.isBlank()) {
|
||||
DiagnosisRun run = diagnosisRunRepository.findBySessionIdAndRunId(sessionId, runId)
|
||||
.orElseThrow(() -> buildRunLookupException(sessionId, runId));
|
||||
return buildRunTraceResponse(sessionId, run);
|
||||
}
|
||||
|
||||
return diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId)
|
||||
.map(run -> buildRunTraceResponse(sessionId, run))
|
||||
.orElseGet(() -> buildLegacyTraceResponse(sessionId));
|
||||
}
|
||||
|
||||
public List<DiagnosisTraceResponse.RunSummary> listRunSummaries(String sessionId) {
|
||||
List<DiagnosisRun> runs = diagnosisRunRepository.findBySessionIdOrderByCreatedAtDescIdDesc(sessionId);
|
||||
if (runs.isEmpty() && chatSessionRepository.findBySessionId(sessionId).isEmpty()) {
|
||||
throw new SessionNotFoundException(sessionId);
|
||||
}
|
||||
return runs.stream()
|
||||
.map(this::toRunSummary)
|
||||
.toList();
|
||||
}
|
||||
|
||||
private RuntimeException buildRunLookupException(String sessionId, String runId) {
|
||||
if (diagnosisRunRepository.findByRunId(runId).isPresent()) {
|
||||
return new IllegalArgumentException("runId does not belong to sessionId: " + runId);
|
||||
}
|
||||
return new SessionNotFoundException(sessionId, "Run not found: " + runId);
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse buildRunTraceResponse(String sessionId, DiagnosisRun run) {
|
||||
List<AgentStep> steps = orderStepsForTrace(agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId()));
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findByRunIdOrderByIdAsc(run.getRunId());
|
||||
DiagnosisTraceResponse.RunTrace runTrace = toRunTrace(run);
|
||||
|
||||
return DiagnosisTraceResponse.builder()
|
||||
.runId(run.getRunId())
|
||||
.chatSession(chatSessionRepository.findBySessionId(sessionId)
|
||||
.map(this::toChatSessionTrace)
|
||||
.orElse(null))
|
||||
.session(toSessionTrace(runTrace))
|
||||
.run(runTrace)
|
||||
.steps(steps.stream().map(this::toAgentStepTrace).toList())
|
||||
.toolInvocations(toolInvocations.stream().map(this::toToolInvocationTrace).toList())
|
||||
.summary(toSummary(runTrace, steps, toolInvocations))
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse buildLegacyTraceResponse(String sessionId) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElseThrow(() -> new SessionNotFoundException(sessionId));
|
||||
List<AgentStep> steps = orderStepsForTrace(agentStepRepository.findBySessionId(sessionId));
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||
|
||||
return DiagnosisTraceResponse.builder()
|
||||
.runId(null)
|
||||
.chatSession(chatSessionRepository.findBySessionId(sessionId)
|
||||
.map(this::toChatSessionTrace)
|
||||
.orElse(null))
|
||||
.session(toSessionTrace(session))
|
||||
.run(null)
|
||||
.steps(steps.stream().map(this::toAgentStepTrace).toList())
|
||||
.toolInvocations(toolInvocations.stream().map(this::toToolInvocationTrace).toList())
|
||||
.summary(toSummary(session, steps, toolInvocations))
|
||||
@@ -52,6 +114,18 @@ public class DiagnosisTraceService {
|
||||
.toList();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.ChatSessionTrace toChatSessionTrace(ChatSession chatSession) {
|
||||
return DiagnosisTraceResponse.ChatSessionTrace.builder()
|
||||
.id(chatSession.getId())
|
||||
.sessionId(chatSession.getSessionId())
|
||||
.status(chatSession.getStatus())
|
||||
.messagePairCount(chatSession.getMessagePairCount())
|
||||
.createdAt(chatSession.getCreatedAt())
|
||||
.lastActiveAt(chatSession.getLastActiveAt())
|
||||
.expiresAt(chatSession.getExpiresAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.SessionTrace toSessionTrace(DiagnosisSession session) {
|
||||
return DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.id(session.getId())
|
||||
@@ -72,10 +146,52 @@ public class DiagnosisTraceService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.SessionTrace toSessionTrace(DiagnosisTraceResponse.RunTrace run) {
|
||||
return DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.id(run.getId())
|
||||
.sessionId(run.getSessionId())
|
||||
.query(run.getQuery())
|
||||
.status(run.getStatus())
|
||||
.agentFlow(run.getAgentFlow())
|
||||
.totalDurationMs(run.getTotalDurationMs())
|
||||
.totalTokenCount(run.getTotalTokenCount())
|
||||
.stepCount(run.getStepCount())
|
||||
.toolCallCount(run.getToolCallCount())
|
||||
.answer(run.getAnswer())
|
||||
.selfEvaluationRaw(run.getSelfEvaluationRaw())
|
||||
.selfEvaluation(run.getSelfEvaluation())
|
||||
.feedback(run.getFeedback())
|
||||
.createdAt(run.getCreatedAt())
|
||||
.updatedAt(run.getUpdatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.RunTrace toRunTrace(DiagnosisRun run) {
|
||||
return DiagnosisTraceResponse.RunTrace.builder()
|
||||
.id(run.getId())
|
||||
.runId(run.getRunId())
|
||||
.sessionId(run.getSessionId())
|
||||
.query(run.getQuery())
|
||||
.status(run.getStatus())
|
||||
.agentFlow(run.getAgentFlow())
|
||||
.totalDurationMs(run.getTotalDurationMs())
|
||||
.totalTokenCount(run.getTotalTokenCount())
|
||||
.stepCount(run.getStepCount())
|
||||
.toolCallCount(run.getToolCallCount())
|
||||
.answer(run.getAnswer())
|
||||
.selfEvaluationRaw(run.getSelfEvaluation())
|
||||
.selfEvaluation(parseJsonObject(run.getSelfEvaluation()))
|
||||
.feedback(run.getFeedback())
|
||||
.createdAt(run.getCreatedAt())
|
||||
.updatedAt(run.getUpdatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.AgentStepTrace toAgentStepTrace(AgentStep step) {
|
||||
return DiagnosisTraceResponse.AgentStepTrace.builder()
|
||||
.id(step.getId())
|
||||
.sessionId(step.getSessionId())
|
||||
.runId(step.getRunId())
|
||||
.stepIndex(step.getStepIndex())
|
||||
.agentName(step.getAgentName())
|
||||
.modelInput(step.getModelInput())
|
||||
@@ -92,6 +208,7 @@ public class DiagnosisTraceService {
|
||||
return DiagnosisTraceResponse.ToolInvocationTrace.builder()
|
||||
.id(invocation.getId())
|
||||
.sessionId(invocation.getSessionId())
|
||||
.runId(invocation.getRunId())
|
||||
.stepId(invocation.getStepId())
|
||||
.toolName(invocation.getToolName())
|
||||
.inputParamsRaw(invocation.getInputParams())
|
||||
@@ -113,6 +230,21 @@ public class DiagnosisTraceService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.RunSummary toRunSummary(DiagnosisRun run) {
|
||||
return DiagnosisTraceResponse.RunSummary.builder()
|
||||
.runId(run.getRunId())
|
||||
.sessionId(run.getSessionId())
|
||||
.query(run.getQuery())
|
||||
.status(run.getStatus())
|
||||
.agentFlow(run.getAgentFlow())
|
||||
.answerPreview(preview(run.getAnswer(), 160))
|
||||
.stepCount(run.getStepCount())
|
||||
.toolCallCount(run.getToolCallCount())
|
||||
.createdAt(run.getCreatedAt())
|
||||
.updatedAt(run.getUpdatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.TraceSummary toSummary(
|
||||
DiagnosisSession session,
|
||||
List<AgentStep> steps,
|
||||
@@ -120,6 +252,7 @@ public class DiagnosisTraceService {
|
||||
) {
|
||||
Map<String, Object> selfEvaluation = parseJsonObject(session.getSelfEvaluation());
|
||||
return DiagnosisTraceResponse.TraceSummary.builder()
|
||||
.resolvedRunId(null)
|
||||
.persistedStepCount(defaultInt(session.getStepCount()))
|
||||
.returnedStepCount(steps.size())
|
||||
.persistedToolCallCount(defaultInt(session.getToolCallCount()))
|
||||
@@ -130,6 +263,24 @@ public class DiagnosisTraceService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.TraceSummary toSummary(
|
||||
DiagnosisTraceResponse.RunTrace run,
|
||||
List<AgentStep> steps,
|
||||
List<ToolInvocation> toolInvocations
|
||||
) {
|
||||
Map<String, Object> selfEvaluation = run.getSelfEvaluation();
|
||||
return DiagnosisTraceResponse.TraceSummary.builder()
|
||||
.resolvedRunId(run.getRunId())
|
||||
.persistedStepCount(defaultInt(run.getStepCount()))
|
||||
.returnedStepCount(steps.size())
|
||||
.persistedToolCallCount(defaultInt(run.getToolCallCount()))
|
||||
.returnedToolCallCount(toolInvocations.size())
|
||||
.hasVerifierEvaluation(selfEvaluation != null && selfEvaluation.containsKey("verifier_evaluation"))
|
||||
.hasAiOpsRuleEvaluation(selfEvaluation != null && selfEvaluation.containsKey("aiops_rule_evaluation"))
|
||||
.hasFeedback(run.getFeedback() != null && !run.getFeedback().isBlank())
|
||||
.build();
|
||||
}
|
||||
|
||||
private Map<String, Object> parseJsonObject(String json) {
|
||||
if (json == null || json.isBlank()) {
|
||||
return null;
|
||||
@@ -144,4 +295,14 @@ public class DiagnosisTraceService {
|
||||
private int defaultInt(Integer value) {
|
||||
return value == null ? 0 : value;
|
||||
}
|
||||
|
||||
private String preview(String text, int maxLength) {
|
||||
if (text == null) {
|
||||
return null;
|
||||
}
|
||||
if (text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,7 +1,9 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import org.slf4j.Logger;
|
||||
@@ -29,6 +31,9 @@ public class EvaluationService {
|
||||
@Autowired
|
||||
private DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisRunRepository diagnosisRunRepository;
|
||||
|
||||
@Autowired
|
||||
private ToolInvocationRepository toolInvocationRepository;
|
||||
|
||||
@@ -40,7 +45,7 @@ public class EvaluationService {
|
||||
diagnosisSessionRepository.findBySessionId(sessionId).ifPresent(session -> {
|
||||
try {
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findBySessionId(sessionId);
|
||||
Map<String, Object> ruleEvaluation = evaluateWithRules(session, toolInvocations);
|
||||
Map<String, Object> ruleEvaluation = evaluateWithRules(session.getStatus(), toolInvocations);
|
||||
String merged = selfEvaluationMergeService.mergeRuleEvaluation(session.getSelfEvaluation(), ruleEvaluation);
|
||||
session.setSelfEvaluation(merged);
|
||||
diagnosisSessionRepository.save(session);
|
||||
@@ -51,14 +56,30 @@ public class EvaluationService {
|
||||
});
|
||||
}
|
||||
|
||||
@Async
|
||||
public void evaluateRun(String runId, String answer) {
|
||||
diagnosisRunRepository.findByRunId(runId).ifPresent(run -> {
|
||||
try {
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findByRunIdOrderByIdAsc(runId);
|
||||
Map<String, Object> ruleEvaluation = evaluateWithRules(run.getStatus(), toolInvocations);
|
||||
String merged = selfEvaluationMergeService.mergeRuleEvaluation(run.getSelfEvaluation(), ruleEvaluation);
|
||||
run.setSelfEvaluation(merged);
|
||||
diagnosisRunRepository.save(run);
|
||||
logger.info("证据评分已写入: runId={}, result={}", runId, merged);
|
||||
} catch (Exception e) {
|
||||
logger.error("评分失败: runId={}", runId, e);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// -------------------------------------------------------------------------
|
||||
// 规则引擎(事实层)
|
||||
// -------------------------------------------------------------------------
|
||||
|
||||
private Map<String, Object> evaluateWithRules(DiagnosisSession session, List<ToolInvocation> invocations) {
|
||||
private Map<String, Object> evaluateWithRules(String status, List<ToolInvocation> invocations) {
|
||||
List<Map<String, Object>> factors = new ArrayList<>();
|
||||
|
||||
if ("FAILED".equals(session.getStatus())) {
|
||||
if ("FAILED".equals(status)) {
|
||||
factors.add(factor("execution_failed", -100, "执行失败"));
|
||||
return buildResult(0, factors);
|
||||
}
|
||||
|
||||
@@ -61,7 +61,25 @@ public class ExecutorGatekeeperService {
|
||||
GatekeeperResult result = new GatekeeperResult(ruleCatalog);
|
||||
validateSchema(structuredOutput, parseStatus, result);
|
||||
if (structuredOutput != null) {
|
||||
validateInvocationRefs(sessionId, structuredOutput, result);
|
||||
List<ToolInvocation> invocations = sessionId == null || sessionId.isBlank()
|
||||
? List.of()
|
||||
: toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||
validateInvocationRefs("session_id", sessionId, invocations, structuredOutput, result);
|
||||
importWarnings(structuredOutput, result);
|
||||
}
|
||||
return result.toMap();
|
||||
}
|
||||
|
||||
public Map<String, Object> validateRun(String runId,
|
||||
Map<String, Object> structuredOutput,
|
||||
Map<String, Object> parseStatus) {
|
||||
GatekeeperResult result = new GatekeeperResult(ruleCatalog);
|
||||
validateSchema(structuredOutput, parseStatus, result);
|
||||
if (structuredOutput != null) {
|
||||
List<ToolInvocation> invocations = runId == null || runId.isBlank()
|
||||
? List.of()
|
||||
: toolInvocationRepository.findByRunIdOrderByIdAsc(runId);
|
||||
validateInvocationRefs("run_id", runId, invocations, structuredOutput, result);
|
||||
importWarnings(structuredOutput, result);
|
||||
}
|
||||
return result.toMap();
|
||||
@@ -134,15 +152,18 @@ public class ExecutorGatekeeperService {
|
||||
requireArray(structuredOutput, "missing_info", result);
|
||||
}
|
||||
|
||||
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_id",
|
||||
private void validateInvocationRefs(String scopeName,
|
||||
String scopeId,
|
||||
List<ToolInvocation> invocations,
|
||||
Map<String, Object> structuredOutput,
|
||||
GatekeeperResult result) {
|
||||
if (scopeId == null || scopeId.isBlank()) {
|
||||
result.fail(RULE_INVOCATION_REF, scopeName, scopeName + " is required to validate source_invocation_id",
|
||||
SEVERITY_LOW_CONFID);
|
||||
return;
|
||||
}
|
||||
|
||||
Map<Long, ToolInvocation> validInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)
|
||||
.stream()
|
||||
Map<Long, ToolInvocation> validInvocations = invocations.stream()
|
||||
.filter(invocation -> invocation.getId() != null)
|
||||
.collect(Collectors.toMap(ToolInvocation::getId, Function.identity(), (left, right) -> left));
|
||||
Object claimsValue = structuredOutput.get("claims");
|
||||
|
||||
@@ -1,7 +1,9 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.dto.FeedbackResponse;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -19,20 +21,81 @@ public class FeedbackService {
|
||||
@Autowired
|
||||
private DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisRunRepository diagnosisRunRepository;
|
||||
|
||||
@Autowired
|
||||
private CaseLibraryService caseLibraryService;
|
||||
|
||||
public FeedbackResponse submitFeedback(String sessionId, String feedback) {
|
||||
return submitFeedback(sessionId, null, feedback);
|
||||
}
|
||||
|
||||
public FeedbackResponse submitFeedback(String sessionId, String runId, String feedback) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return FeedbackResponse.builder().success(false).message("sessionId 不能为空").build();
|
||||
}
|
||||
if (!FEEDBACK_USEFUL.equals(feedback) && !FEEDBACK_NOT_USEFUL.equals(feedback)) {
|
||||
return FeedbackResponse.builder().success(false)
|
||||
.message("feedback 只能是 useful 或 not_useful").build();
|
||||
return FeedbackResponse.builder()
|
||||
.success(false)
|
||||
.message("feedback 只能是 useful 或 not_useful")
|
||||
.build();
|
||||
}
|
||||
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElse(null);
|
||||
if (runId != null && !runId.isBlank()) {
|
||||
return submitRunFeedback(sessionId, runId, feedback, false);
|
||||
}
|
||||
|
||||
return diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId)
|
||||
.map(run -> submitRunFeedback(sessionId, run.getRunId(), feedback, true))
|
||||
.orElseGet(() -> submitLegacySessionFeedback(sessionId, feedback));
|
||||
}
|
||||
|
||||
private FeedbackResponse submitRunFeedback(String sessionId,
|
||||
String runId,
|
||||
String feedback,
|
||||
boolean fallbackToLatestRun) {
|
||||
DiagnosisRun run = diagnosisRunRepository.findBySessionIdAndRunId(sessionId, runId).orElse(null);
|
||||
if (run == null) {
|
||||
if (diagnosisRunRepository.findByRunId(runId).isPresent()) {
|
||||
return FeedbackResponse.builder()
|
||||
.success(false)
|
||||
.message("runId does not belong to sessionId")
|
||||
.runId(runId)
|
||||
.fallbackToLatestRun(fallbackToLatestRun)
|
||||
.build();
|
||||
}
|
||||
return FeedbackResponse.builder()
|
||||
.success(false)
|
||||
.message("run 不存在")
|
||||
.runId(runId)
|
||||
.fallbackToLatestRun(fallbackToLatestRun)
|
||||
.build();
|
||||
}
|
||||
|
||||
run.setFeedback(feedback);
|
||||
|
||||
String caseId = null;
|
||||
if (FEEDBACK_USEFUL.equals(feedback)) {
|
||||
var caseLibrary = caseLibraryService.createFromRun(run);
|
||||
caseId = caseLibrary.getCaseId();
|
||||
}
|
||||
|
||||
diagnosisRunRepository.save(run);
|
||||
logger.info("反馈已记录: sessionId={}, runId={}, feedback={}, fallbackToLatestRun={}, caseId={}",
|
||||
sessionId, runId, feedback, fallbackToLatestRun, caseId);
|
||||
|
||||
return FeedbackResponse.builder()
|
||||
.success(true)
|
||||
.message("反馈已记录")
|
||||
.caseId(caseId)
|
||||
.runId(runId)
|
||||
.fallbackToLatestRun(fallbackToLatestRun)
|
||||
.build();
|
||||
}
|
||||
|
||||
private FeedbackResponse submitLegacySessionFeedback(String sessionId, String feedback) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId).orElse(null);
|
||||
if (session == null) {
|
||||
return FeedbackResponse.builder().success(false).message("会话不存在").build();
|
||||
}
|
||||
@@ -46,12 +109,13 @@ public class FeedbackService {
|
||||
}
|
||||
|
||||
diagnosisSessionRepository.save(session);
|
||||
logger.info("反馈已记录: sessionId={}, feedback={}, caseId={}", sessionId, feedback, caseId);
|
||||
logger.info("历史反馈已记录: sessionId={}, feedback={}, caseId={}", sessionId, feedback, caseId);
|
||||
|
||||
return FeedbackResponse.builder()
|
||||
.success(true)
|
||||
.message("反馈已记录")
|
||||
.caseId(caseId)
|
||||
.fallbackToLatestRun(false)
|
||||
.build();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -48,6 +48,9 @@ public class ToolInvocationRecorder {
|
||||
if (invocation.getSessionId() == null || invocation.getSessionId().isBlank()) {
|
||||
invocation.setSessionId(SessionContextHolder.getSessionId());
|
||||
}
|
||||
if (invocation.getRunId() == null || invocation.getRunId().isBlank()) {
|
||||
invocation.setRunId(SessionContextHolder.getRunId());
|
||||
}
|
||||
if (invocation.getSessionId() == null || invocation.getSessionId().isBlank()) {
|
||||
log.debug("Skip tool_invocation without sessionId: tool={}", invocation.getToolName());
|
||||
return;
|
||||
|
||||
@@ -42,6 +42,19 @@ public class ToolTraceSummaryService {
|
||||
}
|
||||
|
||||
List<ToolInvocation> invocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||
return buildVerifierTraceSummary(invocations, executorFinalAnswer);
|
||||
}
|
||||
|
||||
public List<Map<String, Object>> buildVerifierTraceSummaryForRun(String runId, String executorFinalAnswer) {
|
||||
if (runId == null || runId.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
List<ToolInvocation> invocations = toolInvocationRepository.findByRunIdOrderByIdAsc(runId);
|
||||
return buildVerifierTraceSummary(invocations, executorFinalAnswer);
|
||||
}
|
||||
|
||||
private List<Map<String, Object>> buildVerifierTraceSummary(List<ToolInvocation> invocations, String executorFinalAnswer) {
|
||||
if (invocations.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
@@ -14,16 +14,31 @@ package com.superbiz.agent.util;
|
||||
public class SessionContextHolder {
|
||||
|
||||
private static final ThreadLocal<String> SESSION_ID = new ThreadLocal<>();
|
||||
private static final ThreadLocal<String> RUN_ID = new ThreadLocal<>();
|
||||
|
||||
public static void setContext(String sessionId, String runId) {
|
||||
setSessionId(sessionId);
|
||||
setRunId(runId);
|
||||
}
|
||||
|
||||
public static void setSessionId(String sessionId) {
|
||||
SESSION_ID.set(sessionId);
|
||||
}
|
||||
|
||||
public static void setRunId(String runId) {
|
||||
RUN_ID.set(runId);
|
||||
}
|
||||
|
||||
public static String getSessionId() {
|
||||
return SESSION_ID.get();
|
||||
}
|
||||
|
||||
public static String getRunId() {
|
||||
return RUN_ID.get();
|
||||
}
|
||||
|
||||
public static void clear() {
|
||||
SESSION_ID.remove();
|
||||
RUN_ID.remove();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
-- V011: split conversation metadata from diagnosis execution runs.
|
||||
|
||||
CREATE TABLE chat_session (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
session_id VARCHAR(64) UNIQUE NOT NULL COMMENT 'Chat session id / conversation context id',
|
||||
status VARCHAR(16) DEFAULT 'ACTIVE' COMMENT 'ACTIVE/EXPIRED/CLOSED',
|
||||
message_pair_count INT DEFAULT 0 COMMENT 'Cached Redis message pair count snapshot',
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
last_active_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
expires_at DATETIME COMMENT 'Directory metadata only; Redis message history may expire independently',
|
||||
|
||||
INDEX idx_chat_session_last_active (last_active_at),
|
||||
INDEX idx_chat_session_status (status),
|
||||
INDEX idx_chat_session_expires_at (expires_at)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Chat session metadata table';
|
||||
|
||||
CREATE TABLE diagnosis_run (
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
run_id VARCHAR(64) UNIQUE NOT NULL COMMENT 'Diagnosis run id, format run-uuid',
|
||||
session_id VARCHAR(64) NOT NULL COMMENT 'Owner chat_session.session_id',
|
||||
|
||||
query TEXT NOT NULL COMMENT 'User question or AIOps alert summary for this run',
|
||||
status VARCHAR(16) DEFAULT 'PENDING' COMMENT 'PENDING/RUNNING/SUCCESS/FAILED',
|
||||
agent_flow VARCHAR(32) COMMENT 'CHAT / AI_OPS',
|
||||
|
||||
answer LONGTEXT COMMENT 'Final answer/report for this run',
|
||||
self_evaluation JSON COMMENT 'Run-scoped self evaluation payload',
|
||||
feedback VARCHAR(16) COMMENT 'User feedback for this run',
|
||||
|
||||
total_duration_ms INT COMMENT 'Run duration in milliseconds',
|
||||
total_token_count INT COMMENT 'Run token count',
|
||||
step_count INT COMMENT 'Run agent step count',
|
||||
tool_call_count INT COMMENT 'Run tool invocation count',
|
||||
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
||||
|
||||
INDEX idx_diagnosis_run_session_created (session_id, created_at, id),
|
||||
INDEX idx_diagnosis_run_session_run (session_id, run_id),
|
||||
INDEX idx_diagnosis_run_status (status),
|
||||
INDEX idx_diagnosis_run_agent_flow (agent_flow)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='Diagnosis execution run table';
|
||||
|
||||
ALTER TABLE agent_step
|
||||
ADD COLUMN run_id VARCHAR(64) NULL COMMENT '关联 diagnosis_run.run_id' AFTER session_id,
|
||||
ADD INDEX idx_agent_step_run_step (run_id, step_index);
|
||||
|
||||
ALTER TABLE tool_invocation
|
||||
ADD COLUMN run_id VARCHAR(64) NULL COMMENT '关联 diagnosis_run.run_id' AFTER session_id,
|
||||
ADD INDEX idx_tool_invocation_run_id (run_id, id);
|
||||
|
||||
INSERT INTO chat_session (
|
||||
session_id,
|
||||
status,
|
||||
message_pair_count,
|
||||
created_at,
|
||||
last_active_at,
|
||||
expires_at
|
||||
)
|
||||
SELECT
|
||||
ds.session_id,
|
||||
'ACTIVE',
|
||||
0,
|
||||
ds.created_at,
|
||||
COALESCE(ds.updated_at, ds.created_at),
|
||||
NULL
|
||||
FROM diagnosis_session ds;
|
||||
|
||||
INSERT INTO diagnosis_run (
|
||||
run_id,
|
||||
session_id,
|
||||
query,
|
||||
status,
|
||||
agent_flow,
|
||||
answer,
|
||||
self_evaluation,
|
||||
feedback,
|
||||
total_duration_ms,
|
||||
total_token_count,
|
||||
step_count,
|
||||
tool_call_count,
|
||||
created_at,
|
||||
updated_at
|
||||
)
|
||||
SELECT
|
||||
CONCAT('run-', UUID()),
|
||||
ds.session_id,
|
||||
ds.query,
|
||||
ds.status,
|
||||
ds.agent_flow,
|
||||
ds.answer,
|
||||
ds.self_evaluation,
|
||||
ds.feedback,
|
||||
ds.total_duration_ms,
|
||||
ds.total_token_count,
|
||||
ds.step_count,
|
||||
ds.tool_call_count,
|
||||
ds.created_at,
|
||||
ds.updated_at
|
||||
FROM diagnosis_session ds;
|
||||
|
||||
UPDATE agent_step ast
|
||||
JOIN diagnosis_run dr ON dr.session_id = ast.session_id
|
||||
SET ast.run_id = dr.run_id
|
||||
WHERE ast.run_id IS NULL;
|
||||
|
||||
UPDATE tool_invocation ti
|
||||
JOIN diagnosis_run dr ON dr.session_id = ti.session_id
|
||||
SET ti.run_id = dr.run_id
|
||||
WHERE ti.run_id IS NULL;
|
||||
|
||||
@@ -9,6 +9,7 @@ class SuperBizAgentApp {
|
||||
this.chatHistories = this.loadChatHistories(); // 所有历史对话
|
||||
this.isCurrentChatFromHistory = false; // 标记当前对话是否是从历史记录加载的
|
||||
this.lastTraceSessionId = this.loadLastTraceSessionId();
|
||||
this.lastTraceRunId = this.loadLastTraceRunId();
|
||||
|
||||
this.initializeElements();
|
||||
this.bindEvents();
|
||||
@@ -310,13 +311,14 @@ class SuperBizAgentApp {
|
||||
id: this.sessionId,
|
||||
title: title,
|
||||
messages: [...this.currentChatHistory],
|
||||
lastRunId: this.sessionId === this.lastTraceSessionId ? this.lastTraceRunId : '',
|
||||
createdAt: new Date().toISOString(),
|
||||
updatedAt: new Date().toISOString()
|
||||
};
|
||||
|
||||
// 添加到历史记录列表的开头
|
||||
this.chatHistories.unshift(chatHistory);
|
||||
this.rememberTraceSessionId(this.sessionId);
|
||||
this.rememberTraceTarget(this.sessionId, null);
|
||||
|
||||
// 限制历史记录数量(最多保存50条)
|
||||
if (this.chatHistories.length > 50) {
|
||||
@@ -344,6 +346,9 @@ class SuperBizAgentApp {
|
||||
const history = this.chatHistories[existingIndex];
|
||||
history.messages = [...this.currentChatHistory];
|
||||
history.updatedAt = new Date().toISOString();
|
||||
if (this.sessionId === this.lastTraceSessionId && this.lastTraceRunId) {
|
||||
history.lastRunId = this.lastTraceRunId;
|
||||
}
|
||||
|
||||
// 如果标题需要更新(第一条消息改变了)
|
||||
const firstUserMessage = this.currentChatHistory.find(msg => msg.type === 'user');
|
||||
@@ -444,7 +449,7 @@ class SuperBizAgentApp {
|
||||
|
||||
// 加载历史对话
|
||||
this.sessionId = history.id;
|
||||
this.rememberTraceSessionId(this.sessionId);
|
||||
this.rememberTraceTarget(this.sessionId, history.lastRunId || null);
|
||||
this.currentChatHistory = [...history.messages];
|
||||
this.isCurrentChatFromHistory = true; // 标记为从历史记录加载
|
||||
|
||||
@@ -487,16 +492,39 @@ class SuperBizAgentApp {
|
||||
}
|
||||
}
|
||||
|
||||
loadLastTraceRunId() {
|
||||
try {
|
||||
return localStorage.getItem('lastTraceRunId') || '';
|
||||
} catch (e) {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
rememberTraceSessionId(sessionId) {
|
||||
this.rememberTraceTarget(sessionId, null);
|
||||
}
|
||||
|
||||
rememberTraceTarget(sessionId, runId) {
|
||||
if (!sessionId) {
|
||||
return;
|
||||
}
|
||||
this.lastTraceSessionId = sessionId;
|
||||
this.lastTraceRunId = runId || '';
|
||||
try {
|
||||
localStorage.setItem('lastTraceSessionId', sessionId);
|
||||
if (runId) {
|
||||
localStorage.setItem('lastTraceRunId', runId);
|
||||
} else {
|
||||
localStorage.removeItem('lastTraceRunId');
|
||||
}
|
||||
} catch (e) {
|
||||
// localStorage may be unavailable in private or restricted contexts.
|
||||
}
|
||||
const history = this.chatHistories.find(item => item && item.id === sessionId);
|
||||
if (history && runId) {
|
||||
history.lastRunId = runId;
|
||||
this.saveChatHistories();
|
||||
}
|
||||
this.updateTraceWorkbenchLink();
|
||||
}
|
||||
|
||||
@@ -508,9 +536,30 @@ class SuperBizAgentApp {
|
||||
const sessionId = this.lastTraceSessionId
|
||||
|| (this.currentChatHistory.length > 0 ? this.sessionId : '')
|
||||
|| (recentHistory ? recentHistory.id : '');
|
||||
this.traceWorkbenchLink.href = sessionId
|
||||
? `trace.html?sessionId=${encodeURIComponent(sessionId)}`
|
||||
: 'trace.html';
|
||||
if (!sessionId) {
|
||||
this.traceWorkbenchLink.href = 'trace.html';
|
||||
return;
|
||||
}
|
||||
const params = new URLSearchParams({ sessionId });
|
||||
if (this.lastTraceRunId && sessionId === this.lastTraceSessionId) {
|
||||
params.set('runId', this.lastTraceRunId);
|
||||
}
|
||||
this.traceWorkbenchLink.href = `trace.html?${params.toString()}`;
|
||||
}
|
||||
|
||||
rememberRunMetadata(sseMessage) {
|
||||
if (!sseMessage) {
|
||||
return;
|
||||
}
|
||||
const payload = sseMessage.data && typeof sseMessage.data === 'object' ? sseMessage.data : {};
|
||||
const sessionId = sseMessage.sessionId || payload.sessionId;
|
||||
const runId = sseMessage.runId || payload.runId;
|
||||
if (!sessionId) {
|
||||
return;
|
||||
}
|
||||
this.lastSessionId = sessionId;
|
||||
this.lastRunId = runId || '';
|
||||
this.rememberTraceTarget(sessionId, runId || null);
|
||||
}
|
||||
|
||||
// 切换模式下拉菜单
|
||||
@@ -675,10 +724,11 @@ class SuperBizAgentApp {
|
||||
const chatResponse = data.data;
|
||||
|
||||
if (chatResponse && chatResponse.success) {
|
||||
// 保存后端返回的 sessionId,用于 feedback 提交
|
||||
// 保存后端返回的 sessionId/runId,用于 feedback 提交和 Trace 精确定位
|
||||
if (chatResponse.sessionId) {
|
||||
this.lastSessionId = chatResponse.sessionId;
|
||||
this.rememberTraceSessionId(chatResponse.sessionId);
|
||||
this.lastRunId = chatResponse.runId || '';
|
||||
this.rememberTraceTarget(chatResponse.sessionId, chatResponse.runId || null);
|
||||
}
|
||||
// 成功:添加实际响应消息(即使 answer 为空也显示)
|
||||
const answer = chatResponse.answer || '(无回复内容)';
|
||||
@@ -784,7 +834,9 @@ class SuperBizAgentApp {
|
||||
console.log('[SSE调试] 解析JSON成功:', sseMessage);
|
||||
|
||||
if (sseMessage && typeof sseMessage.type === 'string') {
|
||||
if (sseMessage.type === 'content') {
|
||||
if (sseMessage.type === 'metadata') {
|
||||
this.rememberRunMetadata(sseMessage);
|
||||
} else if (sseMessage.type === 'content') {
|
||||
const content = sseMessage.data || '';
|
||||
fullResponse += content;
|
||||
console.log('[SSE调试] 添加内容:', content);
|
||||
@@ -897,7 +949,7 @@ class SuperBizAgentApp {
|
||||
|
||||
// assistant 消息末尾加反馈栏(流式消息完成后由 handleStreamComplete 添加)
|
||||
if (type === 'assistant' && !isStreaming) {
|
||||
messageContentWrapper.appendChild(this.createFeedbackBar(this.lastSessionId));
|
||||
messageContentWrapper.appendChild(this.createFeedbackBar(this.lastSessionId, this.lastRunId || ''));
|
||||
}
|
||||
|
||||
messageDiv.appendChild(messageContentWrapper);
|
||||
@@ -919,7 +971,7 @@ class SuperBizAgentApp {
|
||||
}
|
||||
|
||||
// 创建反馈栏,sessionId 闭包绑定,避免多轮对话时错位
|
||||
createFeedbackBar(sessionId) {
|
||||
createFeedbackBar(sessionId, runId = '') {
|
||||
const bar = document.createElement('div');
|
||||
bar.className = 'feedback-bar';
|
||||
|
||||
@@ -933,8 +985,8 @@ class SuperBizAgentApp {
|
||||
notUsefulBtn.title = '无用';
|
||||
notUsefulBtn.innerHTML = `<svg viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg"><path d="M10 15v4a3 3 0 0 0 3 3l4-9V2H5.72a2 2 0 0 0-2 1.7l-1.38 9a2 2 0 0 0 2 2.3H10z" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"/><path d="M17 2h2.67A2.31 2.31 0 0 1 22 4v7a2.31 2.31 0 0 1-2.33 2H17" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"/></svg>`;
|
||||
|
||||
usefulBtn.addEventListener('click', () => this.submitFeedback('useful', bar, sessionId));
|
||||
notUsefulBtn.addEventListener('click', () => this.submitFeedback('not_useful', bar, sessionId));
|
||||
usefulBtn.addEventListener('click', () => this.submitFeedback('useful', bar, sessionId, runId));
|
||||
notUsefulBtn.addEventListener('click', () => this.submitFeedback('not_useful', bar, sessionId, runId));
|
||||
|
||||
bar.appendChild(usefulBtn);
|
||||
bar.appendChild(notUsefulBtn);
|
||||
@@ -942,7 +994,7 @@ class SuperBizAgentApp {
|
||||
}
|
||||
|
||||
// 提交反馈
|
||||
async submitFeedback(feedback, barElement, sessionId) {
|
||||
async submitFeedback(feedback, barElement, sessionId, runId = '') {
|
||||
if (!sessionId) return;
|
||||
|
||||
barElement.querySelectorAll('.feedback-btn').forEach(btn => btn.disabled = true);
|
||||
@@ -951,7 +1003,7 @@ class SuperBizAgentApp {
|
||||
const response = await fetch(`${this.apiBaseUrl}/feedback`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ sessionId, feedback })
|
||||
body: JSON.stringify({ sessionId, runId: runId || undefined, feedback })
|
||||
});
|
||||
const data = await response.json();
|
||||
if (data.success) {
|
||||
@@ -1055,7 +1107,7 @@ class SuperBizAgentApp {
|
||||
}
|
||||
// 流式完成后追加反馈栏
|
||||
if (messageContentWrapper && !messageContentWrapper.querySelector('.feedback-bar')) {
|
||||
messageContentWrapper.appendChild(this.createFeedbackBar(this.lastSessionId));
|
||||
messageContentWrapper.appendChild(this.createFeedbackBar(this.lastSessionId, this.lastRunId || ''));
|
||||
}
|
||||
}
|
||||
// 保存流式消息到历史记录
|
||||
@@ -1273,9 +1325,11 @@ class SuperBizAgentApp {
|
||||
for (const jsonStr of matches) {
|
||||
try {
|
||||
const sseMessage = JSON.parse(jsonStr);
|
||||
if (sseMessage.type === 'session') {
|
||||
if (sseMessage.type === 'metadata') {
|
||||
this.rememberRunMetadata(sseMessage);
|
||||
} else if (sseMessage.type === 'session') {
|
||||
this.lastSessionId = sseMessage.data;
|
||||
this.rememberTraceSessionId(sseMessage.data);
|
||||
this.rememberTraceTarget(sseMessage.data, null);
|
||||
} else if (sseMessage.type === 'content') {
|
||||
fullResponse += sseMessage.data || '';
|
||||
} else if (sseMessage.type === 'done') {
|
||||
@@ -1306,9 +1360,11 @@ class SuperBizAgentApp {
|
||||
try {
|
||||
const sseMessage = JSON.parse(rawData);
|
||||
if (sseMessage && sseMessage.type) {
|
||||
if (sseMessage.type === 'session') {
|
||||
if (sseMessage.type === 'metadata') {
|
||||
this.rememberRunMetadata(sseMessage);
|
||||
} else if (sseMessage.type === 'session') {
|
||||
this.lastSessionId = sseMessage.data;
|
||||
this.rememberTraceSessionId(sseMessage.data);
|
||||
this.rememberTraceTarget(sseMessage.data, null);
|
||||
} else if (sseMessage.type === 'content') {
|
||||
fullResponse += sseMessage.data || '';
|
||||
if (loadingMessageElement) {
|
||||
|
||||
@@ -36,7 +36,7 @@ class TraceWorkbench {
|
||||
bindEvents() {
|
||||
this.sessionForm.addEventListener('submit', (event) => {
|
||||
event.preventDefault();
|
||||
this.loadTrace(this.sessionIdInput.value.trim());
|
||||
this.loadTrace(this.sessionIdInput.value.trim(), '');
|
||||
});
|
||||
|
||||
this.sessionIdInput.addEventListener('focus', () => {
|
||||
@@ -66,9 +66,10 @@ class TraceWorkbench {
|
||||
bootstrapFromUrl() {
|
||||
const params = new URLSearchParams(window.location.search);
|
||||
const sessionId = params.get('sessionId') || this.getLastTraceSessionId();
|
||||
const runId = params.get('runId') || '';
|
||||
if (sessionId) {
|
||||
this.sessionIdInput.value = sessionId;
|
||||
this.loadTrace(sessionId);
|
||||
this.loadTrace(sessionId, runId);
|
||||
return;
|
||||
}
|
||||
this.renderEmpty();
|
||||
@@ -107,6 +108,7 @@ class TraceWorkbench {
|
||||
id: history.id,
|
||||
title: history.title || '未命名对话',
|
||||
updatedAt: history.updatedAt || history.createdAt || '',
|
||||
runId: history.lastRunId || '',
|
||||
source: 'history'
|
||||
}));
|
||||
}
|
||||
@@ -237,7 +239,7 @@ class TraceWorkbench {
|
||||
chooseSession(sessionId) {
|
||||
this.sessionIdInput.value = sessionId;
|
||||
this.closeSessionOptions();
|
||||
this.loadTrace(sessionId);
|
||||
this.loadTrace(sessionId, '');
|
||||
}
|
||||
|
||||
shortSessionId(sessionId) {
|
||||
@@ -264,18 +266,23 @@ class TraceWorkbench {
|
||||
});
|
||||
}
|
||||
|
||||
async loadTrace(sessionId) {
|
||||
async loadTrace(sessionId, runId = '') {
|
||||
if (!sessionId) {
|
||||
this.setState('请先输入会话 ID。', 'error');
|
||||
return;
|
||||
}
|
||||
|
||||
this.setState(`正在加载 ${sessionId}...`, 'loading');
|
||||
this.updateUrl(sessionId);
|
||||
const runSuffix = runId ? ` / ${runId}` : '';
|
||||
this.setState(`正在加载 ${sessionId}${runSuffix}...`, 'loading');
|
||||
this.updateUrl(sessionId, runId);
|
||||
this.loadButton.disabled = true;
|
||||
|
||||
try {
|
||||
const response = await fetch(`/api/diagnosis/${encodeURIComponent(sessionId)}/trace`);
|
||||
const url = new URL(`/api/diagnosis/${encodeURIComponent(sessionId)}/trace`, window.location.origin);
|
||||
if (runId) {
|
||||
url.searchParams.set('runId', runId);
|
||||
}
|
||||
const response = await fetch(url.toString());
|
||||
const data = await this.handleResponse(response);
|
||||
this.trace = data;
|
||||
this.selectedToolId = data.toolInvocations && data.toolInvocations.length
|
||||
@@ -284,11 +291,14 @@ class TraceWorkbench {
|
||||
this.activeFilter = 'all';
|
||||
try {
|
||||
localStorage.setItem('lastTraceSessionId', sessionId);
|
||||
if (data.runId) {
|
||||
localStorage.setItem('lastTraceRunId', data.runId);
|
||||
}
|
||||
} catch (error) {
|
||||
// ignore storage failures
|
||||
}
|
||||
this.renderTrace();
|
||||
this.setState(`已加载 ${sessionId}。`, '');
|
||||
this.setState(`已加载 ${sessionId}${data.runId ? ` / ${data.runId}` : ''}。`, '');
|
||||
} catch (error) {
|
||||
this.trace = null;
|
||||
this.renderEmpty();
|
||||
@@ -324,9 +334,14 @@ class TraceWorkbench {
|
||||
return value;
|
||||
}
|
||||
|
||||
updateUrl(sessionId) {
|
||||
updateUrl(sessionId, runId = '') {
|
||||
const url = new URL(window.location.href);
|
||||
url.searchParams.set('sessionId', sessionId);
|
||||
if (runId) {
|
||||
url.searchParams.set('runId', runId);
|
||||
} else {
|
||||
url.searchParams.delete('runId');
|
||||
}
|
||||
window.history.replaceState({}, '', url.toString());
|
||||
}
|
||||
|
||||
@@ -399,6 +414,9 @@ class TraceWorkbench {
|
||||
if (session.sessionId) {
|
||||
subtitleParts.push(`会话 ${session.sessionId}`);
|
||||
}
|
||||
if (this.trace && this.trace.runId) {
|
||||
subtitleParts.push(`运行 ${this.trace.runId}`);
|
||||
}
|
||||
if (session.status) {
|
||||
subtitleParts.push(`状态 ${this.displayStatus(session.status)}`);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.superbiz.agent.service.ChatService;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.verifyNoInteractions;
|
||||
|
||||
class ChatControllerTest {
|
||||
|
||||
@Test
|
||||
void blankChatRequestReturnsErrorBeforeCreatingRun() {
|
||||
ChatController controller = new ChatController();
|
||||
ChatService chatService = mock(ChatService.class);
|
||||
ReflectionTestUtils.setField(controller, "chatService", chatService);
|
||||
|
||||
ChatController.ChatRequest request = new ChatController.ChatRequest();
|
||||
request.setId("invalid-chat-session");
|
||||
request.setQuestion(" ");
|
||||
|
||||
ResponseEntity<ChatController.ApiResponse<ChatController.ChatResponse>> response = controller.chat(request);
|
||||
|
||||
ChatController.ChatResponse body = response.getBody().getData();
|
||||
assertFalse(body.isSuccess());
|
||||
assertEquals("问题内容不能为空", body.getErrorMessage());
|
||||
verifyNoInteractions(chatService);
|
||||
}
|
||||
|
||||
@Test
|
||||
void aiOpsMetadataMessageCarriesSessionAndRunId() {
|
||||
ChatController.SseMessage message = ChatController.SseMessage.metadata("session-1", "run-1");
|
||||
|
||||
assertEquals("metadata", message.getType());
|
||||
assertEquals("session-1", message.getSessionId());
|
||||
assertEquals("run-1", message.getRunId());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.superbiz.agent.dto.FeedbackRequest;
|
||||
import com.superbiz.agent.dto.FeedbackResponse;
|
||||
import com.superbiz.agent.service.FeedbackService;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class FeedbackControllerTest {
|
||||
|
||||
@Test
|
||||
void submitFeedbackPassesRunIdToService() {
|
||||
FeedbackController controller = new FeedbackController();
|
||||
FeedbackService feedbackService = mock(FeedbackService.class);
|
||||
ReflectionTestUtils.setField(controller, "feedbackService", feedbackService);
|
||||
|
||||
FeedbackRequest request = new FeedbackRequest();
|
||||
request.setSessionId("session-1");
|
||||
request.setRunId("run-1");
|
||||
request.setFeedback("useful");
|
||||
|
||||
when(feedbackService.submitFeedback("session-1", "run-1", "useful"))
|
||||
.thenReturn(FeedbackResponse.builder()
|
||||
.success(true)
|
||||
.runId("run-1")
|
||||
.fallbackToLatestRun(false)
|
||||
.build());
|
||||
|
||||
ResponseEntity<FeedbackResponse> response = controller.submitFeedback(request);
|
||||
|
||||
assertEquals(200, response.getStatusCode().value());
|
||||
assertTrue(response.getBody().isSuccess());
|
||||
assertEquals("run-1", response.getBody().getRunId());
|
||||
verify(feedbackService).submitFeedback("session-1", "run-1", "useful");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
package com.superbiz.agent.repository;
|
||||
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.boot.test.autoconfigure.jdbc.AutoConfigureTestDatabase;
|
||||
import org.springframework.boot.test.autoconfigure.orm.jpa.DataJpaTest;
|
||||
import org.springframework.test.context.TestPropertySource;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.UUID;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
|
||||
@DataJpaTest
|
||||
@AutoConfigureTestDatabase(replace = AutoConfigureTestDatabase.Replace.NONE)
|
||||
@TestPropertySource(properties = {
|
||||
"spring.flyway.enabled=true",
|
||||
"spring.jpa.hibernate.ddl-auto=validate",
|
||||
"spring.jpa.show-sql=true"
|
||||
})
|
||||
class DiagnosisRunRepositoryTest {
|
||||
|
||||
@Autowired
|
||||
private ChatSessionRepository chatSessionRepository;
|
||||
|
||||
@Autowired
|
||||
private DiagnosisRunRepository diagnosisRunRepository;
|
||||
|
||||
@Autowired
|
||||
private AgentStepRepository agentStepRepository;
|
||||
|
||||
@Autowired
|
||||
private ToolInvocationRepository toolInvocationRepository;
|
||||
|
||||
@Test
|
||||
void saveAndFindLatestRunBySessionId() {
|
||||
String sessionId = "test-session-" + UUID.randomUUID();
|
||||
DiagnosisRun firstRun = saveRun(sessionId, "run-" + UUID.randomUUID(), "first question");
|
||||
DiagnosisRun secondRun = saveRun(sessionId, "run-" + UUID.randomUUID(), "second question");
|
||||
|
||||
DiagnosisRun latest = diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId)
|
||||
.orElseThrow();
|
||||
|
||||
assertEquals(secondRun.getRunId(), latest.getRunId());
|
||||
assertEquals(firstRun.getRunId(), diagnosisRunRepository.findByRunId(firstRun.getRunId()).orElseThrow().getRunId());
|
||||
assertTrue(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, secondRun.getRunId()).isPresent());
|
||||
assertEquals(2, diagnosisRunRepository.findBySessionIdOrderByCreatedAtDescIdDesc(sessionId).size());
|
||||
}
|
||||
|
||||
@Test
|
||||
void stepAndToolCanBeQueriedByRunId() {
|
||||
String sessionId = "test-session-" + UUID.randomUUID();
|
||||
String runId = "run-" + UUID.randomUUID();
|
||||
saveRun(sessionId, runId, "run scoped trace");
|
||||
|
||||
agentStepRepository.save(AgentStep.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.stepIndex(1)
|
||||
.agentName("executor")
|
||||
.hasToolCall(true)
|
||||
.build());
|
||||
agentStepRepository.save(AgentStep.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.stepIndex(0)
|
||||
.agentName("planner")
|
||||
.hasToolCall(false)
|
||||
.build());
|
||||
|
||||
toolInvocationRepository.save(ToolInvocation.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.toolName("lookup_knowledge")
|
||||
.inputParams("{\"query\":\"payment timeout\"}")
|
||||
.success(true)
|
||||
.build());
|
||||
|
||||
List<AgentStep> steps = agentStepRepository.findByRunIdOrderByStepIndex(runId);
|
||||
List<ToolInvocation> tools = toolInvocationRepository.findByRunIdOrderByIdAsc(runId);
|
||||
|
||||
assertEquals(2, steps.size());
|
||||
assertEquals("planner", steps.get(0).getAgentName());
|
||||
assertEquals(1, tools.size());
|
||||
assertEquals(runId, tools.get(0).getRunId());
|
||||
assertEquals(2, agentStepRepository.countByRunId(runId));
|
||||
assertEquals(1, toolInvocationRepository.countByRunId(runId));
|
||||
}
|
||||
|
||||
private DiagnosisRun saveRun(String sessionId, String runId, String query) {
|
||||
chatSessionRepository.findBySessionId(sessionId)
|
||||
.orElseGet(() -> chatSessionRepository.save(ChatSession.builder()
|
||||
.sessionId(sessionId)
|
||||
.status("ACTIVE")
|
||||
.messagePairCount(0)
|
||||
.build()));
|
||||
|
||||
return diagnosisRunRepository.save(DiagnosisRun.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.query(query)
|
||||
.status("SUCCESS")
|
||||
.agentFlow("CHAT")
|
||||
.answer("answer for " + query)
|
||||
.build());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,13 +1,20 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||
import com.superbiz.agent.config.AiOpsPromptProperties;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.AIOpsRequest;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.ai.chat.model.ChatModel;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.util.Optional;
|
||||
@@ -19,17 +26,22 @@ import static org.mockito.Mockito.*;
|
||||
class AiOpsServiceTest {
|
||||
|
||||
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||
private final DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
|
||||
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
|
||||
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||
private final AiOpsPromptProperties promptProperties = mock(AiOpsPromptProperties.class);
|
||||
private final AiOpsService service = new AiOpsService();
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
|
||||
ReflectionTestUtils.setField(service, "diagnosisRunRepository", diagnosisRunRepository);
|
||||
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
|
||||
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
|
||||
ReflectionTestUtils.setField(service, "aiOpsRuleEvaluationService", new AiOpsRuleEvaluationService());
|
||||
ReflectionTestUtils.setField(service, "selfEvaluationMergeService", new SelfEvaluationMergeService());
|
||||
ReflectionTestUtils.setField(service, "promptProperties", promptProperties);
|
||||
when(promptProperties.getSupervisor()).thenReturn("supervisor prompt");
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -139,19 +151,48 @@ class AiOpsServiceTest {
|
||||
}
|
||||
|
||||
@Test
|
||||
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
void persistFinalReportUpdatesDiagnosisRunAnswerAndEvaluationByRun() {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.sessionId("aiops-session-001")
|
||||
.runId("run-aiops-001")
|
||||
.query("AI Ops alert analysis")
|
||||
.status("SUCCESS")
|
||||
.agentFlow("AI_OPS")
|
||||
.build();
|
||||
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
|
||||
ToolInvocation invocation = ToolInvocation.builder()
|
||||
.sessionId("aiops-session-001")
|
||||
.runId("run-aiops-001")
|
||||
.toolName("query_logs")
|
||||
.success(true)
|
||||
.build();
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("aiops-session-001", "run-aiops-001"))
|
||||
.thenReturn(Optional.of(run));
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-aiops-001")).thenReturn(List.of(invocation));
|
||||
|
||||
service.persistFinalReport("aiops-session-001", "# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.");
|
||||
service.persistFinalReport("aiops-session-001", "run-aiops-001",
|
||||
"# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", null);
|
||||
|
||||
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
|
||||
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", run.getAnswer());
|
||||
assertTrue(run.getSelfEvaluation().contains("aiops_rule_evaluation"));
|
||||
verify(toolInvocationRepository).findByRunIdOrderByIdAsc("run-aiops-001");
|
||||
verify(diagnosisRunRepository).save(run);
|
||||
verify(diagnosisSessionRepository, never()).save(any());
|
||||
}
|
||||
|
||||
@Test
|
||||
void legacyPersistFinalReportStillUpdatesHistoricalDiagnosisSession() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId("legacy-aiops-session")
|
||||
.query("AI Ops alert analysis")
|
||||
.status("SUCCESS")
|
||||
.agentFlow("AI_OPS")
|
||||
.build();
|
||||
when(diagnosisSessionRepository.findBySessionId("legacy-aiops-session")).thenReturn(Optional.of(session));
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("legacy-aiops-session")).thenReturn(List.of());
|
||||
|
||||
service.persistFinalReport("legacy-aiops-session", "# 告警分析报告\nLegacy analysis with evidence summary.");
|
||||
|
||||
assertEquals("# 告警分析报告\nLegacy analysis with evidence summary.", session.getAnswer());
|
||||
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
|
||||
verify(diagnosisSessionRepository).save(session);
|
||||
}
|
||||
@@ -160,32 +201,68 @@ class AiOpsServiceTest {
|
||||
void persistFinalReportSkipsBlankInput() {
|
||||
service.persistFinalReport("aiops-session-001", " ");
|
||||
|
||||
verifyNoInteractions(diagnosisSessionRepository);
|
||||
verifyNoInteractions(diagnosisSessionRepository, diagnosisRunRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void backfillSessionMetricsUsesRealToolInvocationCount() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
void backfillRunMetricsUsesRunScopedRows() {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.sessionId("aiops-session-002")
|
||||
.runId("run-aiops-002")
|
||||
.build();
|
||||
AgentStep stepWithTool = AgentStep.builder()
|
||||
.sessionId("aiops-session-002")
|
||||
.runId("run-aiops-002")
|
||||
.hasToolCall(true)
|
||||
.tokenCount(10)
|
||||
.build();
|
||||
AgentStep stepWithoutTool = AgentStep.builder()
|
||||
.sessionId("aiops-session-002")
|
||||
.runId("run-aiops-002")
|
||||
.hasToolCall(false)
|
||||
.tokenCount(20)
|
||||
.build();
|
||||
when(agentStepRepository.findBySessionIdOrderByStepIndex("aiops-session-002"))
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex("run-aiops-002"))
|
||||
.thenReturn(List.of(stepWithTool, stepWithoutTool));
|
||||
when(toolInvocationRepository.countBySessionId("aiops-session-002")).thenReturn(11L);
|
||||
when(toolInvocationRepository.countByRunId("run-aiops-002")).thenReturn(11L);
|
||||
|
||||
ReflectionTestUtils.invokeMethod(service, "backfillSessionMetrics", session);
|
||||
ReflectionTestUtils.invokeMethod(service, "backfillRunMetrics", run);
|
||||
|
||||
assertEquals(2, session.getStepCount());
|
||||
assertEquals(30, session.getTotalTokenCount());
|
||||
assertEquals(11, session.getToolCallCount());
|
||||
assertEquals(2, run.getStepCount());
|
||||
assertEquals(30, run.getTotalTokenCount());
|
||||
assertEquals(11, run.getToolCallCount());
|
||||
verify(agentStepRepository).findByRunIdOrderByStepIndex("run-aiops-002");
|
||||
verify(toolInvocationRepository).countByRunId("run-aiops-002");
|
||||
}
|
||||
|
||||
@Test
|
||||
void sameAiOpsSessionCanStartDistinctRuns() {
|
||||
AIOpsRequest request = new AIOpsRequest();
|
||||
request.setAlertName("HighCPUUsage");
|
||||
|
||||
ReflectionTestUtils.invokeMethod(service, "startDiagnosisRun", "same-session", "run-aiops-a", request);
|
||||
ReflectionTestUtils.invokeMethod(service, "startDiagnosisRun", "same-session", "run-aiops-b", request);
|
||||
|
||||
verify(diagnosisRunRepository).save(argThat(run ->
|
||||
"same-session".equals(run.getSessionId())
|
||||
&& "run-aiops-a".equals(run.getRunId())
|
||||
&& "AI_OPS".equals(run.getAgentFlow())
|
||||
&& "RUNNING".equals(run.getStatus())));
|
||||
verify(diagnosisRunRepository).save(argThat(run ->
|
||||
"same-session".equals(run.getSessionId())
|
||||
&& "run-aiops-b".equals(run.getRunId())
|
||||
&& "AI_OPS".equals(run.getAgentFlow())
|
||||
&& "RUNNING".equals(run.getStatus())));
|
||||
}
|
||||
|
||||
@Test
|
||||
void buildSupervisorAgentSetsPlannerAsMainAgent() {
|
||||
ChatModel chatModel = mock(ChatModel.class);
|
||||
ReactAgent planner = mock(ReactAgent.class);
|
||||
ReactAgent executor = mock(ReactAgent.class);
|
||||
|
||||
SupervisorAgent supervisor = service.buildSupervisorAgent(chatModel, planner, executor);
|
||||
|
||||
assertSame(planner, supervisor.getMainAgent());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.CaseLibrary;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.enums.SourceType;
|
||||
import com.superbiz.agent.repository.CaseLibraryRepository;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.mockito.ArgumentCaptor;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.util.Optional;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotNull;
|
||||
import static org.mockito.ArgumentMatchers.any;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.never;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class CaseLibraryServiceTest {
|
||||
|
||||
private final CaseLibraryRepository caseLibraryRepository = mock(CaseLibraryRepository.class);
|
||||
private final CaseLibraryService service = new CaseLibraryService();
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
ReflectionTestUtils.setField(service, "caseLibraryRepository", caseLibraryRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void createFromRunStoresRunIdAsDiagnosisIdAndUsesRunContent() {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.sessionId("session-1")
|
||||
.runId("run-1")
|
||||
.query("payment timeout")
|
||||
.answer("redis timeout caused payment latency")
|
||||
.build();
|
||||
when(caseLibraryRepository.findByDiagnosisId("run-1")).thenReturn(Optional.empty());
|
||||
when(caseLibraryRepository.save(any(CaseLibrary.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
|
||||
CaseLibrary saved = service.createFromRun(run);
|
||||
|
||||
assertEquals("run-1", saved.getDiagnosisId());
|
||||
assertEquals("payment timeout", saved.getTitle());
|
||||
assertEquals("redis timeout caused payment latency", saved.getRootCause());
|
||||
assertEquals("redis timeout caused payment latency", saved.getSolution());
|
||||
assertEquals(SourceType.AUTO, saved.getSourceType());
|
||||
assertNotNull(saved.getCaseId());
|
||||
}
|
||||
|
||||
@Test
|
||||
void createFromRunReusesExistingCaseForIdempotency() {
|
||||
DiagnosisRun run = DiagnosisRun.builder()
|
||||
.runId("run-existing")
|
||||
.query("query")
|
||||
.answer("answer")
|
||||
.build();
|
||||
CaseLibrary existing = CaseLibrary.builder()
|
||||
.caseId("case-existing")
|
||||
.diagnosisId("run-existing")
|
||||
.build();
|
||||
when(caseLibraryRepository.findByDiagnosisId("run-existing")).thenReturn(Optional.of(existing));
|
||||
|
||||
CaseLibrary saved = service.createFromRun(run);
|
||||
|
||||
assertEquals("case-existing", saved.getCaseId());
|
||||
verify(caseLibraryRepository, never()).save(any());
|
||||
}
|
||||
|
||||
@Test
|
||||
void createFromSessionPreservesLegacySessionIdSemantics() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId("legacy-session")
|
||||
.query("legacy query")
|
||||
.answer("legacy answer")
|
||||
.build();
|
||||
when(caseLibraryRepository.findByDiagnosisId("legacy-session")).thenReturn(Optional.empty());
|
||||
when(caseLibraryRepository.save(any(CaseLibrary.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
|
||||
service.createFromSession(session);
|
||||
|
||||
ArgumentCaptor<CaseLibrary> captor = ArgumentCaptor.forClass(CaseLibrary.class);
|
||||
verify(caseLibraryRepository).save(captor.capture());
|
||||
assertEquals("legacy-session", captor.getValue().getDiagnosisId());
|
||||
assertEquals("legacy answer", captor.getValue().getRootCause());
|
||||
}
|
||||
}
|
||||
@@ -7,10 +7,12 @@ import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ChatSessionRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||
@@ -31,11 +33,15 @@ import java.util.concurrent.atomic.AtomicInteger;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNotEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertSame;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.ArgumentMatchers.any;
|
||||
import static org.mockito.ArgumentMatchers.anyString;
|
||||
import static org.mockito.ArgumentMatchers.eq;
|
||||
import static org.mockito.ArgumentMatchers.isNull;
|
||||
import static org.mockito.Mockito.atLeast;
|
||||
import static org.mockito.Mockito.atLeastOnce;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.when;
|
||||
@@ -58,8 +64,73 @@ class ChatServiceSequentialAgentTest {
|
||||
assertTrue(result.answer().contains("连接池 active 达到上限"));
|
||||
assertFalse(result.answer().contains("\"answer_version\""));
|
||||
assertEquals("sequential-test-session", result.sessionId());
|
||||
assertTrue(result.runId().startsWith("run-"));
|
||||
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "chat_composer"), chatModel.agentCalls);
|
||||
assertTrue(chatModel.sawVerifierPrompt);
|
||||
|
||||
ChatSessionRepository chatSessionRepository =
|
||||
(ChatSessionRepository) ReflectionTestUtils.getField(chatService, "chatSessionRepository");
|
||||
DiagnosisRunRepository diagnosisRunRepository =
|
||||
(DiagnosisRunRepository) ReflectionTestUtils.getField(chatService, "diagnosisRunRepository");
|
||||
EvaluationService evaluationService =
|
||||
(EvaluationService) ReflectionTestUtils.getField(chatService, "evaluationService");
|
||||
|
||||
ArgumentCaptor<ChatSession> chatSessionCaptor = ArgumentCaptor.forClass(ChatSession.class);
|
||||
verify(chatSessionRepository, atLeastOnce()).save(chatSessionCaptor.capture());
|
||||
assertEquals("sequential-test-session", chatSessionCaptor.getValue().getSessionId());
|
||||
|
||||
ArgumentCaptor<DiagnosisRun> runCaptor = ArgumentCaptor.forClass(DiagnosisRun.class);
|
||||
verify(diagnosisRunRepository, atLeastOnce()).save(runCaptor.capture());
|
||||
DiagnosisRun savedRun = runCaptor.getValue();
|
||||
assertEquals(result.runId(), savedRun.getRunId());
|
||||
assertEquals("sequential-test-session", savedRun.getSessionId());
|
||||
assertEquals("SUCCESS", savedRun.getStatus());
|
||||
assertEquals(result.answer(), savedRun.getAnswer());
|
||||
verify(evaluationService).evaluateRun(eq(result.runId()), eq(result.answer()));
|
||||
}
|
||||
|
||||
@Test
|
||||
void executeChatComplexCreatesDistinctRunsForSameSessionAcrossTurns() throws Exception {
|
||||
ChatService chatService = createChatService();
|
||||
ScriptedChatModel firstRoundModel = new ScriptedChatModel();
|
||||
ScriptedChatModel secondRoundModel = new ScriptedChatModel();
|
||||
String sessionId = "sequential-same-session";
|
||||
|
||||
ChatService.ChatResult first = chatService.executeChatComplex(
|
||||
firstRoundModel,
|
||||
new ToolCallback[0],
|
||||
"第一轮:请分析支付超时",
|
||||
List.of(),
|
||||
sessionId
|
||||
);
|
||||
ChatService.ChatResult second = chatService.executeChatComplex(
|
||||
secondRoundModel,
|
||||
new ToolCallback[0],
|
||||
"第二轮:基于上一轮结论列出缺失证据",
|
||||
List.of(
|
||||
Map.of("role", "user", "content", "第一轮:请分析支付超时"),
|
||||
Map.of("role", "assistant", "content", first.answer())
|
||||
),
|
||||
sessionId
|
||||
);
|
||||
|
||||
assertEquals(sessionId, first.sessionId());
|
||||
assertEquals(sessionId, second.sessionId());
|
||||
assertNotEquals(first.runId(), second.runId());
|
||||
|
||||
DiagnosisRunRepository diagnosisRunRepository =
|
||||
(DiagnosisRunRepository) ReflectionTestUtils.getField(chatService, "diagnosisRunRepository");
|
||||
ArgumentCaptor<DiagnosisRun> runCaptor = ArgumentCaptor.forClass(DiagnosisRun.class);
|
||||
verify(diagnosisRunRepository, atLeast(2)).save(runCaptor.capture());
|
||||
|
||||
List<String> savedRunIds = runCaptor.getAllValues().stream()
|
||||
.filter(run -> sessionId.equals(run.getSessionId()))
|
||||
.map(DiagnosisRun::getRunId)
|
||||
.distinct()
|
||||
.toList();
|
||||
assertEquals(2, savedRunIds.size());
|
||||
assertTrue(savedRunIds.contains(first.runId()));
|
||||
assertTrue(savedRunIds.contains(second.runId()));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -645,9 +716,12 @@ class ChatServiceSequentialAgentTest {
|
||||
private ChatService createChatService() {
|
||||
ChatService chatService = new ChatService();
|
||||
|
||||
DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||
when(diagnosisSessionRepository.findBySessionId(anyString())).thenReturn(Optional.empty());
|
||||
when(diagnosisSessionRepository.save(any(DiagnosisSession.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
ChatSessionRepository chatSessionRepository = mock(ChatSessionRepository.class);
|
||||
when(chatSessionRepository.findBySessionId(anyString())).thenReturn(Optional.empty());
|
||||
when(chatSessionRepository.save(any(ChatSession.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
|
||||
DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
|
||||
when(diagnosisRunRepository.save(any(DiagnosisRun.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
|
||||
AtomicInteger stepId = new AtomicInteger(1);
|
||||
AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
|
||||
@@ -660,13 +734,20 @@ class ChatServiceSequentialAgentTest {
|
||||
});
|
||||
when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep()));
|
||||
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex(anyString())).thenReturn(List.of());
|
||||
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||
when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L);
|
||||
when(toolInvocationRepository.countByRunId(anyString())).thenReturn(0L);
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
|
||||
.id(101L)
|
||||
.toolName("query_metrics")
|
||||
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||
.build()));
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
|
||||
.id(101L)
|
||||
.toolName("query_metrics")
|
||||
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||
.build()));
|
||||
|
||||
EvaluationService evaluationService = mock(EvaluationService.class);
|
||||
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
|
||||
@@ -674,6 +755,7 @@ class ChatServiceSequentialAgentTest {
|
||||
when(knowledgeDomainService.buildKnowledgeMap()).thenReturn("");
|
||||
ToolTraceSummaryService toolTraceSummaryService = mock(ToolTraceSummaryService.class);
|
||||
when(toolTraceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
|
||||
when(toolTraceSummaryService.buildVerifierTraceSummaryForRun(anyString(), anyString())).thenReturn(List.of());
|
||||
SelfEvaluationMergeService selfEvaluationMergeService = mock(SelfEvaluationMergeService.class);
|
||||
when(selfEvaluationMergeService.mergeVerifierEvaluation(any(), any())).thenReturn("{}");
|
||||
ExecutorGatekeeperService executorGatekeeperService = new ExecutorGatekeeperService(toolInvocationRepository);
|
||||
@@ -681,7 +763,8 @@ class ChatServiceSequentialAgentTest {
|
||||
ReflectionTestUtils.setField(chatService, "dateTimeTools", new DateTimeTools());
|
||||
ReflectionTestUtils.setField(chatService, "lookupKnowledgeTool", new LookupKnowledgeTool());
|
||||
ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class)));
|
||||
ReflectionTestUtils.setField(chatService, "diagnosisSessionRepository", diagnosisSessionRepository);
|
||||
ReflectionTestUtils.setField(chatService, "chatSessionRepository", chatSessionRepository);
|
||||
ReflectionTestUtils.setField(chatService, "diagnosisRunRepository", diagnosisRunRepository);
|
||||
ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository);
|
||||
ReflectionTestUtils.setField(chatService, "toolInvocationRepository", toolInvocationRepository);
|
||||
ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService);
|
||||
|
||||
@@ -2,11 +2,15 @@ package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.ChatSession;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||
import com.superbiz.agent.exception.SessionNotFoundException;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.ChatSessionRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import org.junit.jupiter.api.Test;
|
||||
@@ -15,148 +19,245 @@ import java.time.LocalDateTime;
|
||||
import java.util.List;
|
||||
import java.util.Optional;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.*;
|
||||
import static org.mockito.Mockito.*;
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.ArgumentMatchers.any;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.never;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.verifyNoInteractions;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class DiagnosisTraceServiceTest {
|
||||
|
||||
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||
private final ChatSessionRepository chatSessionRepository = mock(ChatSessionRepository.class);
|
||||
private final DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
|
||||
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
|
||||
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||
private final DiagnosisTraceService service = new DiagnosisTraceService(
|
||||
diagnosisSessionRepository,
|
||||
chatSessionRepository,
|
||||
diagnosisRunRepository,
|
||||
agentStepRepository,
|
||||
toolInvocationRepository,
|
||||
new ObjectMapper()
|
||||
);
|
||||
|
||||
@Test
|
||||
void getTraceAggregatesSessionStepsAndTools() {
|
||||
String sessionId = "trace-session-001";
|
||||
LocalDateTime now = LocalDateTime.of(2026, 7, 3, 14, 30);
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.id(1L)
|
||||
.sessionId(sessionId)
|
||||
.query("payment timeout")
|
||||
.status("SUCCESS")
|
||||
.agentFlow("COMPLEX")
|
||||
.totalDurationMs(1200)
|
||||
.totalTokenCount(300)
|
||||
.stepCount(2)
|
||||
.toolCallCount(1)
|
||||
.answer("restart payment gateway pool")
|
||||
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"},\"aiops_rule_evaluation\":{\"verdict\":\"WARN\"}}")
|
||||
.feedback("useful")
|
||||
.createdAt(now)
|
||||
.updatedAt(now)
|
||||
.build();
|
||||
AgentStep step = AgentStep.builder()
|
||||
.id(10L)
|
||||
.sessionId(sessionId)
|
||||
.stepIndex(1)
|
||||
.agentName("chat_executor")
|
||||
.modelInput("input")
|
||||
.modelOutput("output")
|
||||
.thought("executor finished")
|
||||
.hasToolCall(true)
|
||||
.durationMs(500)
|
||||
.tokenCount(100)
|
||||
.createdAt(now)
|
||||
.build();
|
||||
ToolInvocation invocation = ToolInvocation.builder()
|
||||
.id(20L)
|
||||
.sessionId(sessionId)
|
||||
.stepId(10L)
|
||||
.toolName("lookup_knowledge")
|
||||
.inputParams("{\"query\":\"ERR_TIMEOUT\"}")
|
||||
.outputPreview("payment timeout doc")
|
||||
.outputLength(19)
|
||||
.retrievalLayer("L0")
|
||||
.l0MatchCount(1)
|
||||
.l1MatchCount(0)
|
||||
.isTruncated(false)
|
||||
.relevanceLevel("HIGHLY_RELEVANT")
|
||||
.dedupReason("FIRST_HIT")
|
||||
.retrievalDetails("{\"documents\":[\"payment-errors.md\"]}")
|
||||
.durationMs(80)
|
||||
.success(true)
|
||||
.createdAt(now)
|
||||
.build();
|
||||
void getTraceWithoutRunIdResolvesLatestRun() {
|
||||
String sessionId = "trace-session-latest";
|
||||
LocalDateTime base = LocalDateTime.of(2026, 7, 10, 10, 0);
|
||||
DiagnosisRun latest = run(2L, sessionId, "run-latest", "second question", base.plusMinutes(1));
|
||||
AgentStep step = step(20L, sessionId, "run-latest", 0, "composer", base.plusMinutes(1));
|
||||
ToolInvocation invocation = invocation(30L, sessionId, "run-latest", "query_metrics", base.plusMinutes(1));
|
||||
|
||||
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.of(session));
|
||||
when(agentStepRepository.findBySessionId(sessionId)).thenReturn(List.of(step));
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)).thenReturn(List.of(invocation));
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId))
|
||||
.thenReturn(Optional.of(latest));
|
||||
when(chatSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.of(chatSession(sessionId)));
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex("run-latest")).thenReturn(List.of(step));
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-latest")).thenReturn(List.of(invocation));
|
||||
|
||||
DiagnosisTraceResponse response = service.getTrace(sessionId);
|
||||
|
||||
assertEquals(sessionId, response.getSession().getSessionId());
|
||||
assertEquals("payment timeout", response.getSession().getQuery());
|
||||
assertEquals("PASS", ((java.util.Map<?, ?>) response.getSession()
|
||||
.getSelfEvaluation()
|
||||
.get("verifier_evaluation")).get("verdict"));
|
||||
assertEquals(1, response.getSteps().size());
|
||||
assertEquals("chat_executor", response.getSteps().get(0).getAgentName());
|
||||
assertEquals(1, response.getToolInvocations().size());
|
||||
assertEquals("ERR_TIMEOUT", response.getToolInvocations().get(0).getInputParams().get("query"));
|
||||
assertEquals(2, response.getSummary().getPersistedStepCount());
|
||||
assertEquals(1, response.getSummary().getReturnedStepCount());
|
||||
assertEquals(1, response.getSummary().getPersistedToolCallCount());
|
||||
assertEquals(1, response.getSummary().getReturnedToolCallCount());
|
||||
assertTrue(response.getSummary().isHasVerifierEvaluation());
|
||||
assertTrue(response.getSummary().isHasAiOpsRuleEvaluation());
|
||||
assertTrue(response.getSummary().isHasFeedback());
|
||||
assertEquals("run-latest", response.getRunId());
|
||||
assertEquals("run-latest", response.getRun().getRunId());
|
||||
assertEquals("run-latest", response.getSummary().getResolvedRunId());
|
||||
assertEquals("second question", response.getSession().getQuery());
|
||||
assertEquals("run-latest", response.getSteps().get(0).getRunId());
|
||||
assertEquals("run-latest", response.getToolInvocations().get(0).getRunId());
|
||||
assertEquals(1, response.getChatSession().getMessagePairCount());
|
||||
}
|
||||
|
||||
@Test
|
||||
void getTraceOrdersStepsByCreationTimeAndIdNotPerAgentStepIndex() {
|
||||
String sessionId = "trace-session-ordered";
|
||||
LocalDateTime base = LocalDateTime.of(2026, 7, 7, 0, 0);
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.id(1L)
|
||||
.sessionId(sessionId)
|
||||
.query("mysql pool issue")
|
||||
.status("SUCCESS")
|
||||
.stepCount(4)
|
||||
.toolCallCount(0)
|
||||
.createdAt(base)
|
||||
.updatedAt(base)
|
||||
.build();
|
||||
AgentStep planner = step(1L, sessionId, 0, "planner", base.plusSeconds(1));
|
||||
AgentStep executor0 = step(2L, sessionId, 0, "executor", base.plusSeconds(2));
|
||||
AgentStep executor1 = step(3L, sessionId, 1, "executor", base.plusSeconds(3));
|
||||
AgentStep verifier = step(4L, sessionId, 0, "verifier", base.plusSeconds(4));
|
||||
void getTraceWithRunIdReturnsExactFirstRun() {
|
||||
String sessionId = "trace-session-exact";
|
||||
DiagnosisRun first = run(1L, sessionId, "run-first", "first question",
|
||||
LocalDateTime.of(2026, 7, 10, 10, 0));
|
||||
AgentStep firstStep = step(10L, sessionId, "run-first", 0, "planner", first.getCreatedAt());
|
||||
|
||||
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.of(session));
|
||||
when(agentStepRepository.findBySessionId(sessionId))
|
||||
.thenReturn(List.of(planner, executor0, verifier, executor1));
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)).thenReturn(List.of());
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, "run-first"))
|
||||
.thenReturn(Optional.of(first));
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex("run-first")).thenReturn(List.of(firstStep));
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-first")).thenReturn(List.of());
|
||||
|
||||
DiagnosisTraceResponse response = service.getTrace(sessionId);
|
||||
DiagnosisTraceResponse response = service.getTrace(sessionId, "run-first");
|
||||
|
||||
assertEquals(List.of("planner", "executor", "executor", "verifier"),
|
||||
assertEquals("run-first", response.getRunId());
|
||||
assertEquals("first question", response.getRun().getQuery());
|
||||
assertEquals(List.of("planner"),
|
||||
response.getSteps().stream().map(DiagnosisTraceResponse.AgentStepTrace::getAgentName).toList());
|
||||
assertEquals(List.of(0, 0, 1, 0),
|
||||
response.getSteps().stream().map(DiagnosisTraceResponse.AgentStepTrace::getStepIndex).toList());
|
||||
verify(diagnosisRunRepository, never()).findFirstBySessionIdOrderByCreatedAtDescIdDesc(any());
|
||||
}
|
||||
|
||||
@Test
|
||||
void getTraceWithRunIdReturnsExactSecondRun() {
|
||||
String sessionId = "trace-session-exact";
|
||||
DiagnosisRun second = run(2L, sessionId, "run-second", "second question",
|
||||
LocalDateTime.of(2026, 7, 10, 10, 1));
|
||||
ToolInvocation secondTool = invocation(20L, sessionId, "run-second", "query_logs", second.getCreatedAt());
|
||||
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId(sessionId, "run-second"))
|
||||
.thenReturn(Optional.of(second));
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex("run-second")).thenReturn(List.of());
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-second")).thenReturn(List.of(secondTool));
|
||||
|
||||
DiagnosisTraceResponse response = service.getTrace(sessionId, "run-second");
|
||||
|
||||
assertEquals("run-second", response.getRunId());
|
||||
assertEquals("second question", response.getSession().getQuery());
|
||||
assertEquals(List.of("query_logs"),
|
||||
response.getToolInvocations().stream().map(DiagnosisTraceResponse.ToolInvocationTrace::getToolName).toList());
|
||||
}
|
||||
|
||||
@Test
|
||||
void getTraceRejectsRunFromAnotherSession() {
|
||||
String runId = "run-other-session";
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("path-session", runId)).thenReturn(Optional.empty());
|
||||
when(diagnosisRunRepository.findByRunId(runId)).thenReturn(Optional.of(run(
|
||||
1L, "actual-session", runId, "query", LocalDateTime.of(2026, 7, 10, 10, 0))));
|
||||
|
||||
IllegalArgumentException error = assertThrows(IllegalArgumentException.class,
|
||||
() -> service.getTrace("path-session", runId));
|
||||
|
||||
assertTrue(error.getMessage().contains("does not belong"));
|
||||
verifyNoInteractions(agentStepRepository, toolInvocationRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void getTraceThrowsWhenSessionMissing() {
|
||||
String sessionId = "missing-session";
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId))
|
||||
.thenReturn(Optional.empty());
|
||||
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.empty());
|
||||
|
||||
assertThrows(SessionNotFoundException.class, () -> service.getTrace(sessionId));
|
||||
|
||||
verify(diagnosisRunRepository).findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId);
|
||||
verify(diagnosisSessionRepository).findBySessionId(sessionId);
|
||||
verifyNoInteractions(agentStepRepository, toolInvocationRepository);
|
||||
}
|
||||
|
||||
private AgentStep step(Long id, String sessionId, int stepIndex, String agentName, LocalDateTime createdAt) {
|
||||
@Test
|
||||
void getTraceIsReadOnly() {
|
||||
String sessionId = "trace-session-readonly";
|
||||
DiagnosisRun latest = run(3L, sessionId, "run-readonly", "readonly question",
|
||||
LocalDateTime.of(2026, 7, 10, 10, 2));
|
||||
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId))
|
||||
.thenReturn(Optional.of(latest));
|
||||
when(agentStepRepository.findByRunIdOrderByStepIndex("run-readonly")).thenReturn(List.of());
|
||||
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-readonly")).thenReturn(List.of());
|
||||
|
||||
service.getTrace(sessionId);
|
||||
|
||||
verify(chatSessionRepository, never()).save(any());
|
||||
verify(diagnosisRunRepository, never()).save(any());
|
||||
verify(diagnosisSessionRepository, never()).save(any());
|
||||
verify(agentStepRepository, never()).save(any());
|
||||
verify(toolInvocationRepository, never()).save(any());
|
||||
}
|
||||
|
||||
@Test
|
||||
void listRunSummariesDoesNotExpandTraceDetails() {
|
||||
String sessionId = "trace-session-runs";
|
||||
DiagnosisRun second = run(2L, sessionId, "run-second", "second question",
|
||||
LocalDateTime.of(2026, 7, 10, 10, 1));
|
||||
second.setAnswer("answer ".repeat(40));
|
||||
DiagnosisRun first = run(1L, sessionId, "run-first", "first question",
|
||||
LocalDateTime.of(2026, 7, 10, 10, 0));
|
||||
|
||||
when(diagnosisRunRepository.findBySessionIdOrderByCreatedAtDescIdDesc(sessionId))
|
||||
.thenReturn(List.of(second, first));
|
||||
|
||||
List<DiagnosisTraceResponse.RunSummary> summaries = service.listRunSummaries(sessionId);
|
||||
|
||||
assertEquals(List.of("run-second", "run-first"),
|
||||
summaries.stream().map(DiagnosisTraceResponse.RunSummary::getRunId).toList());
|
||||
assertEquals("second question", summaries.get(0).getQuery());
|
||||
assertTrue(summaries.get(0).getAnswerPreview().length() <= 160);
|
||||
verifyNoInteractions(agentStepRepository, toolInvocationRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void legacyTraceFallbackKeepsHistoricalSessionReadable() {
|
||||
String sessionId = "legacy-session";
|
||||
DiagnosisSession legacy = DiagnosisSession.builder()
|
||||
.id(1L)
|
||||
.sessionId(sessionId)
|
||||
.query("legacy question")
|
||||
.status("SUCCESS")
|
||||
.stepCount(0)
|
||||
.toolCallCount(0)
|
||||
.createdAt(LocalDateTime.of(2026, 7, 10, 9, 0))
|
||||
.updatedAt(LocalDateTime.of(2026, 7, 10, 9, 1))
|
||||
.build();
|
||||
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc(sessionId))
|
||||
.thenReturn(Optional.empty());
|
||||
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.of(legacy));
|
||||
when(agentStepRepository.findBySessionId(sessionId)).thenReturn(List.of());
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)).thenReturn(List.of());
|
||||
|
||||
DiagnosisTraceResponse response = service.getTrace(sessionId);
|
||||
|
||||
assertNull(response.getRunId());
|
||||
assertNull(response.getRun());
|
||||
assertEquals("legacy question", response.getSession().getQuery());
|
||||
}
|
||||
|
||||
private DiagnosisRun run(Long id, String sessionId, String runId, String query, LocalDateTime createdAt) {
|
||||
return DiagnosisRun.builder()
|
||||
.id(id)
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.query(query)
|
||||
.status("SUCCESS")
|
||||
.agentFlow("CHAT")
|
||||
.answer("answer for " + runId)
|
||||
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"}}")
|
||||
.feedback("useful")
|
||||
.stepCount(1)
|
||||
.toolCallCount(1)
|
||||
.createdAt(createdAt)
|
||||
.updatedAt(createdAt.plusSeconds(1))
|
||||
.build();
|
||||
}
|
||||
|
||||
private ChatSession chatSession(String sessionId) {
|
||||
return ChatSession.builder()
|
||||
.id(99L)
|
||||
.sessionId(sessionId)
|
||||
.status("ACTIVE")
|
||||
.messagePairCount(1)
|
||||
.createdAt(LocalDateTime.of(2026, 7, 10, 9, 0))
|
||||
.lastActiveAt(LocalDateTime.of(2026, 7, 10, 10, 0))
|
||||
.build();
|
||||
}
|
||||
|
||||
private AgentStep step(Long id, String sessionId, String runId, int stepIndex, String agentName, LocalDateTime createdAt) {
|
||||
return AgentStep.builder()
|
||||
.id(id)
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.stepIndex(stepIndex)
|
||||
.agentName(agentName)
|
||||
.createdAt(createdAt)
|
||||
.build();
|
||||
}
|
||||
|
||||
private ToolInvocation invocation(Long id, String sessionId, String runId, String toolName, LocalDateTime createdAt) {
|
||||
return ToolInvocation.builder()
|
||||
.id(id)
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.toolName(toolName)
|
||||
.inputParams("{\"query\":\"timeout\"}")
|
||||
.retrievalDetails("{\"evidence_refs\":[]}")
|
||||
.success(true)
|
||||
.createdAt(createdAt)
|
||||
.build();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -11,10 +11,29 @@ import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class ExecutorGatekeeperServiceTest {
|
||||
|
||||
@Test
|
||||
void validateRunUsesRunScopedToolRows() {
|
||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||
when(repository.findByRunIdOrderByIdAsc("run-gatekeeper-1")).thenReturn(List.of(
|
||||
invocation(101L, "query_metrics", "$.alerts[0]",
|
||||
"HighCPUUsage firing, service=payment-service, current=92%")
|
||||
));
|
||||
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||
|
||||
Map<String, Object> result = service.validateRun("run-gatekeeper-1",
|
||||
validOutput(101L, "query_metrics", "$.alerts[0]",
|
||||
"HighCPUUsage firing, service=payment-service, current=92%"),
|
||||
Map.of("status", "valid"));
|
||||
|
||||
assertEquals("pass", result.get("status"));
|
||||
verify(repository).findByRunIdOrderByIdAsc("run-gatekeeper-1");
|
||||
}
|
||||
|
||||
@Test
|
||||
void ruleCatalogLoadsDefaultMetadata() {
|
||||
GatekeeperRuleCatalog catalog = GatekeeperRuleCatalog.loadDefault(new com.fasterxml.jackson.databind.ObjectMapper());
|
||||
|
||||
@@ -0,0 +1,145 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.CaseLibrary;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisRun;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.dto.FeedbackResponse;
|
||||
import com.superbiz.agent.repository.DiagnosisRunRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.springframework.test.util.ReflectionTestUtils;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.Optional;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
import static org.mockito.ArgumentMatchers.any;
|
||||
import static org.mockito.Mockito.mock;
|
||||
import static org.mockito.Mockito.never;
|
||||
import static org.mockito.Mockito.verify;
|
||||
import static org.mockito.Mockito.verifyNoInteractions;
|
||||
import static org.mockito.Mockito.when;
|
||||
|
||||
class FeedbackServiceTest {
|
||||
|
||||
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||
private final DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
|
||||
private final CaseLibraryService caseLibraryService = mock(CaseLibraryService.class);
|
||||
private final FeedbackService service = new FeedbackService();
|
||||
|
||||
@BeforeEach
|
||||
void setUp() {
|
||||
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
|
||||
ReflectionTestUtils.setField(service, "diagnosisRunRepository", diagnosisRunRepository);
|
||||
ReflectionTestUtils.setField(service, "caseLibraryService", caseLibraryService);
|
||||
}
|
||||
|
||||
@Test
|
||||
void submitFeedbackWithRunIdUpdatesSpecifiedRun() {
|
||||
DiagnosisRun run = run("session-1", "run-1");
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("session-1", "run-1"))
|
||||
.thenReturn(Optional.of(run));
|
||||
|
||||
FeedbackResponse response = service.submitFeedback("session-1", "run-1", "not_useful");
|
||||
|
||||
assertTrue(response.isSuccess());
|
||||
assertEquals("run-1", response.getRunId());
|
||||
assertFalse(response.isFallbackToLatestRun());
|
||||
assertEquals("not_useful", run.getFeedback());
|
||||
verify(diagnosisRunRepository).save(run);
|
||||
verifyNoInteractions(caseLibraryService);
|
||||
verifyNoInteractions(diagnosisSessionRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void submitFeedbackWithoutRunIdFallsBackToLatestRunObservably() {
|
||||
DiagnosisRun latest = run("session-1", "run-latest");
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc("session-1"))
|
||||
.thenReturn(Optional.of(latest));
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("session-1", "run-latest"))
|
||||
.thenReturn(Optional.of(latest));
|
||||
|
||||
FeedbackResponse response = service.submitFeedback("session-1", null, "not_useful");
|
||||
|
||||
assertTrue(response.isSuccess());
|
||||
assertEquals("run-latest", response.getRunId());
|
||||
assertTrue(response.isFallbackToLatestRun());
|
||||
assertEquals("not_useful", latest.getFeedback());
|
||||
verify(diagnosisRunRepository).save(latest);
|
||||
verify(diagnosisSessionRepository, never()).findBySessionId(any());
|
||||
}
|
||||
|
||||
@Test
|
||||
void submitFeedbackRejectsRunFromAnotherSession() {
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("session-1", "run-other"))
|
||||
.thenReturn(Optional.empty());
|
||||
when(diagnosisRunRepository.findByRunId("run-other"))
|
||||
.thenReturn(Optional.of(run("session-2", "run-other")));
|
||||
|
||||
FeedbackResponse response = service.submitFeedback("session-1", "run-other", "useful");
|
||||
|
||||
assertFalse(response.isSuccess());
|
||||
assertEquals("run-other", response.getRunId());
|
||||
verify(diagnosisRunRepository, never()).save(any());
|
||||
verifyNoInteractions(caseLibraryService);
|
||||
verifyNoInteractions(diagnosisSessionRepository);
|
||||
}
|
||||
|
||||
@Test
|
||||
void usefulFeedbackCreatesCaseFromRun() {
|
||||
DiagnosisRun run = run("session-1", "run-useful");
|
||||
CaseLibrary caseLibrary = CaseLibrary.builder().caseId("case-1").build();
|
||||
when(diagnosisRunRepository.findBySessionIdAndRunId("session-1", "run-useful"))
|
||||
.thenReturn(Optional.of(run));
|
||||
when(caseLibraryService.createFromRun(run)).thenReturn(caseLibrary);
|
||||
|
||||
FeedbackResponse response = service.submitFeedback("session-1", "run-useful", "useful");
|
||||
|
||||
assertTrue(response.isSuccess());
|
||||
assertEquals("case-1", response.getCaseId());
|
||||
assertEquals("run-useful", response.getRunId());
|
||||
verify(caseLibraryService).createFromRun(run);
|
||||
verify(diagnosisRunRepository).save(run);
|
||||
}
|
||||
|
||||
@Test
|
||||
void legacySessionFallbackPreservesOldDataCompatibility() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId("legacy-session")
|
||||
.query("legacy query")
|
||||
.answer("legacy answer")
|
||||
.build();
|
||||
CaseLibrary caseLibrary = CaseLibrary.builder().caseId("legacy-case").build();
|
||||
when(diagnosisRunRepository.findFirstBySessionIdOrderByCreatedAtDescIdDesc("legacy-session"))
|
||||
.thenReturn(Optional.empty());
|
||||
when(diagnosisSessionRepository.findBySessionId("legacy-session")).thenReturn(Optional.of(session));
|
||||
when(caseLibraryService.createFromSession(session)).thenReturn(caseLibrary);
|
||||
|
||||
FeedbackResponse response = service.submitFeedback("legacy-session", null, "useful");
|
||||
|
||||
assertTrue(response.isSuccess());
|
||||
assertNull(response.getRunId());
|
||||
assertFalse(response.isFallbackToLatestRun());
|
||||
assertEquals("legacy-case", response.getCaseId());
|
||||
assertEquals("useful", session.getFeedback());
|
||||
verify(diagnosisSessionRepository).save(session);
|
||||
verify(caseLibraryService).createFromSession(session);
|
||||
}
|
||||
|
||||
private DiagnosisRun run(String sessionId, String runId) {
|
||||
return DiagnosisRun.builder()
|
||||
.sessionId(sessionId)
|
||||
.runId(runId)
|
||||
.query("query for " + runId)
|
||||
.answer("answer for " + runId)
|
||||
.status("SUCCESS")
|
||||
.agentFlow("CHAT")
|
||||
.createdAt(LocalDateTime.of(2026, 7, 10, 10, 0))
|
||||
.updatedAt(LocalDateTime.of(2026, 7, 10, 10, 1))
|
||||
.build();
|
||||
}
|
||||
}
|
||||
@@ -26,6 +26,36 @@ class ToolInvocationRecorderTest {
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
@Test
|
||||
void recordEvidenceToolWritesRunIdFromExecutionContext() {
|
||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||
SessionContextHolder.setContext("recorder-run-session", "run-recorder-1");
|
||||
|
||||
try {
|
||||
recorder.recordEvidenceTool(
|
||||
"query_metrics",
|
||||
Map.of("query", "active_prometheus_alerts"),
|
||||
"{\"success\":true,\"alerts\":[]}",
|
||||
true,
|
||||
12,
|
||||
null,
|
||||
"prometheus_alerts",
|
||||
ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE,
|
||||
Map.of("metric_family", "prometheus_alerts")
|
||||
);
|
||||
} finally {
|
||||
SessionContextHolder.clear();
|
||||
}
|
||||
|
||||
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
|
||||
verify(repository).save(captor.capture());
|
||||
ToolInvocation saved = captor.getValue();
|
||||
assertEquals("recorder-run-session", saved.getSessionId());
|
||||
assertEquals("run-recorder-1", saved.getRunId());
|
||||
}
|
||||
|
||||
@Test
|
||||
void recordEvidenceToolPreservesNoEvidenceSemantics() throws Exception {
|
||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||
|
||||
@@ -15,6 +15,32 @@ import static org.mockito.Mockito.when;
|
||||
|
||||
class ToolTraceSummaryServiceTest {
|
||||
|
||||
@Test
|
||||
void buildVerifierTraceSummaryForRunUsesRunScopedToolRows() {
|
||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||
when(repository.findByRunIdOrderByIdAsc("run-summary-1")).thenReturn(List.of(
|
||||
ToolInvocation.builder()
|
||||
.id(101L)
|
||||
.sessionId("session-1")
|
||||
.runId("run-summary-1")
|
||||
.toolName("query_metrics")
|
||||
.inputParams("{\"query\":\"active_prometheus_alerts\"}")
|
||||
.outputPreview("active=50 max=50")
|
||||
.retrievalDetails("{\"retrieved_domains\":[\"prometheus_alerts\"],\"evidence_status\":\"supported\"}")
|
||||
.success(true)
|
||||
.build()
|
||||
));
|
||||
|
||||
ToolTraceSummaryService service = new ToolTraceSummaryService(repository);
|
||||
|
||||
List<Map<String, Object>> summaries = service.buildVerifierTraceSummaryForRun(
|
||||
"run-summary-1", "active=50 max=50");
|
||||
|
||||
assertEquals(1, summaries.size());
|
||||
assertEquals("query_metrics", summaries.get(0).get("tool_name"));
|
||||
assertEquals(List.of(101L), summaries.get(0).get("source_invocation_ids"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void buildVerifierTraceSummaryTreatsNoEvidenceAsGapWithoutLosingSuccessfulEvidence() {
|
||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||
|
||||
Reference in New Issue
Block a user