feat: add traceable scoped AIOps diagnosis
This commit is contained in:
@@ -5,6 +5,8 @@
|
|||||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||||
|
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||||
|
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||||
|
|||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# Acceptance: aiops-alert-scope-control
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
- [x] Payload-mode prompt focuses the final report on the supplied alert.
|
||||||
|
- [x] No-payload prompt requires active-alert discovery first.
|
||||||
|
- [x] Targeted tests pass.
|
||||||
|
- [x] Compile passes.
|
||||||
|
- [x] OpenSpec validates.
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- Prompt-only scope control may still require runtime observation.
|
||||||
|
- AIOps Verifier remains deferred.
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
# Brief: aiops-alert-scope-control
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Make AIOps scope explicit:
|
||||||
|
|
||||||
|
- Payload present -> targeted diagnosis for the supplied alert.
|
||||||
|
- Payload absent -> automatic active-alert discovery and diagnosis.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- In scope:
|
||||||
|
- `AiOpsService.buildTaskPrompt(...)` scope rules.
|
||||||
|
- Focused tests.
|
||||||
|
- Demo acceptance wording.
|
||||||
|
- Out of scope:
|
||||||
|
- Verifier integration.
|
||||||
|
- Java-side filtering of tool results.
|
||||||
|
- API shape changes.
|
||||||
|
- Database changes.
|
||||||
|
|
||||||
|
## Related OpenSpec
|
||||||
|
|
||||||
|
`openspec/changes/aiops-alert-scope-control/`
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# Decisions: aiops-alert-scope-control
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
|
||||||
|
- Slug: `aiops-alert-scope-control`
|
||||||
|
- Scale: standard-light.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- AIOps traceability is implemented and verified.
|
||||||
|
- Mock Prometheus returns multiple active alerts.
|
||||||
|
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
|
||||||
|
|
||||||
|
## Grill Question Pool
|
||||||
|
|
||||||
|
| # | Dimension | Question | Mode | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
|
||||||
|
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
|
||||||
|
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
|
||||||
|
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
|
||||||
|
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
|
||||||
|
|
||||||
|
## Evidence-Driven Conclusions
|
||||||
|
|
||||||
|
| Conclusion | Evidence Source | Result |
|
||||||
|
|---|---|---|
|
||||||
|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
|
||||||
|
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
|
||||||
|
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
|
||||||
|
|
||||||
|
## GitNexus
|
||||||
|
|
||||||
|
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Payload mode is detected when any alert field is present.
|
||||||
|
- Payload mode final report must focus on the supplied alert.
|
||||||
|
- No-payload mode must first call `queryPrometheusAlerts`.
|
||||||
|
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# Evidence: aiops-alert-scope-control
|
||||||
|
|
||||||
|
## Local Impact Analysis
|
||||||
|
|
||||||
|
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
|
||||||
|
- No controller, DTO, repository, or database changes are required.
|
||||||
|
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
|
||||||
|
|
||||||
|
## Verification Results
|
||||||
|
|
||||||
|
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
|
||||||
|
- `mvn -q -DskipTests compile` passed.
|
||||||
|
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
|
||||||
|
|
||||||
|
## Runtime Verification
|
||||||
|
|
||||||
|
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
|
||||||
|
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
|
||||||
|
- `diagnosis_session` persisted:
|
||||||
|
- `agent_flow = AI_OPS`
|
||||||
|
- `status = SUCCESS`
|
||||||
|
- `total_duration_ms = 69875`
|
||||||
|
- `step_count = 5`
|
||||||
|
- `tool_call_count = 8`
|
||||||
|
- Tool invocation counts:
|
||||||
|
- `query_metrics = 1`
|
||||||
|
- `lookup_knowledge = 1`
|
||||||
|
- `query_logs = 6`
|
||||||
|
- Report scope check:
|
||||||
|
- `告警根因分析 - HighCPUUsage` exists.
|
||||||
|
- `告警根因分析 - HighMemoryUsage` does not exist.
|
||||||
|
- `告警根因分析 - SlowResponse` does not exist.
|
||||||
|
- `相关风险告警` exists.
|
||||||
|
|
||||||
|
## Runtime Fix
|
||||||
|
|
||||||
|
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
|
||||||
|
- `maximum-pool-size: 5`
|
||||||
|
- `minimum-idle: 1`
|
||||||
|
- `connection-timeout: 10000`
|
||||||
|
- `validation-timeout: 5000`
|
||||||
|
- `idle-timeout: 60000`
|
||||||
|
- `max-lifetime: 120000`
|
||||||
|
- `keepalive-time: 30000`
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
# Acceptance: aiops-traceable-diagnosis-entry
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
|
||||||
|
- [x] Targeted AIOps service tests pass.
|
||||||
|
- [x] Compile verification passes.
|
||||||
|
- [x] Demo docs describe AIOps request -> session id -> trace query.
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Accepted for implementation scope.
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- AIOps Verifier integration is deferred.
|
||||||
|
- Runtime still depends on configured model and infrastructure.
|
||||||
|
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
# Brief: aiops-traceable-diagnosis-entry
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- In scope:
|
||||||
|
- Optional AIOps alert request body.
|
||||||
|
- Stable request/session id propagation.
|
||||||
|
- Persisted AIOps query summary and final answer.
|
||||||
|
- SSE session id event.
|
||||||
|
- Demo documentation and focused tests.
|
||||||
|
- Out of scope:
|
||||||
|
- Full AIOps and ChatService unification.
|
||||||
|
- AIOps Verifier integration.
|
||||||
|
- Database schema changes.
|
||||||
|
- Sensitive configuration cleanup.
|
||||||
|
- Fully offline runtime.
|
||||||
|
|
||||||
|
## Related OpenSpec
|
||||||
|
|
||||||
|
`openspec/changes/aiops-traceable-diagnosis-entry/`
|
||||||
@@ -0,0 +1,68 @@
|
|||||||
|
# Decisions: aiops-traceable-diagnosis-entry
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
|
||||||
|
- Slug: `aiops-traceable-diagnosis-entry`
|
||||||
|
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
|
||||||
|
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
|
||||||
|
|
||||||
|
## Grill Question Pool
|
||||||
|
|
||||||
|
| # | Dimension | Question | Mode | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
|
||||||
|
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
|
||||||
|
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
|
||||||
|
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
|
||||||
|
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
|
||||||
|
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
|
||||||
|
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
|
||||||
|
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
|
||||||
|
|
||||||
|
## Evidence-Driven Conclusions
|
||||||
|
|
||||||
|
| Conclusion | Evidence Source | Result |
|
||||||
|
|---|---|---|
|
||||||
|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
|
||||||
|
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
|
||||||
|
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
|
||||||
|
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
|
||||||
|
|
||||||
|
## User-Interview Confirmations
|
||||||
|
|
||||||
|
| Topic | User Words | Decision |
|
||||||
|
|---|---|---|
|
||||||
|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
|
||||||
|
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
|
||||||
|
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Keep `/api/ai_ops` and make its body optional.
|
||||||
|
- Emit `SseMessage.type=session` before long-running analysis starts.
|
||||||
|
- Store AIOps request summary in `diagnosis_session.query`.
|
||||||
|
- Store final report in `diagnosis_session.answer`.
|
||||||
|
- Defer AIOps Verifier to a later change so this slice stays focused.
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/ai_ops
|
||||||
|
-> optional AIOpsRequest
|
||||||
|
-> resolve sessionId
|
||||||
|
-> create diagnosis_session(agentFlow=AI_OPS)
|
||||||
|
-> run ai_ops_supervisor(planner, executor)
|
||||||
|
-> AgentLoggingHook persists steps
|
||||||
|
-> tools persist invocations under SessionContextHolder
|
||||||
|
-> extract final report
|
||||||
|
-> persist answer
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace replays the run
|
||||||
|
```
|
||||||
|
|
||||||
|
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
# Evidence: aiops-traceable-diagnosis-entry
|
||||||
|
|
||||||
|
## Local Impact Analysis
|
||||||
|
|
||||||
|
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
|
||||||
|
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
|
||||||
|
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
|
||||||
|
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
|
||||||
|
|
||||||
|
## GitNexus
|
||||||
|
|
||||||
|
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
|
||||||
|
|
||||||
|
## Expected Verification
|
||||||
|
|
||||||
|
- Focused unit tests for AIOps request/session/report helper behavior.
|
||||||
|
- Compile verification.
|
||||||
|
- OpenSpec validation if CLI is available.
|
||||||
|
|
||||||
|
## Verification Results
|
||||||
|
|
||||||
|
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
|
||||||
|
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
|
||||||
|
- `mvn -q -DskipTests compile`: passed.
|
||||||
|
|
||||||
|
## Demo Alignment
|
||||||
|
|
||||||
|
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
|
||||||
|
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
|
||||||
|
|
||||||
|
## Metric Alignment Follow-up
|
||||||
|
|
||||||
|
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
|
||||||
|
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
|
||||||
|
- Targeted verification:
|
||||||
|
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
|
||||||
|
- `mvn -q -DskipTests compile`: passed.
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
---
|
||||||
|
title: AIOps 告警排障 Runbook
|
||||||
|
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
|
||||||
|
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
|
||||||
|
category: troubleshooting
|
||||||
|
---
|
||||||
|
|
||||||
|
# AIOps 告警排障 Runbook
|
||||||
|
|
||||||
|
## 1. 告警输入处理原则
|
||||||
|
|
||||||
|
AIOps 诊断入口有两种触发方式:
|
||||||
|
|
||||||
|
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
|
||||||
|
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
|
||||||
|
|
||||||
|
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
|
||||||
|
|
||||||
|
## 2. Mock 告警与日志主题映射
|
||||||
|
|
||||||
|
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
|
||||||
|
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
|
||||||
|
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
|
||||||
|
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
|
||||||
|
|
||||||
|
## 3. HighCPUUsage / payment-service 排障步骤
|
||||||
|
|
||||||
|
### 3.1 现象确认
|
||||||
|
|
||||||
|
先确认 Prometheus 活动告警中是否存在:
|
||||||
|
|
||||||
|
- `alert_name = HighCPUUsage`
|
||||||
|
- `service = payment-service`
|
||||||
|
- CPU 使用率超过 80%
|
||||||
|
- 状态为 firing
|
||||||
|
|
||||||
|
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
|
||||||
|
|
||||||
|
### 3.2 指标与日志取证
|
||||||
|
|
||||||
|
推荐工具调用顺序:
|
||||||
|
|
||||||
|
1. `queryPrometheusAlerts`:确认当前活动告警。
|
||||||
|
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
|
||||||
|
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
|
||||||
|
|
||||||
|
### 3.3 根因判断
|
||||||
|
|
||||||
|
可接受的根因结论必须至少满足一项:
|
||||||
|
|
||||||
|
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
|
||||||
|
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
|
||||||
|
- 告警持续时间与日志时间线一致。
|
||||||
|
|
||||||
|
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
|
||||||
|
|
||||||
|
## 4. 处理建议
|
||||||
|
|
||||||
|
### 临时止血
|
||||||
|
|
||||||
|
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
|
||||||
|
- 对高耗时接口开启限流或降级非核心功能。
|
||||||
|
- 如果近期有发布,检查变更窗口并准备回滚。
|
||||||
|
|
||||||
|
### 根因修复
|
||||||
|
|
||||||
|
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
|
||||||
|
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
|
||||||
|
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
|
||||||
|
|
||||||
|
## 5. 报告要求
|
||||||
|
|
||||||
|
告警分析报告至少包含:
|
||||||
|
|
||||||
|
- 活跃告警清单。
|
||||||
|
- 告警根因分析。
|
||||||
|
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
|
||||||
|
- 已执行或建议执行的处理方案。
|
||||||
|
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
|
||||||
@@ -78,6 +78,43 @@ Expected result:
|
|||||||
- `success` is `true`.
|
- `success` is `true`.
|
||||||
- A later trace query shows `data.session.feedback` as `useful`.
|
- A later trace query shows `data.session.feedback` as `useful`.
|
||||||
|
|
||||||
|
## 4. Run AIOps Alert Diagnosis
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||||
|
$aiopsBody = @{
|
||||||
|
sessionId = $aiopsSessionId
|
||||||
|
alertName = "HighCPUUsage"
|
||||||
|
service = "payment-service"
|
||||||
|
severity = "P1"
|
||||||
|
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||||
|
timeRange = "last_15m"
|
||||||
|
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-WebRequest `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/ai_ops" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $aiopsBody
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected result:
|
||||||
|
|
||||||
|
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
|
||||||
|
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
|
||||||
|
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
|
||||||
|
- `data.session.answer` contains the final alert analysis report when a report is generated.
|
||||||
|
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
|
||||||
|
|
||||||
|
Query the AIOps trace:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||||
|
```
|
||||||
|
|
||||||
## Demo Story
|
## Demo Story
|
||||||
|
|
||||||
The important interview story is:
|
The important interview story is:
|
||||||
@@ -92,3 +129,14 @@ one session id
|
|||||||
-> feedback
|
-> feedback
|
||||||
-> trace API for replay and audit
|
-> trace API for replay and audit
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The AIOps story uses the same audit spine:
|
||||||
|
|
||||||
|
```text
|
||||||
|
one session id
|
||||||
|
-> alert payload
|
||||||
|
-> AIOps planner/executor execution
|
||||||
|
-> evidence tools
|
||||||
|
-> alert analysis report
|
||||||
|
-> trace API for replay and audit
|
||||||
|
```
|
||||||
|
|||||||
@@ -0,0 +1,38 @@
|
|||||||
|
# AIOps Alert Acceptance Case
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
|
||||||
|
|
||||||
|
## Input
|
||||||
|
|
||||||
|
- Session id: `mvp-demo-aiops-payment-cpu-001`
|
||||||
|
- Endpoint: `POST /api/ai_ops`
|
||||||
|
- Profile: `mvp-demo`
|
||||||
|
- Alert:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"sessionId": "mvp-demo-aiops-payment-cpu-001",
|
||||||
|
"alertName": "HighCPUUsage",
|
||||||
|
"service": "payment-service",
|
||||||
|
"severity": "P1",
|
||||||
|
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
|
||||||
|
"timeRange": "last_15m",
|
||||||
|
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Acceptance Criteria
|
||||||
|
|
||||||
|
1. The SSE stream emits a `session` message containing the requested session id.
|
||||||
|
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
|
||||||
|
3. The persisted session query contains the alert name, service, severity, time range, and description.
|
||||||
|
4. If a final report is generated, `diagnosis_session.answer` contains that report.
|
||||||
|
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
|
||||||
|
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- This slice does not add a Verifier Agent to AIOps.
|
||||||
|
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-04
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
The AIOps endpoint has two natural modes:
|
||||||
|
|
||||||
|
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
|
||||||
|
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
|
||||||
|
|
||||||
|
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Make AIOps payload mode single-alert focused.
|
||||||
|
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
|
||||||
|
- Keep the change prompt-only and low risk.
|
||||||
|
- Add tests for prompt scope rules.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- Do not add a Verifier Agent.
|
||||||
|
- Do not force tool calls in Java code.
|
||||||
|
- Do not change `/api/ai_ops` request/response contracts.
|
||||||
|
- Do not modify mock alert data.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
| Decision | Choice | Alternative Considered | Rationale |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
|
||||||
|
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
|
||||||
|
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
|
||||||
|
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
|
||||||
|
|
||||||
|
## Prompt Rules
|
||||||
|
|
||||||
|
Payload mode MUST instruct the agent:
|
||||||
|
|
||||||
|
- Treat supplied payload as the primary and only report target.
|
||||||
|
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
|
||||||
|
- Do not create root-cause sections for unrelated active alerts.
|
||||||
|
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
|
||||||
|
|
||||||
|
Auto-discovery mode MUST instruct the agent:
|
||||||
|
|
||||||
|
- First call `queryPrometheusAlerts`.
|
||||||
|
- Select P0/P1 or longest-running firing alerts.
|
||||||
|
- Analyze one or more active alerts based on severity and evidence.
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
|
||||||
|
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
|
||||||
|
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
|
||||||
|
- Preserve full active-alert discovery when no payload is supplied.
|
||||||
|
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
|
||||||
|
- Update tests and demo acceptance wording to lock the new behavior.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
|
||||||
|
- Affected API: no endpoint or request/response shape change.
|
||||||
|
- Affected persistence: no schema change.
|
||||||
|
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
|
||||||
+25
@@ -0,0 +1,25 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: AIOps payload mode focuses on supplied alert
|
||||||
|
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||||
|
|
||||||
|
#### Scenario: Request includes alertName and service
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||||
|
- **THEN** the AIOps task prompt identifies payload mode
|
||||||
|
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||||
|
|
||||||
|
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||||
|
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||||
|
|
||||||
|
#### Scenario: Request body is omitted
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||||
|
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||||
|
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||||
|
|
||||||
|
### Requirement: Payload mode may use active alerts as supporting context
|
||||||
|
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||||
|
|
||||||
|
#### Scenario: Prometheus returns multiple active alerts
|
||||||
|
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||||
|
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||||
|
- **AND** the final report target remains the supplied alert
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
## 1. Flow Records
|
||||||
|
|
||||||
|
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
|
||||||
|
- [x] 1.2 Record local impact analysis and GitNexus skip context.
|
||||||
|
|
||||||
|
## 2. Prompt Scope Control
|
||||||
|
|
||||||
|
- [x] 2.1 Add payload detection helper in `AiOpsService`.
|
||||||
|
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
|
||||||
|
|
||||||
|
## 3. Tests And Docs
|
||||||
|
|
||||||
|
- [x] 3.1 Add tests for payload-mode prompt rules.
|
||||||
|
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
|
||||||
|
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
|
||||||
|
|
||||||
|
## 4. Verification
|
||||||
|
|
||||||
|
- [x] 4.1 Run targeted tests.
|
||||||
|
- [x] 4.2 Run compile verification.
|
||||||
|
- [x] 4.3 Run OpenSpec validation.
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-04
|
||||||
@@ -0,0 +1,90 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
|
||||||
|
|
||||||
|
- `ChatController.aiOps()` accepts no request body.
|
||||||
|
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
|
||||||
|
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
|
||||||
|
|
||||||
|
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
|
||||||
|
- Preserve backward compatibility for callers that post with no request body.
|
||||||
|
- Return the resolved `sessionId` through SSE.
|
||||||
|
- Persist the final report into the existing diagnosis session.
|
||||||
|
- Keep the AIOps path observable through the existing trace API.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- Do not merge AIOps into `ChatService`.
|
||||||
|
- Do not add a new Verifier Agent to AIOps in this slice.
|
||||||
|
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- Do not add database migrations.
|
||||||
|
- Do not clean up sensitive configuration.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
| Decision | Choice | Alternative Considered | Rationale |
|
||||||
|
|---|---|---|---|
|
||||||
|
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
|
||||||
|
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
|
||||||
|
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
|
||||||
|
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
|
||||||
|
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
|
||||||
|
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- Level: L3 API behavior extension.
|
||||||
|
- Endpoint: `POST /api/ai_ops`
|
||||||
|
- Compatibility: callers may still omit a body. New callers may send:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"sessionId": "mvp-demo-aiops-payment-latency-001",
|
||||||
|
"alertName": "payment-service-latency-high",
|
||||||
|
"service": "payment-service",
|
||||||
|
"severity": "P1",
|
||||||
|
"description": "支付服务 P95 延迟升高并伴随超时错误",
|
||||||
|
"timeRange": "last_15m"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The SSE stream emits a first content message containing the resolved session id:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId: mvp-demo-aiops-payment-latency-001
|
||||||
|
```
|
||||||
|
|
||||||
|
## Data Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/ai_ops
|
||||||
|
-> ChatController resolves request body and tools
|
||||||
|
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
|
||||||
|
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
|
||||||
|
-> set SessionContextHolder(sessionId)
|
||||||
|
-> ai_ops_supervisor -> planner_agent -> executor_agent
|
||||||
|
-> persist agent_step and tool_invocation through existing hooks/tools
|
||||||
|
-> extract final report
|
||||||
|
-> persist diagnosis_session.answer/status/counts
|
||||||
|
-> caller queries GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
|
||||||
|
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
|
||||||
|
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
|
||||||
|
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
- No database migration.
|
||||||
|
- Deploy with application restart.
|
||||||
|
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
|
||||||
|
- Resolve a stable session id from the request or generate one when omitted.
|
||||||
|
- Persist the AIOps alert query and final report into `diagnosis_session`.
|
||||||
|
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
|
||||||
|
- Document the AIOps demo path beside the existing MVP demo trace flow.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
|
||||||
|
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
|
||||||
|
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
|
||||||
|
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
|
||||||
+41
@@ -0,0 +1,41 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: AIOps endpoint accepts optional alert input
|
||||||
|
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||||
|
|
||||||
|
#### Scenario: Caller supplies alert input
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||||
|
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||||
|
|
||||||
|
#### Scenario: Caller omits alert input
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||||
|
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||||
|
|
||||||
|
### Requirement: AIOps session id is traceable
|
||||||
|
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||||
|
|
||||||
|
#### Scenario: Request includes session id
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||||
|
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||||
|
- **AND** the SSE stream includes the same session id
|
||||||
|
|
||||||
|
#### Scenario: Request omits session id
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||||
|
- **THEN** the system generates a session id
|
||||||
|
- **AND** the SSE stream includes the generated session id
|
||||||
|
|
||||||
|
### Requirement: AIOps report is persisted for trace replay
|
||||||
|
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||||
|
|
||||||
|
#### Scenario: AIOps report is generated
|
||||||
|
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||||
|
- **THEN** the corresponding diagnosis session is marked successful
|
||||||
|
- **AND** `diagnosis_session.answer` stores the final report
|
||||||
|
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||||
|
|
||||||
|
### Requirement: AIOps trace uses existing evidence tables
|
||||||
|
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||||
|
|
||||||
|
#### Scenario: AIOps uses evidence tools
|
||||||
|
- **WHEN** the AIOps flow calls available evidence tools
|
||||||
|
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
## 1. Flow Records
|
||||||
|
|
||||||
|
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
|
||||||
|
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
|
||||||
|
|
||||||
|
## 2. AIOps API And Service
|
||||||
|
|
||||||
|
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
|
||||||
|
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
|
||||||
|
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
|
||||||
|
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
|
||||||
|
|
||||||
|
## 3. Demo Documentation
|
||||||
|
|
||||||
|
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
|
||||||
|
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
|
||||||
|
|
||||||
|
## 4. Verification
|
||||||
|
|
||||||
|
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
|
||||||
|
- [x] 4.2 Run targeted tests.
|
||||||
|
- [x] 4.3 Run compile verification.
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
# aiops-alert-scope-control Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
|
||||||
|
## Requirements
|
||||||
|
### Requirement: AIOps payload mode focuses on supplied alert
|
||||||
|
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||||
|
|
||||||
|
#### Scenario: Request includes alertName and service
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||||
|
- **THEN** the AIOps task prompt identifies payload mode
|
||||||
|
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||||
|
|
||||||
|
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||||
|
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||||
|
|
||||||
|
#### Scenario: Request body is omitted
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||||
|
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||||
|
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||||
|
|
||||||
|
### Requirement: Payload mode may use active alerts as supporting context
|
||||||
|
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||||
|
|
||||||
|
#### Scenario: Prometheus returns multiple active alerts
|
||||||
|
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||||
|
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||||
|
- **AND** the final report target remains the supplied alert
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# aiops-traceable-diagnosis-entry Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
|
||||||
|
## Requirements
|
||||||
|
### Requirement: AIOps endpoint accepts optional alert input
|
||||||
|
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||||
|
|
||||||
|
#### Scenario: Caller supplies alert input
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||||
|
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||||
|
|
||||||
|
#### Scenario: Caller omits alert input
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||||
|
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||||
|
|
||||||
|
### Requirement: AIOps session id is traceable
|
||||||
|
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||||
|
|
||||||
|
#### Scenario: Request includes session id
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||||
|
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||||
|
- **AND** the SSE stream includes the same session id
|
||||||
|
|
||||||
|
#### Scenario: Request omits session id
|
||||||
|
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||||
|
- **THEN** the system generates a session id
|
||||||
|
- **AND** the SSE stream includes the generated session id
|
||||||
|
|
||||||
|
### Requirement: AIOps report is persisted for trace replay
|
||||||
|
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||||
|
|
||||||
|
#### Scenario: AIOps report is generated
|
||||||
|
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||||
|
- **THEN** the corresponding diagnosis session is marked successful
|
||||||
|
- **AND** `diagnosis_session.answer` stores the final report
|
||||||
|
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||||
|
|
||||||
|
### Requirement: AIOps trace uses existing evidence tables
|
||||||
|
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||||
|
|
||||||
|
#### Scenario: AIOps uses evidence tools
|
||||||
|
- **WHEN** the AIOps flow calls available evidence tools
|
||||||
|
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||||
@@ -4,6 +4,7 @@ import com.alibaba.cloud.ai.graph.OverAllState;
|
|||||||
import lombok.Getter;
|
import lombok.Getter;
|
||||||
import lombok.Setter;
|
import lombok.Setter;
|
||||||
import com.superbiz.agent.domain.model.SessionContext;
|
import com.superbiz.agent.domain.model.SessionContext;
|
||||||
|
import com.superbiz.agent.dto.AIOpsRequest;
|
||||||
import com.superbiz.agent.service.AiOpsService;
|
import com.superbiz.agent.service.AiOpsService;
|
||||||
import com.superbiz.agent.service.ChatService;
|
import com.superbiz.agent.service.ChatService;
|
||||||
import com.superbiz.agent.service.session.SessionManager;
|
import com.superbiz.agent.service.session.SessionManager;
|
||||||
@@ -212,21 +213,23 @@ public class ChatController {
|
|||||||
* 无需用户输入,自动执行告警分析流程
|
* 无需用户输入,自动执行告警分析流程
|
||||||
*/
|
*/
|
||||||
@PostMapping(value = "/ai_ops", produces = "text/event-stream;charset=UTF-8")
|
@PostMapping(value = "/ai_ops", produces = "text/event-stream;charset=UTF-8")
|
||||||
public SseEmitter aiOps() {
|
public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) {
|
||||||
SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢)
|
SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢)
|
||||||
|
String sessionId = aiOpsService.resolveSessionId(request);
|
||||||
|
|
||||||
executor.execute(() -> {
|
executor.execute(() -> {
|
||||||
try {
|
try {
|
||||||
logger.info("收到 AI 智能运维请求 - 启动多 Agent 协作流程");
|
logger.info("收到 AI 智能运维请求 - SessionId: {}, 启动多 Agent 协作流程", sessionId);
|
||||||
|
|
||||||
ChatModel chatModel = chatService.getChatModel();
|
ChatModel chatModel = chatService.getChatModel();
|
||||||
|
|
||||||
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
|
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
|
||||||
|
|
||||||
|
emitter.send(SseEmitter.event().name("message").data(SseMessage.session(sessionId), MediaType.APPLICATION_JSON));
|
||||||
emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n")));
|
emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n")));
|
||||||
|
|
||||||
// 调用 AiOpsService 执行分析流程
|
// 调用 AiOpsService 执行分析流程
|
||||||
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks);
|
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId);
|
||||||
|
|
||||||
if (overAllStateOptional.isEmpty()) {
|
if (overAllStateOptional.isEmpty()) {
|
||||||
emitter.send(SseEmitter.event().name("message")
|
emitter.send(SseEmitter.event().name("message")
|
||||||
@@ -245,6 +248,7 @@ public class ChatController {
|
|||||||
if (finalReportOptional.isPresent()) {
|
if (finalReportOptional.isPresent()) {
|
||||||
String finalReportText = finalReportOptional.get();
|
String finalReportText = finalReportOptional.get();
|
||||||
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
|
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
|
||||||
|
aiOpsService.persistFinalReport(sessionId, finalReportText);
|
||||||
|
|
||||||
// 发送分隔线
|
// 发送分隔线
|
||||||
emitter.send(SseEmitter.event().name("message")
|
emitter.send(SseEmitter.event().name("message")
|
||||||
@@ -443,6 +447,13 @@ public class ChatController {
|
|||||||
return message;
|
return message;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public static SseMessage session(String sessionId) {
|
||||||
|
SseMessage message = new SseMessage();
|
||||||
|
message.setType("session");
|
||||||
|
message.setData(sessionId);
|
||||||
|
return message;
|
||||||
|
}
|
||||||
|
|
||||||
public static SseMessage error(String errorMessage) {
|
public static SseMessage error(String errorMessage) {
|
||||||
SseMessage message = new SseMessage();
|
SseMessage message = new SseMessage();
|
||||||
message.setType("error");
|
message.setType("error");
|
||||||
|
|||||||
@@ -8,6 +8,36 @@ import lombok.Data;
|
|||||||
@Data
|
@Data
|
||||||
public class AIOpsRequest {
|
public class AIOpsRequest {
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 诊断会话 ID;为空时后端自动生成。
|
||||||
|
*/
|
||||||
|
private String sessionId;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 告警名称。
|
||||||
|
*/
|
||||||
|
private String alertName;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 受影响服务。
|
||||||
|
*/
|
||||||
|
private String service;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 告警等级,例如 P0/P1/P2。
|
||||||
|
*/
|
||||||
|
private String severity;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 告警描述。
|
||||||
|
*/
|
||||||
|
private String description;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 排查时间范围,例如 last_15m。
|
||||||
|
*/
|
||||||
|
private String timeRange;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 用户请求描述
|
* 用户请求描述
|
||||||
*/
|
*/
|
||||||
|
|||||||
@@ -22,6 +22,11 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
|
|||||||
*/
|
*/
|
||||||
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
|
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
|
||||||
|
|
||||||
|
/**
|
||||||
|
* 根据会话ID统计真实工具调用次数
|
||||||
|
*/
|
||||||
|
long countBySessionId(String sessionId);
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 根据工具名查询所有调用
|
* 根据工具名查询所有调用
|
||||||
*/
|
*/
|
||||||
|
|||||||
@@ -10,11 +10,12 @@ import com.superbiz.agent.agent.tool.InternalDocsTools;
|
|||||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||||
import com.superbiz.agent.domain.entity.AgentStep;
|
import com.superbiz.agent.domain.entity.AgentStep;
|
||||||
import com.superbiz.agent.domain.entity.AgentStep;
|
|
||||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
|
import com.superbiz.agent.dto.AIOpsRequest;
|
||||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||||
import com.superbiz.agent.repository.AgentStepRepository;
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
import com.superbiz.agent.util.SessionContextHolder;
|
import com.superbiz.agent.util.SessionContextHolder;
|
||||||
import org.slf4j.Logger;
|
import org.slf4j.Logger;
|
||||||
import org.slf4j.LoggerFactory;
|
import org.slf4j.LoggerFactory;
|
||||||
@@ -62,6 +63,9 @@ public class AiOpsService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private AgentStepRepository agentStepRepository;
|
private AgentStepRepository agentStepRepository;
|
||||||
|
|
||||||
|
@Autowired
|
||||||
|
private ToolInvocationRepository toolInvocationRepository;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 执行 AI Ops 告警分析流程
|
* 执行 AI Ops 告警分析流程
|
||||||
*
|
*
|
||||||
@@ -71,22 +75,22 @@ public class AiOpsService {
|
|||||||
* @throws GraphRunnerException 如果 Agent 执行失败
|
* @throws GraphRunnerException 如果 Agent 执行失败
|
||||||
*/
|
*/
|
||||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
||||||
|
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
||||||
|
}
|
||||||
|
|
||||||
|
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||||
|
AIOpsRequest request, String sessionId) throws GraphRunnerException {
|
||||||
logger.info("开始执行 AI Ops 多 Agent 协作流程");
|
logger.info("开始执行 AI Ops 多 Agent 协作流程");
|
||||||
|
|
||||||
String sessionId = UUID.randomUUID().toString().substring(0, 8);
|
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
||||||
long startTime = System.currentTimeMillis();
|
long startTime = System.currentTimeMillis();
|
||||||
|
|
||||||
// 创建诊断会话
|
// 创建或更新诊断会话
|
||||||
DiagnosisSession session = DiagnosisSession.builder()
|
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
||||||
.sessionId(sessionId)
|
|
||||||
.query("AI Ops 告警分析")
|
|
||||||
.status("RUNNING")
|
|
||||||
.agentFlow("AI_OPS")
|
|
||||||
.build();
|
|
||||||
diagnosisSessionRepository.save(session);
|
diagnosisSessionRepository.save(session);
|
||||||
|
|
||||||
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
||||||
SessionContextHolder.setSessionId(sessionId);
|
SessionContextHolder.setSessionId(resolvedSessionId);
|
||||||
|
|
||||||
try {
|
try {
|
||||||
// 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook)
|
// 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook)
|
||||||
@@ -102,7 +106,7 @@ public class AiOpsService {
|
|||||||
.subAgents(List.of(plannerAgent, executorAgent))
|
.subAgents(List.of(plannerAgent, executorAgent))
|
||||||
.build();
|
.build();
|
||||||
|
|
||||||
String taskPrompt = "你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。";
|
String taskPrompt = buildTaskPrompt(request);
|
||||||
|
|
||||||
logger.info("调用 Supervisor Agent 开始编排...");
|
logger.info("调用 Supervisor Agent 开始编排...");
|
||||||
|
|
||||||
@@ -158,6 +162,87 @@ public class AiOpsService {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
public String resolveSessionId(AIOpsRequest request) {
|
||||||
|
if (request != null && !isBlank(request.getSessionId())) {
|
||||||
|
return request.getSessionId().trim();
|
||||||
|
}
|
||||||
|
return UUID.randomUUID().toString();
|
||||||
|
}
|
||||||
|
|
||||||
|
public void persistFinalReport(String sessionId, String finalReport) {
|
||||||
|
if (isBlank(sessionId) || isBlank(finalReport)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
diagnosisSessionRepository.findBySessionId(sessionId.trim()).ifPresent(session -> {
|
||||||
|
session.setAnswer(finalReport);
|
||||||
|
diagnosisSessionRepository.save(session);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
String buildQuerySummary(AIOpsRequest request) {
|
||||||
|
if (request == null) {
|
||||||
|
return "AI Ops 告警分析";
|
||||||
|
}
|
||||||
|
|
||||||
|
StringBuilder summary = new StringBuilder("AI Ops 告警分析");
|
||||||
|
appendField(summary, "告警", request.getAlertName());
|
||||||
|
appendField(summary, "服务", request.getService());
|
||||||
|
appendField(summary, "等级", request.getSeverity());
|
||||||
|
appendField(summary, "时间范围", request.getTimeRange());
|
||||||
|
appendField(summary, "描述", request.getDescription());
|
||||||
|
appendField(summary, "请求", request.getUserRequest());
|
||||||
|
return summary.toString();
|
||||||
|
}
|
||||||
|
|
||||||
|
boolean hasAlertPayload(AIOpsRequest request) {
|
||||||
|
if (request == null) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return !isBlank(request.getAlertName())
|
||||||
|
|| !isBlank(request.getService())
|
||||||
|
|| !isBlank(request.getSeverity())
|
||||||
|
|| !isBlank(request.getDescription())
|
||||||
|
|| !isBlank(request.getTimeRange());
|
||||||
|
}
|
||||||
|
|
||||||
|
String buildTaskPrompt(AIOpsRequest request) {
|
||||||
|
StringBuilder prompt = new StringBuilder();
|
||||||
|
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
|
||||||
|
prompt.append("\n\n本次告警输入:\n");
|
||||||
|
prompt.append(buildQuerySummary(request));
|
||||||
|
if (hasAlertPayload(request)) {
|
||||||
|
prompt.append("\n\nAIOps scope mode: PAYLOAD_TARGETED\n");
|
||||||
|
prompt.append("- The request includes an alert payload. Treat the supplied alert payload as the primary and only main diagnosis target.\n");
|
||||||
|
prompt.append("- The final report must focus on the supplied alert fields such as alertName, service, severity, description, and timeRange.\n");
|
||||||
|
prompt.append("- You may call queryPrometheusAlerts only to verify whether the supplied alert is still active or to identify related risk/context.\n");
|
||||||
|
prompt.append("- If queryPrometheusAlerts returns unrelated active alerts, do not create full root-cause or remediation sections for them.\n");
|
||||||
|
prompt.append("- Mention unrelated active alerts only briefly in a Related Risk section when they help explain the supplied alert.\n");
|
||||||
|
} else {
|
||||||
|
prompt.append("\n\nAIOps scope mode: AUTO_DISCOVERY\n");
|
||||||
|
prompt.append("- The request does not include alert payload fields. First call queryPrometheusAlerts to discover current active/firing alerts.\n");
|
||||||
|
prompt.append("- Prefer P0/P1 alerts or the longest-running firing alerts, then diagnose one or more alerts based on severity and evidence.\n");
|
||||||
|
prompt.append("- Use metrics, logs, and knowledge-base evidence before producing the final alert analysis report.\n");
|
||||||
|
}
|
||||||
|
return prompt.toString();
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisSession startDiagnosisSession(String sessionId, AIOpsRequest request) {
|
||||||
|
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||||
|
.orElseGet(() -> DiagnosisSession.builder()
|
||||||
|
.sessionId(sessionId)
|
||||||
|
.agentFlow("AI_OPS")
|
||||||
|
.build());
|
||||||
|
session.setQuery(buildQuerySummary(request));
|
||||||
|
session.setStatus("RUNNING");
|
||||||
|
session.setAgentFlow("AI_OPS");
|
||||||
|
session.setAnswer(null);
|
||||||
|
session.setTotalDurationMs(null);
|
||||||
|
session.setTotalTokenCount(null);
|
||||||
|
session.setStepCount(null);
|
||||||
|
session.setToolCallCount(null);
|
||||||
|
return session;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 构建 Planner Agent
|
* 构建 Planner Agent
|
||||||
*/
|
*/
|
||||||
@@ -205,25 +290,33 @@ public class AiOpsService {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** 从 agent_step 汇总指标回填 diagnosis_session */
|
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
|
||||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||||
try {
|
try {
|
||||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||||
if (steps.isEmpty()) return;
|
|
||||||
|
|
||||||
int totalTokens = 0;
|
int totalTokens = 0;
|
||||||
int stepCount = 0;
|
int stepCount = 0;
|
||||||
int toolCallCount = 0;
|
|
||||||
for (AgentStep s : steps) {
|
for (AgentStep s : steps) {
|
||||||
stepCount++;
|
stepCount++;
|
||||||
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
||||||
if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++;
|
|
||||||
}
|
}
|
||||||
|
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
|
||||||
session.setTotalTokenCount(totalTokens);
|
session.setTotalTokenCount(totalTokens);
|
||||||
session.setStepCount(stepCount);
|
session.setStepCount(stepCount);
|
||||||
session.setToolCallCount(toolCallCount);
|
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private void appendField(StringBuilder builder, String label, String value) {
|
||||||
|
if (!isBlank(value)) {
|
||||||
|
builder.append("\n- ").append(label).append(": ").append(value.trim());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean isBlank(String value) {
|
||||||
|
return value == null || value.trim().isEmpty();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -18,6 +18,7 @@ import com.superbiz.agent.hook.TokenUsageHolder;
|
|||||||
import com.superbiz.agent.hook.VerifierInputHook;
|
import com.superbiz.agent.hook.VerifierInputHook;
|
||||||
import com.superbiz.agent.repository.AgentStepRepository;
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||||
import com.superbiz.agent.util.QuestionComplexity;
|
import com.superbiz.agent.util.QuestionComplexity;
|
||||||
@@ -86,6 +87,9 @@ public class ChatService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private AgentStepRepository agentStepRepository;
|
private AgentStepRepository agentStepRepository;
|
||||||
|
|
||||||
|
@Autowired
|
||||||
|
private ToolInvocationRepository toolInvocationRepository;
|
||||||
|
|
||||||
@Autowired
|
@Autowired
|
||||||
private EvaluationService evaluationService;
|
private EvaluationService evaluationService;
|
||||||
|
|
||||||
@@ -838,25 +842,22 @@ public class ChatService {
|
|||||||
) {
|
) {
|
||||||
}
|
}
|
||||||
|
|
||||||
/** 从 agent_step 汇总 token、步数等指标回填 diagnosis_session */
|
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
|
||||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||||
try {
|
try {
|
||||||
List<com.superbiz.agent.domain.entity.AgentStep> steps =
|
List<com.superbiz.agent.domain.entity.AgentStep> steps =
|
||||||
agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||||
|
|
||||||
if (steps.isEmpty()) return;
|
|
||||||
|
|
||||||
int totalTokens = 0;
|
int totalTokens = 0;
|
||||||
int stepCount = 0;
|
int stepCount = 0;
|
||||||
int toolCallCount = 0;
|
|
||||||
for (var s : steps) {
|
for (var s : steps) {
|
||||||
stepCount++;
|
stepCount++;
|
||||||
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
|
||||||
if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++;
|
|
||||||
}
|
}
|
||||||
|
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
|
||||||
session.setTotalTokenCount(totalTokens);
|
session.setTotalTokenCount(totalTokens);
|
||||||
session.setStepCount(stepCount);
|
session.setStepCount(stepCount);
|
||||||
session.setToolCallCount(toolCallCount);
|
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -47,6 +47,14 @@ spring:
|
|||||||
username: root
|
username: root
|
||||||
password: '!Fucker123..'
|
password: '!Fucker123..'
|
||||||
driver-class-name: com.mysql.cj.jdbc.Driver
|
driver-class-name: com.mysql.cj.jdbc.Driver
|
||||||
|
hikari:
|
||||||
|
maximum-pool-size: 5
|
||||||
|
minimum-idle: 1
|
||||||
|
connection-timeout: 10000
|
||||||
|
validation-timeout: 5000
|
||||||
|
idle-timeout: 60000
|
||||||
|
max-lifetime: 120000
|
||||||
|
keepalive-time: 30000
|
||||||
|
|
||||||
# =====================================================
|
# =====================================================
|
||||||
# JPA 配置
|
# JPA 配置
|
||||||
|
|||||||
@@ -0,0 +1,169 @@
|
|||||||
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.superbiz.agent.domain.entity.AgentStep;
|
||||||
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
|
import com.superbiz.agent.dto.AIOpsRequest;
|
||||||
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
|
import org.junit.jupiter.api.BeforeEach;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.springframework.test.util.ReflectionTestUtils;
|
||||||
|
|
||||||
|
import java.util.Optional;
|
||||||
|
import java.util.List;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
import static org.mockito.Mockito.*;
|
||||||
|
|
||||||
|
class AiOpsServiceTest {
|
||||||
|
|
||||||
|
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||||
|
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
|
||||||
|
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||||
|
private final AiOpsService service = new AiOpsService();
|
||||||
|
|
||||||
|
@BeforeEach
|
||||||
|
void setUp() {
|
||||||
|
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
|
||||||
|
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
|
||||||
|
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void resolveSessionIdUsesRequestValueWhenPresent() {
|
||||||
|
AIOpsRequest request = new AIOpsRequest();
|
||||||
|
request.setSessionId(" aiops-demo-session ");
|
||||||
|
|
||||||
|
assertEquals("aiops-demo-session", service.resolveSessionId(request));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void resolveSessionIdGeneratesWhenMissing() {
|
||||||
|
String sessionId = service.resolveSessionId(null);
|
||||||
|
|
||||||
|
assertNotNull(sessionId);
|
||||||
|
assertFalse(sessionId.isBlank());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void buildQuerySummaryUsesAlertFieldsAndUserRequestFallback() {
|
||||||
|
AIOpsRequest request = new AIOpsRequest();
|
||||||
|
request.setAlertName("payment-service-latency-high");
|
||||||
|
request.setService("payment-service");
|
||||||
|
request.setSeverity("P1");
|
||||||
|
request.setTimeRange("last_15m");
|
||||||
|
request.setDescription("P95 latency is high");
|
||||||
|
request.setUserRequest("结合日志和指标排查支付超时");
|
||||||
|
|
||||||
|
String summary = service.buildQuerySummary(request);
|
||||||
|
|
||||||
|
assertTrue(summary.contains("AI Ops 告警分析"));
|
||||||
|
assertTrue(summary.contains("告警: payment-service-latency-high"));
|
||||||
|
assertTrue(summary.contains("服务: payment-service"));
|
||||||
|
assertTrue(summary.contains("等级: P1"));
|
||||||
|
assertTrue(summary.contains("时间范围: last_15m"));
|
||||||
|
assertTrue(summary.contains("描述: P95 latency is high"));
|
||||||
|
assertTrue(summary.contains("请求: 结合日志和指标排查支付超时"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void hasAlertPayloadIgnoresUserRequestOnly() {
|
||||||
|
AIOpsRequest request = new AIOpsRequest();
|
||||||
|
request.setUserRequest("please discover active alerts");
|
||||||
|
|
||||||
|
assertFalse(service.hasAlertPayload(request));
|
||||||
|
|
||||||
|
request.setAlertName("HighCPUUsage");
|
||||||
|
|
||||||
|
assertTrue(service.hasAlertPayload(request));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void buildTaskPromptUsesPayloadTargetedModeWhenAlertFieldsExist() {
|
||||||
|
AIOpsRequest request = new AIOpsRequest();
|
||||||
|
request.setAlertName("HighCPUUsage");
|
||||||
|
request.setService("payment-service");
|
||||||
|
request.setSeverity("P1");
|
||||||
|
request.setTimeRange("last_15m");
|
||||||
|
request.setDescription("CPU usage is above 80%");
|
||||||
|
|
||||||
|
String prompt = service.buildTaskPrompt(request);
|
||||||
|
|
||||||
|
assertTrue(prompt.contains("AIOps scope mode: PAYLOAD_TARGETED"));
|
||||||
|
assertTrue(prompt.contains("primary and only main diagnosis target"));
|
||||||
|
assertTrue(prompt.contains("queryPrometheusAlerts only to verify"));
|
||||||
|
assertTrue(prompt.contains("do not create full root-cause or remediation sections"));
|
||||||
|
assertTrue(prompt.contains("Related Risk"));
|
||||||
|
assertTrue(prompt.contains("告警: HighCPUUsage"));
|
||||||
|
assertTrue(prompt.contains("服务: payment-service"));
|
||||||
|
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void buildTaskPromptUsesAutoDiscoveryModeWhenAlertPayloadIsMissing() {
|
||||||
|
String nullRequestPrompt = service.buildTaskPrompt(null);
|
||||||
|
|
||||||
|
assertTrue(nullRequestPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||||
|
assertTrue(nullRequestPrompt.contains("First call queryPrometheusAlerts"));
|
||||||
|
assertTrue(nullRequestPrompt.contains("current active/firing alerts"));
|
||||||
|
assertFalse(nullRequestPrompt.contains("AIOps scope mode: PAYLOAD_TARGETED"));
|
||||||
|
|
||||||
|
AIOpsRequest userRequestOnly = new AIOpsRequest();
|
||||||
|
userRequestOnly.setUserRequest("check what is firing now");
|
||||||
|
|
||||||
|
String userRequestOnlyPrompt = service.buildTaskPrompt(userRequestOnly);
|
||||||
|
|
||||||
|
assertTrue(userRequestOnlyPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||||
|
assertTrue(userRequestOnlyPrompt.contains("First call queryPrometheusAlerts"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
|
||||||
|
DiagnosisSession session = DiagnosisSession.builder()
|
||||||
|
.sessionId("aiops-session-001")
|
||||||
|
.query("AI Ops 告警分析")
|
||||||
|
.status("SUCCESS")
|
||||||
|
.agentFlow("AI_OPS")
|
||||||
|
.build();
|
||||||
|
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
|
||||||
|
|
||||||
|
service.persistFinalReport("aiops-session-001", "# 告警分析报告");
|
||||||
|
|
||||||
|
assertEquals("# 告警分析报告", session.getAnswer());
|
||||||
|
verify(diagnosisSessionRepository).save(session);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void persistFinalReportSkipsBlankInput() {
|
||||||
|
service.persistFinalReport("aiops-session-001", " ");
|
||||||
|
|
||||||
|
verifyNoInteractions(diagnosisSessionRepository);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void backfillSessionMetricsUsesRealToolInvocationCount() {
|
||||||
|
DiagnosisSession session = DiagnosisSession.builder()
|
||||||
|
.sessionId("aiops-session-002")
|
||||||
|
.build();
|
||||||
|
AgentStep stepWithTool = AgentStep.builder()
|
||||||
|
.sessionId("aiops-session-002")
|
||||||
|
.hasToolCall(true)
|
||||||
|
.tokenCount(10)
|
||||||
|
.build();
|
||||||
|
AgentStep stepWithoutTool = AgentStep.builder()
|
||||||
|
.sessionId("aiops-session-002")
|
||||||
|
.hasToolCall(false)
|
||||||
|
.tokenCount(20)
|
||||||
|
.build();
|
||||||
|
when(agentStepRepository.findBySessionIdOrderByStepIndex("aiops-session-002"))
|
||||||
|
.thenReturn(List.of(stepWithTool, stepWithoutTool));
|
||||||
|
when(toolInvocationRepository.countBySessionId("aiops-session-002")).thenReturn(11L);
|
||||||
|
|
||||||
|
ReflectionTestUtils.invokeMethod(service, "backfillSessionMetrics", session);
|
||||||
|
|
||||||
|
assertEquals(2, session.getStepCount());
|
||||||
|
assertEquals(30, session.getTotalTokenCount());
|
||||||
|
assertEquals(11, session.getToolCallCount());
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -7,6 +7,7 @@ import com.superbiz.agent.domain.entity.AgentStep;
|
|||||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
import com.superbiz.agent.repository.AgentStepRepository;
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||||
import org.junit.jupiter.api.Test;
|
import org.junit.jupiter.api.Test;
|
||||||
@@ -143,6 +144,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
});
|
});
|
||||||
when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep()));
|
when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep()));
|
||||||
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
|
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
|
||||||
|
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||||
|
when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L);
|
||||||
|
|
||||||
EvaluationService evaluationService = mock(EvaluationService.class);
|
EvaluationService evaluationService = mock(EvaluationService.class);
|
||||||
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
|
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
|
||||||
@@ -158,6 +161,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class)));
|
ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class)));
|
||||||
ReflectionTestUtils.setField(chatService, "diagnosisSessionRepository", diagnosisSessionRepository);
|
ReflectionTestUtils.setField(chatService, "diagnosisSessionRepository", diagnosisSessionRepository);
|
||||||
ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository);
|
ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository);
|
||||||
|
ReflectionTestUtils.setField(chatService, "toolInvocationRepository", toolInvocationRepository);
|
||||||
ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService);
|
ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService);
|
||||||
ReflectionTestUtils.setField(chatService, "retrievedDocTracker", retrievedDocTracker);
|
ReflectionTestUtils.setField(chatService, "retrievedDocTracker", retrievedDocTracker);
|
||||||
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);
|
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);
|
||||||
|
|||||||
Reference in New Issue
Block a user