diff --git a/devflow/index.md b/devflow/index.md index b06943b..09e4822 100644 --- a/devflow/index.md +++ b/devflow/index.md @@ -5,6 +5,8 @@ | 日期 | slug | 领域 | 关键词 | 状态 | |---|---|---|---|---| | 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived | +| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived | +| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived | | 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived | | 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived | | 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived | diff --git a/devflow/projects/2026-07-04-aiops-alert-scope-control/acceptance.md b/devflow/projects/2026-07-04-aiops-alert-scope-control/acceptance.md new file mode 100644 index 0000000..9fbfce8 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-alert-scope-control/acceptance.md @@ -0,0 +1,14 @@ +# Acceptance: aiops-alert-scope-control + +## Verification + +- [x] Payload-mode prompt focuses the final report on the supplied alert. +- [x] No-payload prompt requires active-alert discovery first. +- [x] Targeted tests pass. +- [x] Compile passes. +- [x] OpenSpec validates. + +## Known Limits + +- Prompt-only scope control may still require runtime observation. +- AIOps Verifier remains deferred. diff --git a/devflow/projects/2026-07-04-aiops-alert-scope-control/brief.md b/devflow/projects/2026-07-04-aiops-alert-scope-control/brief.md new file mode 100644 index 0000000..39a87c9 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-alert-scope-control/brief.md @@ -0,0 +1,28 @@ +# Brief: aiops-alert-scope-control + +## Background + +After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts. + +## Goal + +Make AIOps scope explicit: + +- Payload present -> targeted diagnosis for the supplied alert. +- Payload absent -> automatic active-alert discovery and diagnosis. + +## Scope + +- In scope: + - `AiOpsService.buildTaskPrompt(...)` scope rules. + - Focused tests. + - Demo acceptance wording. +- Out of scope: + - Verifier integration. + - Java-side filtering of tool results. + - API shape changes. + - Database changes. + +## Related OpenSpec + +`openspec/changes/aiops-alert-scope-control/` diff --git a/devflow/projects/2026-07-04-aiops-alert-scope-control/decisions.md b/devflow/projects/2026-07-04-aiops-alert-scope-control/decisions.md new file mode 100644 index 0000000..c95fdf6 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-alert-scope-control/decisions.md @@ -0,0 +1,42 @@ +# Decisions: aiops-alert-scope-control + +## Clarify + +- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts. +- Slug: `aiops-alert-scope-control` +- Scale: standard-light. + +## Context + +- AIOps traceability is implemented and verified. +- Mock Prometheus returns multiple active alerts. +- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`. + +## Grill Question Pool + +| # | Dimension | Question | Mode | Status | +|---|---|---|---|---| +| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. | +| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. | +| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. | +| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. | +| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. | + +## Evidence-Driven Conclusions + +| Conclusion | Evidence Source | Result | +|---|---|---| +| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. | +| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. | +| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. | + +## GitNexus + +GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead. + +## Key Decisions + +- Payload mode is detected when any alert field is present. +- Payload mode final report must focus on the supplied alert. +- No-payload mode must first call `queryPrometheusAlerts`. +- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections. diff --git a/devflow/projects/2026-07-04-aiops-alert-scope-control/evidence.md b/devflow/projects/2026-07-04-aiops-alert-scope-control/evidence.md new file mode 100644 index 0000000..11bfe11 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-alert-scope-control/evidence.md @@ -0,0 +1,44 @@ +# Evidence: aiops-alert-scope-control + +## Local Impact Analysis + +- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`. +- No controller, DTO, repository, or database changes are required. +- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules. + +## Verification Results + +- `mvn -q "-Dtest=AiOpsServiceTest" test` passed. +- `mvn -q -DskipTests compile` passed. +- `openspec.cmd validate aiops-alert-scope-control --strict` passed. + +## Runtime Verification + +- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`. +- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`. +- `diagnosis_session` persisted: + - `agent_flow = AI_OPS` + - `status = SUCCESS` + - `total_duration_ms = 69875` + - `step_count = 5` + - `tool_call_count = 8` +- Tool invocation counts: + - `query_metrics = 1` + - `lookup_knowledge = 1` + - `query_logs = 6` +- Report scope check: + - `告警根因分析 - HighCPUUsage` exists. + - `告警根因分析 - HighMemoryUsage` does not exist. + - `告警根因分析 - SlowResponse` does not exist. + - `相关风险告警` exists. + +## Runtime Fix + +- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections: + - `maximum-pool-size: 5` + - `minimum-idle: 1` + - `connection-timeout: 10000` + - `validation-timeout: 5000` + - `idle-timeout: 60000` + - `max-lifetime: 120000` + - `keepalive-time: 30000` diff --git a/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/acceptance.md b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/acceptance.md new file mode 100644 index 0000000..dec5c82 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/acceptance.md @@ -0,0 +1,18 @@ +# Acceptance: aiops-traceable-diagnosis-entry + +## Verification + +- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`. +- [x] Targeted AIOps service tests pass. +- [x] Compile verification passes. +- [x] Demo docs describe AIOps request -> session id -> trace query. + +## Result + +Accepted for implementation scope. + +## Known Limits + +- AIOps Verifier integration is deferred. +- Runtime still depends on configured model and infrastructure. +- Full browser/SSE runtime verification is not guaranteed in this coding pass. diff --git a/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/brief.md b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/brief.md new file mode 100644 index 0000000..5a10c15 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/brief.md @@ -0,0 +1,28 @@ +# Brief: aiops-traceable-diagnosis-entry + +## Background + +The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers. + +## Goal + +Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow. + +## Scope + +- In scope: + - Optional AIOps alert request body. + - Stable request/session id propagation. + - Persisted AIOps query summary and final answer. + - SSE session id event. + - Demo documentation and focused tests. +- Out of scope: + - Full AIOps and ChatService unification. + - AIOps Verifier integration. + - Database schema changes. + - Sensitive configuration cleanup. + - Fully offline runtime. + +## Related OpenSpec + +`openspec/changes/aiops-traceable-diagnosis-entry/` diff --git a/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/decisions.md b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/decisions.md new file mode 100644 index 0000000..154d3f6 --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/decisions.md @@ -0,0 +1,68 @@ +# Decisions: aiops-traceable-diagnosis-entry + +## Clarify + +- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP. +- Slug: `aiops-traceable-diagnosis-entry` +- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure. + +## Context + +- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`. +- `chat-verifier-agent` made the chat path stronger than the older AIOps path. +- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace. + +## Grill Question Pool + +| # | Dimension | Question | Mode | Status | +|---|---|---|---|---| +| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story | +| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body | +| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback | +| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` | +| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion | +| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up | +| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior | +| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision | + +## Evidence-Driven Conclusions + +| Conclusion | Evidence Source | Result | +|---|---|---| +| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. | +| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. | +| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. | +| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. | + +## User-Interview Confirmations + +| Topic | User Words | Decision | +|---|---|---| +| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. | +| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. | +| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. | + +## Key Decisions + +- Keep `/api/ai_ops` and make its body optional. +- Emit `SseMessage.type=session` before long-running analysis starts. +- Store AIOps request summary in `diagnosis_session.query`. +- Store final report in `diagnosis_session.answer`. +- Defer AIOps Verifier to a later change so this slice stays focused. + +## Architecture Audit + +```text +POST /api/ai_ops + -> optional AIOpsRequest + -> resolve sessionId + -> create diagnosis_session(agentFlow=AI_OPS) + -> run ai_ops_supervisor(planner, executor) + -> AgentLoggingHook persists steps + -> tools persist invocations under SessionContextHolder + -> extract final report + -> persist answer + -> GET /api/diagnosis/{sessionId}/trace replays the run +``` + +Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema. diff --git a/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/evidence.md b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/evidence.md new file mode 100644 index 0000000..3a61e8b --- /dev/null +++ b/devflow/projects/2026-07-04-aiops-traceable-diagnosis-entry/evidence.md @@ -0,0 +1,37 @@ +# Evidence: aiops-traceable-diagnosis-entry + +## Local Impact Analysis + +- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`. +- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`. +- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it. +- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`. + +## GitNexus + +GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead. + +## Expected Verification + +- Focused unit tests for AIOps request/session/report helper behavior. +- Compile verification. +- OpenSpec validation if CLI is available. + +## Verification Results + +- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed. +- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access. +- `mvn -q -DskipTests compile`: passed. + +## Demo Alignment + +- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance. +- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence. + +## Metric Alignment Follow-up + +- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records. +- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`. +- Targeted verification: + - `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed. + - `mvn -q -DskipTests compile`: passed. diff --git a/knowledge_base/troubleshooting/aiops-alert-runbook.md b/knowledge_base/troubleshooting/aiops-alert-runbook.md new file mode 100644 index 0000000..a89fed4 --- /dev/null +++ b/knowledge_base/troubleshooting/aiops-alert-runbook.md @@ -0,0 +1,81 @@ +--- +title: AIOps 告警排障 Runbook +keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs] +summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。 +category: troubleshooting +--- + +# AIOps 告警排障 Runbook + +## 1. 告警输入处理原则 + +AIOps 诊断入口有两种触发方式: + +- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。 +- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。 + +最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。 + +## 2. Mock 告警与日志主题映射 + +| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 | +|---|---|---|---| +| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` | +| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` | +| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` | +| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` | + +## 3. HighCPUUsage / payment-service 排障步骤 + +### 3.1 现象确认 + +先确认 Prometheus 活动告警中是否存在: + +- `alert_name = HighCPUUsage` +- `service = payment-service` +- CPU 使用率超过 80% +- 状态为 firing + +如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。 + +### 3.2 指标与日志取证 + +推荐工具调用顺序: + +1. `queryPrometheusAlerts`:确认当前活动告警。 +2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。 +3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。 + +### 3.3 根因判断 + +可接受的根因结论必须至少满足一项: + +- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。 +- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。 +- 告警持续时间与日志时间线一致。 + +如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。 + +## 4. 处理建议 + +### 临时止血 + +- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。 +- 对高耗时接口开启限流或降级非核心功能。 +- 如果近期有发布,检查变更窗口并准备回滚。 + +### 根因修复 + +- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。 +- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。 +- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。 + +## 5. 报告要求 + +告警分析报告至少包含: + +- 活跃告警清单。 +- 告警根因分析。 +- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。 +- 已执行或建议执行的处理方案。 +- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。 diff --git a/mvp/demo/README.md b/mvp/demo/README.md index 3974c12..7b2dca3 100644 --- a/mvp/demo/README.md +++ b/mvp/demo/README.md @@ -78,6 +78,43 @@ Expected result: - `success` is `true`. - A later trace query shows `data.session.feedback` as `useful`. +## 4. Run AIOps Alert Diagnosis + +```powershell +$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001" +$aiopsBody = @{ + sessionId = $aiopsSessionId + alertName = "HighCPUUsage" + service = "payment-service" + severity = "P1" + description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。" + timeRange = "last_15m" + userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。" +} | ConvertTo-Json + +Invoke-WebRequest ` + -Method Post ` + -Uri "http://localhost:9900/api/ai_ops" ` + -ContentType "application/json" ` + -Body $aiopsBody +``` + +Expected result: + +- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`. +- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload. +- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`. +- `data.session.answer` contains the final alert analysis report when a report is generated. +- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them. + +Query the AIOps trace: + +```powershell +Invoke-RestMethod ` + -Method Get ` + -Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace" +``` + ## Demo Story The important interview story is: @@ -92,3 +129,14 @@ one session id -> feedback -> trace API for replay and audit ``` + +The AIOps story uses the same audit spine: + +```text +one session id +-> alert payload +-> AIOps planner/executor execution +-> evidence tools +-> alert analysis report +-> trace API for replay and audit +``` diff --git a/mvp/demo/aiops-alert-acceptance.md b/mvp/demo/aiops-alert-acceptance.md new file mode 100644 index 0000000..97129ae --- /dev/null +++ b/mvp/demo/aiops-alert-acceptance.md @@ -0,0 +1,38 @@ +# AIOps Alert Acceptance Case + +## Goal + +Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry. + +## Input + +- Session id: `mvp-demo-aiops-payment-cpu-001` +- Endpoint: `POST /api/ai_ops` +- Profile: `mvp-demo` +- Alert: + +```json +{ + "sessionId": "mvp-demo-aiops-payment-cpu-001", + "alertName": "HighCPUUsage", + "service": "payment-service", + "severity": "P1", + "description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。", + "timeRange": "last_15m", + "userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。" +} +``` + +## Acceptance Criteria + +1. The SSE stream emits a `session` message containing the requested session id. +2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`. +3. The persisted session query contains the alert name, service, severity, time range, and description. +4. If a final report is generated, `diagnosis_session.answer` contains that report. +5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations. +6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections. + +## Known Limits + +- This slice does not add a Verifier Agent to AIOps. +- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration. diff --git a/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/.openspec.yaml b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/.openspec.yaml new file mode 100644 index 0000000..d86f152 --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-07-04 diff --git a/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/design.md b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/design.md new file mode 100644 index 0000000..ba8dfcd --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/design.md @@ -0,0 +1,54 @@ +## Context + +The AIOps endpoint has two natural modes: + +- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert. +- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them. + +The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied. + +## Goals / Non-Goals + +**Goals:** + +- Make AIOps payload mode single-alert focused. +- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior. +- Keep the change prompt-only and low risk. +- Add tests for prompt scope rules. + +**Non-Goals:** + +- Do not add a Verifier Agent. +- Do not force tool calls in Java code. +- Do not change `/api/ai_ops` request/response contracts. +- Do not modify mock alert data. + +## Decisions + +| Decision | Choice | Alternative Considered | Rationale | +|---|---|---|---| +| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. | +| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. | +| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. | +| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. | + +## Prompt Rules + +Payload mode MUST instruct the agent: + +- Treat supplied payload as the primary and only report target. +- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk. +- Do not create root-cause sections for unrelated active alerts. +- Report unrelated alerts only in a brief "关联风险" note if they appear relevant. + +Auto-discovery mode MUST instruct the agent: + +- First call `queryPrometheusAlerts`. +- Select P0/P1 or longest-running firing alerts. +- Analyze one or more active alerts based on severity and evidence. + +## Risks / Trade-offs + +- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace. +- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections. +- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier. diff --git a/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/proposal.md b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/proposal.md new file mode 100644 index 0000000..0d9347c --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/proposal.md @@ -0,0 +1,27 @@ +## Why + +Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery. + +## What Changes + +- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert. +- Preserve full active-alert discovery when no payload is supplied. +- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts. +- Update tests and demo acceptance wording to lock the new behavior. + +## Capabilities + +### New Capabilities + +- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode. + +### Modified Capabilities + +- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis. + +## Impact + +- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records. +- Affected API: no endpoint or request/response shape change. +- Affected persistence: no schema change. +- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat. diff --git a/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/specs/aiops-alert-scope-control/spec.md b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/specs/aiops-alert-scope-control/spec.md new file mode 100644 index 0000000..5a6f535 --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/specs/aiops-alert-scope-control/spec.md @@ -0,0 +1,25 @@ +## ADDED Requirements + +### Requirement: AIOps payload mode focuses on supplied alert +When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert. + +#### Scenario: Request includes alertName and service +- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service` +- **THEN** the AIOps task prompt identifies payload mode +- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts + +### Requirement: AIOps auto-discovery mode queries active alerts first +When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts. + +#### Scenario: Request body is omitted +- **WHEN** a caller posts to `/api/ai_ops` without alert fields +- **THEN** the AIOps task prompt identifies auto-discovery mode +- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first + +### Requirement: Payload mode may use active alerts as supporting context +Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections. + +#### Scenario: Prometheus returns multiple active alerts +- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts +- **THEN** the prompt permits mentioning those alerts only as related risk or context +- **AND** the final report target remains the supplied alert diff --git a/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/tasks.md b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/tasks.md new file mode 100644 index 0000000..7393ebc --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-alert-scope-control/tasks.md @@ -0,0 +1,21 @@ +## 1. Flow Records + +- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records. +- [x] 1.2 Record local impact analysis and GitNexus skip context. + +## 2. Prompt Scope Control + +- [x] 2.1 Add payload detection helper in `AiOpsService`. +- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules. + +## 3. Tests And Docs + +- [x] 3.1 Add tests for payload-mode prompt rules. +- [x] 3.2 Add tests for no-payload auto-discovery prompt rules. +- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode. + +## 4. Verification + +- [x] 4.1 Run targeted tests. +- [x] 4.2 Run compile verification. +- [x] 4.3 Run OpenSpec validation. diff --git a/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/.openspec.yaml b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/.openspec.yaml new file mode 100644 index 0000000..d86f152 --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-07-04 diff --git a/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/design.md b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/design.md new file mode 100644 index 0000000..b536f8f --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/design.md @@ -0,0 +1,90 @@ +## Context + +The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story: + +- `ChatController.aiOps()` accepts no request body. +- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally. +- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`. +- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`. + +The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence. + +## Goals / Non-Goals + +**Goals:** + +- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point. +- Preserve backward compatibility for callers that post with no request body. +- Return the resolved `sessionId` through SSE. +- Persist the final report into the existing diagnosis session. +- Keep the AIOps path observable through the existing trace API. + +**Non-Goals:** + +- Do not merge AIOps into `ChatService`. +- Do not add a new Verifier Agent to AIOps in this slice. +- Do not change `GET /api/diagnosis/{sessionId}/trace`. +- Do not add database migrations. +- Do not clean up sensitive configuration. + +## Decisions + +| Decision | Choice | Alternative Considered | Rationale | +|---|---|---|---| +| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. | +| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. | +| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. | +| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. | +| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. | +| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. | + +## Interface Impact + +- Level: L3 API behavior extension. +- Endpoint: `POST /api/ai_ops` +- Compatibility: callers may still omit a body. New callers may send: + +```json +{ + "sessionId": "mvp-demo-aiops-payment-latency-001", + "alertName": "payment-service-latency-high", + "service": "payment-service", + "severity": "P1", + "description": "支付服务 P95 延迟升高并伴随超时错误", + "timeRange": "last_15m" +} +``` + +The SSE stream emits a first content message containing the resolved session id: + +```text +sessionId: mvp-demo-aiops-payment-latency-001 +``` + +## Data Flow + +```text +POST /api/ai_ops + -> ChatController resolves request body and tools + -> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request) + -> create diagnosis_session(agentFlow=AI_OPS, query=) + -> set SessionContextHolder(sessionId) + -> ai_ops_supervisor -> planner_agent -> executor_agent + -> persist agent_step and tool_invocation through existing hooks/tools + -> extract final report + -> persist diagnosis_session.answer/status/counts + -> caller queries GET /api/diagnosis/{sessionId}/trace +``` + +## Risks / Trade-offs + +- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability. +- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report. +- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code. +- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope. + +## Migration Plan + +- No database migration. +- Deploy with application restart. +- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid. diff --git a/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/proposal.md b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/proposal.md new file mode 100644 index 0000000..9e5f029 --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/proposal.md @@ -0,0 +1,29 @@ +## Why + +The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path. + +## What Changes + +- Allow `/api/ai_ops` to accept an optional alert diagnosis request body. +- Resolve a stable session id from the request or generate one when omitted. +- Persist the AIOps alert query and final report into `diagnosis_session`. +- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`. +- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice. +- Document the AIOps demo path beside the existing MVP demo trace flow. + +## Capabilities + +### New Capabilities + +- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API. + +### Modified Capabilities + +- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint. + +## Impact + +- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records. +- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`. +- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts. +- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice. diff --git a/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/specs/aiops-traceable-diagnosis-entry/spec.md b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/specs/aiops-traceable-diagnosis-entry/spec.md new file mode 100644 index 0000000..079ce00 --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/specs/aiops-traceable-diagnosis-entry/spec.md @@ -0,0 +1,41 @@ +## ADDED Requirements + +### Requirement: AIOps endpoint accepts optional alert input +The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request. + +#### Scenario: Caller supplies alert input +- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range +- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query + +#### Scenario: Caller omits alert input +- **WHEN** a caller posts to `/api/ai_ops` without a body +- **THEN** the system still starts the default AIOps alert-analysis flow + +### Requirement: AIOps session id is traceable +The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller. + +#### Scenario: Request includes session id +- **WHEN** a caller posts to `/api/ai_ops` with `sessionId` +- **THEN** the created `diagnosis_session.session_id` equals that value +- **AND** the SSE stream includes the same session id + +#### Scenario: Request omits session id +- **WHEN** a caller posts to `/api/ai_ops` without `sessionId` +- **THEN** the system generates a session id +- **AND** the SSE stream includes the generated session id + +### Requirement: AIOps report is persisted for trace replay +The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available. + +#### Scenario: AIOps report is generated +- **WHEN** the AIOps planner/executor flow returns a final report +- **THEN** the corresponding diagnosis session is marked successful +- **AND** `diagnosis_session.answer` stores the final report +- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer + +### Requirement: AIOps trace uses existing evidence tables +The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence. + +#### Scenario: AIOps uses evidence tools +- **WHEN** the AIOps flow calls available evidence tools +- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id diff --git a/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/tasks.md b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/tasks.md new file mode 100644 index 0000000..ff9e11b --- /dev/null +++ b/openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry/tasks.md @@ -0,0 +1,22 @@ +## 1. Flow Records + +- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`. +- [x] 1.2 Record user-approved GitNexus skip and local impact analysis. + +## 2. AIOps API And Service + +- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields. +- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE. +- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary. +- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`. + +## 3. Demo Documentation + +- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`. +- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`. + +## 4. Verification + +- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical. +- [x] 4.2 Run targeted tests. +- [x] 4.3 Run compile verification. diff --git a/openspec/specs/aiops-alert-scope-control/spec.md b/openspec/specs/aiops-alert-scope-control/spec.md new file mode 100644 index 0000000..46c50e4 --- /dev/null +++ b/openspec/specs/aiops-alert-scope-control/spec.md @@ -0,0 +1,28 @@ +# aiops-alert-scope-control Specification + +## Purpose +TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive. +## Requirements +### Requirement: AIOps payload mode focuses on supplied alert +When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert. + +#### Scenario: Request includes alertName and service +- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service` +- **THEN** the AIOps task prompt identifies payload mode +- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts + +### Requirement: AIOps auto-discovery mode queries active alerts first +When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts. + +#### Scenario: Request body is omitted +- **WHEN** a caller posts to `/api/ai_ops` without alert fields +- **THEN** the AIOps task prompt identifies auto-discovery mode +- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first + +### Requirement: Payload mode may use active alerts as supporting context +Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections. + +#### Scenario: Prometheus returns multiple active alerts +- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts +- **THEN** the prompt permits mentioning those alerts only as related risk or context +- **AND** the final report target remains the supplied alert diff --git a/openspec/specs/aiops-traceable-diagnosis-entry/spec.md b/openspec/specs/aiops-traceable-diagnosis-entry/spec.md new file mode 100644 index 0000000..505b956 --- /dev/null +++ b/openspec/specs/aiops-traceable-diagnosis-entry/spec.md @@ -0,0 +1,44 @@ +# aiops-traceable-diagnosis-entry Specification + +## Purpose +TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive. +## Requirements +### Requirement: AIOps endpoint accepts optional alert input +The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request. + +#### Scenario: Caller supplies alert input +- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range +- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query + +#### Scenario: Caller omits alert input +- **WHEN** a caller posts to `/api/ai_ops` without a body +- **THEN** the system still starts the default AIOps alert-analysis flow + +### Requirement: AIOps session id is traceable +The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller. + +#### Scenario: Request includes session id +- **WHEN** a caller posts to `/api/ai_ops` with `sessionId` +- **THEN** the created `diagnosis_session.session_id` equals that value +- **AND** the SSE stream includes the same session id + +#### Scenario: Request omits session id +- **WHEN** a caller posts to `/api/ai_ops` without `sessionId` +- **THEN** the system generates a session id +- **AND** the SSE stream includes the generated session id + +### Requirement: AIOps report is persisted for trace replay +The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available. + +#### Scenario: AIOps report is generated +- **WHEN** the AIOps planner/executor flow returns a final report +- **THEN** the corresponding diagnosis session is marked successful +- **AND** `diagnosis_session.answer` stores the final report +- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer + +### Requirement: AIOps trace uses existing evidence tables +The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence. + +#### Scenario: AIOps uses evidence tools +- **WHEN** the AIOps flow calls available evidence tools +- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id diff --git a/src/main/java/com/superbiz/agent/controller/ChatController.java b/src/main/java/com/superbiz/agent/controller/ChatController.java index 4f058f2..9b65160 100644 --- a/src/main/java/com/superbiz/agent/controller/ChatController.java +++ b/src/main/java/com/superbiz/agent/controller/ChatController.java @@ -4,6 +4,7 @@ import com.alibaba.cloud.ai.graph.OverAllState; import lombok.Getter; import lombok.Setter; import com.superbiz.agent.domain.model.SessionContext; +import com.superbiz.agent.dto.AIOpsRequest; import com.superbiz.agent.service.AiOpsService; import com.superbiz.agent.service.ChatService; import com.superbiz.agent.service.session.SessionManager; @@ -212,21 +213,23 @@ public class ChatController { * 无需用户输入,自动执行告警分析流程 */ @PostMapping(value = "/ai_ops", produces = "text/event-stream;charset=UTF-8") - public SseEmitter aiOps() { + public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) { SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢) + String sessionId = aiOpsService.resolveSessionId(request); executor.execute(() -> { try { - logger.info("收到 AI 智能运维请求 - 启动多 Agent 协作流程"); + logger.info("收到 AI 智能运维请求 - SessionId: {}, 启动多 Agent 协作流程", sessionId); ChatModel chatModel = chatService.getChatModel(); ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0]; + emitter.send(SseEmitter.event().name("message").data(SseMessage.session(sessionId), MediaType.APPLICATION_JSON)); emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n"))); // 调用 AiOpsService 执行分析流程 - Optional overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks); + Optional overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId); if (overAllStateOptional.isEmpty()) { emitter.send(SseEmitter.event().name("message") @@ -245,6 +248,7 @@ public class ChatController { if (finalReportOptional.isPresent()) { String finalReportText = finalReportOptional.get(); logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length()); + aiOpsService.persistFinalReport(sessionId, finalReportText); // 发送分隔线 emitter.send(SseEmitter.event().name("message") @@ -443,6 +447,13 @@ public class ChatController { return message; } + public static SseMessage session(String sessionId) { + SseMessage message = new SseMessage(); + message.setType("session"); + message.setData(sessionId); + return message; + } + public static SseMessage error(String errorMessage) { SseMessage message = new SseMessage(); message.setType("error"); diff --git a/src/main/java/com/superbiz/agent/dto/AIOpsRequest.java b/src/main/java/com/superbiz/agent/dto/AIOpsRequest.java index fcc0b4d..333bdde 100644 --- a/src/main/java/com/superbiz/agent/dto/AIOpsRequest.java +++ b/src/main/java/com/superbiz/agent/dto/AIOpsRequest.java @@ -7,6 +7,36 @@ import lombok.Data; */ @Data public class AIOpsRequest { + + /** + * 诊断会话 ID;为空时后端自动生成。 + */ + private String sessionId; + + /** + * 告警名称。 + */ + private String alertName; + + /** + * 受影响服务。 + */ + private String service; + + /** + * 告警等级,例如 P0/P1/P2。 + */ + private String severity; + + /** + * 告警描述。 + */ + private String description; + + /** + * 排查时间范围,例如 last_15m。 + */ + private String timeRange; /** * 用户请求描述 diff --git a/src/main/java/com/superbiz/agent/repository/ToolInvocationRepository.java b/src/main/java/com/superbiz/agent/repository/ToolInvocationRepository.java index b48343e..a682b26 100644 --- a/src/main/java/com/superbiz/agent/repository/ToolInvocationRepository.java +++ b/src/main/java/com/superbiz/agent/repository/ToolInvocationRepository.java @@ -22,6 +22,11 @@ public interface ToolInvocationRepository extends JpaRepository findBySessionIdOrderByIdAsc(String sessionId); + /** + * 根据会话ID统计真实工具调用次数 + */ + long countBySessionId(String sessionId); + /** * 根据工具名查询所有调用 */ diff --git a/src/main/java/com/superbiz/agent/service/AiOpsService.java b/src/main/java/com/superbiz/agent/service/AiOpsService.java index 523adba..ec96032 100644 --- a/src/main/java/com/superbiz/agent/service/AiOpsService.java +++ b/src/main/java/com/superbiz/agent/service/AiOpsService.java @@ -10,11 +10,12 @@ import com.superbiz.agent.agent.tool.InternalDocsTools; import com.superbiz.agent.agent.tool.QueryLogsTools; import com.superbiz.agent.agent.tool.QueryMetricsTools; import com.superbiz.agent.domain.entity.AgentStep; -import com.superbiz.agent.domain.entity.AgentStep; import com.superbiz.agent.domain.entity.DiagnosisSession; +import com.superbiz.agent.dto.AIOpsRequest; import com.superbiz.agent.hook.AgentLoggingHook; import com.superbiz.agent.repository.AgentStepRepository; import com.superbiz.agent.repository.DiagnosisSessionRepository; +import com.superbiz.agent.repository.ToolInvocationRepository; import com.superbiz.agent.util.SessionContextHolder; import org.slf4j.Logger; import org.slf4j.LoggerFactory; @@ -62,6 +63,9 @@ public class AiOpsService { @Autowired private AgentStepRepository agentStepRepository; + @Autowired + private ToolInvocationRepository toolInvocationRepository; + /** * 执行 AI Ops 告警分析流程 * @@ -71,22 +75,22 @@ public class AiOpsService { * @throws GraphRunnerException 如果 Agent 执行失败 */ public Optional executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException { + return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null)); + } + + public Optional executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks, + AIOpsRequest request, String sessionId) throws GraphRunnerException { logger.info("开始执行 AI Ops 多 Agent 协作流程"); - String sessionId = UUID.randomUUID().toString().substring(0, 8); + String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim(); long startTime = System.currentTimeMillis(); - // 创建诊断会话 - DiagnosisSession session = DiagnosisSession.builder() - .sessionId(sessionId) - .query("AI Ops 告警分析") - .status("RUNNING") - .agentFlow("AI_OPS") - .build(); + // 创建或更新诊断会话 + DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request); diagnosisSessionRepository.save(session); // 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId) - SessionContextHolder.setSessionId(sessionId); + SessionContextHolder.setSessionId(resolvedSessionId); try { // 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook) @@ -102,7 +106,7 @@ public class AiOpsService { .subAgents(List.of(plannerAgent, executorAgent)) .build(); - String taskPrompt = "你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。"; + String taskPrompt = buildTaskPrompt(request); logger.info("调用 Supervisor Agent 开始编排..."); @@ -158,6 +162,87 @@ public class AiOpsService { } } + public String resolveSessionId(AIOpsRequest request) { + if (request != null && !isBlank(request.getSessionId())) { + return request.getSessionId().trim(); + } + return UUID.randomUUID().toString(); + } + + public void persistFinalReport(String sessionId, String finalReport) { + if (isBlank(sessionId) || isBlank(finalReport)) { + return; + } + diagnosisSessionRepository.findBySessionId(sessionId.trim()).ifPresent(session -> { + session.setAnswer(finalReport); + diagnosisSessionRepository.save(session); + }); + } + + String buildQuerySummary(AIOpsRequest request) { + if (request == null) { + return "AI Ops 告警分析"; + } + + StringBuilder summary = new StringBuilder("AI Ops 告警分析"); + appendField(summary, "告警", request.getAlertName()); + appendField(summary, "服务", request.getService()); + appendField(summary, "等级", request.getSeverity()); + appendField(summary, "时间范围", request.getTimeRange()); + appendField(summary, "描述", request.getDescription()); + appendField(summary, "请求", request.getUserRequest()); + return summary.toString(); + } + + boolean hasAlertPayload(AIOpsRequest request) { + if (request == null) { + return false; + } + return !isBlank(request.getAlertName()) + || !isBlank(request.getService()) + || !isBlank(request.getSeverity()) + || !isBlank(request.getDescription()) + || !isBlank(request.getTimeRange()); + } + + String buildTaskPrompt(AIOpsRequest request) { + StringBuilder prompt = new StringBuilder(); + prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。"); + prompt.append("\n\n本次告警输入:\n"); + prompt.append(buildQuerySummary(request)); + if (hasAlertPayload(request)) { + prompt.append("\n\nAIOps scope mode: PAYLOAD_TARGETED\n"); + prompt.append("- The request includes an alert payload. Treat the supplied alert payload as the primary and only main diagnosis target.\n"); + prompt.append("- The final report must focus on the supplied alert fields such as alertName, service, severity, description, and timeRange.\n"); + prompt.append("- You may call queryPrometheusAlerts only to verify whether the supplied alert is still active or to identify related risk/context.\n"); + prompt.append("- If queryPrometheusAlerts returns unrelated active alerts, do not create full root-cause or remediation sections for them.\n"); + prompt.append("- Mention unrelated active alerts only briefly in a Related Risk section when they help explain the supplied alert.\n"); + } else { + prompt.append("\n\nAIOps scope mode: AUTO_DISCOVERY\n"); + prompt.append("- The request does not include alert payload fields. First call queryPrometheusAlerts to discover current active/firing alerts.\n"); + prompt.append("- Prefer P0/P1 alerts or the longest-running firing alerts, then diagnose one or more alerts based on severity and evidence.\n"); + prompt.append("- Use metrics, logs, and knowledge-base evidence before producing the final alert analysis report.\n"); + } + return prompt.toString(); + } + + private DiagnosisSession startDiagnosisSession(String sessionId, AIOpsRequest request) { + DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId) + .orElseGet(() -> DiagnosisSession.builder() + .sessionId(sessionId) + .agentFlow("AI_OPS") + .build()); + session.setQuery(buildQuerySummary(request)); + session.setStatus("RUNNING"); + session.setAgentFlow("AI_OPS"); + session.setAnswer(null); + session.setTotalDurationMs(null); + session.setTotalTokenCount(null); + session.setStepCount(null); + session.setToolCallCount(null); + return session; + } + /** * 构建 Planner Agent */ @@ -205,25 +290,33 @@ public class AiOpsService { } } - /** 从 agent_step 汇总指标回填 diagnosis_session */ + /** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */ private void backfillSessionMetrics(DiagnosisSession session) { try { List steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId()); - if (steps.isEmpty()) return; int totalTokens = 0; int stepCount = 0; - int toolCallCount = 0; for (AgentStep s : steps) { stepCount++; if (s.getTokenCount() != null) totalTokens += s.getTokenCount(); - if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++; } + long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId()); session.setTotalTokenCount(totalTokens); session.setStepCount(stepCount); - session.setToolCallCount(toolCallCount); + session.setToolCallCount(Math.toIntExact(toolCallCount)); } catch (Exception e) { logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e); } } + + private void appendField(StringBuilder builder, String label, String value) { + if (!isBlank(value)) { + builder.append("\n- ").append(label).append(": ").append(value.trim()); + } + } + + private boolean isBlank(String value) { + return value == null || value.trim().isEmpty(); + } } diff --git a/src/main/java/com/superbiz/agent/service/ChatService.java b/src/main/java/com/superbiz/agent/service/ChatService.java index 50415ac..b3c1a5a 100644 --- a/src/main/java/com/superbiz/agent/service/ChatService.java +++ b/src/main/java/com/superbiz/agent/service/ChatService.java @@ -18,6 +18,7 @@ import com.superbiz.agent.hook.TokenUsageHolder; import com.superbiz.agent.hook.VerifierInputHook; import com.superbiz.agent.repository.AgentStepRepository; import com.superbiz.agent.repository.DiagnosisSessionRepository; +import com.superbiz.agent.repository.ToolInvocationRepository; import com.superbiz.agent.tool.LookupKnowledgeTool; import com.superbiz.agent.tool.RetrievedDocTracker; import com.superbiz.agent.util.QuestionComplexity; @@ -86,6 +87,9 @@ public class ChatService { @Autowired private AgentStepRepository agentStepRepository; + @Autowired + private ToolInvocationRepository toolInvocationRepository; + @Autowired private EvaluationService evaluationService; @@ -838,25 +842,22 @@ public class ChatService { ) { } - /** 从 agent_step 汇总 token、步数等指标回填 diagnosis_session */ + /** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */ private void backfillSessionMetrics(DiagnosisSession session) { try { List steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId()); - if (steps.isEmpty()) return; - int totalTokens = 0; int stepCount = 0; - int toolCallCount = 0; for (var s : steps) { stepCount++; if (s.getTokenCount() != null) totalTokens += s.getTokenCount(); - if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++; } + long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId()); session.setTotalTokenCount(totalTokens); session.setStepCount(stepCount); - session.setToolCallCount(toolCallCount); + session.setToolCallCount(Math.toIntExact(toolCallCount)); } catch (Exception e) { logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e); } diff --git a/src/main/resources/application.yml b/src/main/resources/application.yml index 1df7156..312ba48 100644 --- a/src/main/resources/application.yml +++ b/src/main/resources/application.yml @@ -47,6 +47,14 @@ spring: username: root password: '!Fucker123..' driver-class-name: com.mysql.cj.jdbc.Driver + hikari: + maximum-pool-size: 5 + minimum-idle: 1 + connection-timeout: 10000 + validation-timeout: 5000 + idle-timeout: 60000 + max-lifetime: 120000 + keepalive-time: 30000 # ===================================================== # JPA 配置 diff --git a/src/test/java/com/superbiz/agent/service/AiOpsServiceTest.java b/src/test/java/com/superbiz/agent/service/AiOpsServiceTest.java new file mode 100644 index 0000000..2a8b2d5 --- /dev/null +++ b/src/test/java/com/superbiz/agent/service/AiOpsServiceTest.java @@ -0,0 +1,169 @@ +package com.superbiz.agent.service; + +import com.superbiz.agent.domain.entity.AgentStep; +import com.superbiz.agent.domain.entity.DiagnosisSession; +import com.superbiz.agent.dto.AIOpsRequest; +import com.superbiz.agent.repository.AgentStepRepository; +import com.superbiz.agent.repository.DiagnosisSessionRepository; +import com.superbiz.agent.repository.ToolInvocationRepository; +import org.junit.jupiter.api.BeforeEach; +import org.junit.jupiter.api.Test; +import org.springframework.test.util.ReflectionTestUtils; + +import java.util.Optional; +import java.util.List; + +import static org.junit.jupiter.api.Assertions.*; +import static org.mockito.Mockito.*; + +class AiOpsServiceTest { + + private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class); + private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class); + private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class); + private final AiOpsService service = new AiOpsService(); + + @BeforeEach + void setUp() { + ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository); + ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository); + ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository); + } + + @Test + void resolveSessionIdUsesRequestValueWhenPresent() { + AIOpsRequest request = new AIOpsRequest(); + request.setSessionId(" aiops-demo-session "); + + assertEquals("aiops-demo-session", service.resolveSessionId(request)); + } + + @Test + void resolveSessionIdGeneratesWhenMissing() { + String sessionId = service.resolveSessionId(null); + + assertNotNull(sessionId); + assertFalse(sessionId.isBlank()); + } + + @Test + void buildQuerySummaryUsesAlertFieldsAndUserRequestFallback() { + AIOpsRequest request = new AIOpsRequest(); + request.setAlertName("payment-service-latency-high"); + request.setService("payment-service"); + request.setSeverity("P1"); + request.setTimeRange("last_15m"); + request.setDescription("P95 latency is high"); + request.setUserRequest("结合日志和指标排查支付超时"); + + String summary = service.buildQuerySummary(request); + + assertTrue(summary.contains("AI Ops 告警分析")); + assertTrue(summary.contains("告警: payment-service-latency-high")); + assertTrue(summary.contains("服务: payment-service")); + assertTrue(summary.contains("等级: P1")); + assertTrue(summary.contains("时间范围: last_15m")); + assertTrue(summary.contains("描述: P95 latency is high")); + assertTrue(summary.contains("请求: 结合日志和指标排查支付超时")); + } + + @Test + void hasAlertPayloadIgnoresUserRequestOnly() { + AIOpsRequest request = new AIOpsRequest(); + request.setUserRequest("please discover active alerts"); + + assertFalse(service.hasAlertPayload(request)); + + request.setAlertName("HighCPUUsage"); + + assertTrue(service.hasAlertPayload(request)); + } + + @Test + void buildTaskPromptUsesPayloadTargetedModeWhenAlertFieldsExist() { + AIOpsRequest request = new AIOpsRequest(); + request.setAlertName("HighCPUUsage"); + request.setService("payment-service"); + request.setSeverity("P1"); + request.setTimeRange("last_15m"); + request.setDescription("CPU usage is above 80%"); + + String prompt = service.buildTaskPrompt(request); + + assertTrue(prompt.contains("AIOps scope mode: PAYLOAD_TARGETED")); + assertTrue(prompt.contains("primary and only main diagnosis target")); + assertTrue(prompt.contains("queryPrometheusAlerts only to verify")); + assertTrue(prompt.contains("do not create full root-cause or remediation sections")); + assertTrue(prompt.contains("Related Risk")); + assertTrue(prompt.contains("告警: HighCPUUsage")); + assertTrue(prompt.contains("服务: payment-service")); + assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY")); + } + + @Test + void buildTaskPromptUsesAutoDiscoveryModeWhenAlertPayloadIsMissing() { + String nullRequestPrompt = service.buildTaskPrompt(null); + + assertTrue(nullRequestPrompt.contains("AIOps scope mode: AUTO_DISCOVERY")); + assertTrue(nullRequestPrompt.contains("First call queryPrometheusAlerts")); + assertTrue(nullRequestPrompt.contains("current active/firing alerts")); + assertFalse(nullRequestPrompt.contains("AIOps scope mode: PAYLOAD_TARGETED")); + + AIOpsRequest userRequestOnly = new AIOpsRequest(); + userRequestOnly.setUserRequest("check what is firing now"); + + String userRequestOnlyPrompt = service.buildTaskPrompt(userRequestOnly); + + assertTrue(userRequestOnlyPrompt.contains("AIOps scope mode: AUTO_DISCOVERY")); + assertTrue(userRequestOnlyPrompt.contains("First call queryPrometheusAlerts")); + } + + @Test + void persistFinalReportUpdatesDiagnosisSessionAnswer() { + DiagnosisSession session = DiagnosisSession.builder() + .sessionId("aiops-session-001") + .query("AI Ops 告警分析") + .status("SUCCESS") + .agentFlow("AI_OPS") + .build(); + when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session)); + + service.persistFinalReport("aiops-session-001", "# 告警分析报告"); + + assertEquals("# 告警分析报告", session.getAnswer()); + verify(diagnosisSessionRepository).save(session); + } + + @Test + void persistFinalReportSkipsBlankInput() { + service.persistFinalReport("aiops-session-001", " "); + + verifyNoInteractions(diagnosisSessionRepository); + } + + @Test + void backfillSessionMetricsUsesRealToolInvocationCount() { + DiagnosisSession session = DiagnosisSession.builder() + .sessionId("aiops-session-002") + .build(); + AgentStep stepWithTool = AgentStep.builder() + .sessionId("aiops-session-002") + .hasToolCall(true) + .tokenCount(10) + .build(); + AgentStep stepWithoutTool = AgentStep.builder() + .sessionId("aiops-session-002") + .hasToolCall(false) + .tokenCount(20) + .build(); + when(agentStepRepository.findBySessionIdOrderByStepIndex("aiops-session-002")) + .thenReturn(List.of(stepWithTool, stepWithoutTool)); + when(toolInvocationRepository.countBySessionId("aiops-session-002")).thenReturn(11L); + + ReflectionTestUtils.invokeMethod(service, "backfillSessionMetrics", session); + + assertEquals(2, session.getStepCount()); + assertEquals(30, session.getTotalTokenCount()); + assertEquals(11, session.getToolCallCount()); + } +} diff --git a/src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java b/src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java index a3c5dad..c97aee2 100644 --- a/src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java +++ b/src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java @@ -7,6 +7,7 @@ import com.superbiz.agent.domain.entity.AgentStep; import com.superbiz.agent.domain.entity.DiagnosisSession; import com.superbiz.agent.repository.AgentStepRepository; import com.superbiz.agent.repository.DiagnosisSessionRepository; +import com.superbiz.agent.repository.ToolInvocationRepository; import com.superbiz.agent.tool.LookupKnowledgeTool; import com.superbiz.agent.tool.RetrievedDocTracker; import org.junit.jupiter.api.Test; @@ -143,6 +144,8 @@ class ChatServiceSequentialAgentTest { }); when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep())); when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of()); + ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class); + when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L); EvaluationService evaluationService = mock(EvaluationService.class); RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class); @@ -158,6 +161,7 @@ class ChatServiceSequentialAgentTest { ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class))); ReflectionTestUtils.setField(chatService, "diagnosisSessionRepository", diagnosisSessionRepository); ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository); + ReflectionTestUtils.setField(chatService, "toolInvocationRepository", toolInvocationRepository); ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService); ReflectionTestUtils.setField(chatService, "retrievedDocTracker", retrievedDocTracker); ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);