feat: add traceable scoped AIOps diagnosis

This commit is contained in:
aruo
2026-07-04 22:57:28 +08:00
parent 246c99b954
commit 23ee05c7c3
32 changed files with 1179 additions and 25 deletions
+2
View File
@@ -5,6 +5,8 @@
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.
@@ -0,0 +1,81 @@
---
title: AIOps 告警排障 Runbook
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
category: troubleshooting
---
# AIOps 告警排障 Runbook
## 1. 告警输入处理原则
AIOps 诊断入口有两种触发方式:
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
## 2. Mock 告警与日志主题映射
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|---|---|---|---|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
## 3. HighCPUUsage / payment-service 排障步骤
### 3.1 现象确认
先确认 Prometheus 活动告警中是否存在:
- `alert_name = HighCPUUsage`
- `service = payment-service`
- CPU 使用率超过 80%
- 状态为 firing
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
### 3.2 指标与日志取证
推荐工具调用顺序:
1. `queryPrometheusAlerts`:确认当前活动告警。
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
### 3.3 根因判断
可接受的根因结论必须至少满足一项:
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
- 告警持续时间与日志时间线一致。
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
## 4. 处理建议
### 临时止血
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
- 对高耗时接口开启限流或降级非核心功能。
- 如果近期有发布,检查变更窗口并准备回滚。
### 根因修复
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
## 5. 报告要求
告警分析报告至少包含:
- 活跃告警清单。
- 告警根因分析。
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
- 已执行或建议执行的处理方案。
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
+48
View File
@@ -78,6 +78,43 @@ Expected result:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
## 4. Run AIOps Alert Diagnosis
```powershell
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
Expected result:
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
- `data.session.answer` contains the final alert analysis report when a report is generated.
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
Query the AIOps trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
## Demo Story
The important interview story is:
@@ -92,3 +129,14 @@ one session id
-> feedback
-> trace API for replay and audit
```
The AIOps story uses the same audit spine:
```text
one session id
-> alert payload
-> AIOps planner/executor execution
-> evidence tools
-> alert analysis report
-> trace API for replay and audit
```
+38
View File
@@ -0,0 +1,38 @@
# AIOps Alert Acceptance Case
## Goal
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
## Input
- Session id: `mvp-demo-aiops-payment-cpu-001`
- Endpoint: `POST /api/ai_ops`
- Profile: `mvp-demo`
- Alert:
```json
{
"sessionId": "mvp-demo-aiops-payment-cpu-001",
"alertName": "HighCPUUsage",
"service": "payment-service",
"severity": "P1",
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
"timeRange": "last_15m",
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
}
```
## Acceptance Criteria
1. The SSE stream emits a `session` message containing the requested session id.
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
3. The persisted session query contains the alert name, service, severity, time range, and description.
4. If a final report is generated, `diagnosis_session.answer` contains that report.
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
## Known Limits
- This slice does not add a Verifier Agent to AIOps.
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The AIOps endpoint has two natural modes:
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
## Goals / Non-Goals
**Goals:**
- Make AIOps payload mode single-alert focused.
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
- Keep the change prompt-only and low risk.
- Add tests for prompt scope rules.
**Non-Goals:**
- Do not add a Verifier Agent.
- Do not force tool calls in Java code.
- Do not change `/api/ai_ops` request/response contracts.
- Do not modify mock alert data.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
## Prompt Rules
Payload mode MUST instruct the agent:
- Treat supplied payload as the primary and only report target.
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
- Do not create root-cause sections for unrelated active alerts.
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
Auto-discovery mode MUST instruct the agent:
- First call `queryPrometheusAlerts`.
- Select P0/P1 or longest-running firing alerts.
- Analyze one or more active alerts based on severity and evidence.
## Risks / Trade-offs
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
@@ -0,0 +1,27 @@
## Why
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
## What Changes
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
- Preserve full active-alert discovery when no payload is supplied.
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
- Update tests and demo acceptance wording to lock the new behavior.
## Capabilities
### New Capabilities
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
## Impact
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
- Affected API: no endpoint or request/response shape change.
- Affected persistence: no schema change.
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
@@ -0,0 +1,25 @@
## ADDED Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,21 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
- [x] 1.2 Record local impact analysis and GitNexus skip context.
## 2. Prompt Scope Control
- [x] 2.1 Add payload detection helper in `AiOpsService`.
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
## 3. Tests And Docs
- [x] 3.1 Add tests for payload-mode prompt rules.
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
## 4. Verification
- [x] 4.1 Run targeted tests.
- [x] 4.2 Run compile verification.
- [x] 4.3 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,90 @@
## Context
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
- `ChatController.aiOps()` accepts no request body.
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
## Goals / Non-Goals
**Goals:**
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
- Preserve backward compatibility for callers that post with no request body.
- Return the resolved `sessionId` through SSE.
- Persist the final report into the existing diagnosis session.
- Keep the AIOps path observable through the existing trace API.
**Non-Goals:**
- Do not merge AIOps into `ChatService`.
- Do not add a new Verifier Agent to AIOps in this slice.
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
- Do not add database migrations.
- Do not clean up sensitive configuration.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
## Interface Impact
- Level: L3 API behavior extension.
- Endpoint: `POST /api/ai_ops`
- Compatibility: callers may still omit a body. New callers may send:
```json
{
"sessionId": "mvp-demo-aiops-payment-latency-001",
"alertName": "payment-service-latency-high",
"service": "payment-service",
"severity": "P1",
"description": "支付服务 P95 延迟升高并伴随超时错误",
"timeRange": "last_15m"
}
```
The SSE stream emits a first content message containing the resolved session id:
```text
sessionId: mvp-demo-aiops-payment-latency-001
```
## Data Flow
```text
POST /api/ai_ops
-> ChatController resolves request body and tools
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
-> set SessionContextHolder(sessionId)
-> ai_ops_supervisor -> planner_agent -> executor_agent
-> persist agent_step and tool_invocation through existing hooks/tools
-> extract final report
-> persist diagnosis_session.answer/status/counts
-> caller queries GET /api/diagnosis/{sessionId}/trace
```
## Risks / Trade-offs
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
## Migration Plan
- No database migration.
- Deploy with application restart.
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
@@ -0,0 +1,29 @@
## Why
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
## What Changes
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
- Resolve a stable session id from the request or generate one when omitted.
- Persist the AIOps alert query and final report into `diagnosis_session`.
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
- Document the AIOps demo path beside the existing MVP demo trace flow.
## Capabilities
### New Capabilities
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
### Modified Capabilities
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
## Impact
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
@@ -0,0 +1,41 @@
## ADDED Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
@@ -0,0 +1,22 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
## 2. AIOps API And Service
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
## 3. Demo Documentation
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
## 4. Verification
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
- [x] 4.2 Run targeted tests.
- [x] 4.3 Run compile verification.
@@ -0,0 +1,28 @@
# aiops-alert-scope-control Specification
## Purpose
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
## Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,44 @@
# aiops-traceable-diagnosis-entry Specification
## Purpose
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
## Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
@@ -4,6 +4,7 @@ import com.alibaba.cloud.ai.graph.OverAllState;
import lombok.Getter;
import lombok.Setter;
import com.superbiz.agent.domain.model.SessionContext;
import com.superbiz.agent.dto.AIOpsRequest;
import com.superbiz.agent.service.AiOpsService;
import com.superbiz.agent.service.ChatService;
import com.superbiz.agent.service.session.SessionManager;
@@ -212,21 +213,23 @@ public class ChatController {
* 无需用户输入,自动执行告警分析流程
*/
@PostMapping(value = "/ai_ops", produces = "text/event-stream;charset=UTF-8")
public SseEmitter aiOps() {
public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) {
SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢)
String sessionId = aiOpsService.resolveSessionId(request);
executor.execute(() -> {
try {
logger.info("收到 AI 智能运维请求 - 启动多 Agent 协作流程");
logger.info("收到 AI 智能运维请求 - SessionId: {}, 启动多 Agent 协作流程", sessionId);
ChatModel chatModel = chatService.getChatModel();
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
emitter.send(SseEmitter.event().name("message").data(SseMessage.session(sessionId), MediaType.APPLICATION_JSON));
emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n")));
// 调用 AiOpsService 执行分析流程
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks);
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId);
if (overAllStateOptional.isEmpty()) {
emitter.send(SseEmitter.event().name("message")
@@ -245,6 +248,7 @@ public class ChatController {
if (finalReportOptional.isPresent()) {
String finalReportText = finalReportOptional.get();
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
aiOpsService.persistFinalReport(sessionId, finalReportText);
// 发送分隔线
emitter.send(SseEmitter.event().name("message")
@@ -443,6 +447,13 @@ public class ChatController {
return message;
}
public static SseMessage session(String sessionId) {
SseMessage message = new SseMessage();
message.setType("session");
message.setData(sessionId);
return message;
}
public static SseMessage error(String errorMessage) {
SseMessage message = new SseMessage();
message.setType("error");
@@ -7,6 +7,36 @@ import lombok.Data;
*/
@Data
public class AIOpsRequest {
/**
* 诊断会话 ID;为空时后端自动生成。
*/
private String sessionId;
/**
* 告警名称。
*/
private String alertName;
/**
* 受影响服务。
*/
private String service;
/**
* 告警等级,例如 P0/P1/P2。
*/
private String severity;
/**
* 告警描述。
*/
private String description;
/**
* 排查时间范围,例如 last_15m。
*/
private String timeRange;
/**
* 用户请求描述
@@ -22,6 +22,11 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
*/
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
/**
* 根据会话ID统计真实工具调用次数
*/
long countBySessionId(String sessionId);
/**
* 根据工具名查询所有调用
*/
@@ -10,11 +10,12 @@ import com.superbiz.agent.agent.tool.InternalDocsTools;
import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.dto.AIOpsRequest;
import com.superbiz.agent.hook.AgentLoggingHook;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.util.SessionContextHolder;
import org.slf4j.Logger;
import org.slf4j.LoggerFactory;
@@ -62,6 +63,9 @@ public class AiOpsService {
@Autowired
private AgentStepRepository agentStepRepository;
@Autowired
private ToolInvocationRepository toolInvocationRepository;
/**
* 执行 AI Ops 告警分析流程
*
@@ -71,22 +75,22 @@ public class AiOpsService {
* @throws GraphRunnerException 如果 Agent 执行失败
*/
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
}
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
AIOpsRequest request, String sessionId) throws GraphRunnerException {
logger.info("开始执行 AI Ops 多 Agent 协作流程");
String sessionId = UUID.randomUUID().toString().substring(0, 8);
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
long startTime = System.currentTimeMillis();
// 创建诊断会话
DiagnosisSession session = DiagnosisSession.builder()
.sessionId(sessionId)
.query("AI Ops 告警分析")
.status("RUNNING")
.agentFlow("AI_OPS")
.build();
// 创建或更新诊断会话
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
diagnosisSessionRepository.save(session);
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
SessionContextHolder.setSessionId(sessionId);
SessionContextHolder.setSessionId(resolvedSessionId);
try {
// 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook)
@@ -102,7 +106,7 @@ public class AiOpsService {
.subAgents(List.of(plannerAgent, executorAgent))
.build();
String taskPrompt = "你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。";
String taskPrompt = buildTaskPrompt(request);
logger.info("调用 Supervisor Agent 开始编排...");
@@ -158,6 +162,87 @@ public class AiOpsService {
}
}
public String resolveSessionId(AIOpsRequest request) {
if (request != null && !isBlank(request.getSessionId())) {
return request.getSessionId().trim();
}
return UUID.randomUUID().toString();
}
public void persistFinalReport(String sessionId, String finalReport) {
if (isBlank(sessionId) || isBlank(finalReport)) {
return;
}
diagnosisSessionRepository.findBySessionId(sessionId.trim()).ifPresent(session -> {
session.setAnswer(finalReport);
diagnosisSessionRepository.save(session);
});
}
String buildQuerySummary(AIOpsRequest request) {
if (request == null) {
return "AI Ops 告警分析";
}
StringBuilder summary = new StringBuilder("AI Ops 告警分析");
appendField(summary, "告警", request.getAlertName());
appendField(summary, "服务", request.getService());
appendField(summary, "等级", request.getSeverity());
appendField(summary, "时间范围", request.getTimeRange());
appendField(summary, "描述", request.getDescription());
appendField(summary, "请求", request.getUserRequest());
return summary.toString();
}
boolean hasAlertPayload(AIOpsRequest request) {
if (request == null) {
return false;
}
return !isBlank(request.getAlertName())
|| !isBlank(request.getService())
|| !isBlank(request.getSeverity())
|| !isBlank(request.getDescription())
|| !isBlank(request.getTimeRange());
}
String buildTaskPrompt(AIOpsRequest request) {
StringBuilder prompt = new StringBuilder();
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
prompt.append("\n\n本次告警输入:\n");
prompt.append(buildQuerySummary(request));
if (hasAlertPayload(request)) {
prompt.append("\n\nAIOps scope mode: PAYLOAD_TARGETED\n");
prompt.append("- The request includes an alert payload. Treat the supplied alert payload as the primary and only main diagnosis target.\n");
prompt.append("- The final report must focus on the supplied alert fields such as alertName, service, severity, description, and timeRange.\n");
prompt.append("- You may call queryPrometheusAlerts only to verify whether the supplied alert is still active or to identify related risk/context.\n");
prompt.append("- If queryPrometheusAlerts returns unrelated active alerts, do not create full root-cause or remediation sections for them.\n");
prompt.append("- Mention unrelated active alerts only briefly in a Related Risk section when they help explain the supplied alert.\n");
} else {
prompt.append("\n\nAIOps scope mode: AUTO_DISCOVERY\n");
prompt.append("- The request does not include alert payload fields. First call queryPrometheusAlerts to discover current active/firing alerts.\n");
prompt.append("- Prefer P0/P1 alerts or the longest-running firing alerts, then diagnose one or more alerts based on severity and evidence.\n");
prompt.append("- Use metrics, logs, and knowledge-base evidence before producing the final alert analysis report.\n");
}
return prompt.toString();
}
private DiagnosisSession startDiagnosisSession(String sessionId, AIOpsRequest request) {
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
.orElseGet(() -> DiagnosisSession.builder()
.sessionId(sessionId)
.agentFlow("AI_OPS")
.build());
session.setQuery(buildQuerySummary(request));
session.setStatus("RUNNING");
session.setAgentFlow("AI_OPS");
session.setAnswer(null);
session.setTotalDurationMs(null);
session.setTotalTokenCount(null);
session.setStepCount(null);
session.setToolCallCount(null);
return session;
}
/**
* 构建 Planner Agent
*/
@@ -205,25 +290,33 @@ public class AiOpsService {
}
}
/** 从 agent_step 汇总指标回填 diagnosis_session */
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
private void backfillSessionMetrics(DiagnosisSession session) {
try {
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
if (steps.isEmpty()) return;
int totalTokens = 0;
int stepCount = 0;
int toolCallCount = 0;
for (AgentStep s : steps) {
stepCount++;
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++;
}
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
session.setTotalTokenCount(totalTokens);
session.setStepCount(stepCount);
session.setToolCallCount(toolCallCount);
session.setToolCallCount(Math.toIntExact(toolCallCount));
} catch (Exception e) {
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
}
}
private void appendField(StringBuilder builder, String label, String value) {
if (!isBlank(value)) {
builder.append("\n- ").append(label).append(": ").append(value.trim());
}
}
private boolean isBlank(String value) {
return value == null || value.trim().isEmpty();
}
}
@@ -18,6 +18,7 @@ import com.superbiz.agent.hook.TokenUsageHolder;
import com.superbiz.agent.hook.VerifierInputHook;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.tool.RetrievedDocTracker;
import com.superbiz.agent.util.QuestionComplexity;
@@ -86,6 +87,9 @@ public class ChatService {
@Autowired
private AgentStepRepository agentStepRepository;
@Autowired
private ToolInvocationRepository toolInvocationRepository;
@Autowired
private EvaluationService evaluationService;
@@ -838,25 +842,22 @@ public class ChatService {
) {
}
/** 从 agent_step 汇总 token、步数等指标回填 diagnosis_session */
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
private void backfillSessionMetrics(DiagnosisSession session) {
try {
List<com.superbiz.agent.domain.entity.AgentStep> steps =
agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
if (steps.isEmpty()) return;
int totalTokens = 0;
int stepCount = 0;
int toolCallCount = 0;
for (var s : steps) {
stepCount++;
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
if (Boolean.TRUE.equals(s.getHasToolCall())) toolCallCount++;
}
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
session.setTotalTokenCount(totalTokens);
session.setStepCount(stepCount);
session.setToolCallCount(toolCallCount);
session.setToolCallCount(Math.toIntExact(toolCallCount));
} catch (Exception e) {
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
}
+8
View File
@@ -47,6 +47,14 @@ spring:
username: root
password: '!Fucker123..'
driver-class-name: com.mysql.cj.jdbc.Driver
hikari:
maximum-pool-size: 5
minimum-idle: 1
connection-timeout: 10000
validation-timeout: 5000
idle-timeout: 60000
max-lifetime: 120000
keepalive-time: 30000
# =====================================================
# JPA 配置
@@ -0,0 +1,169 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.dto.AIOpsRequest;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.springframework.test.util.ReflectionTestUtils;
import java.util.Optional;
import java.util.List;
import static org.junit.jupiter.api.Assertions.*;
import static org.mockito.Mockito.*;
class AiOpsServiceTest {
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
private final AiOpsService service = new AiOpsService();
@BeforeEach
void setUp() {
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
}
@Test
void resolveSessionIdUsesRequestValueWhenPresent() {
AIOpsRequest request = new AIOpsRequest();
request.setSessionId(" aiops-demo-session ");
assertEquals("aiops-demo-session", service.resolveSessionId(request));
}
@Test
void resolveSessionIdGeneratesWhenMissing() {
String sessionId = service.resolveSessionId(null);
assertNotNull(sessionId);
assertFalse(sessionId.isBlank());
}
@Test
void buildQuerySummaryUsesAlertFieldsAndUserRequestFallback() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("payment-service-latency-high");
request.setService("payment-service");
request.setSeverity("P1");
request.setTimeRange("last_15m");
request.setDescription("P95 latency is high");
request.setUserRequest("结合日志和指标排查支付超时");
String summary = service.buildQuerySummary(request);
assertTrue(summary.contains("AI Ops 告警分析"));
assertTrue(summary.contains("告警: payment-service-latency-high"));
assertTrue(summary.contains("服务: payment-service"));
assertTrue(summary.contains("等级: P1"));
assertTrue(summary.contains("时间范围: last_15m"));
assertTrue(summary.contains("描述: P95 latency is high"));
assertTrue(summary.contains("请求: 结合日志和指标排查支付超时"));
}
@Test
void hasAlertPayloadIgnoresUserRequestOnly() {
AIOpsRequest request = new AIOpsRequest();
request.setUserRequest("please discover active alerts");
assertFalse(service.hasAlertPayload(request));
request.setAlertName("HighCPUUsage");
assertTrue(service.hasAlertPayload(request));
}
@Test
void buildTaskPromptUsesPayloadTargetedModeWhenAlertFieldsExist() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("HighCPUUsage");
request.setService("payment-service");
request.setSeverity("P1");
request.setTimeRange("last_15m");
request.setDescription("CPU usage is above 80%");
String prompt = service.buildTaskPrompt(request);
assertTrue(prompt.contains("AIOps scope mode: PAYLOAD_TARGETED"));
assertTrue(prompt.contains("primary and only main diagnosis target"));
assertTrue(prompt.contains("queryPrometheusAlerts only to verify"));
assertTrue(prompt.contains("do not create full root-cause or remediation sections"));
assertTrue(prompt.contains("Related Risk"));
assertTrue(prompt.contains("告警: HighCPUUsage"));
assertTrue(prompt.contains("服务: payment-service"));
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
}
@Test
void buildTaskPromptUsesAutoDiscoveryModeWhenAlertPayloadIsMissing() {
String nullRequestPrompt = service.buildTaskPrompt(null);
assertTrue(nullRequestPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
assertTrue(nullRequestPrompt.contains("First call queryPrometheusAlerts"));
assertTrue(nullRequestPrompt.contains("current active/firing alerts"));
assertFalse(nullRequestPrompt.contains("AIOps scope mode: PAYLOAD_TARGETED"));
AIOpsRequest userRequestOnly = new AIOpsRequest();
userRequestOnly.setUserRequest("check what is firing now");
String userRequestOnlyPrompt = service.buildTaskPrompt(userRequestOnly);
assertTrue(userRequestOnlyPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
assertTrue(userRequestOnlyPrompt.contains("First call queryPrometheusAlerts"));
}
@Test
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
DiagnosisSession session = DiagnosisSession.builder()
.sessionId("aiops-session-001")
.query("AI Ops 告警分析")
.status("SUCCESS")
.agentFlow("AI_OPS")
.build();
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
service.persistFinalReport("aiops-session-001", "# 告警分析报告");
assertEquals("# 告警分析报告", session.getAnswer());
verify(diagnosisSessionRepository).save(session);
}
@Test
void persistFinalReportSkipsBlankInput() {
service.persistFinalReport("aiops-session-001", " ");
verifyNoInteractions(diagnosisSessionRepository);
}
@Test
void backfillSessionMetricsUsesRealToolInvocationCount() {
DiagnosisSession session = DiagnosisSession.builder()
.sessionId("aiops-session-002")
.build();
AgentStep stepWithTool = AgentStep.builder()
.sessionId("aiops-session-002")
.hasToolCall(true)
.tokenCount(10)
.build();
AgentStep stepWithoutTool = AgentStep.builder()
.sessionId("aiops-session-002")
.hasToolCall(false)
.tokenCount(20)
.build();
when(agentStepRepository.findBySessionIdOrderByStepIndex("aiops-session-002"))
.thenReturn(List.of(stepWithTool, stepWithoutTool));
when(toolInvocationRepository.countBySessionId("aiops-session-002")).thenReturn(11L);
ReflectionTestUtils.invokeMethod(service, "backfillSessionMetrics", session);
assertEquals(2, session.getStepCount());
assertEquals(30, session.getTotalTokenCount());
assertEquals(11, session.getToolCallCount());
}
}
@@ -7,6 +7,7 @@ import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import com.superbiz.agent.tool.RetrievedDocTracker;
import org.junit.jupiter.api.Test;
@@ -143,6 +144,8 @@ class ChatServiceSequentialAgentTest {
});
when(agentStepRepository.findById(any())).thenReturn(Optional.of(new AgentStep()));
when(agentStepRepository.findBySessionIdOrderByStepIndex(anyString())).thenReturn(List.of());
ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
when(toolInvocationRepository.countBySessionId(anyString())).thenReturn(0L);
EvaluationService evaluationService = mock(EvaluationService.class);
RetrievedDocTracker retrievedDocTracker = mock(RetrievedDocTracker.class);
@@ -158,6 +161,7 @@ class ChatServiceSequentialAgentTest {
ReflectionTestUtils.setField(chatService, "queryLogsTools", new QueryLogsTools(mock(ToolInvocationRecorder.class)));
ReflectionTestUtils.setField(chatService, "diagnosisSessionRepository", diagnosisSessionRepository);
ReflectionTestUtils.setField(chatService, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(chatService, "toolInvocationRepository", toolInvocationRepository);
ReflectionTestUtils.setField(chatService, "evaluationService", evaluationService);
ReflectionTestUtils.setField(chatService, "retrievedDocTracker", retrievedDocTracker);
ReflectionTestUtils.setField(chatService, "knowledgeDomainService", knowledgeDomainService);