feat: add traceable scoped AIOps diagnosis

This commit is contained in:
aruo
2026-07-04 22:57:28 +08:00
parent 246c99b954
commit 23ee05c7c3
32 changed files with 1179 additions and 25 deletions
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.