feat: add traceable scoped AIOps diagnosis

This commit is contained in:
aruo
2026-07-04 22:57:28 +08:00
parent 246c99b954
commit 23ee05c7c3
32 changed files with 1179 additions and 25 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The AIOps endpoint has two natural modes:
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
## Goals / Non-Goals
**Goals:**
- Make AIOps payload mode single-alert focused.
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
- Keep the change prompt-only and low risk.
- Add tests for prompt scope rules.
**Non-Goals:**
- Do not add a Verifier Agent.
- Do not force tool calls in Java code.
- Do not change `/api/ai_ops` request/response contracts.
- Do not modify mock alert data.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
## Prompt Rules
Payload mode MUST instruct the agent:
- Treat supplied payload as the primary and only report target.
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
- Do not create root-cause sections for unrelated active alerts.
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
Auto-discovery mode MUST instruct the agent:
- First call `queryPrometheusAlerts`.
- Select P0/P1 or longest-running firing alerts.
- Analyze one or more active alerts based on severity and evidence.
## Risks / Trade-offs
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
@@ -0,0 +1,27 @@
## Why
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
## What Changes
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
- Preserve full active-alert discovery when no payload is supplied.
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
- Update tests and demo acceptance wording to lock the new behavior.
## Capabilities
### New Capabilities
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
## Impact
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
- Affected API: no endpoint or request/response shape change.
- Affected persistence: no schema change.
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
@@ -0,0 +1,25 @@
## ADDED Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,21 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
- [x] 1.2 Record local impact analysis and GitNexus skip context.
## 2. Prompt Scope Control
- [x] 2.1 Add payload detection helper in `AiOpsService`.
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
## 3. Tests And Docs
- [x] 3.1 Add tests for payload-mode prompt rules.
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
## 4. Verification
- [x] 4.1 Run targeted tests.
- [x] 4.2 Run compile verification.
- [x] 4.3 Run OpenSpec validation.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,90 @@
## Context
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
- `ChatController.aiOps()` accepts no request body.
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
## Goals / Non-Goals
**Goals:**
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
- Preserve backward compatibility for callers that post with no request body.
- Return the resolved `sessionId` through SSE.
- Persist the final report into the existing diagnosis session.
- Keep the AIOps path observable through the existing trace API.
**Non-Goals:**
- Do not merge AIOps into `ChatService`.
- Do not add a new Verifier Agent to AIOps in this slice.
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
- Do not add database migrations.
- Do not clean up sensitive configuration.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
## Interface Impact
- Level: L3 API behavior extension.
- Endpoint: `POST /api/ai_ops`
- Compatibility: callers may still omit a body. New callers may send:
```json
{
"sessionId": "mvp-demo-aiops-payment-latency-001",
"alertName": "payment-service-latency-high",
"service": "payment-service",
"severity": "P1",
"description": "支付服务 P95 延迟升高并伴随超时错误",
"timeRange": "last_15m"
}
```
The SSE stream emits a first content message containing the resolved session id:
```text
sessionId: mvp-demo-aiops-payment-latency-001
```
## Data Flow
```text
POST /api/ai_ops
-> ChatController resolves request body and tools
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
-> set SessionContextHolder(sessionId)
-> ai_ops_supervisor -> planner_agent -> executor_agent
-> persist agent_step and tool_invocation through existing hooks/tools
-> extract final report
-> persist diagnosis_session.answer/status/counts
-> caller queries GET /api/diagnosis/{sessionId}/trace
```
## Risks / Trade-offs
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
## Migration Plan
- No database migration.
- Deploy with application restart.
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
@@ -0,0 +1,29 @@
## Why
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
## What Changes
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
- Resolve a stable session id from the request or generate one when omitted.
- Persist the AIOps alert query and final report into `diagnosis_session`.
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
- Document the AIOps demo path beside the existing MVP demo trace flow.
## Capabilities
### New Capabilities
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
### Modified Capabilities
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
## Impact
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
@@ -0,0 +1,41 @@
## ADDED Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
@@ -0,0 +1,22 @@
## 1. Flow Records
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
## 2. AIOps API And Service
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
## 3. Demo Documentation
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
## 4. Verification
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
- [x] 4.2 Run targeted tests.
- [x] 4.3 Run compile verification.
@@ -0,0 +1,28 @@
# aiops-alert-scope-control Specification
## Purpose
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
## Requirements
### Requirement: AIOps payload mode focuses on supplied alert
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
#### Scenario: Request includes alertName and service
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
- **THEN** the AIOps task prompt identifies payload mode
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
### Requirement: AIOps auto-discovery mode queries active alerts first
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
#### Scenario: Request body is omitted
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
- **THEN** the AIOps task prompt identifies auto-discovery mode
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
### Requirement: Payload mode may use active alerts as supporting context
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
#### Scenario: Prometheus returns multiple active alerts
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
- **THEN** the prompt permits mentioning those alerts only as related risk or context
- **AND** the final report target remains the supplied alert
@@ -0,0 +1,44 @@
# aiops-traceable-diagnosis-entry Specification
## Purpose
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
## Requirements
### Requirement: AIOps endpoint accepts optional alert input
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
#### Scenario: Caller supplies alert input
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
#### Scenario: Caller omits alert input
- **WHEN** a caller posts to `/api/ai_ops` without a body
- **THEN** the system still starts the default AIOps alert-analysis flow
### Requirement: AIOps session id is traceable
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
#### Scenario: Request includes session id
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
- **THEN** the created `diagnosis_session.session_id` equals that value
- **AND** the SSE stream includes the same session id
#### Scenario: Request omits session id
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
- **THEN** the system generates a session id
- **AND** the SSE stream includes the generated session id
### Requirement: AIOps report is persisted for trace replay
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
#### Scenario: AIOps report is generated
- **WHEN** the AIOps planner/executor flow returns a final report
- **THEN** the corresponding diagnosis session is marked successful
- **AND** `diagnosis_session.answer` stores the final report
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
### Requirement: AIOps trace uses existing evidence tables
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
#### Scenario: AIOps uses evidence tools
- **WHEN** the AIOps flow calls available evidence tools
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id