feat: add traceable scoped AIOps diagnosis
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint has two natural modes:
|
||||
|
||||
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
|
||||
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
|
||||
|
||||
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make AIOps payload mode single-alert focused.
|
||||
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
|
||||
- Keep the change prompt-only and low risk.
|
||||
- Add tests for prompt scope rules.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add a Verifier Agent.
|
||||
- Do not force tool calls in Java code.
|
||||
- Do not change `/api/ai_ops` request/response contracts.
|
||||
- Do not modify mock alert data.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
|
||||
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
|
||||
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
|
||||
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
|
||||
|
||||
## Prompt Rules
|
||||
|
||||
Payload mode MUST instruct the agent:
|
||||
|
||||
- Treat supplied payload as the primary and only report target.
|
||||
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
|
||||
- Do not create root-cause sections for unrelated active alerts.
|
||||
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
|
||||
|
||||
Auto-discovery mode MUST instruct the agent:
|
||||
|
||||
- First call `queryPrometheusAlerts`.
|
||||
- Select P0/P1 or longest-running firing alerts.
|
||||
- Analyze one or more active alerts based on severity and evidence.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
|
||||
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
|
||||
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
|
||||
- Preserve full active-alert discovery when no payload is supplied.
|
||||
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
|
||||
- Update tests and demo acceptance wording to lock the new behavior.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
|
||||
- Affected API: no endpoint or request/response shape change.
|
||||
- Affected persistence: no schema change.
|
||||
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps payload mode focuses on supplied alert
|
||||
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||
|
||||
#### Scenario: Request includes alertName and service
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||
- **THEN** the AIOps task prompt identifies payload mode
|
||||
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||
|
||||
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||
|
||||
#### Scenario: Request body is omitted
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||
|
||||
### Requirement: Payload mode may use active alerts as supporting context
|
||||
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||
|
||||
#### Scenario: Prometheus returns multiple active alerts
|
||||
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||
- **AND** the final report target remains the supplied alert
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
|
||||
- [x] 1.2 Record local impact analysis and GitNexus skip context.
|
||||
|
||||
## 2. Prompt Scope Control
|
||||
|
||||
- [x] 2.1 Add payload detection helper in `AiOpsService`.
|
||||
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
|
||||
|
||||
## 3. Tests And Docs
|
||||
|
||||
- [x] 3.1 Add tests for payload-mode prompt rules.
|
||||
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
|
||||
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run targeted tests.
|
||||
- [x] 4.2 Run compile verification.
|
||||
- [x] 4.3 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,90 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
|
||||
|
||||
- `ChatController.aiOps()` accepts no request body.
|
||||
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
|
||||
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
|
||||
|
||||
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
|
||||
- Preserve backward compatibility for callers that post with no request body.
|
||||
- Return the resolved `sessionId` through SSE.
|
||||
- Persist the final report into the existing diagnosis session.
|
||||
- Keep the AIOps path observable through the existing trace API.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not merge AIOps into `ChatService`.
|
||||
- Do not add a new Verifier Agent to AIOps in this slice.
|
||||
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Do not add database migrations.
|
||||
- Do not clean up sensitive configuration.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
|
||||
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
|
||||
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
|
||||
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
|
||||
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
|
||||
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L3 API behavior extension.
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Compatibility: callers may still omit a body. New callers may send:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-latency-001",
|
||||
"alertName": "payment-service-latency-high",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "支付服务 P95 延迟升高并伴随超时错误",
|
||||
"timeRange": "last_15m"
|
||||
}
|
||||
```
|
||||
|
||||
The SSE stream emits a first content message containing the resolved session id:
|
||||
|
||||
```text
|
||||
sessionId: mvp-demo-aiops-payment-latency-001
|
||||
```
|
||||
|
||||
## Data Flow
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> ChatController resolves request body and tools
|
||||
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
|
||||
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
|
||||
-> set SessionContextHolder(sessionId)
|
||||
-> ai_ops_supervisor -> planner_agent -> executor_agent
|
||||
-> persist agent_step and tool_invocation through existing hooks/tools
|
||||
-> extract final report
|
||||
-> persist diagnosis_session.answer/status/counts
|
||||
-> caller queries GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
|
||||
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
|
||||
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
|
||||
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No database migration.
|
||||
- Deploy with application restart.
|
||||
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
|
||||
- Resolve a stable session id from the request or generate one when omitted.
|
||||
- Persist the AIOps alert query and final report into `diagnosis_session`.
|
||||
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
|
||||
- Document the AIOps demo path beside the existing MVP demo trace flow.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
|
||||
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
|
||||
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
|
||||
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps endpoint accepts optional alert input
|
||||
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||
|
||||
#### Scenario: Caller supplies alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||
|
||||
#### Scenario: Caller omits alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||
|
||||
### Requirement: AIOps session id is traceable
|
||||
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||
|
||||
#### Scenario: Request includes session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||
- **AND** the SSE stream includes the same session id
|
||||
|
||||
#### Scenario: Request omits session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||
- **THEN** the system generates a session id
|
||||
- **AND** the SSE stream includes the generated session id
|
||||
|
||||
### Requirement: AIOps report is persisted for trace replay
|
||||
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||
|
||||
#### Scenario: AIOps report is generated
|
||||
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||
- **THEN** the corresponding diagnosis session is marked successful
|
||||
- **AND** `diagnosis_session.answer` stores the final report
|
||||
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||
|
||||
### Requirement: AIOps trace uses existing evidence tables
|
||||
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||
|
||||
#### Scenario: AIOps uses evidence tools
|
||||
- **WHEN** the AIOps flow calls available evidence tools
|
||||
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
|
||||
|
||||
## 2. AIOps API And Service
|
||||
|
||||
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
|
||||
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
|
||||
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
|
||||
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
|
||||
|
||||
## 3. Demo Documentation
|
||||
|
||||
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
|
||||
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
|
||||
- [x] 4.2 Run targeted tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,28 @@
|
||||
# aiops-alert-scope-control Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change aiops-alert-scope-control. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: AIOps payload mode focuses on supplied alert
|
||||
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||
|
||||
#### Scenario: Request includes alertName and service
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||
- **THEN** the AIOps task prompt identifies payload mode
|
||||
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||
|
||||
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||
|
||||
#### Scenario: Request body is omitted
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||
|
||||
### Requirement: Payload mode may use active alerts as supporting context
|
||||
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||
|
||||
#### Scenario: Prometheus returns multiple active alerts
|
||||
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||
- **AND** the final report target remains the supplied alert
|
||||
@@ -0,0 +1,44 @@
|
||||
# aiops-traceable-diagnosis-entry Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change aiops-traceable-diagnosis-entry. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: AIOps endpoint accepts optional alert input
|
||||
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||
|
||||
#### Scenario: Caller supplies alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||
|
||||
#### Scenario: Caller omits alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||
|
||||
### Requirement: AIOps session id is traceable
|
||||
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||
|
||||
#### Scenario: Request includes session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||
- **AND** the SSE stream includes the same session id
|
||||
|
||||
#### Scenario: Request omits session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||
- **THEN** the system generates a session id
|
||||
- **AND** the SSE stream includes the generated session id
|
||||
|
||||
### Requirement: AIOps report is persisted for trace replay
|
||||
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||
|
||||
#### Scenario: AIOps report is generated
|
||||
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||
- **THEN** the corresponding diagnosis session is marked successful
|
||||
- **AND** `diagnosis_session.answer` stores the final report
|
||||
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||
|
||||
### Requirement: AIOps trace uses existing evidence tables
|
||||
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||
|
||||
#### Scenario: AIOps uses evidence tools
|
||||
- **WHEN** the AIOps flow calls available evidence tools
|
||||
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||
Reference in New Issue
Block a user