feat: add traceable scoped AIOps diagnosis
This commit is contained in:
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint has two natural modes:
|
||||
|
||||
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
|
||||
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
|
||||
|
||||
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make AIOps payload mode single-alert focused.
|
||||
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
|
||||
- Keep the change prompt-only and low risk.
|
||||
- Add tests for prompt scope rules.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add a Verifier Agent.
|
||||
- Do not force tool calls in Java code.
|
||||
- Do not change `/api/ai_ops` request/response contracts.
|
||||
- Do not modify mock alert data.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
|
||||
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
|
||||
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
|
||||
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
|
||||
|
||||
## Prompt Rules
|
||||
|
||||
Payload mode MUST instruct the agent:
|
||||
|
||||
- Treat supplied payload as the primary and only report target.
|
||||
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
|
||||
- Do not create root-cause sections for unrelated active alerts.
|
||||
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
|
||||
|
||||
Auto-discovery mode MUST instruct the agent:
|
||||
|
||||
- First call `queryPrometheusAlerts`.
|
||||
- Select P0/P1 or longest-running firing alerts.
|
||||
- Analyze one or more active alerts based on severity and evidence.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
|
||||
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
|
||||
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
|
||||
Reference in New Issue
Block a user