Add diagnosis playbook skills

This commit is contained in:
aruo
2026-07-06 08:35:54 +08:00
parent 88e0a6c944
commit 6ccfd33ec5
23 changed files with 1002 additions and 75 deletions
@@ -0,0 +1,38 @@
---
name: diagnose-aiops-alert
description: Diagnose AIOps alert payloads, active Prometheus alerts, alert scope control, HighCPUUsage, HighMemoryUsage, SlowResponse, ServiceUnavailable, and alert-driven incident reports. Use in AIOps flows or when the user asks to diagnose current alerts.
---
# AIOps Alert Diagnosis
## Workflow
1. Determine scope mode.
- Payload present: treat the supplied alert as the primary diagnosis target.
- No payload: call `queryPrometheusAlerts` first and choose P0/P1 or the longest-running firing alert.
2. For payload mode, preserve alert name, service, severity, description, and time range in the `lookup_knowledge` query.
3. Confirm active alert state with `queryPrometheusAlerts` when useful, but do not diagnose unrelated alerts as the main target.
4. Query metrics/logs that match the alert type and service.
5. Produce a report that distinguishes confirmed evidence, related risks, and missing evidence.
## Required Evidence
- Alert state from payload or `queryPrometheusAlerts`.
- `lookup_knowledge` when playbook or runbook guidance is needed.
- Logs and metrics aligned to the alert type.
## Stop Conditions
- If payload mode returns unrelated active alerts, mention them only as related risk.
- If three calls in the same direction fail or return no data, stop that direction and report the failure.
- Do not invent metric values, log lines, or remediation execution results.
## Report Rules
- Use the existing alert analysis report structure.
- Keep the supplied alert as the main diagnosis target in payload mode.
- Include confidence and evidence gaps.
## Eval Anchor
RAG cases: `aiops-payment-latency-alert`, `aiops-prometheus-alert-scope`.
@@ -0,0 +1,34 @@
---
name: diagnose-jvm-memory-risk
description: Diagnose JVM memory risk, high heap usage, OOM risk, OutOfMemoryError, frequent Full GC, memory leak, pod OOMKilled, or order-service memory alerts. Use when memory, JVM, heap, GC, OOM, or OOMKilled appears.
---
# JVM Memory Risk Diagnosis
## Workflow
1. Extract affected service, memory threshold, heap size, GC symptoms, pod/container events, and time window.
2. Call `query_metrics` or alert tools for heap usage, memory usage, GC count/time, and active memory alerts.
3. Call `query_logs` for Full GC warnings, OutOfMemoryError, OOMKilled, restart events, or allocation-heavy stack traces.
4. Call `lookup_knowledge` when JVM memory troubleshooting or remediation guidance is needed.
5. Decide whether the supported risk is high memory pressure, confirmed OOM, suspected leak, or insufficient evidence.
## Required Evidence
- `query_metrics` for resource pressure claims.
- `query_logs` for OOM, GC, or restart evidence.
## Stop Conditions
- High memory usage alone is not proof of memory leak.
- OOM risk is stronger when high memory metrics align with Full GC, OOMKilled, or OutOfMemoryError logs.
- If evidence is incomplete, return LOW_CONFID wording and list the missing metrics/logs.
## Report Rules
- Include immediate mitigation, heap/GC investigation, leak investigation, and monitoring recommendations.
- Do not say the issue can be ignored while memory remains above threshold.
## Eval Anchor
Fixed diagnosis case: `jvm-memory-risk`.
@@ -0,0 +1,35 @@
---
name: diagnose-mysql-connection-pool
description: Diagnose MySQL, HikariCP, database connection pool exhaustion, connection acquisition timeout, slow SQL, connection leak, or database saturation issues. Use when the user mentions MySQL pool, HikariCP, connection pool, database timeout, order-service timeout, or connection exhaustion.
---
# MySQL Connection Pool Diagnosis
## Workflow
1. Extract service, database, timeout symptom, and time window.
2. Call `lookup_knowledge` with MySQL, HikariCP, connection pool, and the affected service.
3. Call `query_logs` for connection acquisition timeout, active/max pool counts, waiting threads, leak warnings, slow query, or lock waits.
4. Call `query_metrics` when metrics are available for active connections, idle connections, wait time, DB latency, and error rate.
5. Decide whether the evidence supports pool exhaustion, slow SQL causing saturation, connection leak, or insufficient evidence.
## Required Evidence
- `lookup_knowledge` for pool configuration and diagnosis guidance.
- `query_logs` for concrete pool or SQL symptoms.
- `query_metrics` when making saturation or capacity claims.
## Stop Conditions
- Confirmed pool exhaustion requires log or metric evidence such as active equals max, waiting threads, acquisition timeout, or leak warnings.
- If only request timeout is present without pool evidence, state that the pool hypothesis is unconfirmed.
- If logs show slow SQL but not pool saturation, report slow SQL as the stronger supported cause.
## Report Rules
- Include current evidence, likely root cause, missing evidence, short-term mitigation, and long-term fix.
- Avoid saying "fully confirmed" unless at least two evidence sources align.
## Eval Anchor
Fixed diagnosis case: `mysql-pool-exhausted`.
@@ -0,0 +1,37 @@
---
name: diagnose-payment-timeout
description: Diagnose payment API, payment gateway, ERR_TIMEOUT, gateway timeout, payment-service latency, or payment request timeout issues. Use when the user mentions payment timeout, ERR_TIMEOUT, ERR_GATEWAY_TIMEOUT, slow payment, or payment-service latency.
---
# Payment Timeout Diagnosis
## Workflow
1. Identify the affected payment service, error code, endpoint, and time window from the user request.
2. Call `read_skill` only once for this playbook, then follow the evidence order below.
3. Call `lookup_knowledge` with a narrow query containing payment, timeout, the error code if present, and the affected service.
4. Call `query_logs` for payment-service timeout, downstream dependency timeout, gateway timeout, or request duration above threshold.
5. Call `query_metrics` or alert tools for latency, error rate, saturation, and active alerts when metrics are available.
6. Compare knowledge guidance with logs and metrics before stating a root cause.
## Required Evidence
- `lookup_knowledge` for error-code or payment timeout guidance.
- `query_logs` for concrete timeout or dependency evidence.
- `query_metrics` when the question asks for impact, latency, or current alert state.
## Stop Conditions
- If only knowledge is available and logs/metrics are missing, return LOW_CONFID language.
- If tools fail or return no evidence, state which evidence is missing and do not claim a confirmed root cause.
- Do not repeatedly call `lookup_knowledge` with synonym-only queries after a relevant result.
## Report Rules
- Separate immediate mitigation from long-term remediation.
- Cite the evidence source type for each key conclusion.
- Do not claim payment provider failure unless logs or metrics support an upstream dependency issue.
## Eval Anchor
Fixed diagnosis case: `payment-timeout`.
@@ -0,0 +1,34 @@
---
name: diagnose-redis-timeout
description: Diagnose Redis timeout, Redis connection timeout, cache dependency timeout, Redis cluster unavailable, hot key, network latency, or payment-service Redis dependency failures. Use when Redis or cache timeout appears in the user request, logs, or alert payload.
---
# Redis Timeout Diagnosis
## Workflow
1. Extract affected service, Redis operation, host/cluster, timeout value, and time window.
2. Call `lookup_knowledge` for Redis timeout or cache troubleshooting guidance when knowledge evidence is needed.
3. Call `query_logs` for Redis connection timeout, retry count, host, command latency, hot key, or dependency errors.
4. Call `query_metrics` when available for Redis latency, connection count, CPU, memory, error rate, or network saturation.
5. Distinguish client timeout, Redis saturation, network issue, and missing evidence.
## Required Evidence
- `query_logs` is mandatory for a concrete Redis timeout claim.
- `lookup_knowledge` is recommended for remediation and configuration guidance.
- `query_metrics` is required before claiming Redis resource saturation.
## Stop Conditions
- If only one Redis timeout log exists and no metrics are available, return LOW_CONFID wording.
- If Redis is only mentioned as a possible downstream dependency, do not make it the root cause without supporting logs.
## Report Rules
- State whether the supported issue is client-side timeout, Redis cluster issue, network issue, or unconfirmed.
- Include retry/backoff, timeout tuning, connection pool, and monitoring recommendations only when relevant.
## Eval Anchor
Fixed diagnosis case: `redis-timeout`.
@@ -0,0 +1,33 @@
---
name: diagnose-slow-response
description: Diagnose slow response, high P95/P99 latency, API latency regression, slow request, downstream latency, or user-service response time alerts. Use when the user mentions P99, P95, response time, slow endpoint, latency, or SlowResponse alerts.
---
# Slow Response Diagnosis
## Workflow
1. Extract service, endpoint, latency percentile, threshold, and time window.
2. Call `query_metrics` or alert tools to confirm latency and impact.
3. Call `query_logs` for slow request records, endpoint duration, downstream timing, cache misses, or database query timeout.
4. Call `lookup_knowledge` when process guidance, service-specific runbook, or known failure mode evidence is needed.
5. Classify the supported cause: database slow query, downstream dependency, cache miss, resource saturation, or insufficient evidence.
## Required Evidence
- `query_metrics` for latency or alert confirmation.
- `query_logs` for endpoint-level or dependency-level evidence.
## Stop Conditions
- If metrics show latency but logs do not identify a cause, say impact is confirmed but root cause is not.
- If logs identify slow SQL or dependency latency, use that as a candidate cause and mark confidence based on metric alignment.
## Report Rules
- Include impacted endpoints, observed latency, suspected bottleneck, evidence gaps, and next checks.
- Do not say there is no risk when P95/P99 remains above threshold.
## Eval Anchor
Fixed diagnosis case: `slow-response`.