Add diagnosis playbook skills
This commit is contained in:
@@ -159,7 +159,35 @@ ToolCallbackProvider
|
||||
- retrieved domains。
|
||||
- dedup reason。
|
||||
|
||||
## 6. 与旧版设计的差异
|
||||
## 6. Skill / Playbook 流程
|
||||
|
||||
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||
|
||||
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||
ExecutorHook --> ReadSkill["read_skill"]
|
||||
Plan --> Executor["Executor"]
|
||||
Executor --> ReadSkill
|
||||
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
@@ -167,9 +195,9 @@ ToolCallbackProvider
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||
|
||||
## 7. 后续演进
|
||||
## 8. 后续演进
|
||||
|
||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||
|
||||
@@ -184,4 +212,3 @@ ToolCallbackProvider
|
||||
- 不同故障类型的工具权限明显不同。
|
||||
- Trace 能证明某类问题需要独立的推理策略。
|
||||
- 评测集能覆盖拆分前后的行为差异。
|
||||
|
||||
|
||||
@@ -48,6 +48,13 @@ flowchart TB
|
||||
AlertsTool["queryPrometheusAlerts"]
|
||||
end
|
||||
|
||||
subgraph Skills["Skill / Playbook"]
|
||||
SkillRegistry["SkillRegistry"]
|
||||
PlannerSkillHook["PlannerSkillMetadataHook"]
|
||||
SkillsHook["SkillsAgentHook"]
|
||||
ReadSkill["read_skill"]
|
||||
end
|
||||
|
||||
subgraph RAG["RAG Retrieval"]
|
||||
L0["KnowledgeIndexService"]
|
||||
VectorSearch["VectorSearchService"]
|
||||
@@ -66,6 +73,11 @@ flowchart TB
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
SkillRegistry --> SkillsHook
|
||||
SkillsHook --> Executor
|
||||
Executor --> ReadSkill
|
||||
Agent --> Tools
|
||||
KnowledgeTool --> RAG
|
||||
RAG --> Store
|
||||
@@ -101,6 +113,12 @@ Evidence Tools
|
||||
-> query_metrics
|
||||
-> queryPrometheusAlerts
|
||||
|
||||
Skill / Playbook
|
||||
-> SkillRegistry
|
||||
-> PlannerSkillMetadataHook gives Planner name/description only
|
||||
-> SkillsAgentHook gives Executor read_skill
|
||||
-> Verifier is isolated from skills
|
||||
|
||||
RAG Retrieval
|
||||
-> KnowledgeIndexService
|
||||
-> VectorSearchService
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
committed: true
|
||||
date: 2026-07-05
|
||||
@@ -0,0 +1,72 @@
|
||||
# Decisions: diagnosis-playbook-skills
|
||||
|
||||
## Discover
|
||||
|
||||
- Entry summary: extract high-frequency diagnosis workflows into versionable skills/playbooks and make agents load them progressively.
|
||||
- Slug: `diagnosis-playbook-skills`.
|
||||
- Scale: `standard`.
|
||||
- Capability source: sm-flow built-in protocol for Discover; `grill-with-docs` evidence-driven behavior used by reading project docs and code instead of blocking on user questions.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` hit related projects: `diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`, `evidence-trace-hardening`, `aiops-alert-scope-control`, `session-dedup-knowledge-map`, `executor-action-memory-relevance`.
|
||||
- `devflow/glossary/CONTEXT.md` confirms `ReactAgent`, `ToolCall`, `DiagnosisRecord`, and current historical terminology; current architecture documents supersede old `diagnosis_record` as the primary model.
|
||||
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, fallback-capable playbooks.
|
||||
- `mvp/architecture/harness-quality-gates.md` requires Prompt contract, Tool boundary, Trace persistence, Verifier/rule evaluation, and eval baselines to remain authoritative.
|
||||
- `mvp/eval/cases/diagnosis-cases.json` anchors initial playbooks: payment timeout, MySQL pool exhausted, Redis timeout, slow response, JVM memory risk.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| Dimension | Question | Mode | Resolution |
|
||||
|---|---|---|---|
|
||||
| Terminology | Should these artifacts be called Skill or Playbook? | evidence-driven | Use "diagnosis playbook skills": skills are the runtime mechanism, playbooks are the diagnosis workflow content. |
|
||||
| Boundary | Should skills contain factual knowledge or workflow guidance? | evidence-driven | Skills contain workflow guidance; knowledge facts remain in `knowledge_base/`. |
|
||||
| Acceptance | What proves the change works? | evidence-driven | Unit tests for catalog/tool behavior plus existing diagnosis eval compile/test stability. |
|
||||
| Interface impact | Does this alter external API or DB contracts? | evidence-driven | No external API/DB change; L2 internal interface due new tool/service and agent methodTools change. |
|
||||
|
||||
No user-interview question is blocking because the user explicitly asked to implement the change and prior conversation already confirmed the intended direction.
|
||||
|
||||
## Impact Analysis
|
||||
|
||||
GitNexus MCP tools are not exposed in this environment, so required GitNexus impact analysis could not be run. Local substitute analysis:
|
||||
|
||||
- `ChatService.createReactAgent`, `buildChatPlannerAgent`, `buildChatExecutorAgent`, and `buildMethodToolsArray` are called by Chat controller paths and covered by `ChatServiceSequentialAgentTest` / smoke tests.
|
||||
- `AiOpsService.buildPlannerAgent`, `buildExecutorAgent`, and private `buildMethodToolsArray` affect `POST /api/ai_ops` via `ChatController`.
|
||||
- Risk level: medium. Prompt and tool availability changes may alter agent behavior, but no external API, DTO, DB, or status contract changes.
|
||||
|
||||
## Specify / Commit
|
||||
|
||||
- Cross-artifact alignment:
|
||||
- brief/proposal goal -> proposal: aligned.
|
||||
- proposal scope -> design: aligned.
|
||||
- design decisions -> specs/tasks: aligned.
|
||||
- specs observable behavior -> tasks: aligned.
|
||||
- Interface impact: L2 internal interface.
|
||||
- Commit status: `.committed` created after file integrity and consistency checks.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Reference code:
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||
- Tool pattern: Spring AI Alibaba `SkillsAgentHook` contributes the official `read_skill` ToolCallback from hooks.
|
||||
- Prompt pattern: `SkillsInterceptor` from the hook augments model requests with compact skill metadata.
|
||||
- Test pattern: service tests instantiate classes manually with `ReflectionTestUtils`; new dependencies must be injectable or optional enough for tests.
|
||||
- Dependency check: local `1.1.0.0-RC2` jars do not contain `SkillsAgentHook`; `1.1.2.0` jars contain `com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook`, `ReadSkillTool`, `SkillRegistry`, and `ClasspathSkillRegistry`.
|
||||
|
||||
## Apply
|
||||
|
||||
- Capability source: `openspec-apply-change` guidance was loaded; implementation used local sm-flow/OpenSpec fallback because the work required direct file edits and the OpenSpec CLI was not needed for artifact discovery.
|
||||
- Implemented `src/main/resources/skills/*/SKILL.md` for six diagnosis playbooks.
|
||||
- Upgraded Spring AI Alibaba BOMs to `1.1.2.0`.
|
||||
- Implemented `SkillConfig` with `ClasspathSkillRegistry` loading `classpath:skills`.
|
||||
- Wired `SkillsAgentHook` into Chat single-agent, Chat Planner/Executor, and AIOps Planner/Executor agents.
|
||||
- Removed custom prompt-catalog injection from active service paths; the official `SkillsInterceptor` now handles skill catalog injection.
|
||||
- Kept Chat Verifier prompt unchanged.
|
||||
- Removed the earlier local fallback `SkillCatalogService` and `ReadSkillTool` source files after switching to official Alibaba skills support.
|
||||
- Verification command:
|
||||
- `mvn -q "-Dtest=SkillCatalogServiceTest,ChatServiceSequentialAgentTest,AiOpsServiceTest,DiagnosisTraceEvaluatorTest" test`
|
||||
- Verification result: passed.
|
||||
@@ -0,0 +1,82 @@
|
||||
# Design
|
||||
|
||||
## Architecture
|
||||
|
||||
```text
|
||||
src/main/resources/skills/
|
||||
-> SKILL.md files
|
||||
-> ClasspathSkillRegistry bean
|
||||
-> loads classpath skills
|
||||
-> backs official read_skill
|
||||
-> PlannerSkillMetadataHook
|
||||
-> adds planner-only skill metadata messages
|
||||
-> does not expose read_skill
|
||||
-> SkillsAgentHook
|
||||
-> adds official read_skill ToolCallback for Executor / single-agent Chat
|
||||
-> adds SkillsInterceptor prompt augmentation outside Planner
|
||||
-> ChatService / AiOpsService
|
||||
-> Planner receives metadata only
|
||||
-> Executor and single-agent Chat receive official skill hook
|
||||
-> Verifier remains isolated
|
||||
```
|
||||
|
||||
## Skill Contract
|
||||
|
||||
Each skill folder contains a `SKILL.md` with YAML frontmatter:
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: diagnose-mysql-connection-pool
|
||||
description: ...
|
||||
---
|
||||
```
|
||||
|
||||
The body contains:
|
||||
|
||||
- Trigger conditions.
|
||||
- Required evidence.
|
||||
- Recommended tool order.
|
||||
- Query construction hints.
|
||||
- Stop conditions and low-confidence behavior.
|
||||
- Report requirements.
|
||||
- Eval anchor when one exists.
|
||||
|
||||
## Prompt Injection
|
||||
|
||||
Planner agents receive a project-local `PlannerSkillMetadataHook` message that contains skill names and descriptions only. The message also requires `selected_skill`, `selection_reason`, and an ordered `plan` in the Planner output.
|
||||
|
||||
`SkillsAgentHook` provides `SkillsInterceptor`, which injects the official compact skill section containing skill names, descriptions, and loading instructions into model requests for:
|
||||
|
||||
- Chat Executor prompt.
|
||||
- AIOps Executor prompt.
|
||||
- Single-agent Chat prompt.
|
||||
|
||||
Planner prompts are not augmented by `SkillsAgentHook`, so Planner cannot receive the official `read_skill` tool. Verifier prompt is not augmented.
|
||||
|
||||
## Tool Exposure
|
||||
|
||||
Spring AI Alibaba's official `ReadSkillTool` exposes:
|
||||
|
||||
```java
|
||||
read_skill(skill_name)
|
||||
```
|
||||
|
||||
The tool returns the full `SKILL.md` body for a known skill or a structured error for missing skills.
|
||||
|
||||
`read_skill` is supplied by `SkillsAgentHook`, not by local `methodTools`. Planner selects a skill from metadata and writes the selection into `planner_plan`; Executor reads the selected skill before executing scenario-specific evidence collection.
|
||||
|
||||
## Trace Behavior
|
||||
|
||||
`read_skill` is a guidance tool, not an evidence tool. It does not write `tool_invocation` because the existing eval and verifier treat evidence tools as factual data sources. Actual diagnostic evidence must still come from `lookup_knowledge`, `query_logs`, `query_metrics`, and alert tools.
|
||||
|
||||
## Fallback
|
||||
|
||||
If a skill is not found or cannot be read:
|
||||
|
||||
- The tool returns a structured text error.
|
||||
- The agent must fall back to generic Executor prompt behavior.
|
||||
- It must not invent playbook content.
|
||||
|
||||
## Alibaba Skills Integration
|
||||
|
||||
The implementation uses `spring-ai-alibaba-agent-framework:1.1.2.0`, where `SkillsAgentHook` lives in `com.alibaba.cloud.ai.graph.agent.hook.skills` and `ClasspathSkillRegistry` lives in `com.alibaba.cloud.ai.graph.skills.registry.classpath`.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Diagnosis Playbook Skills
|
||||
|
||||
## Problem
|
||||
|
||||
The MVP diagnosis Agent already has trace persistence, evidence tools, verifier gates, and fixed eval cases, but scenario-specific diagnosis workflows still live in broad prompts and knowledge-base documents. This makes high-frequency fault diagnosis depend too much on the generic Executor prompt and makes it harder to version, review, and reuse diagnostic procedures.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Introduce project-local diagnosis playbook skills using progressive disclosure:
|
||||
|
||||
- Store versionable playbook skills under `src/main/resources/skills/`.
|
||||
- Use Spring AI Alibaba `SkillRegistry` + `SkillsAgentHook` so Planner/Executor agents can load full skill instructions only when a matching diagnosis scenario appears.
|
||||
- Let the official skills interceptor inject the compact skill catalog into eligible agent prompts.
|
||||
- Keep knowledge facts in `knowledge_base/`; skills define workflow, evidence requirements, stop conditions, and report rules.
|
||||
- Keep Verifier isolated from skills. It must continue to validate only existing tool evidence.
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- Payment timeout diagnosis playbook.
|
||||
- MySQL connection pool diagnosis playbook.
|
||||
- Redis timeout diagnosis playbook.
|
||||
- Slow response diagnosis playbook.
|
||||
- JVM memory risk diagnosis playbook.
|
||||
- AIOps alert diagnosis playbook.
|
||||
- Classpath skill registry configuration.
|
||||
- Chat and AIOps Planner/Executor `SkillsAgentHook` wiring.
|
||||
- Focused tests for skill loading/catalog behavior and existing diagnosis eval stability.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Replacing `lookup_knowledge` with implicit advisor retrieval.
|
||||
- Replacing the Chat Verifier contract.
|
||||
- Persisting a new database field for playbook usage.
|
||||
- Creating SubAgents for each playbook.
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, traceable, fallback-capable playbooks.
|
||||
- `mvp/architecture/harness-quality-gates.md` requires evidence tool calls, trace persistence, verifier/rule evaluation, and eval baselines to remain authoritative.
|
||||
- `knowledge_base/` remains the source for factual definitions and troubleshooting knowledge.
|
||||
- `mvp/eval/cases/diagnosis-cases.json` provides the first fixed diagnosis scenarios and evidence-tool expectations.
|
||||
- Spring AI Alibaba `1.1.2.0` provides `SkillsAgentHook`, `ClasspathSkillRegistry`, and the official `read_skill` tool.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
L2 internal interface:
|
||||
|
||||
- Adds an internal `SkillRegistry` bean backed by classpath `skills`.
|
||||
- Adds `SkillsAgentHook` to Chat/AIOps Planner and Executor agents.
|
||||
- Does not change HTTP API, DTOs, database schema, or external response contracts.
|
||||
|
||||
## Risks
|
||||
|
||||
- The hook adds the official `read_skill` tool to eligible agents and may affect tool selection.
|
||||
- Skill instructions could conflict with existing prompt constraints if not scoped carefully.
|
||||
- Tests that instantiate `ChatService` manually must inject or tolerate the new skill tool dependency.
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
# diagnosis-playbook-skills Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
|
||||
|
||||
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
|
||||
|
||||
#### Scenario: Planner or Executor receives available skill metadata
|
||||
|
||||
- **GIVEN** classpath skill folders exist under `skills/`
|
||||
- **WHEN** Chat or AIOps Planner/Executor agents are built
|
||||
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
|
||||
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
|
||||
|
||||
### Requirement: Executor SHALL read full playbook instructions on demand
|
||||
|
||||
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
|
||||
|
||||
#### Scenario: Executor reads an existing skill
|
||||
|
||||
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
|
||||
- **WHEN** the Executor calls `read_skill` with that name
|
||||
- **THEN** the tool SHALL return the full skill instructions
|
||||
- **AND** the result SHALL include the skill name
|
||||
|
||||
#### Scenario: Executor requests an unknown skill
|
||||
|
||||
- **WHEN** the Executor calls `read_skill` with an unknown name
|
||||
- **THEN** the tool SHALL return a bounded error message
|
||||
- **AND** the message SHALL list valid skill names
|
||||
|
||||
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
|
||||
|
||||
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
|
||||
|
||||
#### Scenario: Executor uses a playbook
|
||||
|
||||
- **WHEN** a diagnosis playbook applies to a user issue
|
||||
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
|
||||
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
|
||||
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
|
||||
|
||||
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
|
||||
|
||||
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
|
||||
|
||||
#### Scenario: Fixed diagnosis case has a matching playbook
|
||||
|
||||
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
|
||||
- **THEN** a matching diagnosis skill SHALL exist
|
||||
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
|
||||
@@ -0,0 +1,8 @@
|
||||
# Tasks
|
||||
|
||||
- [x] 1. Add diagnosis skill resources under `src/main/resources/skills/`.
|
||||
- [x] 2. Upgrade Spring AI Alibaba to a version that provides `SkillsAgentHook`.
|
||||
- [x] 3. Add a `ClasspathSkillRegistry` bean for classpath skill resources.
|
||||
- [x] 4. Wire `SkillsAgentHook` into Chat/AIOps single-agent, Planner, and Executor agents while keeping Verifier unchanged.
|
||||
- [x] 5. Add focused unit tests for registry loading and official `read_skill` behavior.
|
||||
- [x] 6. Run focused compile/tests and update this task list.
|
||||
@@ -0,0 +1,56 @@
|
||||
# diagnosis-playbook-skills Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
|
||||
|
||||
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
|
||||
|
||||
#### Scenario: Planner or Executor receives available skill metadata
|
||||
|
||||
- **GIVEN** classpath skill folders exist under `skills/`
|
||||
- **WHEN** Chat or AIOps Planner/Executor agents are built
|
||||
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
|
||||
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
|
||||
|
||||
### Requirement: Executor SHALL read full playbook instructions on demand
|
||||
|
||||
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
|
||||
|
||||
#### Scenario: Executor reads an existing skill
|
||||
|
||||
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
|
||||
- **WHEN** the Executor calls `read_skill` with that name
|
||||
- **THEN** the tool SHALL return the full skill instructions
|
||||
- **AND** the result SHALL include the skill name
|
||||
|
||||
#### Scenario: Executor requests an unknown skill
|
||||
|
||||
- **WHEN** the Executor calls `read_skill` with an unknown name
|
||||
- **THEN** the tool SHALL return a bounded error message
|
||||
- **AND** the message SHALL list valid skill names
|
||||
|
||||
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
|
||||
|
||||
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
|
||||
|
||||
#### Scenario: Executor uses a playbook
|
||||
|
||||
- **WHEN** a diagnosis playbook applies to a user issue
|
||||
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
|
||||
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
|
||||
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
|
||||
|
||||
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
|
||||
|
||||
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
|
||||
|
||||
#### Scenario: Fixed diagnosis case has a matching playbook
|
||||
|
||||
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
|
||||
- **THEN** a matching diagnosis skill SHALL exist
|
||||
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
|
||||
@@ -20,8 +20,8 @@
|
||||
<maven.compiler.target>17</maven.compiler.target>
|
||||
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
|
||||
<spring-ai.version>1.1.7</spring-ai.version>
|
||||
<spring-ai-alibaba.version>1.1.0.0-RC2</spring-ai-alibaba.version>
|
||||
<spring-ai-alibaba-extensions.version>1.1.0.0-RC2</spring-ai-alibaba-extensions.version>
|
||||
<spring-ai-alibaba.version>1.1.2.0</spring-ai-alibaba.version>
|
||||
<spring-ai-alibaba-extensions.version>1.1.2.0</spring-ai-alibaba-extensions.version>
|
||||
</properties>
|
||||
<dependencyManagement>
|
||||
<dependencies>
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
package com.superbiz.agent.config;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.skills.SkillMetadata;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.classpath.ClasspathSkillRegistry;
|
||||
import org.springframework.ai.chat.prompt.SystemPromptTemplate;
|
||||
import org.springframework.context.annotation.Bean;
|
||||
import org.springframework.context.annotation.Configuration;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.util.List;
|
||||
import java.util.Optional;
|
||||
|
||||
@Configuration
|
||||
public class SkillConfig {
|
||||
|
||||
private static final String ACTIVE_SKILL_NAME = "diagnose-mysql-connection-pool";
|
||||
|
||||
@Bean
|
||||
public SkillRegistry skillRegistry() {
|
||||
SkillRegistry classpathRegistry = ClasspathSkillRegistry.builder()
|
||||
.classpathPath("skills")
|
||||
.basePath("target/skills-cache")
|
||||
.build();
|
||||
return new SingleSkillRegistry(classpathRegistry, ACTIVE_SKILL_NAME);
|
||||
}
|
||||
|
||||
private record SingleSkillRegistry(SkillRegistry delegate, String activeSkillName) implements SkillRegistry {
|
||||
|
||||
@Override
|
||||
public List<SkillMetadata> listAll() {
|
||||
return delegate.listAll().stream()
|
||||
.filter(skill -> activeSkillName.equals(skill.getName()))
|
||||
.toList();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String getRegistryType() {
|
||||
return delegate.getRegistryType();
|
||||
}
|
||||
|
||||
@Override
|
||||
public String readSkillContent(String skillName) throws IOException {
|
||||
if (!activeSkillName.equals(skillName)) {
|
||||
throw new IOException("Skill not found: " + skillName);
|
||||
}
|
||||
return delegate.readSkillContent(skillName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public String getSkillLoadInstructions() {
|
||||
return delegate.getSkillLoadInstructions();
|
||||
}
|
||||
|
||||
@Override
|
||||
public SystemPromptTemplate getSystemPromptTemplate() {
|
||||
return delegate.getSystemPromptTemplate();
|
||||
}
|
||||
|
||||
@Override
|
||||
public Optional<SkillMetadata> get(String skillName) {
|
||||
if (!activeSkillName.equals(skillName)) {
|
||||
return Optional.empty();
|
||||
}
|
||||
return delegate.get(skillName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public int size() {
|
||||
return listAll().size();
|
||||
}
|
||||
|
||||
@Override
|
||||
public boolean contains(String skillName) {
|
||||
return activeSkillName.equals(skillName) && delegate.contains(skillName);
|
||||
}
|
||||
|
||||
@Override
|
||||
public void reload() {
|
||||
delegate.reload();
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,101 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.Prioritized;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.alibaba.cloud.ai.graph.skills.SkillMetadata;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.SystemMessage;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Injects planner-visible skill metadata without exposing the full skill loader tool.
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions(HookPosition.BEFORE_MODEL)
|
||||
public class PlannerSkillMetadataHook extends MessagesModelHook {
|
||||
|
||||
private static final String CATALOG_MARKER = "\"skill_catalog\"";
|
||||
|
||||
private final SkillRegistry skillRegistry;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
public PlannerSkillMetadataHook(SkillRegistry skillRegistry) {
|
||||
this.skillRegistry = skillRegistry;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String getName() {
|
||||
return "planner_skill_metadata_hook";
|
||||
}
|
||||
|
||||
@Override
|
||||
public int getOrder() {
|
||||
return Prioritized.HIGHEST_PRECEDENCE;
|
||||
}
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
if (skillRegistry == null || skillRegistry.size() == 0 || hasCatalog(previousMessages)) {
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
try {
|
||||
List<Map<String, String>> skills = skillRegistry.listAll().stream()
|
||||
.map(this::toSkillSummary)
|
||||
.toList();
|
||||
if (skills.isEmpty()) {
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
Map<String, Object> catalog = new LinkedHashMap<>();
|
||||
catalog.put("purpose", "Planner-visible diagnosis skill metadata only.");
|
||||
catalog.put("rules", List.of(
|
||||
"Choose at most one primary skill.",
|
||||
"Do not load full skill instructions in Planner.",
|
||||
"Executor reads the selected skill before evidence collection.",
|
||||
"If no skill matches, set selected_skill to null."
|
||||
));
|
||||
catalog.put("skills", skills);
|
||||
catalog.put("required_planner_output", Map.of(
|
||||
"selected_skill", "skill name or null",
|
||||
"selection_reason", "short reason",
|
||||
"plan", "ordered execution step list"
|
||||
));
|
||||
|
||||
Map<String, Object> payload = Map.of("skill_catalog", catalog);
|
||||
String content = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(payload);
|
||||
|
||||
List<Message> updatedMessages = new ArrayList<>(previousMessages.size() + 1);
|
||||
updatedMessages.add(new SystemMessage(content));
|
||||
updatedMessages.addAll(previousMessages);
|
||||
return new AgentCommand(updatedMessages);
|
||||
} catch (Exception e) {
|
||||
log.warn("Failed to inject planner skill metadata, fallback to original messages", e);
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
}
|
||||
|
||||
private Map<String, String> toSkillSummary(SkillMetadata skill) {
|
||||
Map<String, String> summary = new LinkedHashMap<>();
|
||||
summary.put("name", skill.getName());
|
||||
summary.put("description", skill.getDescription());
|
||||
return summary;
|
||||
}
|
||||
|
||||
private boolean hasCatalog(List<Message> messages) {
|
||||
return messages.stream()
|
||||
.map(Message::getText)
|
||||
.anyMatch(text -> text != null && text.contains(CATALOG_MARKER));
|
||||
}
|
||||
}
|
||||
@@ -4,7 +4,10 @@ import org.springframework.ai.chat.model.ChatModel;
|
||||
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook;
|
||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
@@ -13,6 +16,7 @@ import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.dto.AIOpsRequest;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
@@ -32,8 +36,8 @@ import java.util.Optional;
|
||||
import java.util.UUID;
|
||||
|
||||
/**
|
||||
* AI Ops 智能运维服务
|
||||
* 负责多 Agent 协作的告警分析流程
|
||||
* AI Ops 闂備礁鎼幊妯肩磽濮樿泛绀傛俊顖滅帛娴溿倝鏌熼柇锕€鏋熸俊顖氾躬閺岋繝宕煎┑鍩裤垹鈹?
|
||||
* 闂佽崵濮甸崝妤呭窗閺囥垺鍎楁俊銈勭缁?Agent 闂備礁鎲¢〃鍛崲鐎n剛绀婇柡鍐ㄧ墛閸庡秹鏌涢弴銊ヤ簼闁哥喓鍋ら幃褰掑焵椤掑嫭鏅濋柍褜鍓熷畷瑙勬償閵娿儱鍤戦梺褰掑亰閸橀箖濡堕敂鍓х<?
|
||||
*/
|
||||
@Service
|
||||
public class AiOpsService {
|
||||
@@ -49,7 +53,7 @@ public class AiOpsService {
|
||||
@Autowired
|
||||
private QueryMetricsTools queryMetricsTools;
|
||||
|
||||
@Autowired(required = false) // Mock 模式下才注册
|
||||
@Autowired(required = false) // Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒佺箾閹寸偞灏い鎴濇閺呭爼鎮╁ù瀣亙闂侀潧顭堥崕閬嶅棘閳?
|
||||
private QueryLogsTools queryLogsTools;
|
||||
|
||||
@Autowired
|
||||
@@ -73,13 +77,16 @@ public class AiOpsService {
|
||||
@Autowired
|
||||
private SelfEvaluationMergeService selfEvaluationMergeService;
|
||||
|
||||
@Autowired(required = false)
|
||||
private SkillRegistry skillRegistry;
|
||||
|
||||
/**
|
||||
* 执行 AI Ops 告警分析流程
|
||||
* 闂備礁婀遍悷鎶藉幢閳哄倹鏉?AI Ops 闂備礁鎲$粙鎴︽晝閵娾晩鏁嗛柣鏃傚帶缁€鍡涙煕閳╁喚鐒介柍褜鍓濆Λ鍕箒婵炶揪缍€閵嗏偓闁?
|
||||
*
|
||||
* @param chatModel 大模型实例
|
||||
* @param toolCallbacks 工具回调数组
|
||||
* @return 分析结果状态
|
||||
* @throws GraphRunnerException 如果 Agent 执行失败
|
||||
* @param chatModel 濠电姰鍨归悥銏ゅ礋閳ь剚绗熼埀顒€鐣烽崷顓涘亾閿濆簼绨介柡澶庢閵?
|
||||
* @param toolCallbacks 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柛褎顨呴悙濠囨煟閹邦剙顣虫繛鍫濈埣閺屸剝鎷呴悷閭︽缂?
|
||||
* @return 闂備礁鎲$敮鎺懳涘▎鎾村€甸柣锝呯灱绾惧ジ鏌熼幆褜鍤熷ù婊庡灦閺岋絽螣閸喚鍘梺?
|
||||
* @throws GraphRunnerException 濠电姷顣介埀顒€鍟块埀顒€缍婇幃?Agent 闂備礁婀遍悷鎶藉幢閳哄倹鏉稿┑鐘灪閸庤偐鍒掗崜褎鍠?
|
||||
*/
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
||||
@@ -87,27 +94,27 @@ public class AiOpsService {
|
||||
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
AIOpsRequest request, String sessionId) throws GraphRunnerException {
|
||||
logger.info("开始执行 AI Ops 多 Agent 协作流程");
|
||||
logger.info("Starting AI Ops multi-agent analysis");
|
||||
|
||||
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
// 创建或更新诊断会话
|
||||
// 闂備礁鎲$敮妤冪矙閹寸姷纾介柟鎹愵嚙缁狅綁鏌″鍐ㄥ缂佺虎鍨堕弻锟犲磼濞戞﹩鈧粓鏌i敂鐣屽⒌鐎殿噮鍓熼、妯衡攽閸垻宕堕梺?
|
||||
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
||||
// 闂佽崵濮崇粈浣规櫠娴犲鍋?ThreadLocal 濠电偞鍨堕幐鎼佹晝閿濆洨绠旈柛娑欐綑濡﹢鏌涢妷鈺婃缂佲偓閸戯箰okupKnowledgeTool 闂傚倷绶¢崑鍛┍閾忚宕查柛鎰电厛濞间即鏌曢崼婵堝缂佺媭鍨堕弻?sessionId闂?
|
||||
SessionContextHolder.setSessionId(resolvedSessionId);
|
||||
|
||||
try {
|
||||
// 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook)
|
||||
// 闂備礁鎼鍛偓姘煎墰缁?Planner 闂?Executor Agent闂備焦瀵х粙鎴︽偋婵犲洦鍎婇柍鈺佸暟閳?Agent 闂備礁鎲¢懝鍓р偓姘煎灦瀹曢潧顭ㄩ崨顔芥?Hook闂?
|
||||
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
|
||||
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
|
||||
|
||||
// 构建 Supervisor Agent(不加 Hook)
|
||||
// 闂備礁鎼鍛偓姘煎墰缁?Supervisor Agent闂備焦瀵х粙鎴︽偋閸涱垳绠斿鑸靛姇缁€?Hook闂?
|
||||
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
|
||||
.name("ai_ops_supervisor")
|
||||
.description("负责调度 Planner 与 Executor 的多 Agent 控制器")
|
||||
.description("Coordinates Planner and Executor agents")
|
||||
.model(chatModel)
|
||||
.systemPrompt(promptProperties.getSupervisor())
|
||||
.subAgents(List.of(plannerAgent, executorAgent))
|
||||
@@ -115,19 +122,19 @@ public class AiOpsService {
|
||||
|
||||
String taskPrompt = buildTaskPrompt(request);
|
||||
|
||||
logger.info("调用 Supervisor Agent 开始编排...");
|
||||
logger.info("闂佽崵濮撮鍛村疮娴兼潙鏋?Supervisor Agent 闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢垹绱掓鏍﹂偗妤?..");
|
||||
|
||||
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
// 更新诊断会话
|
||||
// 闂備礁鎼ú銈夋偤閵娾晛钃熷┑鐘插鐎氭艾鈹戦悩鎻掓殲闁绘帟妫勯湁闁稿繘妫挎禍銏ゆ煟?
|
||||
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
// 添加调试代码
|
||||
// 婵犵數鍎戠紞鈧い鏇嗗嫭鍙忛柣鎰仛鐎氼剟鏌涢幇闈涘箻婵¤尙顭堥湁闁绘瑥鎳愰幃濂告煟?
|
||||
if (stateOptional.isPresent()) {
|
||||
OverAllState state = stateOptional.get();
|
||||
logger.debug("Final State Keys: {}", state.data().keySet());
|
||||
@@ -146,25 +153,25 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 从执行结果中提取最终报告文本
|
||||
* 濠电偛顕慨瀵糕偓娑掓櫆閺呭爼鎮剧仦鎯т粧閻庡厜鍋撻柍褜鍓涢崚鎺楀Ω閳轰礁鍤戝┑鐘才堥崑鎾绘煠閸偄鐏存鐐存崌楠炲洭顢楅埀顒傚緤閸ф鐓涢柛顐h壘娴滃墽绱撻崒娆戭槮闁绘锕ラ幈銊╁Χ婢跺﹤绐涙繝鐢靛Т閸燁垶鎮楅鈧弻?
|
||||
*
|
||||
* @param state 执行状态
|
||||
* @return 报告文本(如果存在)
|
||||
* @param state 闂備礁婀遍悷鎶藉幢閳哄倹鏉搁梻浣虹帛椤牓宕洪弽顓炵劦?
|
||||
* @return 闂備胶顢婄紙浼村磿闁秴绠熼柨鐔哄Т濡﹢鏌涢妷锝呭闁圭兘浜堕弻銊モ槈濡厧顣哄銈傛暘閸パ冨殤濠电姴锕ら崯浼村箺閻樼粯鐓曢柨鏃囧吹閸樻粎绱?
|
||||
*/
|
||||
public Optional<String> extractFinalReport(OverAllState state) {
|
||||
logger.info("开始提取最终报告...");
|
||||
logger.info("闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢劗鎲告0浣虹獢鐎规洩缍佸浠嬪Ω閿旇法甯涚紓鍌氬€风粈渚€鎮ф繝鍐╁弿闁靛牆顦?..");
|
||||
|
||||
// 提取 Planner 最终输出(包含完整的告警分析报告)
|
||||
// 闂備礁婀辩划顖炲礉閺嚶颁汗?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢э絿绱掓0婵嗗籍鐎规洘鐟╅幃顔锯偓闈涙憸椤︹晠姊洪崨濠勫ⅹ闁瑰啿閰i獮鍡涘醇閳垛晛浜鹃柣鐔哄濠€浼存煛閸☆厾绉柟顖氬暣瀹曠喖顢曢敐鍛畼闂佽崵濮崑鎾绘煥閺囨浜鹃梺鎼炲妼闁帮絽顕i幖浣哥疀妞ゆ挾鍊幘缁樼厱婵炴垶锕╅悡顓犵磼?
|
||||
Optional<AssistantMessage> plannerFinalOutput = state.value("planner_plan")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast);
|
||||
|
||||
if (plannerFinalOutput.isPresent()) {
|
||||
String reportText = plannerFinalOutput.get().getText();
|
||||
logger.info("成功提取到 Planner 最终报告,长度: {}", reportText.length());
|
||||
logger.info("闂備胶鎳撻悺銊╁礉閺囩喐鍙忔繛鎴欏灩缁犵敻鏌熼柇锕€澧紒鎻掓健閺?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢ф稑鈹戦鍝勨偓婵嗙暦閵婏妇绡€闊洦娲滈ˇ顕€姊婚崒妤€浜鹃梺鍓茬厛閸犳牠顢? {}", reportText.length());
|
||||
return Optional.of(reportText);
|
||||
} else {
|
||||
logger.warn("未能提取到 Planner 最终报告");
|
||||
logger.warn("Unable to extract Planner final report");
|
||||
return Optional.empty();
|
||||
}
|
||||
}
|
||||
@@ -197,16 +204,16 @@ public class AiOpsService {
|
||||
|
||||
String buildQuerySummary(AIOpsRequest request) {
|
||||
if (request == null) {
|
||||
return "AI Ops 告警分析";
|
||||
return "AI Ops alert analysis";
|
||||
}
|
||||
|
||||
StringBuilder summary = new StringBuilder("AI Ops 告警分析");
|
||||
appendField(summary, "告警", request.getAlertName());
|
||||
appendField(summary, "服务", request.getService());
|
||||
appendField(summary, "等级", request.getSeverity());
|
||||
appendField(summary, "时间范围", request.getTimeRange());
|
||||
appendField(summary, "描述", request.getDescription());
|
||||
appendField(summary, "请求", request.getUserRequest());
|
||||
StringBuilder summary = new StringBuilder("AI Ops alert analysis");
|
||||
appendField(summary, "alert", request.getAlertName());
|
||||
appendField(summary, "service", request.getService());
|
||||
appendField(summary, "severity", request.getSeverity());
|
||||
appendField(summary, "timeRange", request.getTimeRange());
|
||||
appendField(summary, "description", request.getDescription());
|
||||
appendField(summary, "request", request.getUserRequest());
|
||||
return summary.toString();
|
||||
}
|
||||
|
||||
@@ -238,8 +245,8 @@ public class AiOpsService {
|
||||
|
||||
String buildTaskPrompt(AIOpsRequest request) {
|
||||
StringBuilder prompt = new StringBuilder();
|
||||
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
|
||||
prompt.append("\n\n本次告警输入:\n");
|
||||
prompt.append("You are an enterprise SRE handling an automated alert diagnosis task. Combine tool evidence, run a plan-execute-replan loop, and output the final alert analysis report. Do not fabricate data; if repeated queries fail, clearly state why the task cannot be completed.");
|
||||
prompt.append("\n\nAlert input:\n");
|
||||
prompt.append(buildQuerySummary(request));
|
||||
if (hasAlertPayload(request)) {
|
||||
String knowledgeQuery = buildKnowledgeRetrievalQuery(request);
|
||||
@@ -278,53 +285,69 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Planner Agent
|
||||
* 闂備礁鎼鍛偓姘煎墰缁?Planner Agent
|
||||
*/
|
||||
private ReactAgent buildPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||
return ReactAgent.builder()
|
||||
.name("planner_agent")
|
||||
.description("负责拆解告警、规划与再规划步骤")
|
||||
.description("Plans alert diagnosis steps")
|
||||
.model(chatModel)
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(toolCallbacks)
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "planner"))
|
||||
.hooks(buildHooks("planner"))
|
||||
.outputKey("planner_plan")
|
||||
.build();
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Executor Agent
|
||||
* 闂備礁鎼鍛偓姘煎墰缁?Executor Agent
|
||||
*/
|
||||
private ReactAgent buildExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||
return ReactAgent.builder()
|
||||
.name("executor_agent")
|
||||
.description("负责执行 Planner 的首个步骤并及时反馈")
|
||||
.description("Executes the current Planner step and reports feedback")
|
||||
.model(chatModel)
|
||||
.systemPrompt(promptProperties.getExecutor())
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(toolCallbacks)
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "executor"))
|
||||
.hooks(buildHooks("executor"))
|
||||
.outputKey("executor_feedback")
|
||||
.build();
|
||||
}
|
||||
|
||||
/**
|
||||
* 动态构建方法工具数组
|
||||
* 根据 cls.mock-enabled 决定是否包含 QueryLogsTools
|
||||
* 工具顺序:知识库查询优先,日志查询次之,弃用工具最后
|
||||
* 闂備礁鎲¢弻锝夊礉瀹ュ鐒垫い鎴f硶閸斿秹鏌f惔顔肩仩妞ゆ洘鐟╅幃婊兾熼懡銈呭箥婵犵數鍋涢ˇ鏉棵洪弽銊ヮ嚤闁圭増婢樼粈鍌炴⒑閸噮鍎愭繛鍫濆缁?
|
||||
* 闂備礁鎼粔鐑斤綖婢跺﹦鏆?cls.mock-enabled 闂備礁鎲¢崝鏇㈠疮閸ф鍋╁Δ锝呭暙閸欏﹥銇勯弽銊ь暡闁稿骸锕弻娑㈠冀瑜庨崳褰掓煙?QueryLogsTools
|
||||
* 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柟绋跨昂娴滄粓鏌涘┑鍡楊伀缁炬澘绉归弻銊モ槈濞嗘劗娈ら梺缁樻惈缁辨洟骞忛悩璇插耿婵°倕鍟惃鎴︽⒑閸濆嫯顫﹂柛搴㈡尦椤㈡艾螖娴e壊鍤ゅ┑鈽嗗灠閹碱偆鏁妷鈺傜叆婵炴垶蓱濠€鐗堜繆椤愮喐娅堢紒鐘崇☉铻栧ù锝呮惈瀵劑鏌i悩鍙夊偍闁搞劍妞介、鏇㈠礂閼测斁鏋欓柣搴到婢у海绮堟径灞稿亾濞堝灝鏋涢柛鐔跺嵆瀵偊濡堕崪浣告櫊闂侀潧顦崕鍗烆嚗閺冨牊鐓涢柛顐h壘娴滈箖姊?
|
||||
*/
|
||||
private Object[] buildMethodToolsArray() {
|
||||
if (queryLogsTools != null) {
|
||||
// Mock 模式:包含 QueryLogsTools
|
||||
// Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒勬⒑閹稿海鈯曢柤鐟板⒔閳ь剙鐏氶敃銏犵暦?QueryLogsTools
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools, queryLogsTools};
|
||||
} else {
|
||||
// 真实模式:不包含 QueryLogsTools(由 MCP 提供日志查询功能)
|
||||
}
|
||||
// Real mode excludes local QueryLogsTools because logs are provided by MCP.
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||
}
|
||||
|
||||
private Hook[] buildHooks(String agentName) {
|
||||
AgentLoggingHook loggingHook = new AgentLoggingHook(agentStepRepository, agentName);
|
||||
if (skillRegistry == null || skillRegistry.size() == 0) {
|
||||
return new Hook[]{loggingHook};
|
||||
}
|
||||
if ("planner".equals(agentName)) {
|
||||
return new Hook[]{
|
||||
new PlannerSkillMetadataHook(skillRegistry),
|
||||
loggingHook
|
||||
};
|
||||
}
|
||||
return new Hook[]{
|
||||
loggingHook,
|
||||
SkillsAgentHook.builder()
|
||||
.skillRegistry(skillRegistry)
|
||||
.build()
|
||||
};
|
||||
}
|
||||
|
||||
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
|
||||
/** 濠?agent_step 闂?tool_invocation 婵犳鍠氶幊鎾趁洪敃鍌氱劦妞ゆ帒鍊荤敮娑㈡倵閸偄鍝虹€殿喕绮欏畷鎯邦槼缂佲偓閳ь剚绻?diagnosis_session */
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
try {
|
||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||
@@ -340,7 +363,7 @@ public class AiOpsService {
|
||||
session.setStepCount(stepCount);
|
||||
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
} catch (Exception e) {
|
||||
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
||||
logger.warn("闂備焦鎮堕崕鎶藉磻濞戙垺鏅查柣鎰綑椤曡鲸鎱ㄥΟ铏癸紞婵☆垰鐗撻弻鐔虹矙閹稿骸顦╅梺缁樼壄缁叉儳顕ラ崟顒佺秶妞ゆ劑鍎? sessionId={}", session.getSessionId(), e);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -4,7 +4,10 @@ import com.alibaba.cloud.ai.graph.OverAllState;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SequentialAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook;
|
||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
@@ -13,6 +16,7 @@ import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
||||
import com.superbiz.agent.hook.TokenUsageHolder;
|
||||
import com.superbiz.agent.hook.VerifierInputHook;
|
||||
@@ -99,6 +103,9 @@ public class ChatService {
|
||||
@Autowired
|
||||
private KnowledgeDomainService knowledgeDomainService;
|
||||
|
||||
@Autowired(required = false)
|
||||
private SkillRegistry skillRegistry;
|
||||
|
||||
@Autowired
|
||||
private ToolTraceSummaryService toolTraceSummaryService;
|
||||
|
||||
@@ -160,6 +167,7 @@ public class ChatService {
|
||||
systemPromptBuilder.append("你是一个专业的智能助手,可以获取当前时间、查询天气信息、搜索内部文档知识库,以及查询 Prometheus 告警信息。\n");
|
||||
systemPromptBuilder.append("当用户询问时间相关问题时,**必须每次都调用 getCurrentDateTime 工具**,因为时间会不断变化。即使历史消息中有时间信息,也不要直接复用,必须重新查询最新时间。\n");
|
||||
systemPromptBuilder.append("当用户需要查询公司内部文档、流程、最佳实践或技术指南时,使用 lookupKnowledgeTool 工具。\n");
|
||||
systemPromptBuilder.append("当用户的问题匹配某个诊断 Skill 时,先调用 read_skill 读取对应流程,再按流程调用证据工具。\n");
|
||||
systemPromptBuilder.append("当用户需要查询 Prometheus 告警、监控指标或系统告警状态时,使用 queryPrometheusAlerts 工具。\n");
|
||||
systemPromptBuilder.append("当用户需要查询腾讯云日志时,请调用腾讯云mcp服务查询,默认查询地域ap-guangzhou,查询时间范围为近一个月。\n\n");
|
||||
|
||||
@@ -271,7 +279,7 @@ public class ChatService {
|
||||
.systemPrompt(systemPrompt)
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(getToolCallbacks())
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "intelligent_assistant"))
|
||||
.hooks(buildHooks("intelligent_assistant"))
|
||||
.build();
|
||||
}
|
||||
|
||||
@@ -516,7 +524,7 @@ public class ChatService {
|
||||
.description("负责拆解问题、规划步骤")
|
||||
.model(chatModel)
|
||||
.systemPrompt(prompt.toString())
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "planner"))
|
||||
.hooks(buildHooks("planner"))
|
||||
.outputKey("planner_plan")
|
||||
.build();
|
||||
}
|
||||
@@ -553,11 +561,30 @@ public class ChatService {
|
||||
.systemPrompt(prompt.toString())
|
||||
.methodTools(buildMethodToolsArray())
|
||||
.tools(toolCallbacks)
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "executor"))
|
||||
.hooks(buildHooks("executor"))
|
||||
.outputKey("executor_feedback")
|
||||
.build();
|
||||
}
|
||||
|
||||
private Hook[] buildHooks(String agentName) {
|
||||
AgentLoggingHook loggingHook = new AgentLoggingHook(agentStepRepository, agentName);
|
||||
if (skillRegistry == null || skillRegistry.size() == 0) {
|
||||
return new Hook[]{loggingHook};
|
||||
}
|
||||
if ("planner".equals(agentName)) {
|
||||
return new Hook[]{
|
||||
new PlannerSkillMetadataHook(skillRegistry),
|
||||
loggingHook
|
||||
};
|
||||
}
|
||||
return new Hook[]{
|
||||
loggingHook,
|
||||
SkillsAgentHook.builder()
|
||||
.skillRegistry(skillRegistry)
|
||||
.build()
|
||||
};
|
||||
}
|
||||
|
||||
private String resolveSessionId(String requestedSessionId) {
|
||||
if (requestedSessionId != null && !requestedSessionId.isBlank()) {
|
||||
return requestedSessionId;
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: diagnose-aiops-alert
|
||||
description: Diagnose AIOps alert payloads, active Prometheus alerts, alert scope control, HighCPUUsage, HighMemoryUsage, SlowResponse, ServiceUnavailable, and alert-driven incident reports. Use in AIOps flows or when the user asks to diagnose current alerts.
|
||||
---
|
||||
|
||||
# AIOps Alert Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Determine scope mode.
|
||||
- Payload present: treat the supplied alert as the primary diagnosis target.
|
||||
- No payload: call `queryPrometheusAlerts` first and choose P0/P1 or the longest-running firing alert.
|
||||
2. For payload mode, preserve alert name, service, severity, description, and time range in the `lookup_knowledge` query.
|
||||
3. Confirm active alert state with `queryPrometheusAlerts` when useful, but do not diagnose unrelated alerts as the main target.
|
||||
4. Query metrics/logs that match the alert type and service.
|
||||
5. Produce a report that distinguishes confirmed evidence, related risks, and missing evidence.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- Alert state from payload or `queryPrometheusAlerts`.
|
||||
- `lookup_knowledge` when playbook or runbook guidance is needed.
|
||||
- Logs and metrics aligned to the alert type.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- If payload mode returns unrelated active alerts, mention them only as related risk.
|
||||
- If three calls in the same direction fail or return no data, stop that direction and report the failure.
|
||||
- Do not invent metric values, log lines, or remediation execution results.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- Use the existing alert analysis report structure.
|
||||
- Keep the supplied alert as the main diagnosis target in payload mode.
|
||||
- Include confidence and evidence gaps.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
RAG cases: `aiops-payment-latency-alert`, `aiops-prometheus-alert-scope`.
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
name: diagnose-jvm-memory-risk
|
||||
description: Diagnose JVM memory risk, high heap usage, OOM risk, OutOfMemoryError, frequent Full GC, memory leak, pod OOMKilled, or order-service memory alerts. Use when memory, JVM, heap, GC, OOM, or OOMKilled appears.
|
||||
---
|
||||
|
||||
# JVM Memory Risk Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Extract affected service, memory threshold, heap size, GC symptoms, pod/container events, and time window.
|
||||
2. Call `query_metrics` or alert tools for heap usage, memory usage, GC count/time, and active memory alerts.
|
||||
3. Call `query_logs` for Full GC warnings, OutOfMemoryError, OOMKilled, restart events, or allocation-heavy stack traces.
|
||||
4. Call `lookup_knowledge` when JVM memory troubleshooting or remediation guidance is needed.
|
||||
5. Decide whether the supported risk is high memory pressure, confirmed OOM, suspected leak, or insufficient evidence.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- `query_metrics` for resource pressure claims.
|
||||
- `query_logs` for OOM, GC, or restart evidence.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- High memory usage alone is not proof of memory leak.
|
||||
- OOM risk is stronger when high memory metrics align with Full GC, OOMKilled, or OutOfMemoryError logs.
|
||||
- If evidence is incomplete, return LOW_CONFID wording and list the missing metrics/logs.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- Include immediate mitigation, heap/GC investigation, leak investigation, and monitoring recommendations.
|
||||
- Do not say the issue can be ignored while memory remains above threshold.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
Fixed diagnosis case: `jvm-memory-risk`.
|
||||
@@ -0,0 +1,35 @@
|
||||
---
|
||||
name: diagnose-mysql-connection-pool
|
||||
description: Diagnose MySQL, HikariCP, database connection pool exhaustion, connection acquisition timeout, slow SQL, connection leak, or database saturation issues. Use when the user mentions MySQL pool, HikariCP, connection pool, database timeout, order-service timeout, or connection exhaustion.
|
||||
---
|
||||
|
||||
# MySQL Connection Pool Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Extract service, database, timeout symptom, and time window.
|
||||
2. Call `lookup_knowledge` with MySQL, HikariCP, connection pool, and the affected service.
|
||||
3. Call `query_logs` for connection acquisition timeout, active/max pool counts, waiting threads, leak warnings, slow query, or lock waits.
|
||||
4. Call `query_metrics` when metrics are available for active connections, idle connections, wait time, DB latency, and error rate.
|
||||
5. Decide whether the evidence supports pool exhaustion, slow SQL causing saturation, connection leak, or insufficient evidence.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- `lookup_knowledge` for pool configuration and diagnosis guidance.
|
||||
- `query_logs` for concrete pool or SQL symptoms.
|
||||
- `query_metrics` when making saturation or capacity claims.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- Confirmed pool exhaustion requires log or metric evidence such as active equals max, waiting threads, acquisition timeout, or leak warnings.
|
||||
- If only request timeout is present without pool evidence, state that the pool hypothesis is unconfirmed.
|
||||
- If logs show slow SQL but not pool saturation, report slow SQL as the stronger supported cause.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- Include current evidence, likely root cause, missing evidence, short-term mitigation, and long-term fix.
|
||||
- Avoid saying "fully confirmed" unless at least two evidence sources align.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
Fixed diagnosis case: `mysql-pool-exhausted`.
|
||||
@@ -0,0 +1,37 @@
|
||||
---
|
||||
name: diagnose-payment-timeout
|
||||
description: Diagnose payment API, payment gateway, ERR_TIMEOUT, gateway timeout, payment-service latency, or payment request timeout issues. Use when the user mentions payment timeout, ERR_TIMEOUT, ERR_GATEWAY_TIMEOUT, slow payment, or payment-service latency.
|
||||
---
|
||||
|
||||
# Payment Timeout Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Identify the affected payment service, error code, endpoint, and time window from the user request.
|
||||
2. Call `read_skill` only once for this playbook, then follow the evidence order below.
|
||||
3. Call `lookup_knowledge` with a narrow query containing payment, timeout, the error code if present, and the affected service.
|
||||
4. Call `query_logs` for payment-service timeout, downstream dependency timeout, gateway timeout, or request duration above threshold.
|
||||
5. Call `query_metrics` or alert tools for latency, error rate, saturation, and active alerts when metrics are available.
|
||||
6. Compare knowledge guidance with logs and metrics before stating a root cause.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- `lookup_knowledge` for error-code or payment timeout guidance.
|
||||
- `query_logs` for concrete timeout or dependency evidence.
|
||||
- `query_metrics` when the question asks for impact, latency, or current alert state.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- If only knowledge is available and logs/metrics are missing, return LOW_CONFID language.
|
||||
- If tools fail or return no evidence, state which evidence is missing and do not claim a confirmed root cause.
|
||||
- Do not repeatedly call `lookup_knowledge` with synonym-only queries after a relevant result.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- Separate immediate mitigation from long-term remediation.
|
||||
- Cite the evidence source type for each key conclusion.
|
||||
- Do not claim payment provider failure unless logs or metrics support an upstream dependency issue.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
Fixed diagnosis case: `payment-timeout`.
|
||||
@@ -0,0 +1,34 @@
|
||||
---
|
||||
name: diagnose-redis-timeout
|
||||
description: Diagnose Redis timeout, Redis connection timeout, cache dependency timeout, Redis cluster unavailable, hot key, network latency, or payment-service Redis dependency failures. Use when Redis or cache timeout appears in the user request, logs, or alert payload.
|
||||
---
|
||||
|
||||
# Redis Timeout Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Extract affected service, Redis operation, host/cluster, timeout value, and time window.
|
||||
2. Call `lookup_knowledge` for Redis timeout or cache troubleshooting guidance when knowledge evidence is needed.
|
||||
3. Call `query_logs` for Redis connection timeout, retry count, host, command latency, hot key, or dependency errors.
|
||||
4. Call `query_metrics` when available for Redis latency, connection count, CPU, memory, error rate, or network saturation.
|
||||
5. Distinguish client timeout, Redis saturation, network issue, and missing evidence.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- `query_logs` is mandatory for a concrete Redis timeout claim.
|
||||
- `lookup_knowledge` is recommended for remediation and configuration guidance.
|
||||
- `query_metrics` is required before claiming Redis resource saturation.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- If only one Redis timeout log exists and no metrics are available, return LOW_CONFID wording.
|
||||
- If Redis is only mentioned as a possible downstream dependency, do not make it the root cause without supporting logs.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- State whether the supported issue is client-side timeout, Redis cluster issue, network issue, or unconfirmed.
|
||||
- Include retry/backoff, timeout tuning, connection pool, and monitoring recommendations only when relevant.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
Fixed diagnosis case: `redis-timeout`.
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
name: diagnose-slow-response
|
||||
description: Diagnose slow response, high P95/P99 latency, API latency regression, slow request, downstream latency, or user-service response time alerts. Use when the user mentions P99, P95, response time, slow endpoint, latency, or SlowResponse alerts.
|
||||
---
|
||||
|
||||
# Slow Response Diagnosis
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Extract service, endpoint, latency percentile, threshold, and time window.
|
||||
2. Call `query_metrics` or alert tools to confirm latency and impact.
|
||||
3. Call `query_logs` for slow request records, endpoint duration, downstream timing, cache misses, or database query timeout.
|
||||
4. Call `lookup_knowledge` when process guidance, service-specific runbook, or known failure mode evidence is needed.
|
||||
5. Classify the supported cause: database slow query, downstream dependency, cache miss, resource saturation, or insufficient evidence.
|
||||
|
||||
## Required Evidence
|
||||
|
||||
- `query_metrics` for latency or alert confirmation.
|
||||
- `query_logs` for endpoint-level or dependency-level evidence.
|
||||
|
||||
## Stop Conditions
|
||||
|
||||
- If metrics show latency but logs do not identify a cause, say impact is confirmed but root cause is not.
|
||||
- If logs identify slow SQL or dependency latency, use that as a candidate cause and mark confidence based on metric alignment.
|
||||
|
||||
## Report Rules
|
||||
|
||||
- Include impacted endpoints, observed latency, suspected bottleneck, evidence gaps, and next checks.
|
||||
- Do not say there is no risk when P95/P99 remains above threshold.
|
||||
|
||||
## Eval Anchor
|
||||
|
||||
Fixed diagnosis case: `slow-response`.
|
||||
@@ -56,17 +56,17 @@ class AiOpsServiceTest {
|
||||
request.setSeverity("P1");
|
||||
request.setTimeRange("last_15m");
|
||||
request.setDescription("P95 latency is high");
|
||||
request.setUserRequest("结合日志和指标排查支付超时");
|
||||
request.setUserRequest("check logs and metrics for payment timeout");
|
||||
|
||||
String summary = service.buildQuerySummary(request);
|
||||
|
||||
assertTrue(summary.contains("AI Ops 告警分析"));
|
||||
assertTrue(summary.contains("告警: payment-service-latency-high"));
|
||||
assertTrue(summary.contains("服务: payment-service"));
|
||||
assertTrue(summary.contains("等级: P1"));
|
||||
assertTrue(summary.contains("时间范围: last_15m"));
|
||||
assertTrue(summary.contains("描述: P95 latency is high"));
|
||||
assertTrue(summary.contains("请求: 结合日志和指标排查支付超时"));
|
||||
assertTrue(summary.contains("AI Ops alert analysis"));
|
||||
assertTrue(summary.contains("alert: payment-service-latency-high"));
|
||||
assertTrue(summary.contains("service: payment-service"));
|
||||
assertTrue(summary.contains("severity: P1"));
|
||||
assertTrue(summary.contains("timeRange: last_15m"));
|
||||
assertTrue(summary.contains("description: P95 latency is high"));
|
||||
assertTrue(summary.contains("request: "));
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -99,8 +99,8 @@ class AiOpsServiceTest {
|
||||
assertTrue(prompt.contains("Related Risk"));
|
||||
assertTrue(prompt.contains("Recommended lookup_knowledge query: HighCPUUsage payment-service P1 CPU usage is above 80% last_15m"));
|
||||
assertTrue(prompt.contains("preserves alertName and service"));
|
||||
assertTrue(prompt.contains("告警: HighCPUUsage"));
|
||||
assertTrue(prompt.contains("服务: payment-service"));
|
||||
assertTrue(prompt.contains("alert: HighCPUUsage"));
|
||||
assertTrue(prompt.contains("service: payment-service"));
|
||||
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||
}
|
||||
|
||||
@@ -112,11 +112,11 @@ class AiOpsServiceTest {
|
||||
request.setSeverity(" ");
|
||||
request.setDescription("P95 latency above threshold");
|
||||
request.setTimeRange("last_10m");
|
||||
request.setUserRequest("结合日志和指标排查");
|
||||
request.setUserRequest("check logs and metrics");
|
||||
|
||||
String query = service.buildKnowledgeRetrievalQuery(request);
|
||||
|
||||
assertEquals("HighLatency payment-service P95 latency above threshold last_10m 结合日志和指标排查", query);
|
||||
assertEquals("HighLatency payment-service P95 latency above threshold last_10m check logs and metrics", query);
|
||||
}
|
||||
|
||||
@Test
|
||||
@@ -142,16 +142,16 @@ class AiOpsServiceTest {
|
||||
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId("aiops-session-001")
|
||||
.query("AI Ops 告警分析")
|
||||
.query("AI Ops alert analysis")
|
||||
.status("SUCCESS")
|
||||
.agentFlow("AI_OPS")
|
||||
.build();
|
||||
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
|
||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
|
||||
|
||||
service.persistFinalReport("aiops-session-001", "# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.");
|
||||
service.persistFinalReport("aiops-session-001", "# 闁告稑锕ㄩ鐔煎礆閸℃鈧粙骞庨妷銉﹀暈\nHighCPUUsage payment-service analysis with evidence summary.");
|
||||
|
||||
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
|
||||
assertEquals("# 闁告稑锕ㄩ鐔煎礆閸℃鈧粙骞庨妷銉﹀暈\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
|
||||
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
|
||||
verify(diagnosisSessionRepository).save(session);
|
||||
}
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.classpath.ClasspathSkillRegistry;
|
||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||
@@ -194,6 +197,56 @@ class ChatServiceSequentialAgentTest {
|
||||
assertSame(queryMetricsTools, methodTools[3]);
|
||||
}
|
||||
|
||||
@Test
|
||||
void createReactAgentInjectsSkillCatalogThroughAlibabaHook() throws Exception {
|
||||
ChatService chatService = createChatService();
|
||||
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
|
||||
.classpathPath("skills")
|
||||
.basePath("target/test-skills-cache")
|
||||
.build();
|
||||
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
|
||||
|
||||
ReactAgent agent = chatService.createReactAgent(chatModel, "BASE_TEST_PROMPT");
|
||||
agent.call("diagnose mysql connection pool exhaustion");
|
||||
|
||||
assertTrue(chatModel.promptText.contains("BASE_TEST_PROMPT"));
|
||||
assertTrue(chatModel.promptText.contains("## Skills System"));
|
||||
assertTrue(chatModel.promptText.contains("diagnose-mysql-connection-pool"));
|
||||
assertTrue(chatModel.promptText.contains("read_skill"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void plannerGetsSkillMetadataAndExecutorGetsReadSkillTool() throws Exception {
|
||||
ChatService chatService = createChatService();
|
||||
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
|
||||
.classpathPath("skills")
|
||||
.basePath("target/test-skills-cache")
|
||||
.build();
|
||||
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
|
||||
|
||||
chatService.executeChatComplex(
|
||||
chatModel,
|
||||
new ToolCallback[0],
|
||||
"diagnose mysql connection pool exhaustion",
|
||||
List.of(),
|
||||
"planner-skill-metadata-session"
|
||||
);
|
||||
|
||||
assertTrue(chatModel.plannerPromptText.contains("\"skill_catalog\""));
|
||||
assertTrue(chatModel.plannerPromptText.contains("diagnose-mysql-connection-pool"));
|
||||
assertTrue(chatModel.plannerPromptText.contains("\"selected_skill\""));
|
||||
assertFalse(chatModel.plannerPromptText.contains("## Skills System"));
|
||||
assertFalse(chatModel.plannerPromptText.contains("read_skill"));
|
||||
|
||||
assertTrue(chatModel.executorPromptText.contains("## Skills System"));
|
||||
assertTrue(chatModel.executorPromptText.contains("diagnose-mysql-connection-pool"));
|
||||
assertTrue(chatModel.executorPromptText.contains("read_skill"));
|
||||
assertFalse(chatModel.verifierPromptText.contains("diagnose-mysql-connection-pool"));
|
||||
assertFalse(chatModel.verifierPromptText.contains("read_skill"));
|
||||
}
|
||||
|
||||
private ChatService createChatService() {
|
||||
ChatService chatService = new ChatService();
|
||||
|
||||
@@ -245,6 +298,9 @@ class ChatServiceSequentialAgentTest {
|
||||
private static final class ScriptedChatModel implements ChatModel {
|
||||
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
|
||||
private String promptText = "";
|
||||
private String plannerPromptText = "";
|
||||
private String executorPromptText = "";
|
||||
private String verifierPromptText = "";
|
||||
private boolean sawVerifierPrompt;
|
||||
private final java.util.List<String> verifierOutputs;
|
||||
private int verifierOutputIndex;
|
||||
@@ -283,12 +339,15 @@ class ChatServiceSequentialAgentTest {
|
||||
String text;
|
||||
if (promptText.contains("PLANNER_TEST_PROMPT")) {
|
||||
agentCalls.add("chat_planner");
|
||||
plannerPromptText = promptText;
|
||||
text = "PLANNER_PLAN";
|
||||
} else if (promptText.contains("EXECUTOR_TEST_PROMPT")) {
|
||||
agentCalls.add("chat_executor");
|
||||
executorPromptText = promptText;
|
||||
text = "EXECUTOR_FINAL_ANSWER";
|
||||
} else if (promptText.contains("VERIFIER_TEST_PROMPT")) {
|
||||
agentCalls.add("chat_verifier");
|
||||
verifierPromptText = promptText;
|
||||
sawVerifierPrompt = true;
|
||||
int index = Math.min(verifierOutputIndex, verifierOutputs.size() - 1);
|
||||
text = verifierOutputs.get(index);
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.skills.ReadSkillTool;
|
||||
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||
import com.superbiz.agent.config.SkillConfig;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class SkillCatalogServiceTest {
|
||||
|
||||
@Test
|
||||
void loadsDiagnosisSkillsFromClasspathRegistry() {
|
||||
SkillRegistry registry = newRegistry();
|
||||
|
||||
assertEquals(1, registry.size());
|
||||
assertTrue(registry.contains("diagnose-mysql-connection-pool"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void readSkillReturnsFullInstructionsFromOfficialTool() {
|
||||
ReadSkillTool tool = new ReadSkillTool(newRegistry());
|
||||
|
||||
String skill = tool.apply(new ReadSkillTool.ReadSkillRequest("diagnose-mysql-connection-pool"), null);
|
||||
|
||||
assertTrue(skill.contains("## Workflow"));
|
||||
assertTrue(skill.contains("query_logs"));
|
||||
assertTrue(skill.contains("Fixed diagnosis case: `mysql-pool-exhausted`"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void readSkillToolReturnsUnknownSkillError() {
|
||||
ReadSkillTool tool = new ReadSkillTool(newRegistry());
|
||||
|
||||
String result = tool.apply(new ReadSkillTool.ReadSkillRequest("missing-skill"), null);
|
||||
|
||||
assertTrue(result.contains("Skill not found: missing-skill"));
|
||||
}
|
||||
|
||||
private SkillRegistry newRegistry() {
|
||||
return new SkillConfig().skillRegistry();
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user