Add diagnosis playbook skills

This commit is contained in:
aruo
2026-07-06 08:35:54 +08:00
parent 88e0a6c944
commit 6ccfd33ec5
23 changed files with 1002 additions and 75 deletions
@@ -0,0 +1,2 @@
committed: true
date: 2026-07-05
@@ -0,0 +1,72 @@
# Decisions: diagnosis-playbook-skills
## Discover
- Entry summary: extract high-frequency diagnosis workflows into versionable skills/playbooks and make agents load them progressively.
- Slug: `diagnosis-playbook-skills`.
- Scale: `standard`.
- Capability source: sm-flow built-in protocol for Discover; `grill-with-docs` evidence-driven behavior used by reading project docs and code instead of blocking on user questions.
## Context
- `devflow/index.md` hit related projects: `diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`, `evidence-trace-hardening`, `aiops-alert-scope-control`, `session-dedup-knowledge-map`, `executor-action-memory-relevance`.
- `devflow/glossary/CONTEXT.md` confirms `ReactAgent`, `ToolCall`, `DiagnosisRecord`, and current historical terminology; current architecture documents supersede old `diagnosis_record` as the primary model.
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, fallback-capable playbooks.
- `mvp/architecture/harness-quality-gates.md` requires Prompt contract, Tool boundary, Trace persistence, Verifier/rule evaluation, and eval baselines to remain authoritative.
- `mvp/eval/cases/diagnosis-cases.json` anchors initial playbooks: payment timeout, MySQL pool exhausted, Redis timeout, slow response, JVM memory risk.
## Grill Question Pool
| Dimension | Question | Mode | Resolution |
|---|---|---|---|
| Terminology | Should these artifacts be called Skill or Playbook? | evidence-driven | Use "diagnosis playbook skills": skills are the runtime mechanism, playbooks are the diagnosis workflow content. |
| Boundary | Should skills contain factual knowledge or workflow guidance? | evidence-driven | Skills contain workflow guidance; knowledge facts remain in `knowledge_base/`. |
| Acceptance | What proves the change works? | evidence-driven | Unit tests for catalog/tool behavior plus existing diagnosis eval compile/test stability. |
| Interface impact | Does this alter external API or DB contracts? | evidence-driven | No external API/DB change; L2 internal interface due new tool/service and agent methodTools change. |
No user-interview question is blocking because the user explicitly asked to implement the change and prior conversation already confirmed the intended direction.
## Impact Analysis
GitNexus MCP tools are not exposed in this environment, so required GitNexus impact analysis could not be run. Local substitute analysis:
- `ChatService.createReactAgent`, `buildChatPlannerAgent`, `buildChatExecutorAgent`, and `buildMethodToolsArray` are called by Chat controller paths and covered by `ChatServiceSequentialAgentTest` / smoke tests.
- `AiOpsService.buildPlannerAgent`, `buildExecutorAgent`, and private `buildMethodToolsArray` affect `POST /api/ai_ops` via `ChatController`.
- Risk level: medium. Prompt and tool availability changes may alter agent behavior, but no external API, DTO, DB, or status contract changes.
## Specify / Commit
- Cross-artifact alignment:
- brief/proposal goal -> proposal: aligned.
- proposal scope -> design: aligned.
- design decisions -> specs/tasks: aligned.
- specs observable behavior -> tasks: aligned.
- Interface impact: L2 internal interface.
- Commit status: `.committed` created after file integrity and consistency checks.
## Pre-apply Research
- Reference code:
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
- Tool pattern: Spring AI Alibaba `SkillsAgentHook` contributes the official `read_skill` ToolCallback from hooks.
- Prompt pattern: `SkillsInterceptor` from the hook augments model requests with compact skill metadata.
- Test pattern: service tests instantiate classes manually with `ReflectionTestUtils`; new dependencies must be injectable or optional enough for tests.
- Dependency check: local `1.1.0.0-RC2` jars do not contain `SkillsAgentHook`; `1.1.2.0` jars contain `com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook`, `ReadSkillTool`, `SkillRegistry`, and `ClasspathSkillRegistry`.
## Apply
- Capability source: `openspec-apply-change` guidance was loaded; implementation used local sm-flow/OpenSpec fallback because the work required direct file edits and the OpenSpec CLI was not needed for artifact discovery.
- Implemented `src/main/resources/skills/*/SKILL.md` for six diagnosis playbooks.
- Upgraded Spring AI Alibaba BOMs to `1.1.2.0`.
- Implemented `SkillConfig` with `ClasspathSkillRegistry` loading `classpath:skills`.
- Wired `SkillsAgentHook` into Chat single-agent, Chat Planner/Executor, and AIOps Planner/Executor agents.
- Removed custom prompt-catalog injection from active service paths; the official `SkillsInterceptor` now handles skill catalog injection.
- Kept Chat Verifier prompt unchanged.
- Removed the earlier local fallback `SkillCatalogService` and `ReadSkillTool` source files after switching to official Alibaba skills support.
- Verification command:
- `mvn -q "-Dtest=SkillCatalogServiceTest,ChatServiceSequentialAgentTest,AiOpsServiceTest,DiagnosisTraceEvaluatorTest" test`
- Verification result: passed.
@@ -0,0 +1,82 @@
# Design
## Architecture
```text
src/main/resources/skills/
-> SKILL.md files
-> ClasspathSkillRegistry bean
-> loads classpath skills
-> backs official read_skill
-> PlannerSkillMetadataHook
-> adds planner-only skill metadata messages
-> does not expose read_skill
-> SkillsAgentHook
-> adds official read_skill ToolCallback for Executor / single-agent Chat
-> adds SkillsInterceptor prompt augmentation outside Planner
-> ChatService / AiOpsService
-> Planner receives metadata only
-> Executor and single-agent Chat receive official skill hook
-> Verifier remains isolated
```
## Skill Contract
Each skill folder contains a `SKILL.md` with YAML frontmatter:
```yaml
---
name: diagnose-mysql-connection-pool
description: ...
---
```
The body contains:
- Trigger conditions.
- Required evidence.
- Recommended tool order.
- Query construction hints.
- Stop conditions and low-confidence behavior.
- Report requirements.
- Eval anchor when one exists.
## Prompt Injection
Planner agents receive a project-local `PlannerSkillMetadataHook` message that contains skill names and descriptions only. The message also requires `selected_skill`, `selection_reason`, and an ordered `plan` in the Planner output.
`SkillsAgentHook` provides `SkillsInterceptor`, which injects the official compact skill section containing skill names, descriptions, and loading instructions into model requests for:
- Chat Executor prompt.
- AIOps Executor prompt.
- Single-agent Chat prompt.
Planner prompts are not augmented by `SkillsAgentHook`, so Planner cannot receive the official `read_skill` tool. Verifier prompt is not augmented.
## Tool Exposure
Spring AI Alibaba's official `ReadSkillTool` exposes:
```java
read_skill(skill_name)
```
The tool returns the full `SKILL.md` body for a known skill or a structured error for missing skills.
`read_skill` is supplied by `SkillsAgentHook`, not by local `methodTools`. Planner selects a skill from metadata and writes the selection into `planner_plan`; Executor reads the selected skill before executing scenario-specific evidence collection.
## Trace Behavior
`read_skill` is a guidance tool, not an evidence tool. It does not write `tool_invocation` because the existing eval and verifier treat evidence tools as factual data sources. Actual diagnostic evidence must still come from `lookup_knowledge`, `query_logs`, `query_metrics`, and alert tools.
## Fallback
If a skill is not found or cannot be read:
- The tool returns a structured text error.
- The agent must fall back to generic Executor prompt behavior.
- It must not invent playbook content.
## Alibaba Skills Integration
The implementation uses `spring-ai-alibaba-agent-framework:1.1.2.0`, where `SkillsAgentHook` lives in `com.alibaba.cloud.ai.graph.agent.hook.skills` and `ClasspathSkillRegistry` lives in `com.alibaba.cloud.ai.graph.skills.registry.classpath`.
@@ -0,0 +1,58 @@
# Diagnosis Playbook Skills
## Problem
The MVP diagnosis Agent already has trace persistence, evidence tools, verifier gates, and fixed eval cases, but scenario-specific diagnosis workflows still live in broad prompts and knowledge-base documents. This makes high-frequency fault diagnosis depend too much on the generic Executor prompt and makes it harder to version, review, and reuse diagnostic procedures.
## Proposed Solution
Introduce project-local diagnosis playbook skills using progressive disclosure:
- Store versionable playbook skills under `src/main/resources/skills/`.
- Use Spring AI Alibaba `SkillRegistry` + `SkillsAgentHook` so Planner/Executor agents can load full skill instructions only when a matching diagnosis scenario appears.
- Let the official skills interceptor inject the compact skill catalog into eligible agent prompts.
- Keep knowledge facts in `knowledge_base/`; skills define workflow, evidence requirements, stop conditions, and report rules.
- Keep Verifier isolated from skills. It must continue to validate only existing tool evidence.
## Scope
In scope:
- Payment timeout diagnosis playbook.
- MySQL connection pool diagnosis playbook.
- Redis timeout diagnosis playbook.
- Slow response diagnosis playbook.
- JVM memory risk diagnosis playbook.
- AIOps alert diagnosis playbook.
- Classpath skill registry configuration.
- Chat and AIOps Planner/Executor `SkillsAgentHook` wiring.
- Focused tests for skill loading/catalog behavior and existing diagnosis eval stability.
Out of scope:
- Replacing `lookup_knowledge` with implicit advisor retrieval.
- Replacing the Chat Verifier contract.
- Persisting a new database field for playbook usage.
- Creating SubAgents for each playbook.
## Context Constraints
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, traceable, fallback-capable playbooks.
- `mvp/architecture/harness-quality-gates.md` requires evidence tool calls, trace persistence, verifier/rule evaluation, and eval baselines to remain authoritative.
- `knowledge_base/` remains the source for factual definitions and troubleshooting knowledge.
- `mvp/eval/cases/diagnosis-cases.json` provides the first fixed diagnosis scenarios and evidence-tool expectations.
- Spring AI Alibaba `1.1.2.0` provides `SkillsAgentHook`, `ClasspathSkillRegistry`, and the official `read_skill` tool.
## Interface Impact
L2 internal interface:
- Adds an internal `SkillRegistry` bean backed by classpath `skills`.
- Adds `SkillsAgentHook` to Chat/AIOps Planner and Executor agents.
- Does not change HTTP API, DTOs, database schema, or external response contracts.
## Risks
- The hook adds the official `read_skill` tool to eligible agents and may affect tool selection.
- Skill instructions could conflict with existing prompt constraints if not scoped carefully.
- Tests that instantiate `ChatService` manually must inject or tolerate the new skill tool dependency.
@@ -0,0 +1,56 @@
# diagnosis-playbook-skills Specification
## Purpose
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
## ADDED Requirements
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
#### Scenario: Planner or Executor receives available skill metadata
- **GIVEN** classpath skill folders exist under `skills/`
- **WHEN** Chat or AIOps Planner/Executor agents are built
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
### Requirement: Executor SHALL read full playbook instructions on demand
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
#### Scenario: Executor reads an existing skill
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
- **WHEN** the Executor calls `read_skill` with that name
- **THEN** the tool SHALL return the full skill instructions
- **AND** the result SHALL include the skill name
#### Scenario: Executor requests an unknown skill
- **WHEN** the Executor calls `read_skill` with an unknown name
- **THEN** the tool SHALL return a bounded error message
- **AND** the message SHALL list valid skill names
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
#### Scenario: Executor uses a playbook
- **WHEN** a diagnosis playbook applies to a user issue
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
#### Scenario: Fixed diagnosis case has a matching playbook
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
- **THEN** a matching diagnosis skill SHALL exist
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
@@ -0,0 +1,8 @@
# Tasks
- [x] 1. Add diagnosis skill resources under `src/main/resources/skills/`.
- [x] 2. Upgrade Spring AI Alibaba to a version that provides `SkillsAgentHook`.
- [x] 3. Add a `ClasspathSkillRegistry` bean for classpath skill resources.
- [x] 4. Wire `SkillsAgentHook` into Chat/AIOps single-agent, Planner, and Executor agents while keeping Verifier unchanged.
- [x] 5. Add focused unit tests for registry loading and official `read_skill` behavior.
- [x] 6. Run focused compile/tests and update this task list.
@@ -0,0 +1,56 @@
# diagnosis-playbook-skills Specification
## Purpose
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
## Requirements
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
#### Scenario: Planner or Executor receives available skill metadata
- **GIVEN** classpath skill folders exist under `skills/`
- **WHEN** Chat or AIOps Planner/Executor agents are built
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
### Requirement: Executor SHALL read full playbook instructions on demand
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
#### Scenario: Executor reads an existing skill
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
- **WHEN** the Executor calls `read_skill` with that name
- **THEN** the tool SHALL return the full skill instructions
- **AND** the result SHALL include the skill name
#### Scenario: Executor requests an unknown skill
- **WHEN** the Executor calls `read_skill` with an unknown name
- **THEN** the tool SHALL return a bounded error message
- **AND** the message SHALL list valid skill names
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
#### Scenario: Executor uses a playbook
- **WHEN** a diagnosis playbook applies to a user issue
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
#### Scenario: Fixed diagnosis case has a matching playbook
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
- **THEN** a matching diagnosis skill SHALL exist
- **AND** the skill SHALL state required evidence tools and low-confidence behavior