Files
2026-07-06 08:35:54 +08:00

2.9 KiB

Diagnosis Playbook Skills

Problem

The MVP diagnosis Agent already has trace persistence, evidence tools, verifier gates, and fixed eval cases, but scenario-specific diagnosis workflows still live in broad prompts and knowledge-base documents. This makes high-frequency fault diagnosis depend too much on the generic Executor prompt and makes it harder to version, review, and reuse diagnostic procedures.

Proposed Solution

Introduce project-local diagnosis playbook skills using progressive disclosure:

  • Store versionable playbook skills under src/main/resources/skills/.
  • Use Spring AI Alibaba SkillRegistry + SkillsAgentHook so Planner/Executor agents can load full skill instructions only when a matching diagnosis scenario appears.
  • Let the official skills interceptor inject the compact skill catalog into eligible agent prompts.
  • Keep knowledge facts in knowledge_base/; skills define workflow, evidence requirements, stop conditions, and report rules.
  • Keep Verifier isolated from skills. It must continue to validate only existing tool evidence.

Scope

In scope:

  • Payment timeout diagnosis playbook.
  • MySQL connection pool diagnosis playbook.
  • Redis timeout diagnosis playbook.
  • Slow response diagnosis playbook.
  • JVM memory risk diagnosis playbook.
  • AIOps alert diagnosis playbook.
  • Classpath skill registry configuration.
  • Chat and AIOps Planner/Executor SkillsAgentHook wiring.
  • Focused tests for skill loading/catalog behavior and existing diagnosis eval stability.

Out of scope:

  • Replacing lookup_knowledge with implicit advisor retrieval.
  • Replacing the Chat Verifier contract.
  • Persisting a new database field for playbook usage.
  • Creating SubAgents for each playbook.

Context Constraints

  • mvp/architecture/evolution-roadmap.md defines Skill/Playbook as P1 and requires eval-backed, traceable, fallback-capable playbooks.
  • mvp/architecture/harness-quality-gates.md requires evidence tool calls, trace persistence, verifier/rule evaluation, and eval baselines to remain authoritative.
  • knowledge_base/ remains the source for factual definitions and troubleshooting knowledge.
  • mvp/eval/cases/diagnosis-cases.json provides the first fixed diagnosis scenarios and evidence-tool expectations.
  • Spring AI Alibaba 1.1.2.0 provides SkillsAgentHook, ClasspathSkillRegistry, and the official read_skill tool.

Interface Impact

L2 internal interface:

  • Adds an internal SkillRegistry bean backed by classpath skills.
  • Adds SkillsAgentHook to Chat/AIOps Planner and Executor agents.
  • Does not change HTTP API, DTOs, database schema, or external response contracts.

Risks

  • The hook adds the official read_skill tool to eligible agents and may affect tool selection.
  • Skill instructions could conflict with existing prompt constraints if not scoped carefully.
  • Tests that instantiate ChatService manually must inject or tolerate the new skill tool dependency.