Compare commits
2
Commits
88e0a6c944
...
3e5c6a159c
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3e5c6a159c | ||
|
|
6ccfd33ec5 |
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: sm-flow
|
name: sm-flow
|
||||||
description: OpenSpec-first 的结构化工程开发协议层 harness。编排 OpenSpec 的完整生命周期,通过阶段、门控、人类对齐和长期记忆,约束 agent 以正确的顺序、条件和标准使用 OpenSpec。用户想把粗略想法、issue、PRD 或已有 research 推进为准确 OpenSpec change,并通过 OpenSpec apply 实现、验证、归档时使用。
|
description: OpenSpec-first 工程流程 harness。仅在用户显式调用 /sm-flow、/sm-flow explore、/sm-flow apply、/sm-flow archive,或明确要求使用 sm-flow 流程时使用;不要根据需求类型自动触发。
|
||||||
---
|
---
|
||||||
|
|
||||||
# SM Flow
|
# SM Flow
|
||||||
@@ -9,6 +9,15 @@ SM Flow 是一个**协议层 harness**——编排 OpenSpec 的完整生命周
|
|||||||
|
|
||||||
sm-flow 会自动维护 `devflow/` 目录作为项目长期记忆。用户不需要手动管理它,sm-flow 会在流程中自动读取和回填。
|
sm-flow 会自动维护 `devflow/` 目录作为项目长期记忆。用户不需要手动管理它,sm-flow 会在流程中自动读取和回填。
|
||||||
|
|
||||||
|
## 触发规则
|
||||||
|
|
||||||
|
只在用户显式调用时使用 sm-flow:
|
||||||
|
|
||||||
|
- 用户输入 `/sm-flow`、`/sm-flow explore`、`/sm-flow apply`、`/sm-flow archive`。
|
||||||
|
- 用户用自然语言明确要求"使用 sm-flow"、"走 sm-flow 流程"或等价表达。
|
||||||
|
|
||||||
|
不要根据需求类型自动触发 sm-flow。即使任务涉及 OpenSpec、跨模块、接口契约、需求澄清或 devflow 归档,只要用户没有显式要求 sm-flow,就按普通工程任务处理。
|
||||||
|
|
||||||
## 四层架构
|
## 四层架构
|
||||||
|
|
||||||
```
|
```
|
||||||
@@ -30,10 +39,10 @@ sm-flow → 编排层(harness):阶段、门控、产物约束、人
|
|||||||
|
|
||||||
1. **OpenSpec 是唯一执行真理源**。apply 阶段必须读取 Committed OpenSpec 文件作为执行依据;对话中的描述不等于产物。Draft OpenSpec 是讨论对象,不是执行许可。
|
1. **OpenSpec 是唯一执行真理源**。apply 阶段必须读取 Committed OpenSpec 文件作为执行依据;对话中的描述不等于产物。Draft OpenSpec 是讨论对象,不是执行许可。
|
||||||
2. **不得跳过 context**。生成 OpenSpec 前,必须先读取相关 devflow 上下文(glossary、ADR、历史项目)。
|
2. **不得跳过 context**。生成 OpenSpec 前,必须先读取相关 devflow 上下文(glossary、ADR、历史项目)。
|
||||||
3. **不得跳过 grill**。即使需求看起来很清楚,至少解决三个高价值澄清或验证问题。
|
3. **不得跳过 grill**。必须按 `references/scales.md` 的当前分档要求完成澄清或验证。
|
||||||
4. **不得跳过 commit**。进入 apply 前,Draft OpenSpec 必须通过 commit 检查成为 Committed OpenSpec。
|
4. **不得跳过 commit**。进入 apply 前,Draft OpenSpec 必须通过 commit 检查成为 Committed OpenSpec。
|
||||||
5. **冲突必须先分类再处理**。OpenSpec 不准(规格遗漏)→ 修正 OpenSpec;代码偏离(实现偏差)→ 修正代码;不确定或涉及设计方向 → 暂停并等待用户确认。
|
5. **冲突必须先分类再处理**。OpenSpec 不准(规格遗漏)→ 修正 OpenSpec;代码偏离(实现偏差)→ 修正代码;不确定或涉及设计方向 → 暂停并等待用户确认。
|
||||||
6. **子 skill 必须显式调用**。每个阶段指定的子 skill 必须显式调用;如果子 skill 不存在,流程失败,不得静默跳过或降级执行。
|
6. **能力来源必须显式声明**。每个阶段先声明使用外部子 skill / OpenSpec CLI / sm-flow 内置协议;外部能力不可用时可使用 `references/fallbacks.md` 的内置协议,但必须标注为 fallback。若外部能力和内置协议都不可用,流程失败。
|
||||||
|
|
||||||
每个阶段的过程约束(question pool、one-at-a-time、cross-artifact 对齐、冲突回写等)和质量约束(可观测产出要求)见 `references/phase-contracts.md` 中对应阶段的退出条件和 checkpoint。
|
每个阶段的过程约束(question pool、one-at-a-time、cross-artifact 对齐、冲突回写等)和质量约束(可观测产出要求)见 `references/phase-contracts.md` 中对应阶段的退出条件和 checkpoint。
|
||||||
|
|
||||||
@@ -48,11 +57,26 @@ sm-flow → 编排层(harness):阶段、门控、产物约束、人
|
|||||||
|
|
||||||
用户也可以用自然语言指定从某个阶段继续,例如"ops-message-support 的 grill 已经做完了,继续"。harness 识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
用户也可以用自然语言指定从某个阶段继续,例如"ops-message-support 的 grill 已经做完了,继续"。harness 识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
||||||
|
|
||||||
|
## 可见 Checkpoint
|
||||||
|
|
||||||
|
内部阶段不是用户 API。对用户汇报进度时,默认只暴露 4 个 checkpoint:
|
||||||
|
|
||||||
|
| Checkpoint | 覆盖内部阶段 | 用户可见含义 |
|
||||||
|
|---|---|---|
|
||||||
|
| Discover | clarify + context + propose + grill | 澄清目标、读取 devflow、形成轻量 proposal、解决关键问题 |
|
||||||
|
| Commit | specify + audit + commit | 补全 OpenSpec、做架构/产物对齐、生成 Committed OpenSpec |
|
||||||
|
| Apply | apply | 基于 Committed OpenSpec 实现和验证 |
|
||||||
|
| Archive | archive | 回填 devflow、汇报验收、询问是否归档 OpenSpec |
|
||||||
|
|
||||||
|
除非用户要求看细节,进度汇报、暂停点和恢复提示应使用 checkpoint 名称,而不是逐个暴露 9 个内部阶段。内部阶段仍按顺序执行,并以 `references/phase-contracts.md` 为准。
|
||||||
|
|
||||||
## 首次加载
|
## 首次加载
|
||||||
|
|
||||||
执行前只读取当前任务需要的 reference 文件:
|
执行前只读取当前任务需要的 reference 文件:
|
||||||
|
|
||||||
- 需要执行阶段时,先读取 `references/phase-contracts.md`;如果当前阶段涉及接口影响分级、分档、启动规则、快速模式或完成标准,再补读 `references/operating-rules.md`。
|
- 需要执行阶段时,先读取 `references/phase-contracts.md`;如果当前阶段涉及接口影响分级、分档、启动规则、快速模式或完成标准,再补读 `references/operating-rules.md`;如果外部 OpenSpec 能力或子 skill 不可用,再补读 `references/fallbacks.md`。
|
||||||
|
- 判断或执行 `micro / standard / complex` 分档时,读取 `references/scales.md`;其它文件不得重复定义分档细节。
|
||||||
|
- 当 checkpoint / gate / fallback / Draft / Committed 等术语含义不清,或需要统一对用户说明时,读取 `references/glossary.md`。
|
||||||
- 创建或更新 PRD、ADR、验收报告、词汇表、复合知识文档时,读取 `references/templates.md`。
|
- 创建或更新 PRD、ADR、验收报告、词汇表、复合知识文档时,读取 `references/templates.md`。
|
||||||
- archive 阶段或需要从 OpenSpec 提取产物时,读取 `references/archive-rules.md`。
|
- archive 阶段或需要从 OpenSpec 提取产物时,读取 `references/archive-rules.md`。
|
||||||
|
|
||||||
|
|||||||
@@ -2,6 +2,49 @@
|
|||||||
|
|
||||||
archive 阶段的目标是把 OpenSpec 产物、实现结果和过程日志转化为持久、可读、可复用的项目记忆。sm-flow 在 clarify → apply 期间只维护 `decisions.md` 作为过程日志,archive 阶段从中提取完整 devflow 档案。
|
archive 阶段的目标是把 OpenSpec 产物、实现结果和过程日志转化为持久、可读、可复用的项目记忆。sm-flow 在 clarify → apply 期间只维护 `decisions.md` 作为过程日志,archive 阶段从中提取完整 devflow 档案。
|
||||||
|
|
||||||
|
## Archive 强制执行顺序
|
||||||
|
|
||||||
|
Archive 阶段必须按以下顺序执行,不得跳过或重排:
|
||||||
|
|
||||||
|
### Step 1: 创建 devflow 档案(必需)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/brief.md`
|
||||||
|
(从 proposal.md 提取:背景、目标、范围、非目标)
|
||||||
|
|
||||||
|
- [ ] 按 `references/scales.md` 的当前分档决定是否创建 `devflow/projects/YYYY-MM-DD-{slug}/evidence.md`
|
||||||
|
(创建时从 decisions.md 提取 evidence-driven 记录)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/decisions.md`
|
||||||
|
(整理为最终版:关键决策、权衡、风险)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/acceptance.md`
|
||||||
|
(记录:静态验证、脚本验证、浏览器/人工验证、未验证)
|
||||||
|
|
||||||
|
### Step 2: 更新索引(必需)
|
||||||
|
|
||||||
|
- [ ] 在 `devflow/index.md` 末尾追加或更新一行:
|
||||||
|
`| YYYY-MM-DD | slug | 领域 | 关键词 | openspec/changes/xxx | {status} |`
|
||||||
|
|
||||||
|
### Step 3: 标记 OpenSpec(必需)
|
||||||
|
|
||||||
|
- [ ] 创建 `openspec/changes/{slug}/.archive-ready` 文件
|
||||||
|
|
||||||
|
### Step 4: 向用户汇报(必需)
|
||||||
|
|
||||||
|
- [ ] 列出创建的 devflow 档案文件路径(验证文件实际存在于磁盘)
|
||||||
|
- [ ] 汇报验证情况(按静态验证、脚本验证、浏览器/人工验证、未验证分类)
|
||||||
|
- [ ] 列出剩余风险或后续事项
|
||||||
|
- [ ] 询问:**是否现在归档 OpenSpec?**
|
||||||
|
|
||||||
|
### Step 5: 用户确认后执行 OpenSpec Archive(可选)
|
||||||
|
|
||||||
|
- [ ] 调用 `openspec-archive-change`
|
||||||
|
- [ ] 记录 archive 结果
|
||||||
|
|
||||||
|
**自检**:在执行 Step 4 前,检查 Step 1-3 是否都完成。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## 目录规则
|
## 目录规则
|
||||||
|
|
||||||
项目档案路径:
|
项目档案路径:
|
||||||
@@ -10,10 +53,10 @@ archive 阶段的目标是把 OpenSpec 产物、实现结果和过程日志转
|
|||||||
devflow/projects/YYYY-MM-DD-{slug}/
|
devflow/projects/YYYY-MM-DD-{slug}/
|
||||||
```
|
```
|
||||||
|
|
||||||
archive 阶段创建以下文件:
|
archive 阶段按 `references/scales.md` 的当前分档创建以下文件:
|
||||||
|
|
||||||
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
||||||
- `evidence.md`:从 decisions.md 中的 evidence-driven 记录提取。
|
- `evidence.md`:从 decisions.md 中的 evidence-driven 记录提取;是否独立创建按 `references/scales.md` 执行。
|
||||||
- `decisions.md`:保持为最终版,整理格式。
|
- `decisions.md`:保持为最终版,整理格式。
|
||||||
- `acceptance.md`:从实现结果和验证结果提取。
|
- `acceptance.md`:从实现结果和验证结果提取。
|
||||||
|
|
||||||
@@ -34,11 +77,7 @@ archive 阶段创建以下文件:
|
|||||||
|
|
||||||
## 产物分档
|
## 产物分档
|
||||||
|
|
||||||
| 分档 | 适用场景 | 必须文件 | 扩展文件 |
|
分档的适用场景和必须文件见 `references/scales.md`。本文件只定义 archive 阶段的创建顺序、提取映射和索引规则。
|
||||||
| --- | --- | --- | --- |
|
|
||||||
| `micro` | 小改动、低风险、需求明确 | `brief.md`、`decisions.md`、`acceptance.md` | 证据少时并入 `brief.md` |
|
|
||||||
| `standard` | 默认模式 | `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md` | 按需 ADR/compound |
|
|
||||||
| `complex` | 高风险、跨模块、需求不清、多人协作 | standard 全部文件 | 按需 `prd.md`、`research.md`、`design.md`、`tasks.md`、`alignment.md` |
|
|
||||||
|
|
||||||
## 提取映射
|
## 提取映射
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# 内置执行协议
|
||||||
|
|
||||||
|
本文件只在外部 OpenSpec CLI 或子 skill 不可用时使用。fallback 不是跳过阶段,而是由 sm-flow 用文件方式完成同等最小产物。每次使用 fallback 都必须写入 `decisions.md` 或 `acceptance.md`,说明能力来源、缺失能力、影响和剩余风险。
|
||||||
|
|
||||||
|
## 通用规则
|
||||||
|
|
||||||
|
- 优先使用外部能力;只有不可用、不可发现或无法在当前环境调用时才使用内置协议。
|
||||||
|
- 不得因为使用 fallback 跳过 context、grill、commit、apply 授权或 archive 确认。
|
||||||
|
- fallback 产物仍写入 `openspec/changes/{slug}/` 和 `devflow/projects/YYYY-MM-DD-{slug}/`。
|
||||||
|
- 如果内置协议也无法满足阶段退出条件,暂停并向用户说明阻塞项。
|
||||||
|
|
||||||
|
## grill 内置协议
|
||||||
|
|
||||||
|
- 建立 question pool,至少覆盖术语、边界、验收;涉及参考实现或项目基础设施时加入技术实现问题。
|
||||||
|
- 将问题标记为 `evidence-driven` 或 `user-interview`。
|
||||||
|
- 先查证 evidence-driven 问题并汇报结论,再逐个询问 user-interview 问题。
|
||||||
|
- 按 `references/scales.md` 的当前分档满足 grill 要求。
|
||||||
|
- 将 question pool、证据结论、用户原话和确认状态写入 `decisions.md`;影响实现的结论回写 `proposal.md`。
|
||||||
|
|
||||||
|
## openspec 提案内置协议
|
||||||
|
|
||||||
|
- 在 `openspec/changes/{slug}/` 创建或更新:
|
||||||
|
- `proposal.md`:问题、方案、范围、非目标、上下文约束、风险。
|
||||||
|
- 设计产物:实现设计、接口影响、关键决策、架构风险;形式按 `references/scales.md` 的当前分档要求执行。
|
||||||
|
- `specs/*/spec.md` 或等价 functional spec:描述用户可观察行为和验收场景。
|
||||||
|
- `tasks.md`:按可执行切片拆分任务,并给每项写可验证验收标准。
|
||||||
|
- 运行 cross-artifact 对齐检查:proposal → 设计产物 → specs → tasks。
|
||||||
|
- 如果发现 gap,先修正 OpenSpec,再进入 commit。
|
||||||
|
|
||||||
|
## audit 内置协议
|
||||||
|
|
||||||
|
- 用 5 句话以内说明模块链路、数据所有权、跨模块依赖、架构风险和是否需要回写 OpenSpec。
|
||||||
|
- 如果风险影响实现,修正设计产物或 `tasks.md`。
|
||||||
|
- 将结论写入 `decisions.md`。
|
||||||
|
|
||||||
|
## openspec apply 内置协议
|
||||||
|
|
||||||
|
- 只依据 Committed OpenSpec 的 specs/tasks 实现;devflow 只作上下文参考。
|
||||||
|
- 开始前检查 `.committed` 文件;缺失则返回 commit。
|
||||||
|
- 如触发 pre-apply checkpoint,先阅读参考实现、grep 项目基础设施模式,并把技术栈清单写入 `decisions.md`。
|
||||||
|
- 按 tasks 的纵向切片实现、验证并更新任务状态。
|
||||||
|
- 发现冲突时按三类处理:OpenSpec 不准则修 OpenSpec,代码偏离则修代码,不确定则暂停等用户确认。
|
||||||
|
|
||||||
|
## openspec archive 内置协议
|
||||||
|
|
||||||
|
- 不删除或移动 OpenSpec change;只标记归档准备状态。
|
||||||
|
- 完成 devflow 回填、更新 `devflow/index.md`、创建 `.archive-ready`。
|
||||||
|
- 向用户汇报已创建文件、验证分类、剩余风险,并询问是否需要真实 OpenSpec archive。
|
||||||
|
- 如果外部 archive 能力仍不可用,在 `acceptance.md` 标记 `accepted-unarchived`。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# 术语表
|
||||||
|
|
||||||
|
本文件统一 sm-flow 协议中的核心词。优先使用这些词,避免同一概念多种说法。
|
||||||
|
|
||||||
|
| 术语 | 含义 | 使用边界 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| sm-flow | 协议层 harness | 编排 OpenSpec 生命周期,不替代 OpenSpec |
|
||||||
|
| OpenSpec | 当前变更的执行真理源 | apply 只能依据 Committed OpenSpec |
|
||||||
|
| devflow | 长期记忆和上下文层 | 提供术语、历史决策、验收记录,不直接指挥实现 |
|
||||||
|
| checkpoint | 用户可见检查点 | 默认只暴露 Discover / Commit / Apply / Archive |
|
||||||
|
| gate | 硬门控 | 不满足就不能进入下一关键动作,如 commit gate |
|
||||||
|
| Draft OpenSpec | 讨论和审计对象 | propose/specify 期间产生,不能直接 apply |
|
||||||
|
| Committed OpenSpec | 已通过 commit gate 的 OpenSpec | apply 的唯一执行依据 |
|
||||||
|
| fallback | 内置执行协议 | 外部 OpenSpec CLI 或子 skill 不可用时使用,必须标注 |
|
||||||
|
| decisions.md | 过程日志 | clarify 到 apply 期间记录问题、证据、决策、冲突和回写 |
|
||||||
|
| .committed | commit gate 标记文件 | 存在才可进入合规 apply |
|
||||||
|
| .archive-ready | archive 准备标记文件 | 表示 devflow 已回填,等待用户确认是否 archive |
|
||||||
|
| Discover | 用户可见 checkpoint | 覆盖 clarify + context + propose + grill |
|
||||||
|
| Commit | 用户可见 checkpoint | 覆盖 specify + audit + commit |
|
||||||
|
| Apply | 用户可见 checkpoint | 覆盖 apply |
|
||||||
|
| Archive | 用户可见 checkpoint | 覆盖 archive |
|
||||||
@@ -27,13 +27,14 @@
|
|||||||
- `/sm-flow apply [change]`:只执行,检查 commit gate → apply。
|
- `/sm-flow apply [change]`:只执行,检查 commit gate → apply。
|
||||||
- `/sm-flow explore`:带上下文的探索模式,不走标准阶段链。
|
- `/sm-flow explore`:带上下文的探索模式,不走标准阶段链。
|
||||||
- `/sm-flow archive [change]`:收尾,回填 devflow + 归档确认。
|
- `/sm-flow archive [change]`:收尾,回填 devflow + 归档确认。
|
||||||
|
- 明确要求"使用 sm-flow"或"走 sm-flow 流程":按显式调用处理。
|
||||||
- 自然语言指定阶段继续:识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
- 自然语言指定阶段继续:识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
||||||
2. 判断启动模式:
|
2. 判断启动模式:
|
||||||
- 完整模式:用户提供粗略想法或初始 PRD。
|
- 完整模式:用户提供粗略想法或初始 PRD。
|
||||||
- Research 模式:用户已有 research,需要转成或修正 OpenSpec。
|
- Research 模式:用户已有 research,需要转成或修正 OpenSpec。
|
||||||
- PRD 文件模式:用户提供已有 PRD 路径。
|
- PRD 文件模式:用户提供已有 PRD 路径。
|
||||||
- 恢复模式:用户希望从某个阶段继续(补做最小前置检查)。
|
- 恢复模式:用户希望从某个阶段继续(补做最小前置检查)。
|
||||||
- 快速模式:小改动,合并 gate(见下文)。
|
- 快速模式:小改动,合并 gate;具体分档规则见 `references/scales.md`。
|
||||||
3. 如果缺少 `devflow/`,初始化:
|
3. 如果缺少 `devflow/`,初始化:
|
||||||
- `devflow/projects/`
|
- `devflow/projects/`
|
||||||
- `devflow/glossary/CONTEXT.md`
|
- `devflow/glossary/CONTEXT.md`
|
||||||
@@ -42,7 +43,25 @@
|
|||||||
5. 检查 OpenSpec 和子 skill 是否可用:
|
5. 检查 OpenSpec 和子 skill 是否可用:
|
||||||
- OpenSpec 能力:`openspec-propose`、`openspec-apply-change`、`openspec-archive-change`。
|
- OpenSpec 能力:`openspec-propose`、`openspec-apply-change`、`openspec-archive-change`。
|
||||||
- 辅助能力:`to-prd`、`grill-with-docs`、`diagnose`、`tdd`、`zoom-out`。
|
- 辅助能力:`to-prd`、`grill-with-docs`、`diagnose`、`tdd`、`zoom-out`。
|
||||||
6. 如果 OpenSpec 不可用,不要直接绕过;使用内置执行协议(见 `references/fallbacks.md`),并在 apply 前向用户说明。
|
6. 如果 OpenSpec 或子 skill 不可用,不要静默跳过;使用内置执行协议(见 `references/fallbacks.md`),并在当前 checkpoint 说明 fallback 来源、影响和剩余风险。
|
||||||
|
|
||||||
|
## 进度汇报
|
||||||
|
|
||||||
|
用户可见进度默认折叠为 4 个 checkpoint:
|
||||||
|
|
||||||
|
| Checkpoint | 内部阶段 |
|
||||||
|
| --- | --- |
|
||||||
|
| Discover | clarify + context + propose + grill |
|
||||||
|
| Commit | specify + audit + commit |
|
||||||
|
| Apply | apply |
|
||||||
|
| Archive | archive |
|
||||||
|
|
||||||
|
汇报规则:
|
||||||
|
|
||||||
|
- 面向用户时优先使用 checkpoint 名称,不逐个汇报 9 个内部阶段。
|
||||||
|
- 内部阶段只在 checkpoint 摘要中作为证据列出,例如"Discover 已完成:读取了 devflow、生成 proposal、解决 2 个问题"。
|
||||||
|
- 只有发生阻塞、冲突、fallback、用户要求继续某个内部阶段,或需要解释恢复位置时,才暴露内部阶段名。
|
||||||
|
- 当前分档的汇报压缩规则见 `references/scales.md`;无论分档如何,都不要把内部阶段名当作用户操作入口。
|
||||||
|
|
||||||
## 项目标识规则
|
## 项目标识规则
|
||||||
|
|
||||||
@@ -63,7 +82,7 @@ Devflow 是 sm-flow 自动维护的项目长期记忆层,不复制 OpenSpec
|
|||||||
**最终档案**(archive 阶段从 decisions.md + OpenSpec 产物提取):
|
**最终档案**(archive 阶段从 decisions.md + OpenSpec 产物提取):
|
||||||
|
|
||||||
- `brief.md`:背景、目标、范围、非目标、分档、关联 OpenSpec change。
|
- `brief.md`:背景、目标、范围、非目标、分档、关联 OpenSpec change。
|
||||||
- `evidence.md`:代码/文档证据、历史决策、evidence-driven 结论和汇报状态。
|
- `evidence.md`:代码/文档证据、历史决策、evidence-driven 结论和汇报状态;分档要求见 `references/scales.md` 和 `references/archive-rules.md`。
|
||||||
- `acceptance.md`:实现结果、验证命令、未验证项、归档状态、后续事项。
|
- `acceptance.md`:实现结果、验证命令、未验证项、归档状态、后续事项。
|
||||||
|
|
||||||
**按需产物**(archive 阶段按需创建):
|
**按需产物**(archive 阶段按需创建):
|
||||||
@@ -75,28 +94,17 @@ Devflow 是 sm-flow 自动维护的项目长期记忆层,不复制 OpenSpec
|
|||||||
- `alignment.md` / `clarifications.md`:仅在 gap 或澄清很多时使用。
|
- `alignment.md` / `clarifications.md`:仅在 gap 或澄清很多时使用。
|
||||||
- `adr/*.md` 和 `compound/*.md`:仅在满足 ADR / compound knowledge 规则时使用。
|
- `adr/*.md` 和 `compound/*.md`:仅在满足 ADR / compound knowledge 规则时使用。
|
||||||
|
|
||||||
**规模分档**:
|
**规模分档**:`micro / standard / complex` 的唯一规则源是 `references/scales.md`。
|
||||||
|
|
||||||
- `micro`:小且低风险,gate 合并(见快速模式),最终档案同 standard。
|
|
||||||
- `standard`:默认模式。
|
|
||||||
- `complex`:高风险、跨模块、需求不清或多人协作时,在 standard 基础上按需增加扩展产物。
|
|
||||||
|
|
||||||
## 快速模式
|
## 快速模式
|
||||||
|
|
||||||
快速模式适用于小而低风险的变更。它合并 gate 而不仅仅是压缩产物:
|
快速模式适用于 `references/scales.md` 定义的 micro 变更。它合并 gate 而不仅仅是压缩产物;具体覆盖规则见 `references/scales.md`。
|
||||||
|
|
||||||
```
|
|
||||||
standard 流程:clarify → context → propose checkpoint → grill → specify → audit checkpoint → commit
|
|
||||||
micro 流程:clarify+context 合并 checkpoint → propose+specify 合并 checkpoint → grill(最少 1 个问题) → commit(简化检查)
|
|
||||||
```
|
|
||||||
|
|
||||||
micro 的定位:**gate 变少但保留最关键的**(grill 最小澄清 + commit gate)。
|
|
||||||
|
|
||||||
无论什么模式,以下内容必须保留:
|
无论什么模式,以下内容必须保留:
|
||||||
|
|
||||||
- context 最小上下文收集:至少检查 glossary 和相关 ADR。
|
- context 最小上下文收集:至少检查 glossary 和相关 ADR。
|
||||||
- grill 最小澄清:至少一个术语问题、一个边界问题、一个验收问题;evidence-driven 结论仍需汇报。
|
- grill 最小澄清:按 `references/scales.md` 当前分档要求执行;evidence-driven 结论仍需汇报。
|
||||||
- commit gate:确认没有未解决用户问题、接口影响已记录、OpenSpec tasks/specs 可执行。
|
- commit gate:确认没有未解决用户问题、接口影响已记录、OpenSpec tasks/specs 可执行;完整性检查按 `references/scales.md` 当前分档要求执行。
|
||||||
- apply 仍由 OpenSpec tasks/specs 驱动执行。
|
- apply 仍由 OpenSpec tasks/specs 驱动执行。
|
||||||
- archive 轻量回填:记录验收结果、OpenSpec 链接和归档状态。
|
- archive 轻量回填:记录验收结果、OpenSpec 链接和归档状态。
|
||||||
|
|
||||||
@@ -104,8 +112,9 @@ micro 的定位:**gate 变少但保留最关键的**(grill 最小澄清 + co
|
|||||||
|
|
||||||
只有同时满足以下条件,流程才算完成:
|
只有同时满足以下条件,流程才算完成:
|
||||||
|
|
||||||
- OpenSpec proposal/design/specs/tasks 已生成或更新到可执行状态。
|
- 用户可见的 Discover、Commit、Apply、Archive checkpoint 已完成,或未完成项已明确标记为暂停/不适用。
|
||||||
|
- OpenSpec proposal、设计产物、specs、tasks 已按当前分档生成或更新到可执行状态。
|
||||||
- 实现或规划工作已完成,且执行依据来自 OpenSpec。
|
- 实现或规划工作已完成,且执行依据来自 OpenSpec。
|
||||||
- 已运行验证,或已记录未运行验证的原因。
|
- 已运行验证,或已记录未运行验证的原因。
|
||||||
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 brief.md、evidence.md、decisions.md、acceptance.md。
|
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 `references/scales.md` 和 `references/archive-rules.md` 要求的当前分档档案。
|
||||||
- 用户知道剩余风险与下一步,并已被询问是否归档 OpenSpec change。
|
- 用户知道剩余风险与下一步,并已被询问是否归档 OpenSpec change。
|
||||||
|
|||||||
@@ -4,6 +4,18 @@
|
|||||||
|
|
||||||
执行顺序:clarify → context → propose → grill → specify → audit → commit → apply → archive。
|
执行顺序:clarify → context → propose → grill → specify → audit → commit → apply → archive。
|
||||||
|
|
||||||
|
## 目录
|
||||||
|
|
||||||
|
- clarify — 入口澄清
|
||||||
|
- context — 上下文收集
|
||||||
|
- propose — 轻量 propose
|
||||||
|
- grill — 人类对齐澄清
|
||||||
|
- specify — 细化 + 对齐
|
||||||
|
- audit — 架构审计
|
||||||
|
- commit — Commit OpenSpec
|
||||||
|
- apply — OpenSpec 执行
|
||||||
|
- archive — 回填 + 归档
|
||||||
|
|
||||||
## clarify — 入口澄清
|
## clarify — 入口澄清
|
||||||
|
|
||||||
**进入条件**:用户提供粗略想法、初始 PRD、已有 research、issue,或要求启动 SM Flow。
|
**进入条件**:用户提供粗略想法、初始 PRD、已有 research、issue,或要求启动 SM Flow。
|
||||||
@@ -13,7 +25,7 @@
|
|||||||
- 如果用户已有 research,先识别它是否已经包含用户价值、技术方案、验收标准和任务拆分。
|
- 如果用户已有 research,先识别它是否已经包含用户价值、技术方案、验收标准和任务拆分。
|
||||||
- 如果输入过于模糊,最多追加三轮聚焦问题。
|
- 如果输入过于模糊,最多追加三轮聚焦问题。
|
||||||
- 当答案会改变 OpenSpec proposal/specs/tasks 时,优先一次只问一个问题。
|
- 当答案会改变 OpenSpec proposal/specs/tasks 时,优先一次只问一个问题。
|
||||||
- 如果需要判断 `micro / standard / complex` 分档,补读 `references/operating-rules.md`。
|
- 如果需要判断 `micro / standard / complex` 分档,补读 `references/scales.md`。
|
||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- 问题可以用 1-2 句话说清楚。
|
- 问题可以用 1-2 句话说清楚。
|
||||||
@@ -36,7 +48,7 @@
|
|||||||
- 读取 `devflow/glossary/CONTEXT.md`,提取相关术语和业务规则。
|
- 读取 `devflow/glossary/CONTEXT.md`,提取相关术语和业务规则。
|
||||||
- 搜索 `devflow/projects/` 中相关 PRD、design、tasks、acceptance 和 ADR。
|
- 搜索 `devflow/projects/` 中相关 PRD、design、tasks、acceptance 和 ADR。
|
||||||
- 搜索 `devflow/compound/` 中可复用 learning、trick、decision、explore。
|
- 搜索 `devflow/compound/` 中可复用 learning、trick、decision、explore。
|
||||||
- 记录哪些上下文会影响 OpenSpec proposal/design/specs/tasks。
|
- 记录哪些上下文会影响 OpenSpec proposal、设计产物、specs 或 tasks。
|
||||||
- 如果发现旧根目录 `CONTEXT.md` 与 `devflow/glossary/CONTEXT.md` 冲突,暂停并向用户汇报。
|
- 如果发现旧根目录 `CONTEXT.md` 与 `devflow/glossary/CONTEXT.md` 冲突,暂停并向用户汇报。
|
||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
@@ -70,18 +82,22 @@
|
|||||||
|
|
||||||
**Human checkpoint**:
|
**Human checkpoint**:
|
||||||
- 向用户简要说明 proposal 范围、关键假设、主要风险、devflow 上下文如何影响方案。
|
- 向用户简要说明 proposal 范围、关键假设、主要风险、devflow 上下文如何影响方案。
|
||||||
- 询问是否继续进入 grill 澄清阶段;用户明确要求"全自动执行"时可跳过等待。
|
- 作为 Discover checkpoint 的中间状态汇报;询问是否继续完成 Discover 的人类澄清部分。用户明确要求"全自动执行"时可跳过等待。
|
||||||
|
|
||||||
## grill — 人类对齐澄清
|
## grill — 人类对齐澄清
|
||||||
|
|
||||||
**进入条件**:propose 已有轻量 proposal.md。
|
**进入条件**:propose 已有轻量 proposal.md。
|
||||||
|
|
||||||
**显式子 skill**:`grill-with-docs`。进入本阶段必须调用 `.agents/skills/grill-with-docs/SKILL.md`。
|
**能力来源**:优先使用 `grill-with-docs`;不可用时使用 `references/fallbacks.md#grill-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
**动作**:
|
**动作**:
|
||||||
- 优先使用 `grill-with-docs`。
|
- 优先使用 `grill-with-docs`。
|
||||||
- 进入 grill 时先建立一个 question pool,并记录到 `decisions.md`:
|
- 进入 grill 时先建立一个 question pool,并记录到 `decisions.md`:
|
||||||
- 默认至少覆盖术语、边界、验收三个维度。
|
- 默认至少覆盖术语、边界、验收三个维度。
|
||||||
|
- **技术实现维度**(新增):当 proposal 提到参考实现、或涉及项目现有基础设施时,增加技术澄清问题:
|
||||||
|
- 参考实现的具体文件路径是什么?
|
||||||
|
- 项目现有的 [请求结构/MQ/缓存/加密/工具类] 标准是什么?
|
||||||
|
- 有哪些技术点需要先调研或新建?
|
||||||
- 如果变更涉及多模块、接口、权限、下游消费者、响应结构或生命周期规则,先把这些维度补进问题池。
|
- 如果变更涉及多模块、接口、权限、下游消费者、响应结构或生命周期规则,先把这些维度补进问题池。
|
||||||
- 逐项标记每个问题的模式:
|
- 逐项标记每个问题的模式:
|
||||||
- `evidence-driven`:问题能通过代码、文档、测试、OpenSpec 或既有 ADR 证明;代理先查证,再向用户汇报证据、结论和是否需要确认。
|
- `evidence-driven`:问题能通过代码、文档、测试、OpenSpec 或既有 ADR 证明;代理先查证,再向用户汇报证据、结论和是否需要确认。
|
||||||
@@ -97,7 +113,7 @@
|
|||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- question pool 已建立并覆盖当前 change 所需维度。
|
- question pool 已建立并覆盖当前 change 所需维度。
|
||||||
- 至少解决三个高价值澄清或验证问题,并记录每个问题属于 `evidence-driven` 还是 `user-interview`。
|
- 已满足 `references/scales.md` 中当前分档的 grill 要求。每个问题都必须记录属于 `evidence-driven` 还是 `user-interview`。
|
||||||
- 所有 evidence-driven 结论已向用户汇报。
|
- 所有 evidence-driven 结论已向用户汇报。
|
||||||
- 所有 user-interview 决策已获得用户确认。
|
- 所有 user-interview 决策已获得用户确认。
|
||||||
- 没有未解决或代理代确认的 user-interview 问题。
|
- 没有未解决或代理代确认的 user-interview 问题。
|
||||||
@@ -113,20 +129,20 @@
|
|||||||
|
|
||||||
**Human checkpoint**:
|
**Human checkpoint**:
|
||||||
- 汇报已解决和未解决的问题、proposal 变更、术语和 ADR 更新。
|
- 汇报已解决和未解决的问题、proposal 变更、术语和 ADR 更新。
|
||||||
- 询问是否继续进入 specify 细化阶段。
|
- 汇报 Discover checkpoint 完成情况,并询问是否继续进入 Commit checkpoint。
|
||||||
|
|
||||||
## specify — 细化 + 对齐
|
## specify — 细化 + 对齐
|
||||||
|
|
||||||
**进入条件**:grill 已退出,需求已通过澄清稳定下来。
|
**进入条件**:grill 已退出,需求已通过澄清稳定下来。
|
||||||
|
|
||||||
**显式子 skill**:`openspec-propose`(基于已稳定的 proposal 补全完整 OpenSpec);`to-prd`(按需生成 PRD)。进入本阶段必须先声明调用方式。
|
**能力来源**:优先使用 `openspec-propose`(基于已稳定的 proposal 补全完整 OpenSpec);按需使用 `to-prd`。进入本阶段必须先声明调用方式;外部能力不可用时使用 `references/fallbacks.md#openspec-提案-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
**动作**:
|
**动作**:
|
||||||
- 基于已稳定的 proposal.md 补全 design.md、specs/、tasks.md:
|
- 基于已稳定的 proposal.md 补全设计产物、specs/、tasks.md:
|
||||||
- 优先调用 `openspec-propose`,输入中明确说明"proposal.md 已存在,本次只需补全 design/specs/tasks"。
|
- 优先调用 `openspec-propose`,输入中明确说明"proposal.md 已存在,本次只需按当前分档补全设计产物/specs/tasks"。
|
||||||
- 如果不可用,执行 `references/fallbacks.md#openspec-提案-降级`。
|
- 如果不可用,执行 `references/fallbacks.md#openspec-提案-内置协议`。
|
||||||
- 如果没有结构化 PRD,按需按 `to-prd` 协议生成 `brief.md`;复杂需求、对外协作或用户明确要求时再生成 `prd.md`。
|
- 如果没有结构化 PRD,按需按 `to-prd` 协议生成 `brief.md`;复杂需求、对外协作或用户明确要求时再生成 `prd.md`。
|
||||||
- `micro` 模式默认不创建独立 PRD,除非用户要求或需求复杂度升级。
|
- 独立 PRD 是否需要按 `references/scales.md` 的当前分档和用户要求判断。
|
||||||
- 用 grill 阶段的 decisions.md 记录增强 OpenSpec 产物:确保 design/specs/tasks 反映所有已确认的决策。
|
- 用 grill 阶段的 decisions.md 记录增强 OpenSpec 产物:确保 design/specs/tasks 反映所有已确认的决策。
|
||||||
- **显式 cross-artifact 对齐检查**——在 checkpoint 中输出对齐检查表:
|
- **显式 cross-artifact 对齐检查**——在 checkpoint 中输出对齐检查表:
|
||||||
- `brief/prd` 中的目标、范围、非目标和验收预期 → `proposal` 是否覆盖。
|
- `brief/prd` 中的目标、范围、非目标和验收预期 → `proposal` 是否覆盖。
|
||||||
@@ -143,14 +159,14 @@
|
|||||||
- 如果发现不一致,优先修正 OpenSpec,而不是只修改 devflow 文档。
|
- 如果发现不一致,优先修正 OpenSpec,而不是只修改 devflow 文档。
|
||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- `design.md`、`specs/`、`tasks.md` 存在且与 proposal 对齐。
|
- OpenSpec 细化产物存在且与 proposal 对齐;产物形态按 `references/scales.md` 的当前分档要求执行。
|
||||||
- `brief.md` 已覆盖背景、目标、范围和非目标;复杂需求存在独立 `prd.md` 或用户明确不需要 PRD。
|
- `brief.md` 已覆盖背景、目标、范围和非目标;复杂需求存在独立 `prd.md` 或用户明确不需要 PRD。
|
||||||
- cross-artifact 对齐检查表已生成(4 行,每行标记已对齐/存在 gap),没有未处理 gap。
|
- cross-artifact 对齐检查表已生成(4 行,每行标记已对齐/存在 gap),没有未处理 gap。
|
||||||
- 涉及接口变更时,已记录接口影响等级和产物要求;不确定项已标记。
|
- 涉及接口变更时,已记录接口影响等级和产物要求;不确定项已标记。
|
||||||
- 所有已知冲突已修正或等待用户决策。
|
- 所有已知冲突已修正或等待用户决策。
|
||||||
|
|
||||||
**输出**:
|
**输出**:
|
||||||
- 完整的 Draft OpenSpec:proposal.md + design.md + specs/ + tasks.md。
|
- Draft OpenSpec:按 `references/scales.md` 的当前分档要求生成 proposal、设计、specs 和 tasks。
|
||||||
- `brief.md`,以及按需创建的 `prd.md`。
|
- `brief.md`,以及按需创建的 `prd.md`。
|
||||||
- cross-artifact 对齐检查表(写入 checkpoint 或 decisions.md)。
|
- cross-artifact 对齐检查表(写入 checkpoint 或 decisions.md)。
|
||||||
- 必要的 OpenSpec 修正。
|
- 必要的 OpenSpec 修正。
|
||||||
@@ -159,7 +175,7 @@
|
|||||||
|
|
||||||
**进入条件**:specify 已退出,完整 OpenSpec 产物已存在。
|
**进入条件**:specify 已退出,完整 OpenSpec 产物已存在。
|
||||||
|
|
||||||
**显式子 skill**:`zoom-out`。进入本阶段必须调用 `.agents/skills/zoom-out/SKILL.md`。
|
**能力来源**:优先使用 `zoom-out`;不可用时使用 `references/fallbacks.md#audit-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
**动作**:
|
**动作**:
|
||||||
- 画出输入 → 处理 → 输出的模块链路。
|
- 画出输入 → 处理 → 输出的模块链路。
|
||||||
@@ -171,20 +187,20 @@
|
|||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- 架构风险已被接受,或流程返回 grill/specify 修正 OpenSpec。
|
- 架构风险已被接受,或流程返回 grill/specify 修正 OpenSpec。
|
||||||
- OpenSpec design/tasks 已反映会影响实现的架构审计结论。
|
- OpenSpec 设计产物/tasks 已反映会影响实现的架构审计结论。
|
||||||
|
|
||||||
**输出**:
|
**输出**:
|
||||||
- 架构审计记录,写入 `decisions.md`;复杂架构审计可拆出 `design.md`。
|
- 架构审计记录,写入 `decisions.md`;复杂架构审计可拆出 `design.md`。
|
||||||
- 必要的 OpenSpec design/tasks 修正。
|
- 必要的 OpenSpec 设计产物/tasks 修正。
|
||||||
|
|
||||||
**Human checkpoint**:
|
**Human checkpoint**:
|
||||||
- 用不超过五句话向用户说明架构风险、OpenSpec 修正点和实现计划。
|
- 用不超过五句话向用户说明架构风险、OpenSpec 修正点和实现计划。
|
||||||
- 询问是否进入 commit。
|
- 作为 Commit checkpoint 的中间状态汇报;询问是否继续完成 commit gate。
|
||||||
|
|
||||||
## commit — Commit OpenSpec
|
## commit — Commit OpenSpec
|
||||||
|
|
||||||
**进入条件**:
|
**进入条件**:
|
||||||
- grill 已解决术语、边界、验收三个维度的高价值问题。
|
- grill 已满足 `references/scales.md` 中当前分档要求。
|
||||||
- 所有 `user-interview` 问题都已获得用户显式确认。
|
- 所有 `user-interview` 问题都已获得用户显式确认。
|
||||||
- audit 已经完成,或快速模式下已记录跳过原因;快速模式定义见 `references/operating-rules.md#快速模式`。
|
- audit 已经完成,或快速模式下已记录跳过原因;快速模式定义见 `references/operating-rules.md#快速模式`。
|
||||||
- Draft OpenSpec 已回写所有会影响实现的澄清、接口影响和架构审计结论。
|
- Draft OpenSpec 已回写所有会影响实现的澄清、接口影响和架构审计结论。
|
||||||
@@ -194,7 +210,7 @@
|
|||||||
- 检查 design 是否记录上下文约束、关键技术决策、架构风险和接口影响。
|
- 检查 design 是否记录上下文约束、关键技术决策、架构风险和接口影响。
|
||||||
- 检查 specs 是否表达外部可观察行为,并覆盖验收口径。
|
- 检查 specs 是否表达外部可观察行为,并覆盖验收口径。
|
||||||
- 检查 tasks 是否是可执行的纵向切片,而不是泛泛描述。
|
- 检查 tasks 是否是可执行的纵向切片,而不是泛泛描述。
|
||||||
- 复核 cross-artifact 对齐:`brief/prd → proposal → design → specs → tasks` 是否闭环,没有把字段、范围项、验收行为或实现切片丢在上游产物里。
|
- 复核 cross-artifact 对齐:`brief/prd → proposal → 设计产物 → specs → tasks` 是否闭环,没有把字段、范围项、验收行为或实现切片丢在上游产物里。
|
||||||
- 检查 `decisions.md` 中所有影响实现的发现,是否已回写到 proposal、design、specs 或 tasks。
|
- 检查 `decisions.md` 中所有影响实现的发现,是否已回写到 proposal、design、specs 或 tasks。
|
||||||
- 接口影响分级定义见 `references/operating-rules.md#接口影响分级`。
|
- 接口影响分级定义见 `references/operating-rules.md#接口影响分级`。
|
||||||
- 检查接口影响是否已按 L1/L4 判级;L3/L4 是否有独立接口文档或等价独立章节。
|
- 检查接口影响是否已按 L1/L4 判级;L3/L4 是否有独立接口文档或等价独立章节。
|
||||||
@@ -203,7 +219,16 @@
|
|||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- Draft OpenSpec 已达到可执行状态,并记录为 Committed OpenSpec。
|
- Draft OpenSpec 已达到可执行状态,并记录为 Committed OpenSpec。
|
||||||
- apply 所需的 proposal、design、specs 和 tasks 均存在且一致;commit checkpoint 必须验证文件实际存在于磁盘,如果任一文件不存在,commit 失败,返回 specify 补写。
|
- **文件完整性检查**(按 `references/scales.md` 的当前分档要求执行):
|
||||||
|
- [ ] proposal 存在,且足以说明问题、建议方案、范围和非目标。
|
||||||
|
- [ ] 设计产物存在,形式符合当前分档要求。
|
||||||
|
- [ ] specs 存在,且表达用户可观察行为。
|
||||||
|
- [ ] tasks 存在,且任务可执行、验收标准可验证。
|
||||||
|
- **一致性检查**(必须通过):
|
||||||
|
- [ ] proposal 中的核心概念在设计产物中有对应设计
|
||||||
|
- [ ] 设计产物中的关键决策在 tasks 中有对应实现任务
|
||||||
|
- [ ] tasks 的验收标准可验证(不是"正确实现""完成功能"这类模糊描述)
|
||||||
|
- **标记文件**:检查通过后,创建 `openspec/changes/{slug}/.committed` 文件标记为 Committed OpenSpec
|
||||||
- 所有 preflight 风险已消除或明确记录为已接受。
|
- 所有 preflight 风险已消除或明确记录为已接受。
|
||||||
|
|
||||||
**输出**:
|
**输出**:
|
||||||
@@ -212,36 +237,83 @@
|
|||||||
|
|
||||||
**Human checkpoint**:
|
**Human checkpoint**:
|
||||||
- 用不超过五句话说明 Committed OpenSpec 的范围、接口影响、剩余风险和执行计划。
|
- 用不超过五句话说明 Committed OpenSpec 的范围、接口影响、剩余风险和执行计划。
|
||||||
- 询问是否进入 apply;除非用户在启动时明确要求"全自动执行",必须等待用户明确说出进入 apply、开始实现、执行修改或等价授权。
|
- 汇报 Commit checkpoint 完成情况,并询问是否进入 Apply checkpoint;除非用户在启动时明确要求"全自动执行",必须等待用户明确说出进入 apply、开始实现、执行修改或等价授权。
|
||||||
- 不得把 grill 的单个决策确认当作本 checkpoint 的授权。
|
- 不得把 grill 的单个决策确认当作本 checkpoint 的授权。
|
||||||
|
|
||||||
## apply — OpenSpec 执行
|
## apply — OpenSpec 执行
|
||||||
|
|
||||||
**进入条件**:
|
**进入条件**:
|
||||||
- `openspec/changes/{slug}/` 中 proposal/design/specs/tasks 已通过 commit,成为 Committed OpenSpec。
|
- `openspec/changes/{slug}/` 中 proposal、设计产物、specs、tasks 已通过 commit,成为 Committed OpenSpec。
|
||||||
|
- **前置门控检查**(硬约束):
|
||||||
|
- 检查 `openspec/changes/{slug}/.committed` 文件是否存在
|
||||||
|
- 如不存在,执行以下流程:
|
||||||
|
1. 汇报:Draft OpenSpec 未通过 commit 检查
|
||||||
|
2. 列出缺失的 checkpoint 项(文件完整性、一致性检查)
|
||||||
|
3. 询问用户:是否补做 commit 检查;如用户要求不补做,则中止 apply 或标记为 `emergency-bypass`,且本次流程不得视为合规 sm-flow apply
|
||||||
- commit 后已获得用户明确的 apply 授权,除非用户在启动时要求"全自动执行"。
|
- commit 后已获得用户明确的 apply 授权,除非用户在启动时要求"全自动执行"。
|
||||||
- devflow 与 OpenSpec 没有未解决冲突。
|
- devflow 与 OpenSpec 没有未解决冲突。
|
||||||
- 没有未解决的 user-interview 问题、未判级接口影响、未汇报 evidence-driven 结论或未接受架构风险。
|
- 没有未解决的 user-interview 问题、未判级接口影响、未汇报 evidence-driven 结论或未接受架构风险。
|
||||||
|
|
||||||
**显式子 skill**:`openspec-apply-change`;遇到 bug/不确定行为时显式调用 `diagnose`;需要测试驱动时显式调用 `tdd`。进入本阶段必须调用指定子 skill,不得静默跳过。
|
**能力来源**:优先使用 `openspec-apply-change`;不可用时使用 `references/fallbacks.md#openspec-apply-内置协议`,并在 `decisions.md` 标注 fallback。遇到 bug/不确定行为时优先使用 `diagnose`;需要测试驱动时优先使用 `tdd`。不可用时执行对应最小协议并记录原因,不得静默跳过。
|
||||||
|
|
||||||
**动作**:
|
**动作**:
|
||||||
|
|
||||||
|
### Pre-apply Checkpoint
|
||||||
|
|
||||||
|
**触发条件**:当 OpenSpec 涉及以下任一情况时必须执行
|
||||||
|
- design 或 tasks 中提到"参考 XXX 实现"
|
||||||
|
- 需要调用项目现有基础设施(MQ/统一请求结构/工具类等)
|
||||||
|
- 技术栈不熟悉或第一次在该项目实现类似功能
|
||||||
|
|
||||||
|
**执行步骤**:
|
||||||
|
1. **阅读所有参考实现**
|
||||||
|
- 从 OpenSpec design 或 tasks 中定位参考实现文件
|
||||||
|
- 如果路径不明确,通过 Grep 搜索关键类名或模式
|
||||||
|
- 理解关键逻辑,提取可复用代码片段和模式
|
||||||
|
|
||||||
|
2. **Grep 关键技术栈**
|
||||||
|
- 请求/响应结构模式(如 `RequestMsg`、`ResponseMsg`、DTO 规范)
|
||||||
|
- 消息队列模式(如 `@KafkaListener`、`@YkMsg`、发送模板)
|
||||||
|
- 统一工具类(如 `XxxUtil`、`XxxHelper`、加密/验签工具)
|
||||||
|
- 异常处理和日志记录标准
|
||||||
|
|
||||||
|
3. **形成技术栈清单并写入 decisions.md**
|
||||||
|
- 项目使用的请求/响应结构标准
|
||||||
|
- MQ 消息定义和发送标准
|
||||||
|
- Consumer 标准位置和写法
|
||||||
|
- 加密/验签/工具类的标准用法
|
||||||
|
- 识别需要新建的工具类或基础设施
|
||||||
|
|
||||||
|
**输出要求**:
|
||||||
|
- 技术栈清单已写入 `decisions.md` 的 "Pre-apply Research" 章节。
|
||||||
|
- 已列出所有参考实现的文件路径。
|
||||||
|
- 已识别需要新建的工具类/基础设施。
|
||||||
|
|
||||||
|
**按风险执行**:执行深度按 `references/scales.md` 的当前分档和实现风险决定;退出判断以清单是否足以指导实现为准。
|
||||||
|
|
||||||
|
### 实现过程
|
||||||
|
|
||||||
- 优先调用 `openspec-apply-change`。
|
- 优先调用 `openspec-apply-change`。
|
||||||
- 执行依据是 OpenSpec specs/tasks;devflow 只能作为上下文参考。
|
- 执行依据是 OpenSpec specs/tasks;devflow 只能作为上下文参考。
|
||||||
- 按 OpenSpec tasks 的纵向切片实现。
|
- 按 OpenSpec tasks 的纵向切片实现。
|
||||||
|
- **分步实现**:建议按 Controller → Service → MQ/异步组件 → Consumer/下游 顺序,每完成一层验证后再继续。
|
||||||
- 进入实现前先汇报本阶段的 capability 来源、当前 task 进度和本轮要推进的切片;否则 apply 不算真正开始。
|
- 进入实现前先汇报本阶段的 capability 来源、当前 task 进度和本轮要推进的切片;否则 apply 不算真正开始。
|
||||||
|
- **首模块完成后对齐检查**:完成第一个接口/模块后,对比 OpenSpec design/tasks,标记"已完成/TODO";核心功能(加密/验签/核心业务逻辑)不允许空实现或纯 TODO 注释。
|
||||||
- 当用户质疑、用户要求修改、代码检查、测试失败或运行行为与 OpenSpec 冲突时,做三类判断:
|
- 当用户质疑、用户要求修改、代码检查、测试失败或运行行为与 OpenSpec 冲突时,做三类判断:
|
||||||
- OpenSpec 不准(规格遗漏、边界未覆盖、验收口径缺失)→ 暂停 apply,修正 OpenSpec 后重新提交。
|
- OpenSpec 不准(规格遗漏、边界未覆盖、验收口径缺失)→ 暂停 apply,修正 OpenSpec 后重新提交。
|
||||||
- 代码偏离(实现没按 OpenSpec 做)→ 修正代码,不改 OpenSpec。
|
- 代码偏离(实现没按 OpenSpec 做)→ 修正代码,不改 OpenSpec。
|
||||||
- 不确定根因、涉及设计方向、用户改变目标或范围 → 暂停并等待用户确认。
|
- 不确定根因、涉及设计方向、用户改变目标或范围 → 暂停并等待用户确认。
|
||||||
- 判断结果、证据、用户确认和 OpenSpec 回写状态必须记录到 `decisions.md`。
|
- 判断结果、证据、用户确认和 OpenSpec 回写状态必须记录到 `decisions.md`。
|
||||||
|
- **快速失败**:连续返工 ≥ 2 次时,暂停并重新执行 pre-apply checkpoint 或向用户汇报。
|
||||||
- 当用户要求、行为复杂或回归风险高时使用 TDD。
|
- 当用户要求、行为复杂或回归风险高时使用 TDD。
|
||||||
- 当测试失败、行为意外或原因不确定时使用 diagnose。
|
- 当测试失败、行为意外或原因不确定时使用 diagnose。
|
||||||
- 如果 diagnose 发现根因是 OpenSpec 不准确,先修正 OpenSpec,再继续 apply。
|
- 如果 diagnose 发现根因是 OpenSpec 不准确,先修正 OpenSpec,再继续 apply。
|
||||||
- 修改文件前遵守仓库指令,例如 `AGENTS.md`。
|
- 修改文件前遵守仓库指令,例如 `AGENTS.md`。
|
||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
|
- 已完成 pre-apply checkpoint(如触发条件满足),技术栈清单已写入 `decisions.md`。
|
||||||
- OpenSpec tasks 已完成,或剩余 tasks 已明确记录。
|
- OpenSpec tasks 已完成,或剩余 tasks 已明确记录。
|
||||||
|
- 核心功能已实现或明确标注"待联调",无纯 TODO 占位。
|
||||||
- 所有实现期冲突已分类并处理;没有未确认的规格遗漏、设计冲突或用户变更。
|
- 所有实现期冲突已分类并处理;没有未确认的规格遗漏、设计冲突或用户变更。
|
||||||
- 已运行验证,或记录了未验证原因。
|
- 已运行验证,或记录了未验证原因。
|
||||||
- 已列出已知限制。
|
- 已列出已知限制。
|
||||||
@@ -255,13 +327,13 @@
|
|||||||
|
|
||||||
**进入条件**:实现或规划工作已经达到可交接状态。
|
**进入条件**:实现或规划工作已经达到可交接状态。
|
||||||
|
|
||||||
**显式子 skill**:`openspec-archive-change` 在用户确认 archive 后调用;archive 回填由 `sm-flow` 执行。必须调用子 skill,不得静默跳过。
|
**能力来源**:`openspec-archive-change` 在用户确认 archive 后优先调用;不可用时使用 `references/fallbacks.md#openspec-archive-内置协议`,并在 `acceptance.md` 标注 fallback。archive 回填由 `sm-flow` 执行。
|
||||||
|
|
||||||
**动作**:
|
**动作**:
|
||||||
- 遵循 `references/archive-rules.md`。
|
- 遵循 `references/archive-rules.md`。
|
||||||
- 从 `decisions.md`(过程日志)+ OpenSpec 产物提炼完整 devflow 档案:
|
- 从 `decisions.md`(过程日志)+ OpenSpec 产物提炼完整 devflow 档案:
|
||||||
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
||||||
- `evidence.md`:从 decisions.md 中的 evidence-driven 记录提取。
|
- `evidence.md`:按 `references/scales.md` 和 `references/archive-rules.md` 的当前分档要求处理。
|
||||||
- `decisions.md`:保持为最终版,整理格式。
|
- `decisions.md`:保持为最终版,整理格式。
|
||||||
- `acceptance.md`:从实现结果和验证结果提取。
|
- `acceptance.md`:从实现结果和验证结果提取。
|
||||||
- 只在复杂场景按需拆出 PRD/research/design/tasks/alignment。
|
- 只在复杂场景按需拆出 PRD/research/design/tasks/alignment。
|
||||||
@@ -271,7 +343,7 @@
|
|||||||
- 询问用户是否要 archive OpenSpec change;不要默认执行归档。
|
- 询问用户是否要 archive OpenSpec change;不要默认执行归档。
|
||||||
|
|
||||||
**退出条件**:
|
**退出条件**:
|
||||||
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 brief.md、evidence.md、decisions.md、acceptance.md;archive checkpoint 必须列出所有已创建的文件路径,验证文件实际存在于磁盘。
|
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 `references/scales.md` 和 `references/archive-rules.md` 要求的当前分档档案;archive checkpoint 必须列出所有已创建的文件路径,验证文件实际存在于磁盘。
|
||||||
- `devflow/index.md` 已包含或更新本项目条目。
|
- `devflow/index.md` 已包含或更新本项目条目。
|
||||||
- 用户已被询问是否 archive OpenSpec change。
|
- 用户已被询问是否 archive OpenSpec change。
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# 分档规则
|
||||||
|
|
||||||
|
本文件是 `micro / standard / complex` 的唯一规则源。其它文件只引用本文件,不重复定义分档细节。
|
||||||
|
|
||||||
|
## standard 基准
|
||||||
|
|
||||||
|
standard 是默认分档,适用于普通功能、明确但有一定实现范围的变更。
|
||||||
|
|
||||||
|
- 用户可见 checkpoint:Discover → Commit → Apply → Archive。
|
||||||
|
- OpenSpec 产物:`proposal.md`、独立 `design.md`、`specs/`、`tasks.md`。
|
||||||
|
- grill:解决术语、边界、验收三个维度的高价值问题。
|
||||||
|
- commit gate:检查 proposal、design、specs、tasks 的完整性和一致性。
|
||||||
|
- devflow 档案:`brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`。
|
||||||
|
|
||||||
|
## micro 覆盖
|
||||||
|
|
||||||
|
micro 适用于小改动、低风险、需求明确的变更。micro 是 standard 的减法,不是跳过流程。
|
||||||
|
|
||||||
|
- checkpoint 可合并:Discover + Commit 可在无阻塞时合并汇报。
|
||||||
|
- micro 内部流程压缩为:clarify+context 合并 checkpoint → 轻量 propose → grill → specify+commit 合并 checkpoint。
|
||||||
|
- context 保留最小收集:至少检查 glossary 和相关 ADR。
|
||||||
|
- grill 保留最小澄清:至少解决一个高价值问题,并记录术语、边界、验收三类是否明确;不明确项必须补问或标记风险。
|
||||||
|
- OpenSpec 仍需要 `proposal.md`、`specs/`、`tasks.md`。
|
||||||
|
- `design.md` 可不独立创建;允许在 `proposal.md` 或 `tasks.md` 中写等价设计小节。
|
||||||
|
- `specs/` 和 `tasks.md` 可轻量,但必须表达可观察行为和可执行任务。
|
||||||
|
- commit gate 仍必须通过,并创建 `.committed`。
|
||||||
|
- devflow 档案至少包含 `brief.md`、`decisions.md`、`acceptance.md`;证据少时可并入 `brief.md` 或 `decisions.md`。
|
||||||
|
- apply 仍只能依据 Committed OpenSpec。
|
||||||
|
- archive 仍要轻量回填 devflow,并询问是否归档 OpenSpec。
|
||||||
|
|
||||||
|
micro 不适用于接口影响不清、跨团队消费者、迁移/回滚、复杂状态机、长期架构决策或需求边界不清的变更;遇到这些情况应升级为 standard 或 complex。
|
||||||
|
|
||||||
|
## complex 增量
|
||||||
|
|
||||||
|
complex 适用于高风险、跨模块、需求不清、多人协作或长期架构影响明显的变更。complex 是 standard 的加法。
|
||||||
|
|
||||||
|
- 需要更完整的 Discover:增加需求澄清、证据查证、范围确认和风险接受。
|
||||||
|
- checkpoint 内可补充关键内部阶段结果,但不要把内部阶段名当作用户操作入口。
|
||||||
|
- 按需创建 `prd.md`、`research.md`、`alignment.md`、接口文档、ADR 或 compound knowledge。
|
||||||
|
- 接口影响、迁移、灰度、回滚、兼容性和消费者边界必须显式记录。
|
||||||
|
- audit 需要覆盖模块链路、数据所有权、生命周期、耦合风险和 ADR 冲突。
|
||||||
|
- archive 在 standard 档案基础上按需提炼长期 design、research、tasks、ADR 和 compound knowledge。
|
||||||
@@ -124,7 +124,7 @@
|
|||||||
|
|
||||||
- 触发来源:用户质疑 / 用户变更 / 代码发现 / 测试失败 / 运行行为
|
- 触发来源:用户质疑 / 用户变更 / 代码发现 / 测试失败 / 运行行为
|
||||||
- 冲突对象:proposal / design / specs / tasks / ADR / 代码行为
|
- 冲突对象:proposal / design / specs / tasks / ADR / 代码行为
|
||||||
- 分类:实现偏差 / 规格遗漏 / 设计冲突 / 用户变更
|
- 分类:OpenSpec 不准 / 代码偏离 / 不确定
|
||||||
|
|
||||||
## 证据
|
## 证据
|
||||||
|
|
||||||
@@ -136,7 +136,7 @@
|
|||||||
|
|
||||||
- 决策:
|
- 决策:
|
||||||
- 是否需要用户确认:是 / 否
|
- 是否需要用户确认:是 / 否
|
||||||
- OpenSpec 回写:不需要 / 已回写 / 待回写
|
- OpenSpec 回写:不需要 / 已回写 / 待回写 / 等待用户确认
|
||||||
- 代码处理:
|
- 代码处理:
|
||||||
- 验证方式:
|
- 验证方式:
|
||||||
```
|
```
|
||||||
@@ -353,8 +353,8 @@ specify 阶段的 checkpoint 必须包含此检查表。每项标记"已对齐"
|
|||||||
| 上游 → 下游 | 检查内容 | 状态 |
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| brief/prd → proposal | 目标、范围、非目标、验收预期是否进入 proposal | 已对齐 / 存在 gap |
|
| brief/prd → proposal | 目标、范围、非目标、验收预期是否进入 proposal | 已对齐 / 存在 gap |
|
||||||
| proposal → design | 范围、约束、关键承诺是否进入 design | 已对齐 / 存在 gap |
|
| proposal → 设计产物 | 范围、约束、关键承诺是否进入 design.md 或等价设计小节 | 已对齐 / 存在 gap |
|
||||||
| design → specs/tasks | 影响实现的约束、接口影响、架构结论是否进入 specs 或 tasks | 已对齐 / 存在 gap |
|
| 设计产物 → specs/tasks | 影响实现的约束、接口影响、架构结论是否进入 specs 或 tasks | 已对齐 / 存在 gap |
|
||||||
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 / 存在 gap |
|
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 / 存在 gap |
|
||||||
|
|
||||||
### Gap 详情(如有)
|
### Gap 详情(如有)
|
||||||
|
|||||||
@@ -0,0 +1,107 @@
|
|||||||
|
---
|
||||||
|
name: sm-flow
|
||||||
|
description: OpenSpec-first 工程流程 harness。仅在用户显式调用 /sm-flow、/sm-flow explore、/sm-flow apply、/sm-flow archive,或明确要求使用 sm-flow 流程时使用;不要根据需求类型自动触发。
|
||||||
|
---
|
||||||
|
|
||||||
|
# SM Flow
|
||||||
|
|
||||||
|
SM Flow 是一个**协议层 harness**——编排 OpenSpec 的完整生命周期。它通过阶段、门控、人类对齐和长期记忆,约束 agent 以正确的顺序、条件和标准使用 OpenSpec。
|
||||||
|
|
||||||
|
sm-flow 会自动维护 `devflow/` 目录作为项目长期记忆。用户不需要手动管理它,sm-flow 会在流程中自动读取和回填。
|
||||||
|
|
||||||
|
## 触发规则
|
||||||
|
|
||||||
|
只在用户显式调用时使用 sm-flow:
|
||||||
|
|
||||||
|
- 用户输入 `/sm-flow`、`/sm-flow explore`、`/sm-flow apply`、`/sm-flow archive`。
|
||||||
|
- 用户用自然语言明确要求"使用 sm-flow"、"走 sm-flow 流程"或等价表达。
|
||||||
|
|
||||||
|
不要根据需求类型自动触发 sm-flow。即使任务涉及 OpenSpec、跨模块、接口契约、需求澄清或 devflow 归档,只要用户没有显式要求 sm-flow,就按普通工程任务处理。
|
||||||
|
|
||||||
|
## 四层架构
|
||||||
|
|
||||||
|
```
|
||||||
|
sm-flow → 编排层(harness):阶段、门控、产物约束、人类对齐
|
||||||
|
OpenSpec → 执行引擎:propose/apply/archive 的能力提供方
|
||||||
|
devflow/ → 记忆层:为编排层提供上下文,接收执行结果的回填
|
||||||
|
code → 实现结果:apply 的产出
|
||||||
|
```
|
||||||
|
|
||||||
|
- OpenSpec 是唯一执行真理源:apply 阶段只能基于 OpenSpec 执行,不能绕过 OpenSpec 直接写代码。
|
||||||
|
- devflow 是上下文真理源:术语、历史决策、验收记录来自 devflow,用于增强 OpenSpec,不替代 OpenSpec。
|
||||||
|
- 如果 devflow 和 OpenSpec 冲突,先汇报冲突、让用户确认、修正 OpenSpec,再继续执行。
|
||||||
|
- propose 阶段产出的 OpenSpec 默认为 **Draft OpenSpec**:它是澄清和审计对象,不是 apply 的执行许可。
|
||||||
|
- 只有通过 commit 检查后的 OpenSpec 才是 **Committed OpenSpec**;apply 只能执行 Committed OpenSpec。
|
||||||
|
|
||||||
|
## 核心规则
|
||||||
|
|
||||||
|
以下 6 条是硬约束,违反即流程失败。其余约束按阶段定义在 `references/phase-contracts.md`。
|
||||||
|
|
||||||
|
1. **OpenSpec 是唯一执行真理源**。apply 阶段必须读取 Committed OpenSpec 文件作为执行依据;对话中的描述不等于产物。Draft OpenSpec 是讨论对象,不是执行许可。
|
||||||
|
2. **不得跳过 context**。生成 OpenSpec 前,必须先读取相关 devflow 上下文(glossary、ADR、历史项目)。
|
||||||
|
3. **不得跳过 grill**。必须按 `references/scales.md` 的当前分档要求完成澄清或验证。
|
||||||
|
4. **不得跳过 commit**。进入 apply 前,Draft OpenSpec 必须通过 commit 检查成为 Committed OpenSpec。
|
||||||
|
5. **冲突必须先分类再处理**。OpenSpec 不准(规格遗漏)→ 修正 OpenSpec;代码偏离(实现偏差)→ 修正代码;不确定或涉及设计方向 → 暂停并等待用户确认。
|
||||||
|
6. **能力来源必须显式声明**。每个阶段先声明使用外部子 skill / OpenSpec CLI / sm-flow 内置协议;外部能力不可用时可使用 `references/fallbacks.md` 的内置协议,但必须标注为 fallback。若外部能力和内置协议都不可用,流程失败。
|
||||||
|
|
||||||
|
每个阶段的过程约束(question pool、one-at-a-time、cross-artifact 对齐、冲突回写等)和质量约束(可观测产出要求)见 `references/phase-contracts.md` 中对应阶段的退出条件和 checkpoint。
|
||||||
|
|
||||||
|
## 用户命令
|
||||||
|
|
||||||
|
| 命令 | 用户意图 | harness 内部行为 |
|
||||||
|
|---|---|---|
|
||||||
|
| `/sm-flow` | 完整流程 | clarify → context → propose → grill → specify → audit → commit → apply → archive |
|
||||||
|
| `/sm-flow explore` | 先想想 | 带上下文的探索模式 |
|
||||||
|
| `/sm-flow apply` | 只执行 | 检查 commit gate → apply |
|
||||||
|
| `/sm-flow archive` | 收尾 | 回填 devflow + 归档确认 |
|
||||||
|
|
||||||
|
用户也可以用自然语言指定从某个阶段继续,例如"ops-message-support 的 grill 已经做完了,继续"。harness 识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
||||||
|
|
||||||
|
## 可见 Checkpoint
|
||||||
|
|
||||||
|
内部阶段不是用户 API。对用户汇报进度时,默认只暴露 4 个 checkpoint:
|
||||||
|
|
||||||
|
| Checkpoint | 覆盖内部阶段 | 用户可见含义 |
|
||||||
|
|---|---|---|
|
||||||
|
| Discover | clarify + context + propose + grill | 澄清目标、读取 devflow、形成轻量 proposal、解决关键问题 |
|
||||||
|
| Commit | specify + audit + commit | 补全 OpenSpec、做架构/产物对齐、生成 Committed OpenSpec |
|
||||||
|
| Apply | apply | 基于 Committed OpenSpec 实现和验证 |
|
||||||
|
| Archive | archive | 回填 devflow、汇报验收、询问是否归档 OpenSpec |
|
||||||
|
|
||||||
|
除非用户要求看细节,进度汇报、暂停点和恢复提示应使用 checkpoint 名称,而不是逐个暴露 9 个内部阶段。内部阶段仍按顺序执行,并以 `references/phase-contracts.md` 为准。
|
||||||
|
|
||||||
|
## 首次加载
|
||||||
|
|
||||||
|
执行前只读取当前任务需要的 reference 文件:
|
||||||
|
|
||||||
|
- 需要执行阶段时,先读取 `references/phase-contracts.md`;如果当前阶段涉及接口影响分级、分档、启动规则、快速模式或完成标准,再补读 `references/operating-rules.md`;如果外部 OpenSpec 能力或子 skill 不可用,再补读 `references/fallbacks.md`。
|
||||||
|
- 判断或执行 `micro / standard / complex` 分档时,读取 `references/scales.md`;其它文件不得重复定义分档细节。
|
||||||
|
- 当 checkpoint / gate / fallback / Draft / Committed 等术语含义不清,或需要统一对用户说明时,读取 `references/glossary.md`。
|
||||||
|
- 创建或更新 PRD、ADR、验收报告、词汇表、复合知识文档时,读取 `references/templates.md`。
|
||||||
|
- archive 阶段或需要从 OpenSpec 提取产物时,读取 `references/archive-rules.md`。
|
||||||
|
|
||||||
|
## 内部阶段
|
||||||
|
|
||||||
|
9 个内部阶段,按执行顺序:
|
||||||
|
|
||||||
|
1. clarify — 入口澄清:接收初始需求,澄清到可生成轻量 proposal。
|
||||||
|
2. context — 上下文收集:读取 devflow 的 glossary、ADR、历史项目、compound knowledge。
|
||||||
|
3. propose — 轻量 propose:只生成 proposal.md,不调用 openspec-propose。
|
||||||
|
4. grill — 人类对齐澄清:evidence-driven 查证 + user-interview one-at-a-time,回写 proposal。
|
||||||
|
5. specify — 细化 + 对齐:基于已稳定的 proposal 补全 design/specs/tasks,做 cross-artifact 对齐。
|
||||||
|
6. audit — 架构审计:审计结果如果影响实现,回写 OpenSpec design/tasks。
|
||||||
|
7. commit — Commit OpenSpec:检查 Draft OpenSpec 是否达到可执行状态,提交为 Committed OpenSpec。
|
||||||
|
8. apply — OpenSpec 执行:基于 Committed OpenSpec 实现代码。
|
||||||
|
9. archive — 回填 + 归档:从 OpenSpec 产物和 decisions.md 提炼长期档案,询问是否归档。
|
||||||
|
|
||||||
|
每个阶段的进入条件、动作、输出和退出标准见 `references/phase-contracts.md`。
|
||||||
|
|
||||||
|
关键阶段的完成判断也以 `references/phase-contracts.md` 为准;如果缺少显式 checkpoint 或能力来源声明,该阶段不得视为已完成。
|
||||||
|
|
||||||
|
## 快速模式
|
||||||
|
|
||||||
|
快速模式的具体约束见 `references/operating-rules.md`。
|
||||||
|
|
||||||
|
## 完成标准
|
||||||
|
|
||||||
|
流程完成标准见 `references/operating-rules.md`。
|
||||||
@@ -0,0 +1,167 @@
|
|||||||
|
# 归档规则
|
||||||
|
|
||||||
|
archive 阶段的目标是把 OpenSpec 产物、实现结果和过程日志转化为持久、可读、可复用的项目记忆。sm-flow 在 clarify → apply 期间只维护 `decisions.md` 作为过程日志,archive 阶段从中提取完整 devflow 档案。
|
||||||
|
|
||||||
|
## Archive 强制执行顺序
|
||||||
|
|
||||||
|
Archive 阶段必须按以下顺序执行,不得跳过或重排:
|
||||||
|
|
||||||
|
### Step 1: 创建 devflow 档案(必需)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/brief.md`
|
||||||
|
(从 proposal.md 提取:背景、目标、范围、非目标)
|
||||||
|
|
||||||
|
- [ ] 按 `references/scales.md` 的当前分档决定是否创建 `devflow/projects/YYYY-MM-DD-{slug}/evidence.md`
|
||||||
|
(创建时从 decisions.md 提取 evidence-driven 记录)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/decisions.md`
|
||||||
|
(整理为最终版:关键决策、权衡、风险)
|
||||||
|
|
||||||
|
- [ ] 创建 `devflow/projects/YYYY-MM-DD-{slug}/acceptance.md`
|
||||||
|
(记录:静态验证、脚本验证、浏览器/人工验证、未验证)
|
||||||
|
|
||||||
|
### Step 2: 更新索引(必需)
|
||||||
|
|
||||||
|
- [ ] 在 `devflow/index.md` 末尾追加或更新一行:
|
||||||
|
`| YYYY-MM-DD | slug | 领域 | 关键词 | openspec/changes/xxx | {status} |`
|
||||||
|
|
||||||
|
### Step 3: 标记 OpenSpec(必需)
|
||||||
|
|
||||||
|
- [ ] 创建 `openspec/changes/{slug}/.archive-ready` 文件
|
||||||
|
|
||||||
|
### Step 4: 向用户汇报(必需)
|
||||||
|
|
||||||
|
- [ ] 列出创建的 devflow 档案文件路径(验证文件实际存在于磁盘)
|
||||||
|
- [ ] 汇报验证情况(按静态验证、脚本验证、浏览器/人工验证、未验证分类)
|
||||||
|
- [ ] 列出剩余风险或后续事项
|
||||||
|
- [ ] 询问:**是否现在归档 OpenSpec?**
|
||||||
|
|
||||||
|
### Step 5: 用户确认后执行 OpenSpec Archive(可选)
|
||||||
|
|
||||||
|
- [ ] 调用 `openspec-archive-change`
|
||||||
|
- [ ] 记录 archive 结果
|
||||||
|
|
||||||
|
**自检**:在执行 Step 4 前,检查 Step 1-3 是否都完成。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 目录规则
|
||||||
|
|
||||||
|
项目档案路径:
|
||||||
|
|
||||||
|
```text
|
||||||
|
devflow/projects/YYYY-MM-DD-{slug}/
|
||||||
|
```
|
||||||
|
|
||||||
|
archive 阶段按 `references/scales.md` 的当前分档创建以下文件:
|
||||||
|
|
||||||
|
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
||||||
|
- `evidence.md`:从 decisions.md 中的 evidence-driven 记录提取;是否独立创建按 `references/scales.md` 执行。
|
||||||
|
- `decisions.md`:保持为最终版,整理格式。
|
||||||
|
- `acceptance.md`:从实现结果和验证结果提取。
|
||||||
|
|
||||||
|
同时维护仓库级索引:
|
||||||
|
|
||||||
|
- `devflow/index.md`
|
||||||
|
|
||||||
|
按需创建以下扩展文件:
|
||||||
|
|
||||||
|
- `prd.md`
|
||||||
|
- `research.md`
|
||||||
|
- `design.md`
|
||||||
|
- `tasks.md`
|
||||||
|
- `alignment.md`
|
||||||
|
- `adr/*.md`
|
||||||
|
|
||||||
|
不要逐字复制完整 OpenSpec 文件,也不要重复 OpenSpec 的 proposal/design/tasks。应提炼 OpenSpec 如何指导执行:背景、证据、用户决策、任务状态、假设、验证结果、风险,以及执行中对 OpenSpec 的修正。
|
||||||
|
|
||||||
|
## 产物分档
|
||||||
|
|
||||||
|
分档的适用场景和必须文件见 `references/scales.md`。本文件只定义 archive 阶段的创建顺序、提取映射和索引规则。
|
||||||
|
|
||||||
|
## 提取映射
|
||||||
|
|
||||||
|
| 来源 | 提取内容 | 写入位置 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `decisions.md`(过程日志) | question pool、evidence-driven 汇报状态、user-interview 确认状态、关键取舍 | `decisions.md`(整理格式为最终版) |
|
||||||
|
| `decisions.md`(过程日志) | evidence-driven 结论、代码/文档证据 | `evidence.md` |
|
||||||
|
| `proposal.md` | 为什么做、做什么、范围、非目标 | `brief.md` |
|
||||||
|
| `design.md` | 技术方案、关键决策、风险;只提炼长期有用内容 | `evidence.md` / 按需 `design.md` |
|
||||||
|
| `specs/**/*.md` | requirement 标题和 scenario 意图 | `brief.md` 或 `acceptance.md` 的验收追踪 |
|
||||||
|
| `tasks.md` | checkbox 状态、剩余工作、执行切片 | `acceptance.md`;复杂项目可拆 `tasks.md` |
|
||||||
|
| 测试/构建输出 | 验证命令、结果、验证类型 | `acceptance.md` |
|
||||||
|
| diagnose 记录 | 根因、修复、回归验证 | `acceptance.md` |
|
||||||
|
| 词汇表更新 | 术语和业务规则 | `devflow/glossary/CONTEXT.md` |
|
||||||
|
| 可复用经验 | 持久工程知识 | `devflow/compound/YYYY-MM-DD-{type}-{slug}.md` |
|
||||||
|
| 项目索引 | 日期、slug、领域、关键词、关联 OpenSpec、状态 | `devflow/index.md` |
|
||||||
|
|
||||||
|
## 索引维护规则
|
||||||
|
|
||||||
|
`devflow/index.md` 是 context 阶段的默认入口,archive 阶段回填时必须维护。
|
||||||
|
|
||||||
|
最小字段:
|
||||||
|
|
||||||
|
| 日期 | slug | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||||
|
| --- | --- | --- | --- | --- | --- |
|
||||||
|
|
||||||
|
规则:
|
||||||
|
|
||||||
|
- 每个 `devflow/projects/YYYY-MM-DD-{slug}/` 默认对应一行索引。
|
||||||
|
- archive 阶段新建或更新项目档案时,必须新增或更新对应行。
|
||||||
|
- 如果项目仍在进行,状态写 `active`;已验收但未 archive 写 `accepted-unarchived`;已 archive 写 `archived`;暂停写 `paused`。
|
||||||
|
- 关键词只放能帮助 context 阶段定位的术语,不复制 brief 内容。
|
||||||
|
- 如果无法准确判断领域或状态,写 `unknown`,并在 `acceptance.md` 记录待补。
|
||||||
|
|
||||||
|
## 验收记录规则
|
||||||
|
|
||||||
|
必须真实记录验证情况,并按类型分类:
|
||||||
|
|
||||||
|
- **静态验证**:语法检查、grep/rg 检查、结构检查、类型检查等不运行完整功能的验证。
|
||||||
|
- **脚本验证**:生成脚本、测试命令、构建命令、自动化检查等可重复命令。
|
||||||
|
- **浏览器/人工验证**:需要用户或代理在界面中点击、观察、确认的行为验证。
|
||||||
|
- **未验证**:未运行的验证必须记录原因、风险和建议补验步骤。
|
||||||
|
|
||||||
|
记录要求:
|
||||||
|
|
||||||
|
- 如果验证通过,记录命令/步骤和覆盖范围。
|
||||||
|
- 如果验证失败,记录失败摘要和是否阻塞验收。
|
||||||
|
- 如果需要人工验证,列出明确步骤,不要用"手动测试一下"这种模糊描述。
|
||||||
|
|
||||||
|
## ADR 规则
|
||||||
|
|
||||||
|
同时满足以下条件时创建 ADR:
|
||||||
|
|
||||||
|
1. 决策难以逆转。
|
||||||
|
2. 缺少上下文会让未来维护者困惑。
|
||||||
|
3. 决策来自真实权衡,而不是简单偏好。
|
||||||
|
|
||||||
|
项目内 ADR 存放于:
|
||||||
|
|
||||||
|
```text
|
||||||
|
devflow/projects/YYYY-MM-DD-{slug}/adr/
|
||||||
|
```
|
||||||
|
|
||||||
|
跨项目可复用决策或经验存放于:
|
||||||
|
|
||||||
|
```text
|
||||||
|
devflow/compound/YYYY-MM-DD-decision-{slug}.md
|
||||||
|
```
|
||||||
|
|
||||||
|
## 归档确认
|
||||||
|
|
||||||
|
OpenSpec archive 是显式 human-in-the-loop 动作。archive 前必须确认 devflow 已经回填 OpenSpec 的关键执行信息:
|
||||||
|
|
||||||
|
- archive 阶段可以建议 archive,但必须先询问用户。
|
||||||
|
- 在用户确认前,不要执行 archive。
|
||||||
|
- 如果用户暂不归档,在 acceptance 中记录原因或状态。
|
||||||
|
- 如果用户确认归档,执行后记录 archive 结果和剩余档案位置。
|
||||||
|
|
||||||
|
## 归档交接
|
||||||
|
|
||||||
|
archive 阶段结束时告诉用户:
|
||||||
|
|
||||||
|
- 创建或更新了哪些档案文件。
|
||||||
|
- `devflow/index.md` 是否已更新。
|
||||||
|
- 运行了哪些验证,并按静态验证、脚本验证、浏览器/人工验证、未验证分类。
|
||||||
|
- 还剩哪些风险或后续事项。
|
||||||
|
- 明确询问:是否现在 archive OpenSpec change?
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# 内置执行协议
|
||||||
|
|
||||||
|
本文件只在外部 OpenSpec CLI 或子 skill 不可用时使用。fallback 不是跳过阶段,而是由 sm-flow 用文件方式完成同等最小产物。每次使用 fallback 都必须写入 `decisions.md` 或 `acceptance.md`,说明能力来源、缺失能力、影响和剩余风险。
|
||||||
|
|
||||||
|
## 通用规则
|
||||||
|
|
||||||
|
- 优先使用外部能力;只有不可用、不可发现或无法在当前环境调用时才使用内置协议。
|
||||||
|
- 不得因为使用 fallback 跳过 context、grill、commit、apply 授权或 archive 确认。
|
||||||
|
- fallback 产物仍写入 `openspec/changes/{slug}/` 和 `devflow/projects/YYYY-MM-DD-{slug}/`。
|
||||||
|
- 如果内置协议也无法满足阶段退出条件,暂停并向用户说明阻塞项。
|
||||||
|
|
||||||
|
## grill 内置协议
|
||||||
|
|
||||||
|
- 建立 question pool,至少覆盖术语、边界、验收;涉及参考实现或项目基础设施时加入技术实现问题。
|
||||||
|
- 将问题标记为 `evidence-driven` 或 `user-interview`。
|
||||||
|
- 先查证 evidence-driven 问题并汇报结论,再逐个询问 user-interview 问题。
|
||||||
|
- 按 `references/scales.md` 的当前分档满足 grill 要求。
|
||||||
|
- 将 question pool、证据结论、用户原话和确认状态写入 `decisions.md`;影响实现的结论回写 `proposal.md`。
|
||||||
|
|
||||||
|
## openspec 提案内置协议
|
||||||
|
|
||||||
|
- 在 `openspec/changes/{slug}/` 创建或更新:
|
||||||
|
- `proposal.md`:问题、方案、范围、非目标、上下文约束、风险。
|
||||||
|
- 设计产物:实现设计、接口影响、关键决策、架构风险;形式按 `references/scales.md` 的当前分档要求执行。
|
||||||
|
- `specs/*/spec.md` 或等价 functional spec:描述用户可观察行为和验收场景。
|
||||||
|
- `tasks.md`:按可执行切片拆分任务,并给每项写可验证验收标准。
|
||||||
|
- 运行 cross-artifact 对齐检查:proposal → 设计产物 → specs → tasks。
|
||||||
|
- 如果发现 gap,先修正 OpenSpec,再进入 commit。
|
||||||
|
|
||||||
|
## audit 内置协议
|
||||||
|
|
||||||
|
- 用 5 句话以内说明模块链路、数据所有权、跨模块依赖、架构风险和是否需要回写 OpenSpec。
|
||||||
|
- 如果风险影响实现,修正设计产物或 `tasks.md`。
|
||||||
|
- 将结论写入 `decisions.md`。
|
||||||
|
|
||||||
|
## openspec apply 内置协议
|
||||||
|
|
||||||
|
- 只依据 Committed OpenSpec 的 specs/tasks 实现;devflow 只作上下文参考。
|
||||||
|
- 开始前检查 `.committed` 文件;缺失则返回 commit。
|
||||||
|
- 如触发 pre-apply checkpoint,先阅读参考实现、grep 项目基础设施模式,并把技术栈清单写入 `decisions.md`。
|
||||||
|
- 按 tasks 的纵向切片实现、验证并更新任务状态。
|
||||||
|
- 发现冲突时按三类处理:OpenSpec 不准则修 OpenSpec,代码偏离则修代码,不确定则暂停等用户确认。
|
||||||
|
|
||||||
|
## openspec archive 内置协议
|
||||||
|
|
||||||
|
- 不删除或移动 OpenSpec change;只标记归档准备状态。
|
||||||
|
- 完成 devflow 回填、更新 `devflow/index.md`、创建 `.archive-ready`。
|
||||||
|
- 向用户汇报已创建文件、验证分类、剩余风险,并询问是否需要真实 OpenSpec archive。
|
||||||
|
- 如果外部 archive 能力仍不可用,在 `acceptance.md` 标记 `accepted-unarchived`。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# 术语表
|
||||||
|
|
||||||
|
本文件统一 sm-flow 协议中的核心词。优先使用这些词,避免同一概念多种说法。
|
||||||
|
|
||||||
|
| 术语 | 含义 | 使用边界 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| sm-flow | 协议层 harness | 编排 OpenSpec 生命周期,不替代 OpenSpec |
|
||||||
|
| OpenSpec | 当前变更的执行真理源 | apply 只能依据 Committed OpenSpec |
|
||||||
|
| devflow | 长期记忆和上下文层 | 提供术语、历史决策、验收记录,不直接指挥实现 |
|
||||||
|
| checkpoint | 用户可见检查点 | 默认只暴露 Discover / Commit / Apply / Archive |
|
||||||
|
| gate | 硬门控 | 不满足就不能进入下一关键动作,如 commit gate |
|
||||||
|
| Draft OpenSpec | 讨论和审计对象 | propose/specify 期间产生,不能直接 apply |
|
||||||
|
| Committed OpenSpec | 已通过 commit gate 的 OpenSpec | apply 的唯一执行依据 |
|
||||||
|
| fallback | 内置执行协议 | 外部 OpenSpec CLI 或子 skill 不可用时使用,必须标注 |
|
||||||
|
| decisions.md | 过程日志 | clarify 到 apply 期间记录问题、证据、决策、冲突和回写 |
|
||||||
|
| .committed | commit gate 标记文件 | 存在才可进入合规 apply |
|
||||||
|
| .archive-ready | archive 准备标记文件 | 表示 devflow 已回填,等待用户确认是否 archive |
|
||||||
|
| Discover | 用户可见 checkpoint | 覆盖 clarify + context + propose + grill |
|
||||||
|
| Commit | 用户可见 checkpoint | 覆盖 specify + audit + commit |
|
||||||
|
| Apply | 用户可见 checkpoint | 覆盖 apply |
|
||||||
|
| Archive | 用户可见 checkpoint | 覆盖 archive |
|
||||||
@@ -0,0 +1,120 @@
|
|||||||
|
# 运行规则
|
||||||
|
|
||||||
|
本文件承载稳定但不必放在顶层 `SKILL.md` 的运行规则。
|
||||||
|
|
||||||
|
## 接口影响分级
|
||||||
|
|
||||||
|
接口影响分级判断的是"记录在哪里、是否需要独立文档",不是判断"是否需要关注"。凡涉及字段、DTO、service 方法、API、事件、回调、数据库契约、命令契约、跨模块调用语义或内部决策逻辑变化,都必须先做分级。
|
||||||
|
|
||||||
|
| 级别 | 判断条件 | 产物要求 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| L1 内部实现 | 不改变任何调用方可观察的接口、字段、状态、错误码、数据范围、排序、过滤、权限结果、状态流转、副作用或文档承诺 | 不需要接口影响文档,只在 OpenSpec tasks 或 acceptance 记录验证 |
|
||||||
|
| L2 内部接口 | 改 DTO、service 方法、内部事件、内部 RPC 或内部判断逻辑,且所有消费者都在同一实现范围内 | 必须记录接口影响范围,可内联到 OpenSpec design/specs/tasks 或 devflow evidence/decisions |
|
||||||
|
| L3 协作接口 | 影响其他模块、其他服务、前端、外部系统、跨团队消费者、数据库契约、消息事件、回调或 SDK | 必须产出独立接口文档或等价独立章节 |
|
||||||
|
| L4 破坏性接口 | 删除字段、改字段语义、改状态机、改错误码、破坏兼容、旧调用方可能失败,或需要迁移、灰度、回滚 | 独立接口文档 + 迁移/回滚说明;必要时创建 ADR |
|
||||||
|
|
||||||
|
判断策略:
|
||||||
|
|
||||||
|
- 如果只是修复 bug,让接口回到原 OpenSpec 或原文档承诺,通常是 L1/L2。
|
||||||
|
- 如果判断逻辑改变了返回数据、错误码、状态、权限结果、排序/过滤、幂等性、时序或副作用,至少按 L3 检查。
|
||||||
|
- 如果旧调用方不改代码会失败、少数据、多数据、状态不同或错误码不同,按 L4 处理。
|
||||||
|
- 如果无法确定调用方边界或兼容性,默认提高一级并作为 `user-interview` 问题等待确认。
|
||||||
|
|
||||||
|
## 启动检查
|
||||||
|
|
||||||
|
1. 识别用户命令意图:
|
||||||
|
- `/sm-flow`(无参数):完整流程,从 clarify 开始。
|
||||||
|
- `/sm-flow apply [change]`:只执行,检查 commit gate → apply。
|
||||||
|
- `/sm-flow explore`:带上下文的探索模式,不走标准阶段链。
|
||||||
|
- `/sm-flow archive [change]`:收尾,回填 devflow + 归档确认。
|
||||||
|
- 明确要求"使用 sm-flow"或"走 sm-flow 流程":按显式调用处理。
|
||||||
|
- 自然语言指定阶段继续:识别意图后,自动补做最小前置检查,然后从指定阶段继续。
|
||||||
|
2. 判断启动模式:
|
||||||
|
- 完整模式:用户提供粗略想法或初始 PRD。
|
||||||
|
- Research 模式:用户已有 research,需要转成或修正 OpenSpec。
|
||||||
|
- PRD 文件模式:用户提供已有 PRD 路径。
|
||||||
|
- 恢复模式:用户希望从某个阶段继续(补做最小前置检查)。
|
||||||
|
- 快速模式:小改动,合并 gate;具体分档规则见 `references/scales.md`。
|
||||||
|
3. 如果缺少 `devflow/`,初始化:
|
||||||
|
- `devflow/projects/`
|
||||||
|
- `devflow/glossary/CONTEXT.md`
|
||||||
|
- `devflow/compound/`
|
||||||
|
4. 如果根目录存在旧 `CONTEXT.md`,且 `devflow/glossary/CONTEXT.md` 不存在或为空,询问用户是迁移还是合并。
|
||||||
|
5. 检查 OpenSpec 和子 skill 是否可用:
|
||||||
|
- OpenSpec 能力:`openspec-propose`、`openspec-apply-change`、`openspec-archive-change`。
|
||||||
|
- 辅助能力:`to-prd`、`grill-with-docs`、`diagnose`、`tdd`、`zoom-out`。
|
||||||
|
6. 如果 OpenSpec 或子 skill 不可用,不要静默跳过;使用内置执行协议(见 `references/fallbacks.md`),并在当前 checkpoint 说明 fallback 来源、影响和剩余风险。
|
||||||
|
|
||||||
|
## 进度汇报
|
||||||
|
|
||||||
|
用户可见进度默认折叠为 4 个 checkpoint:
|
||||||
|
|
||||||
|
| Checkpoint | 内部阶段 |
|
||||||
|
| --- | --- |
|
||||||
|
| Discover | clarify + context + propose + grill |
|
||||||
|
| Commit | specify + audit + commit |
|
||||||
|
| Apply | apply |
|
||||||
|
| Archive | archive |
|
||||||
|
|
||||||
|
汇报规则:
|
||||||
|
|
||||||
|
- 面向用户时优先使用 checkpoint 名称,不逐个汇报 9 个内部阶段。
|
||||||
|
- 内部阶段只在 checkpoint 摘要中作为证据列出,例如"Discover 已完成:读取了 devflow、生成 proposal、解决 2 个问题"。
|
||||||
|
- 只有发生阻塞、冲突、fallback、用户要求继续某个内部阶段,或需要解释恢复位置时,才暴露内部阶段名。
|
||||||
|
- 当前分档的汇报压缩规则见 `references/scales.md`;无论分档如何,都不要把内部阶段名当作用户操作入口。
|
||||||
|
|
||||||
|
## 项目标识规则
|
||||||
|
|
||||||
|
- 整个流程使用同一个 slug。
|
||||||
|
- 优先使用 OpenSpec change name。
|
||||||
|
- 如果还没有,则从功能标题生成 kebab-case slug。
|
||||||
|
- 项目档案目录格式:`devflow/projects/YYYY-MM-DD-{slug}/`。
|
||||||
|
- 如果目录已存在,默认恢复该项目;除非用户明确要求新开一轮。
|
||||||
|
|
||||||
|
## Devflow 产物分层
|
||||||
|
|
||||||
|
Devflow 是 sm-flow 自动维护的项目长期记忆层,不复制 OpenSpec 的执行产物。
|
||||||
|
|
||||||
|
**过程日志**(clarify → apply 期间维护):
|
||||||
|
|
||||||
|
- `decisions.md`:question pool、evidence-driven 汇报状态、user-interview 确认状态、关键取舍、风险接受、OpenSpec 回写记录、冲突分类记录。
|
||||||
|
|
||||||
|
**最终档案**(archive 阶段从 decisions.md + OpenSpec 产物提取):
|
||||||
|
|
||||||
|
- `brief.md`:背景、目标、范围、非目标、分档、关联 OpenSpec change。
|
||||||
|
- `evidence.md`:代码/文档证据、历史决策、evidence-driven 结论和汇报状态;分档要求见 `references/scales.md` 和 `references/archive-rules.md`。
|
||||||
|
- `acceptance.md`:实现结果、验证命令、未验证项、归档状态、后续事项。
|
||||||
|
|
||||||
|
**按需产物**(archive 阶段按需创建):
|
||||||
|
|
||||||
|
- `prd.md`:需求复杂、用户明确要求、或需要对外协作。
|
||||||
|
- `research.md`:存在真实调研、代码考古、竞品/API 对比或复杂方案比较。
|
||||||
|
- `design.md`:不适合放进 OpenSpec design 的长期背景或架构审计摘要。
|
||||||
|
- `tasks.md`:跨会话的人类追踪;执行任务仍属于 OpenSpec。
|
||||||
|
- `alignment.md` / `clarifications.md`:仅在 gap 或澄清很多时使用。
|
||||||
|
- `adr/*.md` 和 `compound/*.md`:仅在满足 ADR / compound knowledge 规则时使用。
|
||||||
|
|
||||||
|
**规模分档**:`micro / standard / complex` 的唯一规则源是 `references/scales.md`。
|
||||||
|
|
||||||
|
## 快速模式
|
||||||
|
|
||||||
|
快速模式适用于 `references/scales.md` 定义的 micro 变更。它合并 gate 而不仅仅是压缩产物;具体覆盖规则见 `references/scales.md`。
|
||||||
|
|
||||||
|
无论什么模式,以下内容必须保留:
|
||||||
|
|
||||||
|
- context 最小上下文收集:至少检查 glossary 和相关 ADR。
|
||||||
|
- grill 最小澄清:按 `references/scales.md` 当前分档要求执行;evidence-driven 结论仍需汇报。
|
||||||
|
- commit gate:确认没有未解决用户问题、接口影响已记录、OpenSpec tasks/specs 可执行;完整性检查按 `references/scales.md` 当前分档要求执行。
|
||||||
|
- apply 仍由 OpenSpec tasks/specs 驱动执行。
|
||||||
|
- archive 轻量回填:记录验收结果、OpenSpec 链接和归档状态。
|
||||||
|
|
||||||
|
## 完成标准
|
||||||
|
|
||||||
|
只有同时满足以下条件,流程才算完成:
|
||||||
|
|
||||||
|
- 用户可见的 Discover、Commit、Apply、Archive checkpoint 已完成,或未完成项已明确标记为暂停/不适用。
|
||||||
|
- OpenSpec proposal、设计产物、specs、tasks 已按当前分档生成或更新到可执行状态。
|
||||||
|
- 实现或规划工作已完成,且执行依据来自 OpenSpec。
|
||||||
|
- 已运行验证,或已记录未运行验证的原因。
|
||||||
|
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 `references/scales.md` 和 `references/archive-rules.md` 要求的当前分档档案。
|
||||||
|
- 用户知道剩余风险与下一步,并已被询问是否归档 OpenSpec change。
|
||||||
@@ -0,0 +1,352 @@
|
|||||||
|
# 阶段契约
|
||||||
|
|
||||||
|
本文件是 SM Flow 的逐阶段执行准则。核心原则:**sm-flow 编排 OpenSpec,OpenSpec 指挥执行,执行结果回填 devflow**。
|
||||||
|
|
||||||
|
执行顺序:clarify → context → propose → grill → specify → audit → commit → apply → archive。
|
||||||
|
|
||||||
|
## 目录
|
||||||
|
|
||||||
|
- clarify — 入口澄清
|
||||||
|
- context — 上下文收集
|
||||||
|
- propose — 轻量 propose
|
||||||
|
- grill — 人类对齐澄清
|
||||||
|
- specify — 细化 + 对齐
|
||||||
|
- audit — 架构审计
|
||||||
|
- commit — Commit OpenSpec
|
||||||
|
- apply — OpenSpec 执行
|
||||||
|
- archive — 回填 + 归档
|
||||||
|
|
||||||
|
## clarify — 入口澄清
|
||||||
|
|
||||||
|
**进入条件**:用户提供粗略想法、初始 PRD、已有 research、issue,或要求启动 SM Flow。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 收集问题、期望结果、目标用户、涉及代码区域、约束条件和可能的非目标。
|
||||||
|
- 如果用户已有 research,先识别它是否已经包含用户价值、技术方案、验收标准和任务拆分。
|
||||||
|
- 如果输入过于模糊,最多追加三轮聚焦问题。
|
||||||
|
- 当答案会改变 OpenSpec proposal/specs/tasks 时,优先一次只问一个问题。
|
||||||
|
- 如果需要判断 `micro / standard / complex` 分档,补读 `references/scales.md`。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- 问题可以用 1-2 句话说清楚。
|
||||||
|
- 期望结果可以用 1-2 句话说清楚。
|
||||||
|
- 已列出已知影响代码或模块;如果未知,也明确标记。
|
||||||
|
- 可以生成 OpenSpec change slug。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 入口摘要。
|
||||||
|
- 初步 slug。
|
||||||
|
- devflow 规模分档:`micro` / `standard` / `complex`。
|
||||||
|
|
||||||
|
## context — 上下文收集
|
||||||
|
|
||||||
|
**进入条件**:clarify 已经有足够信息定位领域、项目或变更方向。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 优先读取 `devflow/index.md`,按日期、slug、领域、关键词和关联 OpenSpec 定位候选项目。
|
||||||
|
- 如果 `devflow/index.md` 不存在,先从 `devflow/projects/` 现有目录初始化轻量索引,再继续本次上下文收集。
|
||||||
|
- 读取 `devflow/glossary/CONTEXT.md`,提取相关术语和业务规则。
|
||||||
|
- 搜索 `devflow/projects/` 中相关 PRD、design、tasks、acceptance 和 ADR。
|
||||||
|
- 搜索 `devflow/compound/` 中可复用 learning、trick、decision、explore。
|
||||||
|
- 记录哪些上下文会影响 OpenSpec proposal、设计产物、specs 或 tasks。
|
||||||
|
- 如果发现旧根目录 `CONTEXT.md` 与 `devflow/glossary/CONTEXT.md` 冲突,暂停并向用户汇报。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- 已形成"OpenSpec 输入上下文摘要"。
|
||||||
|
- 已记录 `devflow/index.md` 的使用状态:已命中 / 已初始化 / 无相关条目。
|
||||||
|
- 已列出相关 ADR 和不能违反的历史决策。
|
||||||
|
- 已列出需要写入或修正 OpenSpec 的上下文点。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 上下文摘要,写入 `decisions.md`(过程日志)。会影响实现的上下文必须标记为"需进入 OpenSpec"。
|
||||||
|
|
||||||
|
## propose — 轻量 propose
|
||||||
|
|
||||||
|
**进入条件**:clarify + context 已经足够生成轻量 proposal。
|
||||||
|
|
||||||
|
**执行者**:sm-flow 内置协议。**不调用 openspec-propose**(完整 OpenSpec 产物留待 specify 阶段生成)。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 创建或识别 `openspec/changes/{slug}/`。
|
||||||
|
- 写入 `proposal.md`,包含:问题、建议方案、范围、非目标、来自 devflow 的上下文约束、风险。
|
||||||
|
- **不生成 design.md、specs/、tasks.md**——这些留待 grill 澄清需求后在 specify 阶段补全。
|
||||||
|
- 用 context 阶段的 devflow 上下文增强 proposal。
|
||||||
|
- 在承诺方案方向前,先检查相关仓库代码。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- `openspec/changes/{slug}/proposal.md` 存在。
|
||||||
|
- 关键假设已显式记录。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- Draft OpenSpec proposal.md(轻量版)。
|
||||||
|
|
||||||
|
**Human checkpoint**:
|
||||||
|
- 向用户简要说明 proposal 范围、关键假设、主要风险、devflow 上下文如何影响方案。
|
||||||
|
- 作为 Discover checkpoint 的中间状态汇报;询问是否继续完成 Discover 的人类澄清部分。用户明确要求"全自动执行"时可跳过等待。
|
||||||
|
|
||||||
|
## grill — 人类对齐澄清
|
||||||
|
|
||||||
|
**进入条件**:propose 已有轻量 proposal.md。
|
||||||
|
|
||||||
|
**能力来源**:优先使用 `grill-with-docs`;不可用时使用 `references/fallbacks.md#grill-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 优先使用 `grill-with-docs`。
|
||||||
|
- 进入 grill 时先建立一个 question pool,并记录到 `decisions.md`:
|
||||||
|
- 默认至少覆盖术语、边界、验收三个维度。
|
||||||
|
- **技术实现维度**(新增):当 proposal 提到参考实现、或涉及项目现有基础设施时,增加技术澄清问题:
|
||||||
|
- 参考实现的具体文件路径是什么?
|
||||||
|
- 项目现有的 [请求结构/MQ/缓存/加密/工具类] 标准是什么?
|
||||||
|
- 有哪些技术点需要先调研或新建?
|
||||||
|
- 如果变更涉及多模块、接口、权限、下游消费者、响应结构或生命周期规则,先把这些维度补进问题池。
|
||||||
|
- 逐项标记每个问题的模式:
|
||||||
|
- `evidence-driven`:问题能通过代码、文档、测试、OpenSpec 或既有 ADR 证明;代理先查证,再向用户汇报证据、结论和是否需要确认。
|
||||||
|
- `user-interview`:问题涉及产品偏好、范围边界、验收口径、风险接受度或价值取舍;必须问用户并等待确认。
|
||||||
|
- evidence-driven 和 user-interview 的推进节奏:先批量查证 evidence-driven 并一次性汇报结论,再逐个处理 user-interview 问题。不要把所有问题攒到最后一起问。
|
||||||
|
- 一次只问一个 `user-interview` 问题。
|
||||||
|
- 每个 `user-interview` 问题必须等待用户显式回答,并在 decisions.md 中记录:问题原文、用户原话、确认状态(已确认/未确认)。未确认的问题不能从 question pool 移除。
|
||||||
|
- 单个 `user-interview` 的确认只能解除该问题本身的阻塞,不能被解释为进入 apply 或修改执行目标文件的授权。
|
||||||
|
- 对接口影响等级、消费者边界或兼容性存在不确定时,必须作为 `user-interview` 问题等待用户确认。
|
||||||
|
- 如果澄清结果影响实现,必须回写 proposal.md。
|
||||||
|
- 术语一旦确认,更新 `devflow/glossary/CONTEXT.md`。
|
||||||
|
- 对难以逆转、依赖上下文、源自真实权衡的决策创建 ADR。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- question pool 已建立并覆盖当前 change 所需维度。
|
||||||
|
- 已满足 `references/scales.md` 中当前分档的 grill 要求。每个问题都必须记录属于 `evidence-driven` 还是 `user-interview`。
|
||||||
|
- 所有 evidence-driven 结论已向用户汇报。
|
||||||
|
- 所有 user-interview 决策已获得用户确认。
|
||||||
|
- 没有未解决或代理代确认的 user-interview 问题。
|
||||||
|
- 没有未判级或未确认的接口影响问题。
|
||||||
|
- 影响实现的结论已回写 proposal.md。
|
||||||
|
- 单个 grill 决策确认不等于 apply 授权;grill 完成后必须停在 commit,等待用户明确要求进入 apply。
|
||||||
|
- question pool、evidence-driven 结论、user-interview 确认必须写入 `decisions.md` 文件,不能只记录在对话中。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 更新后的 proposal.md。
|
||||||
|
- 澄清记录:写入 `decisions.md`。包含 question pool、evidence-driven 汇报状态、user-interview 确认状态。
|
||||||
|
- 更新后的词汇表和 ADR。
|
||||||
|
|
||||||
|
**Human checkpoint**:
|
||||||
|
- 汇报已解决和未解决的问题、proposal 变更、术语和 ADR 更新。
|
||||||
|
- 汇报 Discover checkpoint 完成情况,并询问是否继续进入 Commit checkpoint。
|
||||||
|
|
||||||
|
## specify — 细化 + 对齐
|
||||||
|
|
||||||
|
**进入条件**:grill 已退出,需求已通过澄清稳定下来。
|
||||||
|
|
||||||
|
**能力来源**:优先使用 `openspec-propose`(基于已稳定的 proposal 补全完整 OpenSpec);按需使用 `to-prd`。进入本阶段必须先声明调用方式;外部能力不可用时使用 `references/fallbacks.md#openspec-提案-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 基于已稳定的 proposal.md 补全设计产物、specs/、tasks.md:
|
||||||
|
- 优先调用 `openspec-propose`,输入中明确说明"proposal.md 已存在,本次只需按当前分档补全设计产物/specs/tasks"。
|
||||||
|
- 如果不可用,执行 `references/fallbacks.md#openspec-提案-内置协议`。
|
||||||
|
- 如果没有结构化 PRD,按需按 `to-prd` 协议生成 `brief.md`;复杂需求、对外协作或用户明确要求时再生成 `prd.md`。
|
||||||
|
- 独立 PRD 是否需要按 `references/scales.md` 的当前分档和用户要求判断。
|
||||||
|
- 用 grill 阶段的 decisions.md 记录增强 OpenSpec 产物:确保 design/specs/tasks 反映所有已确认的决策。
|
||||||
|
- **显式 cross-artifact 对齐检查**——在 checkpoint 中输出对齐检查表:
|
||||||
|
- `brief/prd` 中的目标、范围、非目标和验收预期 → `proposal` 是否覆盖。
|
||||||
|
- `proposal` 中的范围、约束和关键承诺 → `design` 是否覆盖。
|
||||||
|
- `design` 中影响实现的约束、接口影响和架构结论 → `specs` 或 `tasks` 是否覆盖。
|
||||||
|
- `specs` 中的可观察行为 → `tasks` 是否覆盖为可执行切片。
|
||||||
|
- 每项标记:已对齐 / 存在 gap。
|
||||||
|
- 检查是否涉及接口影响:
|
||||||
|
- 接口影响分级定义见 `references/operating-rules.md#接口影响分级`。
|
||||||
|
- 是否改变字段、DTO、service 方法、API、事件、回调、数据库契约、命令契约或跨模块调用语义。
|
||||||
|
- 接口内部判断逻辑是否改变调用方可观察行为。
|
||||||
|
- 按 L1/L2/L3/L4 记录接口影响等级;不确定时标记为 `user-interview` 问题。
|
||||||
|
- 如果存在 gap,在进入下一阶段前修复 OpenSpec。
|
||||||
|
- 如果发现不一致,优先修正 OpenSpec,而不是只修改 devflow 文档。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- OpenSpec 细化产物存在且与 proposal 对齐;产物形态按 `references/scales.md` 的当前分档要求执行。
|
||||||
|
- `brief.md` 已覆盖背景、目标、范围和非目标;复杂需求存在独立 `prd.md` 或用户明确不需要 PRD。
|
||||||
|
- cross-artifact 对齐检查表已生成(4 行,每行标记已对齐/存在 gap),没有未处理 gap。
|
||||||
|
- 涉及接口变更时,已记录接口影响等级和产物要求;不确定项已标记。
|
||||||
|
- 所有已知冲突已修正或等待用户决策。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- Draft OpenSpec:按 `references/scales.md` 的当前分档要求生成 proposal、设计、specs 和 tasks。
|
||||||
|
- `brief.md`,以及按需创建的 `prd.md`。
|
||||||
|
- cross-artifact 对齐检查表(写入 checkpoint 或 decisions.md)。
|
||||||
|
- 必要的 OpenSpec 修正。
|
||||||
|
|
||||||
|
## audit — 架构审计
|
||||||
|
|
||||||
|
**进入条件**:specify 已退出,完整 OpenSpec 产物已存在。
|
||||||
|
|
||||||
|
**能力来源**:优先使用 `zoom-out`;不可用时使用 `references/fallbacks.md#audit-内置协议`,并在 `decisions.md` 标注 fallback。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 画出输入 → 处理 → 输出的模块链路。
|
||||||
|
- 识别跨模块依赖、数据所有权、生命周期和耦合风险。
|
||||||
|
- 检查是否与既有架构、ADR、OpenSpec design 冲突。
|
||||||
|
- 用不超过五句话写出架构风险评估。
|
||||||
|
- 如果审计结果影响实现,必须回写 OpenSpec design/tasks;只写入 devflow design 不够。
|
||||||
|
- 审计结论写入 `decisions.md`。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- 架构风险已被接受,或流程返回 grill/specify 修正 OpenSpec。
|
||||||
|
- OpenSpec 设计产物/tasks 已反映会影响实现的架构审计结论。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 架构审计记录,写入 `decisions.md`;复杂架构审计可拆出 `design.md`。
|
||||||
|
- 必要的 OpenSpec 设计产物/tasks 修正。
|
||||||
|
|
||||||
|
**Human checkpoint**:
|
||||||
|
- 用不超过五句话向用户说明架构风险、OpenSpec 修正点和实现计划。
|
||||||
|
- 作为 Commit checkpoint 的中间状态汇报;询问是否继续完成 commit gate。
|
||||||
|
|
||||||
|
## commit — Commit OpenSpec
|
||||||
|
|
||||||
|
**进入条件**:
|
||||||
|
- grill 已满足 `references/scales.md` 中当前分档要求。
|
||||||
|
- 所有 `user-interview` 问题都已获得用户显式确认。
|
||||||
|
- audit 已经完成,或快速模式下已记录跳过原因;快速模式定义见 `references/operating-rules.md#快速模式`。
|
||||||
|
- Draft OpenSpec 已回写所有会影响实现的澄清、接口影响和架构审计结论。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 检查 proposal 是否说明为什么做、做什么、范围和非目标。
|
||||||
|
- 检查 design 是否记录上下文约束、关键技术决策、架构风险和接口影响。
|
||||||
|
- 检查 specs 是否表达外部可观察行为,并覆盖验收口径。
|
||||||
|
- 检查 tasks 是否是可执行的纵向切片,而不是泛泛描述。
|
||||||
|
- 复核 cross-artifact 对齐:`brief/prd → proposal → 设计产物 → specs → tasks` 是否闭环,没有把字段、范围项、验收行为或实现切片丢在上游产物里。
|
||||||
|
- 检查 `decisions.md` 中所有影响实现的发现,是否已回写到 proposal、design、specs 或 tasks。
|
||||||
|
- 接口影响分级定义见 `references/operating-rules.md#接口影响分级`。
|
||||||
|
- 检查接口影响是否已按 L1/L4 判级;L3/L4 是否有独立接口文档或等价独立章节。
|
||||||
|
- 检查没有未汇报的 evidence-driven 结论,没有未确认的 user-interview 问题,没有 devflow/OpenSpec 冲突。
|
||||||
|
- 如果检查失败,返回 propose、grill、specify 或 audit 修正 Draft OpenSpec。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- Draft OpenSpec 已达到可执行状态,并记录为 Committed OpenSpec。
|
||||||
|
- **文件完整性检查**(按 `references/scales.md` 的当前分档要求执行):
|
||||||
|
- [ ] proposal 存在,且足以说明问题、建议方案、范围和非目标。
|
||||||
|
- [ ] 设计产物存在,形式符合当前分档要求。
|
||||||
|
- [ ] specs 存在,且表达用户可观察行为。
|
||||||
|
- [ ] tasks 存在,且任务可执行、验收标准可验证。
|
||||||
|
- **一致性检查**(必须通过):
|
||||||
|
- [ ] proposal 中的核心概念在设计产物中有对应设计
|
||||||
|
- [ ] 设计产物中的关键决策在 tasks 中有对应实现任务
|
||||||
|
- [ ] tasks 的验收标准可验证(不是"正确实现""完成功能"这类模糊描述)
|
||||||
|
- **标记文件**:检查通过后,创建 `openspec/changes/{slug}/.committed` 文件标记为 Committed OpenSpec
|
||||||
|
- 所有 preflight 风险已消除或明确记录为已接受。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- Committed OpenSpec 状态说明。
|
||||||
|
- preflight 检查结果,写入 `decisions.md` 或 `acceptance.md`。
|
||||||
|
|
||||||
|
**Human checkpoint**:
|
||||||
|
- 用不超过五句话说明 Committed OpenSpec 的范围、接口影响、剩余风险和执行计划。
|
||||||
|
- 汇报 Commit checkpoint 完成情况,并询问是否进入 Apply checkpoint;除非用户在启动时明确要求"全自动执行",必须等待用户明确说出进入 apply、开始实现、执行修改或等价授权。
|
||||||
|
- 不得把 grill 的单个决策确认当作本 checkpoint 的授权。
|
||||||
|
|
||||||
|
## apply — OpenSpec 执行
|
||||||
|
|
||||||
|
**进入条件**:
|
||||||
|
- `openspec/changes/{slug}/` 中 proposal、设计产物、specs、tasks 已通过 commit,成为 Committed OpenSpec。
|
||||||
|
- **前置门控检查**(硬约束):
|
||||||
|
- 检查 `openspec/changes/{slug}/.committed` 文件是否存在
|
||||||
|
- 如不存在,执行以下流程:
|
||||||
|
1. 汇报:Draft OpenSpec 未通过 commit 检查
|
||||||
|
2. 列出缺失的 checkpoint 项(文件完整性、一致性检查)
|
||||||
|
3. 询问用户:是否补做 commit 检查;如用户要求不补做,则中止 apply 或标记为 `emergency-bypass`,且本次流程不得视为合规 sm-flow apply
|
||||||
|
- commit 后已获得用户明确的 apply 授权,除非用户在启动时要求"全自动执行"。
|
||||||
|
- devflow 与 OpenSpec 没有未解决冲突。
|
||||||
|
- 没有未解决的 user-interview 问题、未判级接口影响、未汇报 evidence-driven 结论或未接受架构风险。
|
||||||
|
|
||||||
|
**能力来源**:优先使用 `openspec-apply-change`;不可用时使用 `references/fallbacks.md#openspec-apply-内置协议`,并在 `decisions.md` 标注 fallback。遇到 bug/不确定行为时优先使用 `diagnose`;需要测试驱动时优先使用 `tdd`。不可用时执行对应最小协议并记录原因,不得静默跳过。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
|
||||||
|
### Pre-apply Checkpoint
|
||||||
|
|
||||||
|
**触发条件**:当 OpenSpec 涉及以下任一情况时必须执行
|
||||||
|
- design 或 tasks 中提到"参考 XXX 实现"
|
||||||
|
- 需要调用项目现有基础设施(MQ/统一请求结构/工具类等)
|
||||||
|
- 技术栈不熟悉或第一次在该项目实现类似功能
|
||||||
|
|
||||||
|
**执行步骤**:
|
||||||
|
1. **阅读所有参考实现**
|
||||||
|
- 从 OpenSpec design 或 tasks 中定位参考实现文件
|
||||||
|
- 如果路径不明确,通过 Grep 搜索关键类名或模式
|
||||||
|
- 理解关键逻辑,提取可复用代码片段和模式
|
||||||
|
|
||||||
|
2. **Grep 关键技术栈**
|
||||||
|
- 请求/响应结构模式(如 `RequestMsg`、`ResponseMsg`、DTO 规范)
|
||||||
|
- 消息队列模式(如 `@KafkaListener`、`@YkMsg`、发送模板)
|
||||||
|
- 统一工具类(如 `XxxUtil`、`XxxHelper`、加密/验签工具)
|
||||||
|
- 异常处理和日志记录标准
|
||||||
|
|
||||||
|
3. **形成技术栈清单并写入 decisions.md**
|
||||||
|
- 项目使用的请求/响应结构标准
|
||||||
|
- MQ 消息定义和发送标准
|
||||||
|
- Consumer 标准位置和写法
|
||||||
|
- 加密/验签/工具类的标准用法
|
||||||
|
- 识别需要新建的工具类或基础设施
|
||||||
|
|
||||||
|
**输出要求**:
|
||||||
|
- 技术栈清单已写入 `decisions.md` 的 "Pre-apply Research" 章节。
|
||||||
|
- 已列出所有参考实现的文件路径。
|
||||||
|
- 已识别需要新建的工具类/基础设施。
|
||||||
|
|
||||||
|
**按风险执行**:执行深度按 `references/scales.md` 的当前分档和实现风险决定;退出判断以清单是否足以指导实现为准。
|
||||||
|
|
||||||
|
### 实现过程
|
||||||
|
|
||||||
|
- 优先调用 `openspec-apply-change`。
|
||||||
|
- 执行依据是 OpenSpec specs/tasks;devflow 只能作为上下文参考。
|
||||||
|
- 按 OpenSpec tasks 的纵向切片实现。
|
||||||
|
- **分步实现**:建议按 Controller → Service → MQ/异步组件 → Consumer/下游 顺序,每完成一层验证后再继续。
|
||||||
|
- 进入实现前先汇报本阶段的 capability 来源、当前 task 进度和本轮要推进的切片;否则 apply 不算真正开始。
|
||||||
|
- **首模块完成后对齐检查**:完成第一个接口/模块后,对比 OpenSpec design/tasks,标记"已完成/TODO";核心功能(加密/验签/核心业务逻辑)不允许空实现或纯 TODO 注释。
|
||||||
|
- 当用户质疑、用户要求修改、代码检查、测试失败或运行行为与 OpenSpec 冲突时,做三类判断:
|
||||||
|
- OpenSpec 不准(规格遗漏、边界未覆盖、验收口径缺失)→ 暂停 apply,修正 OpenSpec 后重新提交。
|
||||||
|
- 代码偏离(实现没按 OpenSpec 做)→ 修正代码,不改 OpenSpec。
|
||||||
|
- 不确定根因、涉及设计方向、用户改变目标或范围 → 暂停并等待用户确认。
|
||||||
|
- 判断结果、证据、用户确认和 OpenSpec 回写状态必须记录到 `decisions.md`。
|
||||||
|
- **快速失败**:连续返工 ≥ 2 次时,暂停并重新执行 pre-apply checkpoint 或向用户汇报。
|
||||||
|
- 当用户要求、行为复杂或回归风险高时使用 TDD。
|
||||||
|
- 当测试失败、行为意外或原因不确定时使用 diagnose。
|
||||||
|
- 如果 diagnose 发现根因是 OpenSpec 不准确,先修正 OpenSpec,再继续 apply。
|
||||||
|
- 修改文件前遵守仓库指令,例如 `AGENTS.md`。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- 已完成 pre-apply checkpoint(如触发条件满足),技术栈清单已写入 `decisions.md`。
|
||||||
|
- OpenSpec tasks 已完成,或剩余 tasks 已明确记录。
|
||||||
|
- 核心功能已实现或明确标注"待联调",无纯 TODO 占位。
|
||||||
|
- 所有实现期冲突已分类并处理;没有未确认的规格遗漏、设计冲突或用户变更。
|
||||||
|
- 已运行验证,或记录了未验证原因。
|
||||||
|
- 已列出已知限制。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 代码变更、必要测试和实现说明。
|
||||||
|
- 更新后的 OpenSpec task 状态。
|
||||||
|
- 冲突记录写入 `decisions.md`。
|
||||||
|
|
||||||
|
## archive — 回填 + 归档
|
||||||
|
|
||||||
|
**进入条件**:实现或规划工作已经达到可交接状态。
|
||||||
|
|
||||||
|
**能力来源**:`openspec-archive-change` 在用户确认 archive 后优先调用;不可用时使用 `references/fallbacks.md#openspec-archive-内置协议`,并在 `acceptance.md` 标注 fallback。archive 回填由 `sm-flow` 执行。
|
||||||
|
|
||||||
|
**动作**:
|
||||||
|
- 遵循 `references/archive-rules.md`。
|
||||||
|
- 从 `decisions.md`(过程日志)+ OpenSpec 产物提炼完整 devflow 档案:
|
||||||
|
- `brief.md`:从 proposal.md 提取背景、目标、范围、非目标。
|
||||||
|
- `evidence.md`:按 `references/scales.md` 和 `references/archive-rules.md` 的当前分档要求处理。
|
||||||
|
- `decisions.md`:保持为最终版,整理格式。
|
||||||
|
- `acceptance.md`:从实现结果和验证结果提取。
|
||||||
|
- 只在复杂场景按需拆出 PRD/research/design/tasks/alignment。
|
||||||
|
- 写入或更新验收记录,并区分静态验证、脚本验证、浏览器/人工验证、未验证。
|
||||||
|
- 如果本次流程产生可复用经验,写入 compound knowledge。
|
||||||
|
- 更新 `devflow/index.md`,记录日期、slug、领域、关键词、关联 OpenSpec 和状态。
|
||||||
|
- 询问用户是否要 archive OpenSpec change;不要默认执行归档。
|
||||||
|
|
||||||
|
**退出条件**:
|
||||||
|
- `devflow/projects/YYYY-MM-DD-{slug}/` 包含 `references/scales.md` 和 `references/archive-rules.md` 要求的当前分档档案;archive checkpoint 必须列出所有已创建的文件路径,验证文件实际存在于磁盘。
|
||||||
|
- `devflow/index.md` 已包含或更新本项目条目。
|
||||||
|
- 用户已被询问是否 archive OpenSpec change。
|
||||||
|
|
||||||
|
**输出**:
|
||||||
|
- 完整 devflow 档案。
|
||||||
|
- 归档交接清单:创建或更新了哪些文件、验证分类、剩余风险、是否 archive。
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# 分档规则
|
||||||
|
|
||||||
|
本文件是 `micro / standard / complex` 的唯一规则源。其它文件只引用本文件,不重复定义分档细节。
|
||||||
|
|
||||||
|
## standard 基准
|
||||||
|
|
||||||
|
standard 是默认分档,适用于普通功能、明确但有一定实现范围的变更。
|
||||||
|
|
||||||
|
- 用户可见 checkpoint:Discover → Commit → Apply → Archive。
|
||||||
|
- OpenSpec 产物:`proposal.md`、独立 `design.md`、`specs/`、`tasks.md`。
|
||||||
|
- grill:解决术语、边界、验收三个维度的高价值问题。
|
||||||
|
- commit gate:检查 proposal、design、specs、tasks 的完整性和一致性。
|
||||||
|
- devflow 档案:`brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`。
|
||||||
|
|
||||||
|
## micro 覆盖
|
||||||
|
|
||||||
|
micro 适用于小改动、低风险、需求明确的变更。micro 是 standard 的减法,不是跳过流程。
|
||||||
|
|
||||||
|
- checkpoint 可合并:Discover + Commit 可在无阻塞时合并汇报。
|
||||||
|
- micro 内部流程压缩为:clarify+context 合并 checkpoint → 轻量 propose → grill → specify+commit 合并 checkpoint。
|
||||||
|
- context 保留最小收集:至少检查 glossary 和相关 ADR。
|
||||||
|
- grill 保留最小澄清:至少解决一个高价值问题,并记录术语、边界、验收三类是否明确;不明确项必须补问或标记风险。
|
||||||
|
- OpenSpec 仍需要 `proposal.md`、`specs/`、`tasks.md`。
|
||||||
|
- `design.md` 可不独立创建;允许在 `proposal.md` 或 `tasks.md` 中写等价设计小节。
|
||||||
|
- `specs/` 和 `tasks.md` 可轻量,但必须表达可观察行为和可执行任务。
|
||||||
|
- commit gate 仍必须通过,并创建 `.committed`。
|
||||||
|
- devflow 档案至少包含 `brief.md`、`decisions.md`、`acceptance.md`;证据少时可并入 `brief.md` 或 `decisions.md`。
|
||||||
|
- apply 仍只能依据 Committed OpenSpec。
|
||||||
|
- archive 仍要轻量回填 devflow,并询问是否归档 OpenSpec。
|
||||||
|
|
||||||
|
micro 不适用于接口影响不清、跨团队消费者、迁移/回滚、复杂状态机、长期架构决策或需求边界不清的变更;遇到这些情况应升级为 standard 或 complex。
|
||||||
|
|
||||||
|
## complex 增量
|
||||||
|
|
||||||
|
complex 适用于高风险、跨模块、需求不清、多人协作或长期架构影响明显的变更。complex 是 standard 的加法。
|
||||||
|
|
||||||
|
- 需要更完整的 Discover:增加需求澄清、证据查证、范围确认和风险接受。
|
||||||
|
- checkpoint 内可补充关键内部阶段结果,但不要把内部阶段名当作用户操作入口。
|
||||||
|
- 按需创建 `prd.md`、`research.md`、`alignment.md`、接口文档、ADR 或 compound knowledge。
|
||||||
|
- 接口影响、迁移、灰度、回滚、兼容性和消费者边界必须显式记录。
|
||||||
|
- audit 需要覆盖模块链路、数据所有权、生命周期、耦合风险和 ADR 冲突。
|
||||||
|
- archive 在 standard 档案基础上按需提炼长期 design、research、tasks、ADR 和 compound knowledge。
|
||||||
@@ -0,0 +1,385 @@
|
|||||||
|
# 模板
|
||||||
|
|
||||||
|
这些是最小模板。只有在能提升未来可读性时,才增加额外章节。保留 PRD、ADR、OpenSpec、slug 等行业术语,其余说明尽量使用中文。
|
||||||
|
|
||||||
|
## Brief 模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:{goal}
|
||||||
|
- 当前问题:{problem}
|
||||||
|
- 关联 OpenSpec:`openspec/changes/{slug}/`
|
||||||
|
- devflow 分档:micro | standard | complex
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:{in scope}
|
||||||
|
- 本次不做:{out of scope}
|
||||||
|
- 影响区域:{modules/files if known}
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||||
|
- specs 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||||
|
- tasks 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||||
|
```
|
||||||
|
|
||||||
|
## Evidence 模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| {file/doc/test/ADR} | {evidence summary} | {conclusion} | 是 / 否 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- 结论:{conclusion}
|
||||||
|
- 证据:{evidence}
|
||||||
|
- 风险:{risk if any}
|
||||||
|
- 用户确认:需要 / 不需要 / 已确认
|
||||||
|
```
|
||||||
|
|
||||||
|
## Decisions 模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} Decisions
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | {question} | evidence-driven / user-interview | 已解决 / 未解决 |
|
||||||
|
| Q2 | 边界 | {question} | evidence-driven / user-interview | 已解决 / 未解决 |
|
||||||
|
| Q3 | 验收 | {question} | evidence-driven / user-interview | 已解决 / 未解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| {conclusion} | {file/doc/test/ADR} | 已汇报 / 待汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| {question} | {user's exact words} | 已确认 / 未确认 | 已回写 / 不影响 / 待回写 |
|
||||||
|
|
||||||
|
## 关键取舍
|
||||||
|
|
||||||
|
- 决策:{decision}
|
||||||
|
- 原因:{why}
|
||||||
|
- 影响:{impact}
|
||||||
|
- 风险接受:{accepted by whom/when}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 接口影响记录模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 接口影响记录
|
||||||
|
|
||||||
|
## 分级
|
||||||
|
|
||||||
|
- 级别:L1 内部实现 / L2 内部接口 / L3 协作接口 / L4 破坏性接口
|
||||||
|
- 判级原因:{why this level}
|
||||||
|
- 是否需要独立接口文档:是 / 否
|
||||||
|
|
||||||
|
## 变更对象
|
||||||
|
|
||||||
|
- 接口/字段/DTO/事件/回调/数据库契约:
|
||||||
|
- 判断逻辑变化:
|
||||||
|
- 可观察行为变化:返回数据 / 状态 / 错误码 / 权限结果 / 过滤排序 / 幂等性 / 时序 / 副作用 / 无
|
||||||
|
|
||||||
|
## 影响范围
|
||||||
|
|
||||||
|
- 调用方/消费者:
|
||||||
|
- 是否跨模块/跨服务/跨团队:
|
||||||
|
- 旧调用方是否需要改动:
|
||||||
|
|
||||||
|
## 兼容与迁移
|
||||||
|
|
||||||
|
- 是否向后兼容:
|
||||||
|
- 迁移/灰度/回滚要求:
|
||||||
|
- 风险接受:
|
||||||
|
|
||||||
|
## 验收方式
|
||||||
|
|
||||||
|
- 如何证明新行为正确:
|
||||||
|
- 如何证明旧行为未破坏:
|
||||||
|
- 需要用户确认的问题:
|
||||||
|
```
|
||||||
|
|
||||||
|
## 实现期冲突记录模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 实现期冲突记录
|
||||||
|
|
||||||
|
## 冲突摘要
|
||||||
|
|
||||||
|
- 触发来源:用户质疑 / 用户变更 / 代码发现 / 测试失败 / 运行行为
|
||||||
|
- 冲突对象:proposal / design / specs / tasks / ADR / 代码行为
|
||||||
|
- 分类:OpenSpec 不准 / 代码偏离 / 不确定
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
- OpenSpec 依据:
|
||||||
|
- 代码或测试证据:
|
||||||
|
- 用户反馈:
|
||||||
|
|
||||||
|
## 处理
|
||||||
|
|
||||||
|
- 决策:
|
||||||
|
- 是否需要用户确认:是 / 否
|
||||||
|
- OpenSpec 回写:不需要 / 已回写 / 待回写 / 等待用户确认
|
||||||
|
- 代码处理:
|
||||||
|
- 验证方式:
|
||||||
|
```
|
||||||
|
|
||||||
|
## PRD 模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} PRD
|
||||||
|
|
||||||
|
## 问题陈述
|
||||||
|
|
||||||
|
用用户视角描述问题。
|
||||||
|
|
||||||
|
## 解决方案
|
||||||
|
|
||||||
|
用用户视角描述预期解决方案。
|
||||||
|
|
||||||
|
## 用户故事
|
||||||
|
|
||||||
|
1. 作为{角色},我希望{能力},以便{收益}。
|
||||||
|
|
||||||
|
## 实现决策
|
||||||
|
|
||||||
|
- 决策:{decision}
|
||||||
|
- 原因:{why}
|
||||||
|
- 影响:{affected modules or behavior}
|
||||||
|
|
||||||
|
## 测试决策
|
||||||
|
|
||||||
|
- 好测试应该通过{public interface}验证{observable behavior}。
|
||||||
|
- 必须覆盖:{critical paths}
|
||||||
|
- 不测试:{explicit exclusions}
|
||||||
|
|
||||||
|
## 非目标
|
||||||
|
|
||||||
|
- {excluded behavior}
|
||||||
|
|
||||||
|
## 补充说明
|
||||||
|
|
||||||
|
- {open question or useful context}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 词汇表模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# 上下文词汇表
|
||||||
|
|
||||||
|
## 术语
|
||||||
|
|
||||||
|
### {术语}
|
||||||
|
|
||||||
|
- 定义:{precise definition}
|
||||||
|
- 使用场景:{feature/module/context}
|
||||||
|
- 备注:{ambiguities, synonyms, or rejected meanings}
|
||||||
|
|
||||||
|
## 业务规则
|
||||||
|
|
||||||
|
- {rule}: {meaning and source}
|
||||||
|
```
|
||||||
|
|
||||||
|
## ADR 模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# ADR-{编号}: {决策标题}
|
||||||
|
|
||||||
|
**状态**:提议中 | 已接受 | 已废弃
|
||||||
|
**日期**:YYYY-MM-DD
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
是什么情况迫使我们做这个决策?
|
||||||
|
|
||||||
|
## 决策
|
||||||
|
|
||||||
|
我们选择了什么?
|
||||||
|
|
||||||
|
## 替代方案
|
||||||
|
|
||||||
|
| 方案 | 拒绝原因 |
|
||||||
|
| --- | --- |
|
||||||
|
| {option} | {reason} |
|
||||||
|
|
||||||
|
## 后果
|
||||||
|
|
||||||
|
### 正面
|
||||||
|
|
||||||
|
- {benefit}
|
||||||
|
|
||||||
|
### 负面
|
||||||
|
|
||||||
|
- {cost or risk}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 技术调研模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 技术调研
|
||||||
|
|
||||||
|
## 摘要
|
||||||
|
|
||||||
|
- 变更原因:{reason}
|
||||||
|
- 变更范围:{scope}
|
||||||
|
- 主要技术方案:{approach}
|
||||||
|
|
||||||
|
## 源产物
|
||||||
|
|
||||||
|
- OpenSpec change: `openspec/changes/{slug}/`
|
||||||
|
- 关联 PRD: `prd.md` 或 `brief.md`
|
||||||
|
|
||||||
|
## 关键发现
|
||||||
|
|
||||||
|
- {finding}
|
||||||
|
|
||||||
|
## 假设
|
||||||
|
|
||||||
|
- {assumption and validation status}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 设计模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 设计
|
||||||
|
|
||||||
|
## 架构摘要
|
||||||
|
|
||||||
|
描述输入 → 处理 → 输出。
|
||||||
|
|
||||||
|
## 关键决策
|
||||||
|
|
||||||
|
- {decision}: {reason}
|
||||||
|
|
||||||
|
## 模块地图
|
||||||
|
|
||||||
|
| 模块 | 职责 | 备注 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| {module} | {responsibility} | {notes} |
|
||||||
|
|
||||||
|
## 架构审计
|
||||||
|
|
||||||
|
- 风险:{risk}
|
||||||
|
- 缓解:{mitigation}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 任务模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 任务
|
||||||
|
|
||||||
|
## 需求追踪
|
||||||
|
|
||||||
|
| 需求 | 状态 | 备注 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| {requirement} | 已完成 / 待处理 / 部分完成 | {notes} |
|
||||||
|
|
||||||
|
## 实现任务
|
||||||
|
|
||||||
|
- [ ] {task}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 验收模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题} 验收
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受 / 部分接受 / 未接受。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- 命令/检查:`{command or check}`
|
||||||
|
- 结果:{passed/failed/not run}
|
||||||
|
- 备注:{important output or reason not run}
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 命令:`{command}`
|
||||||
|
- 结果:{passed/failed/not run}
|
||||||
|
- 备注:{important output or reason not run}
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 步骤:{manual steps}
|
||||||
|
- 结果:{passed/failed/not run}
|
||||||
|
- 备注:{observations or reason not run}
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- {completed behavior}
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- {limitation}
|
||||||
|
|
||||||
|
## Bug 修复和诊断
|
||||||
|
|
||||||
|
- {bug}: {diagnosis summary and regression coverage}
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:{archive, deploy, review, or follow-up}
|
||||||
|
- OpenSpec 归档确认:{已询问/用户确认归档/用户暂不归档/不适用}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Cross-Artifact 对齐检查表模板
|
||||||
|
|
||||||
|
specify 阶段的 checkpoint 必须包含此检查表。每项标记"已对齐"或"存在 gap"。
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
## Cross-Artifact 对齐检查
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/prd → proposal | 目标、范围、非目标、验收预期是否进入 proposal | 已对齐 / 存在 gap |
|
||||||
|
| proposal → 设计产物 | 范围、约束、关键承诺是否进入 design.md 或等价设计小节 | 已对齐 / 存在 gap |
|
||||||
|
| 设计产物 → specs/tasks | 影响实现的约束、接口影响、架构结论是否进入 specs 或 tasks | 已对齐 / 存在 gap |
|
||||||
|
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 / 存在 gap |
|
||||||
|
|
||||||
|
### Gap 详情(如有)
|
||||||
|
|
||||||
|
- gap 1:{描述哪个字段/约束/行为/切片只停留在上游,未进入下游}
|
||||||
|
- 修复:{如何修正 OpenSpec}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 复合知识模板
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# {标题}
|
||||||
|
|
||||||
|
**类型**:learning | trick | decision | explore
|
||||||
|
**日期**:YYYY-MM-DD
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
这条经验来自哪里?
|
||||||
|
|
||||||
|
## 经验
|
||||||
|
|
||||||
|
未来代理应该复用什么经验?
|
||||||
|
|
||||||
|
## 适用性
|
||||||
|
|
||||||
|
什么时候适用?什么时候不适用?
|
||||||
|
```
|
||||||
@@ -4,6 +4,7 @@
|
|||||||
|
|
||||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||||
|
|||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Acceptance: diagnosis-playbook-skills
|
||||||
|
|
||||||
|
## Implemented
|
||||||
|
|
||||||
|
- Added six classpath diagnosis playbook skills under `src/main/resources/skills/`.
|
||||||
|
- Added classpath skill catalog loading and full skill reading.
|
||||||
|
- Added `read_skill` as a Spring AI tool.
|
||||||
|
- Wired skill catalog and tool into Chat and AIOps Planner/Executor paths.
|
||||||
|
- Kept Chat Verifier isolated from skills.
|
||||||
|
- Added focused tests for skill loading, unknown skill handling, and method tool injection.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Static/unit verification:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q "-Dtest=SkillCatalogServiceTest,ChatServiceSequentialAgentTest,AiOpsServiceTest,DiagnosisTraceEvaluatorTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
Result: passed.
|
||||||
|
|
||||||
|
## Not Verified
|
||||||
|
|
||||||
|
- Live LLM behavior with actual `read_skill` tool calls was not run.
|
||||||
|
- Java2AI `SkillsAgentHook` integration was not attempted because dependency classes were not locally verified.
|
||||||
|
|
||||||
|
## Archive Status
|
||||||
|
|
||||||
|
OpenSpec change not archived yet. User should confirm whether to archive `diagnosis-playbook-skills`.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Brief: diagnosis-playbook-skills
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
The MVP Agent has trace persistence, evidence tools, verifier gates, and fixed diagnosis eval cases, but scenario-specific diagnosis procedures were still embedded in broad prompts and knowledge-base documents.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Introduce versionable diagnosis playbook skills that agents can discover from a compact catalog and read on demand through a `read_skill` tool.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Six initial diagnosis playbooks: payment timeout, MySQL connection pool, Redis timeout, slow response, JVM memory risk, and AIOps alert.
|
||||||
|
- Classpath skill catalog loader.
|
||||||
|
- `read_skill` Spring AI tool.
|
||||||
|
- Prompt catalog injection for Chat and AIOps Planner/Executor paths.
|
||||||
|
- Focused tests.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- No external API or database schema changes.
|
||||||
|
- No replacement of `lookup_knowledge`.
|
||||||
|
- No Verifier skill loading.
|
||||||
|
- No direct dependency on Java2AI `SkillsAgentHook` until local package names are verified.
|
||||||
|
|
||||||
|
## Scale
|
||||||
|
|
||||||
|
standard
|
||||||
|
|
||||||
|
## OpenSpec
|
||||||
|
|
||||||
|
`openspec/changes/diagnosis-playbook-skills/`
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Decisions: diagnosis-playbook-skills
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Use "diagnosis playbook skills" as the canonical term: skill is the runtime loading unit, playbook is the diagnosis workflow content.
|
||||||
|
- Keep factual knowledge in `knowledge_base/`; skills contain workflow, evidence order, stop conditions, and report rules.
|
||||||
|
- Implement a project-local progressive disclosure mechanism first because local Maven cache did not confirm Java2AI skill hook package names.
|
||||||
|
- Keep Verifier unchanged so it only validates existing tool evidence.
|
||||||
|
- Do not persist `read_skill` as evidence in `tool_invocation`; diagnostic facts must still come from evidence tools.
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
L2 internal interface:
|
||||||
|
|
||||||
|
- New `SkillCatalogService`.
|
||||||
|
- New `ReadSkillTool`.
|
||||||
|
- Chat/AIOps internal method tools include `read_skill`.
|
||||||
|
- No HTTP, DTO, database, or external response contract changes.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
- `mvn -q "-Dtest=SkillCatalogServiceTest,ChatServiceSequentialAgentTest,AiOpsServiceTest,DiagnosisTraceEvaluatorTest" test`
|
||||||
|
- Result: passed.
|
||||||
@@ -159,7 +159,35 @@ ToolCallbackProvider
|
|||||||
- retrieved domains。
|
- retrieved domains。
|
||||||
- dedup reason。
|
- dedup reason。
|
||||||
|
|
||||||
## 6. 与旧版设计的差异
|
## 6. Skill / Playbook 流程
|
||||||
|
|
||||||
|
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||||
|
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||||
|
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||||
|
|
||||||
|
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||||
|
ExecutorHook --> ReadSkill["read_skill"]
|
||||||
|
Plan --> Executor["Executor"]
|
||||||
|
Executor --> ReadSkill
|
||||||
|
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||||
|
SkillBody --> Executor
|
||||||
|
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||||
|
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||||
|
Executor --> Verifier["Verifier"]
|
||||||
|
ToolTrace --> Verifier
|
||||||
|
```
|
||||||
|
|
||||||
|
| 角色 | Skill 可见性 | 工具权限 |
|
||||||
|
|---|---|---|
|
||||||
|
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||||
|
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||||
|
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` |
|
||||||
|
|
||||||
|
## 7. 与旧版设计的差异
|
||||||
|
|
||||||
| 旧版设想 | 当前实现 |
|
| 旧版设想 | 当前实现 |
|
||||||
|---|---|
|
|---|---|
|
||||||
@@ -167,9 +195,9 @@ ToolCallbackProvider
|
|||||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||||
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
|
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||||
|
|
||||||
## 7. 后续演进
|
## 8. 后续演进
|
||||||
|
|
||||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||||
|
|
||||||
@@ -184,4 +212,3 @@ ToolCallbackProvider
|
|||||||
- 不同故障类型的工具权限明显不同。
|
- 不同故障类型的工具权限明显不同。
|
||||||
- Trace 能证明某类问题需要独立的推理策略。
|
- Trace 能证明某类问题需要独立的推理策略。
|
||||||
- 评测集能覆盖拆分前后的行为差异。
|
- 评测集能覆盖拆分前后的行为差异。
|
||||||
|
|
||||||
|
|||||||
@@ -48,6 +48,13 @@ flowchart TB
|
|||||||
AlertsTool["queryPrometheusAlerts"]
|
AlertsTool["queryPrometheusAlerts"]
|
||||||
end
|
end
|
||||||
|
|
||||||
|
subgraph Skills["Skill / Playbook"]
|
||||||
|
SkillRegistry["SkillRegistry"]
|
||||||
|
PlannerSkillHook["PlannerSkillMetadataHook"]
|
||||||
|
SkillsHook["SkillsAgentHook"]
|
||||||
|
ReadSkill["read_skill"]
|
||||||
|
end
|
||||||
|
|
||||||
subgraph RAG["RAG Retrieval"]
|
subgraph RAG["RAG Retrieval"]
|
||||||
L0["KnowledgeIndexService"]
|
L0["KnowledgeIndexService"]
|
||||||
VectorSearch["VectorSearchService"]
|
VectorSearch["VectorSearchService"]
|
||||||
@@ -66,6 +73,11 @@ flowchart TB
|
|||||||
API --> App
|
API --> App
|
||||||
ChatService --> Agent
|
ChatService --> Agent
|
||||||
AiOpsService --> Agent
|
AiOpsService --> Agent
|
||||||
|
SkillRegistry --> PlannerSkillHook
|
||||||
|
PlannerSkillHook --> Planner
|
||||||
|
SkillRegistry --> SkillsHook
|
||||||
|
SkillsHook --> Executor
|
||||||
|
Executor --> ReadSkill
|
||||||
Agent --> Tools
|
Agent --> Tools
|
||||||
KnowledgeTool --> RAG
|
KnowledgeTool --> RAG
|
||||||
RAG --> Store
|
RAG --> Store
|
||||||
@@ -101,6 +113,12 @@ Evidence Tools
|
|||||||
-> query_metrics
|
-> query_metrics
|
||||||
-> queryPrometheusAlerts
|
-> queryPrometheusAlerts
|
||||||
|
|
||||||
|
Skill / Playbook
|
||||||
|
-> SkillRegistry
|
||||||
|
-> PlannerSkillMetadataHook gives Planner name/description only
|
||||||
|
-> SkillsAgentHook gives Executor read_skill
|
||||||
|
-> Verifier is isolated from skills
|
||||||
|
|
||||||
RAG Retrieval
|
RAG Retrieval
|
||||||
-> KnowledgeIndexService
|
-> KnowledgeIndexService
|
||||||
-> VectorSearchService
|
-> VectorSearchService
|
||||||
|
|||||||
@@ -0,0 +1,2 @@
|
|||||||
|
committed: true
|
||||||
|
date: 2026-07-05
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
# Decisions: diagnosis-playbook-skills
|
||||||
|
|
||||||
|
## Discover
|
||||||
|
|
||||||
|
- Entry summary: extract high-frequency diagnosis workflows into versionable skills/playbooks and make agents load them progressively.
|
||||||
|
- Slug: `diagnosis-playbook-skills`.
|
||||||
|
- Scale: `standard`.
|
||||||
|
- Capability source: sm-flow built-in protocol for Discover; `grill-with-docs` evidence-driven behavior used by reading project docs and code instead of blocking on user questions.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `devflow/index.md` hit related projects: `diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`, `evidence-trace-hardening`, `aiops-alert-scope-control`, `session-dedup-knowledge-map`, `executor-action-memory-relevance`.
|
||||||
|
- `devflow/glossary/CONTEXT.md` confirms `ReactAgent`, `ToolCall`, `DiagnosisRecord`, and current historical terminology; current architecture documents supersede old `diagnosis_record` as the primary model.
|
||||||
|
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, fallback-capable playbooks.
|
||||||
|
- `mvp/architecture/harness-quality-gates.md` requires Prompt contract, Tool boundary, Trace persistence, Verifier/rule evaluation, and eval baselines to remain authoritative.
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json` anchors initial playbooks: payment timeout, MySQL pool exhausted, Redis timeout, slow response, JVM memory risk.
|
||||||
|
|
||||||
|
## Grill Question Pool
|
||||||
|
|
||||||
|
| Dimension | Question | Mode | Resolution |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Terminology | Should these artifacts be called Skill or Playbook? | evidence-driven | Use "diagnosis playbook skills": skills are the runtime mechanism, playbooks are the diagnosis workflow content. |
|
||||||
|
| Boundary | Should skills contain factual knowledge or workflow guidance? | evidence-driven | Skills contain workflow guidance; knowledge facts remain in `knowledge_base/`. |
|
||||||
|
| Acceptance | What proves the change works? | evidence-driven | Unit tests for catalog/tool behavior plus existing diagnosis eval compile/test stability. |
|
||||||
|
| Interface impact | Does this alter external API or DB contracts? | evidence-driven | No external API/DB change; L2 internal interface due new tool/service and agent methodTools change. |
|
||||||
|
|
||||||
|
No user-interview question is blocking because the user explicitly asked to implement the change and prior conversation already confirmed the intended direction.
|
||||||
|
|
||||||
|
## Impact Analysis
|
||||||
|
|
||||||
|
GitNexus MCP tools are not exposed in this environment, so required GitNexus impact analysis could not be run. Local substitute analysis:
|
||||||
|
|
||||||
|
- `ChatService.createReactAgent`, `buildChatPlannerAgent`, `buildChatExecutorAgent`, and `buildMethodToolsArray` are called by Chat controller paths and covered by `ChatServiceSequentialAgentTest` / smoke tests.
|
||||||
|
- `AiOpsService.buildPlannerAgent`, `buildExecutorAgent`, and private `buildMethodToolsArray` affect `POST /api/ai_ops` via `ChatController`.
|
||||||
|
- Risk level: medium. Prompt and tool availability changes may alter agent behavior, but no external API, DTO, DB, or status contract changes.
|
||||||
|
|
||||||
|
## Specify / Commit
|
||||||
|
|
||||||
|
- Cross-artifact alignment:
|
||||||
|
- brief/proposal goal -> proposal: aligned.
|
||||||
|
- proposal scope -> design: aligned.
|
||||||
|
- design decisions -> specs/tasks: aligned.
|
||||||
|
- specs observable behavior -> tasks: aligned.
|
||||||
|
- Interface impact: L2 internal interface.
|
||||||
|
- Commit status: `.committed` created after file integrity and consistency checks.
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
- Reference code:
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||||
|
- Tool pattern: Spring AI Alibaba `SkillsAgentHook` contributes the official `read_skill` ToolCallback from hooks.
|
||||||
|
- Prompt pattern: `SkillsInterceptor` from the hook augments model requests with compact skill metadata.
|
||||||
|
- Test pattern: service tests instantiate classes manually with `ReflectionTestUtils`; new dependencies must be injectable or optional enough for tests.
|
||||||
|
- Dependency check: local `1.1.0.0-RC2` jars do not contain `SkillsAgentHook`; `1.1.2.0` jars contain `com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook`, `ReadSkillTool`, `SkillRegistry`, and `ClasspathSkillRegistry`.
|
||||||
|
|
||||||
|
## Apply
|
||||||
|
|
||||||
|
- Capability source: `openspec-apply-change` guidance was loaded; implementation used local sm-flow/OpenSpec fallback because the work required direct file edits and the OpenSpec CLI was not needed for artifact discovery.
|
||||||
|
- Implemented `src/main/resources/skills/*/SKILL.md` for six diagnosis playbooks.
|
||||||
|
- Upgraded Spring AI Alibaba BOMs to `1.1.2.0`.
|
||||||
|
- Implemented `SkillConfig` with `ClasspathSkillRegistry` loading `classpath:skills`.
|
||||||
|
- Wired `SkillsAgentHook` into Chat single-agent, Chat Planner/Executor, and AIOps Planner/Executor agents.
|
||||||
|
- Removed custom prompt-catalog injection from active service paths; the official `SkillsInterceptor` now handles skill catalog injection.
|
||||||
|
- Kept Chat Verifier prompt unchanged.
|
||||||
|
- Removed the earlier local fallback `SkillCatalogService` and `ReadSkillTool` source files after switching to official Alibaba skills support.
|
||||||
|
- Verification command:
|
||||||
|
- `mvn -q "-Dtest=SkillCatalogServiceTest,ChatServiceSequentialAgentTest,AiOpsServiceTest,DiagnosisTraceEvaluatorTest" test`
|
||||||
|
- Verification result: passed.
|
||||||
@@ -0,0 +1,82 @@
|
|||||||
|
# Design
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
src/main/resources/skills/
|
||||||
|
-> SKILL.md files
|
||||||
|
-> ClasspathSkillRegistry bean
|
||||||
|
-> loads classpath skills
|
||||||
|
-> backs official read_skill
|
||||||
|
-> PlannerSkillMetadataHook
|
||||||
|
-> adds planner-only skill metadata messages
|
||||||
|
-> does not expose read_skill
|
||||||
|
-> SkillsAgentHook
|
||||||
|
-> adds official read_skill ToolCallback for Executor / single-agent Chat
|
||||||
|
-> adds SkillsInterceptor prompt augmentation outside Planner
|
||||||
|
-> ChatService / AiOpsService
|
||||||
|
-> Planner receives metadata only
|
||||||
|
-> Executor and single-agent Chat receive official skill hook
|
||||||
|
-> Verifier remains isolated
|
||||||
|
```
|
||||||
|
|
||||||
|
## Skill Contract
|
||||||
|
|
||||||
|
Each skill folder contains a `SKILL.md` with YAML frontmatter:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
---
|
||||||
|
name: diagnose-mysql-connection-pool
|
||||||
|
description: ...
|
||||||
|
---
|
||||||
|
```
|
||||||
|
|
||||||
|
The body contains:
|
||||||
|
|
||||||
|
- Trigger conditions.
|
||||||
|
- Required evidence.
|
||||||
|
- Recommended tool order.
|
||||||
|
- Query construction hints.
|
||||||
|
- Stop conditions and low-confidence behavior.
|
||||||
|
- Report requirements.
|
||||||
|
- Eval anchor when one exists.
|
||||||
|
|
||||||
|
## Prompt Injection
|
||||||
|
|
||||||
|
Planner agents receive a project-local `PlannerSkillMetadataHook` message that contains skill names and descriptions only. The message also requires `selected_skill`, `selection_reason`, and an ordered `plan` in the Planner output.
|
||||||
|
|
||||||
|
`SkillsAgentHook` provides `SkillsInterceptor`, which injects the official compact skill section containing skill names, descriptions, and loading instructions into model requests for:
|
||||||
|
|
||||||
|
- Chat Executor prompt.
|
||||||
|
- AIOps Executor prompt.
|
||||||
|
- Single-agent Chat prompt.
|
||||||
|
|
||||||
|
Planner prompts are not augmented by `SkillsAgentHook`, so Planner cannot receive the official `read_skill` tool. Verifier prompt is not augmented.
|
||||||
|
|
||||||
|
## Tool Exposure
|
||||||
|
|
||||||
|
Spring AI Alibaba's official `ReadSkillTool` exposes:
|
||||||
|
|
||||||
|
```java
|
||||||
|
read_skill(skill_name)
|
||||||
|
```
|
||||||
|
|
||||||
|
The tool returns the full `SKILL.md` body for a known skill or a structured error for missing skills.
|
||||||
|
|
||||||
|
`read_skill` is supplied by `SkillsAgentHook`, not by local `methodTools`. Planner selects a skill from metadata and writes the selection into `planner_plan`; Executor reads the selected skill before executing scenario-specific evidence collection.
|
||||||
|
|
||||||
|
## Trace Behavior
|
||||||
|
|
||||||
|
`read_skill` is a guidance tool, not an evidence tool. It does not write `tool_invocation` because the existing eval and verifier treat evidence tools as factual data sources. Actual diagnostic evidence must still come from `lookup_knowledge`, `query_logs`, `query_metrics`, and alert tools.
|
||||||
|
|
||||||
|
## Fallback
|
||||||
|
|
||||||
|
If a skill is not found or cannot be read:
|
||||||
|
|
||||||
|
- The tool returns a structured text error.
|
||||||
|
- The agent must fall back to generic Executor prompt behavior.
|
||||||
|
- It must not invent playbook content.
|
||||||
|
|
||||||
|
## Alibaba Skills Integration
|
||||||
|
|
||||||
|
The implementation uses `spring-ai-alibaba-agent-framework:1.1.2.0`, where `SkillsAgentHook` lives in `com.alibaba.cloud.ai.graph.agent.hook.skills` and `ClasspathSkillRegistry` lives in `com.alibaba.cloud.ai.graph.skills.registry.classpath`.
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
# Diagnosis Playbook Skills
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The MVP diagnosis Agent already has trace persistence, evidence tools, verifier gates, and fixed eval cases, but scenario-specific diagnosis workflows still live in broad prompts and knowledge-base documents. This makes high-frequency fault diagnosis depend too much on the generic Executor prompt and makes it harder to version, review, and reuse diagnostic procedures.
|
||||||
|
|
||||||
|
## Proposed Solution
|
||||||
|
|
||||||
|
Introduce project-local diagnosis playbook skills using progressive disclosure:
|
||||||
|
|
||||||
|
- Store versionable playbook skills under `src/main/resources/skills/`.
|
||||||
|
- Use Spring AI Alibaba `SkillRegistry` + `SkillsAgentHook` so Planner/Executor agents can load full skill instructions only when a matching diagnosis scenario appears.
|
||||||
|
- Let the official skills interceptor inject the compact skill catalog into eligible agent prompts.
|
||||||
|
- Keep knowledge facts in `knowledge_base/`; skills define workflow, evidence requirements, stop conditions, and report rules.
|
||||||
|
- Keep Verifier isolated from skills. It must continue to validate only existing tool evidence.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
In scope:
|
||||||
|
|
||||||
|
- Payment timeout diagnosis playbook.
|
||||||
|
- MySQL connection pool diagnosis playbook.
|
||||||
|
- Redis timeout diagnosis playbook.
|
||||||
|
- Slow response diagnosis playbook.
|
||||||
|
- JVM memory risk diagnosis playbook.
|
||||||
|
- AIOps alert diagnosis playbook.
|
||||||
|
- Classpath skill registry configuration.
|
||||||
|
- Chat and AIOps Planner/Executor `SkillsAgentHook` wiring.
|
||||||
|
- Focused tests for skill loading/catalog behavior and existing diagnosis eval stability.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Replacing `lookup_knowledge` with implicit advisor retrieval.
|
||||||
|
- Replacing the Chat Verifier contract.
|
||||||
|
- Persisting a new database field for playbook usage.
|
||||||
|
- Creating SubAgents for each playbook.
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- `mvp/architecture/evolution-roadmap.md` defines Skill/Playbook as P1 and requires eval-backed, traceable, fallback-capable playbooks.
|
||||||
|
- `mvp/architecture/harness-quality-gates.md` requires evidence tool calls, trace persistence, verifier/rule evaluation, and eval baselines to remain authoritative.
|
||||||
|
- `knowledge_base/` remains the source for factual definitions and troubleshooting knowledge.
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json` provides the first fixed diagnosis scenarios and evidence-tool expectations.
|
||||||
|
- Spring AI Alibaba `1.1.2.0` provides `SkillsAgentHook`, `ClasspathSkillRegistry`, and the official `read_skill` tool.
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
L2 internal interface:
|
||||||
|
|
||||||
|
- Adds an internal `SkillRegistry` bean backed by classpath `skills`.
|
||||||
|
- Adds `SkillsAgentHook` to Chat/AIOps Planner and Executor agents.
|
||||||
|
- Does not change HTTP API, DTOs, database schema, or external response contracts.
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- The hook adds the official `read_skill` tool to eligible agents and may affect tool selection.
|
||||||
|
- Skill instructions could conflict with existing prompt constraints if not scoped carefully.
|
||||||
|
- Tests that instantiate `ChatService` manually must inject or tolerate the new skill tool dependency.
|
||||||
+56
@@ -0,0 +1,56 @@
|
|||||||
|
# diagnosis-playbook-skills Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
|
||||||
|
|
||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
|
||||||
|
|
||||||
|
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
|
||||||
|
|
||||||
|
#### Scenario: Planner or Executor receives available skill metadata
|
||||||
|
|
||||||
|
- **GIVEN** classpath skill folders exist under `skills/`
|
||||||
|
- **WHEN** Chat or AIOps Planner/Executor agents are built
|
||||||
|
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
|
||||||
|
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
|
||||||
|
|
||||||
|
### Requirement: Executor SHALL read full playbook instructions on demand
|
||||||
|
|
||||||
|
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
|
||||||
|
|
||||||
|
#### Scenario: Executor reads an existing skill
|
||||||
|
|
||||||
|
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
|
||||||
|
- **WHEN** the Executor calls `read_skill` with that name
|
||||||
|
- **THEN** the tool SHALL return the full skill instructions
|
||||||
|
- **AND** the result SHALL include the skill name
|
||||||
|
|
||||||
|
#### Scenario: Executor requests an unknown skill
|
||||||
|
|
||||||
|
- **WHEN** the Executor calls `read_skill` with an unknown name
|
||||||
|
- **THEN** the tool SHALL return a bounded error message
|
||||||
|
- **AND** the message SHALL list valid skill names
|
||||||
|
|
||||||
|
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
|
||||||
|
|
||||||
|
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
|
||||||
|
|
||||||
|
#### Scenario: Executor uses a playbook
|
||||||
|
|
||||||
|
- **WHEN** a diagnosis playbook applies to a user issue
|
||||||
|
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
|
||||||
|
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
|
||||||
|
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
|
||||||
|
|
||||||
|
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
|
||||||
|
|
||||||
|
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
|
||||||
|
|
||||||
|
#### Scenario: Fixed diagnosis case has a matching playbook
|
||||||
|
|
||||||
|
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
|
||||||
|
- **THEN** a matching diagnosis skill SHALL exist
|
||||||
|
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
# Tasks
|
||||||
|
|
||||||
|
- [x] 1. Add diagnosis skill resources under `src/main/resources/skills/`.
|
||||||
|
- [x] 2. Upgrade Spring AI Alibaba to a version that provides `SkillsAgentHook`.
|
||||||
|
- [x] 3. Add a `ClasspathSkillRegistry` bean for classpath skill resources.
|
||||||
|
- [x] 4. Wire `SkillsAgentHook` into Chat/AIOps single-agent, Planner, and Executor agents while keeping Verifier unchanged.
|
||||||
|
- [x] 5. Add focused unit tests for registry loading and official `read_skill` behavior.
|
||||||
|
- [x] 6. Run focused compile/tests and update this task list.
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
# diagnosis-playbook-skills Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
|
||||||
|
|
||||||
|
## Requirements
|
||||||
|
|
||||||
|
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
|
||||||
|
|
||||||
|
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
|
||||||
|
|
||||||
|
#### Scenario: Planner or Executor receives available skill metadata
|
||||||
|
|
||||||
|
- **GIVEN** classpath skill folders exist under `skills/`
|
||||||
|
- **WHEN** Chat or AIOps Planner/Executor agents are built
|
||||||
|
- **THEN** their system prompts SHALL include a compact diagnosis skill catalog
|
||||||
|
- **AND** the catalog SHALL include skill names and descriptions only, not full skill bodies
|
||||||
|
|
||||||
|
### Requirement: Executor SHALL read full playbook instructions on demand
|
||||||
|
|
||||||
|
The system SHALL expose a `read_skill` tool to Executor agents for loading a full `SKILL.md` body by skill name.
|
||||||
|
|
||||||
|
#### Scenario: Executor reads an existing skill
|
||||||
|
|
||||||
|
- **GIVEN** a skill named `diagnose-mysql-connection-pool`
|
||||||
|
- **WHEN** the Executor calls `read_skill` with that name
|
||||||
|
- **THEN** the tool SHALL return the full skill instructions
|
||||||
|
- **AND** the result SHALL include the skill name
|
||||||
|
|
||||||
|
#### Scenario: Executor requests an unknown skill
|
||||||
|
|
||||||
|
- **WHEN** the Executor calls `read_skill` with an unknown name
|
||||||
|
- **THEN** the tool SHALL return a bounded error message
|
||||||
|
- **AND** the message SHALL list valid skill names
|
||||||
|
|
||||||
|
### Requirement: Playbook skills SHALL preserve evidence and verifier boundaries
|
||||||
|
|
||||||
|
The system SHALL keep skills as workflow guidance and keep factual evidence collection in existing evidence tools.
|
||||||
|
|
||||||
|
#### Scenario: Executor uses a playbook
|
||||||
|
|
||||||
|
- **WHEN** a diagnosis playbook applies to a user issue
|
||||||
|
- **THEN** the Executor SHALL use the playbook to decide evidence order and stop conditions
|
||||||
|
- **AND** factual claims SHALL still be supported by `lookup_knowledge`, `query_logs`, `query_metrics`, or alert tools
|
||||||
|
- **AND** Chat Verifier SHALL continue to validate only existing `tool_trace_summary`
|
||||||
|
|
||||||
|
### Requirement: Initial playbook set SHALL cover fixed MVP diagnosis cases
|
||||||
|
|
||||||
|
The system SHALL provide playbooks for the existing fixed diagnosis evaluation scenarios.
|
||||||
|
|
||||||
|
#### Scenario: Fixed diagnosis case has a matching playbook
|
||||||
|
|
||||||
|
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
|
||||||
|
- **THEN** a matching diagnosis skill SHALL exist
|
||||||
|
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
|
||||||
@@ -20,8 +20,8 @@
|
|||||||
<maven.compiler.target>17</maven.compiler.target>
|
<maven.compiler.target>17</maven.compiler.target>
|
||||||
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
|
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
|
||||||
<spring-ai.version>1.1.7</spring-ai.version>
|
<spring-ai.version>1.1.7</spring-ai.version>
|
||||||
<spring-ai-alibaba.version>1.1.0.0-RC2</spring-ai-alibaba.version>
|
<spring-ai-alibaba.version>1.1.2.0</spring-ai-alibaba.version>
|
||||||
<spring-ai-alibaba-extensions.version>1.1.0.0-RC2</spring-ai-alibaba-extensions.version>
|
<spring-ai-alibaba-extensions.version>1.1.2.0</spring-ai-alibaba-extensions.version>
|
||||||
</properties>
|
</properties>
|
||||||
<dependencyManagement>
|
<dependencyManagement>
|
||||||
<dependencies>
|
<dependencies>
|
||||||
|
|||||||
@@ -0,0 +1,83 @@
|
|||||||
|
package com.superbiz.agent.config;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.SkillMetadata;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.classpath.ClasspathSkillRegistry;
|
||||||
|
import org.springframework.ai.chat.prompt.SystemPromptTemplate;
|
||||||
|
import org.springframework.context.annotation.Bean;
|
||||||
|
import org.springframework.context.annotation.Configuration;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Optional;
|
||||||
|
|
||||||
|
@Configuration
|
||||||
|
public class SkillConfig {
|
||||||
|
|
||||||
|
private static final String ACTIVE_SKILL_NAME = "diagnose-mysql-connection-pool";
|
||||||
|
|
||||||
|
@Bean
|
||||||
|
public SkillRegistry skillRegistry() {
|
||||||
|
SkillRegistry classpathRegistry = ClasspathSkillRegistry.builder()
|
||||||
|
.classpathPath("skills")
|
||||||
|
.basePath("target/skills-cache")
|
||||||
|
.build();
|
||||||
|
return new SingleSkillRegistry(classpathRegistry, ACTIVE_SKILL_NAME);
|
||||||
|
}
|
||||||
|
|
||||||
|
private record SingleSkillRegistry(SkillRegistry delegate, String activeSkillName) implements SkillRegistry {
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public List<SkillMetadata> listAll() {
|
||||||
|
return delegate.listAll().stream()
|
||||||
|
.filter(skill -> activeSkillName.equals(skill.getName()))
|
||||||
|
.toList();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getRegistryType() {
|
||||||
|
return delegate.getRegistryType();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String readSkillContent(String skillName) throws IOException {
|
||||||
|
if (!activeSkillName.equals(skillName)) {
|
||||||
|
throw new IOException("Skill not found: " + skillName);
|
||||||
|
}
|
||||||
|
return delegate.readSkillContent(skillName);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getSkillLoadInstructions() {
|
||||||
|
return delegate.getSkillLoadInstructions();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public SystemPromptTemplate getSystemPromptTemplate() {
|
||||||
|
return delegate.getSystemPromptTemplate();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public Optional<SkillMetadata> get(String skillName) {
|
||||||
|
if (!activeSkillName.equals(skillName)) {
|
||||||
|
return Optional.empty();
|
||||||
|
}
|
||||||
|
return delegate.get(skillName);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int size() {
|
||||||
|
return listAll().size();
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public boolean contains(String skillName) {
|
||||||
|
return activeSkillName.equals(skillName) && delegate.contains(skillName);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public void reload() {
|
||||||
|
delegate.reload();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
package com.superbiz.agent.hook;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.Prioritized;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.SkillMetadata;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import lombok.extern.slf4j.Slf4j;
|
||||||
|
import org.springframework.ai.chat.messages.Message;
|
||||||
|
import org.springframework.ai.chat.messages.SystemMessage;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Injects planner-visible skill metadata without exposing the full skill loader tool.
|
||||||
|
*/
|
||||||
|
@Slf4j
|
||||||
|
@HookPositions(HookPosition.BEFORE_MODEL)
|
||||||
|
public class PlannerSkillMetadataHook extends MessagesModelHook {
|
||||||
|
|
||||||
|
private static final String CATALOG_MARKER = "\"skill_catalog\"";
|
||||||
|
|
||||||
|
private final SkillRegistry skillRegistry;
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
|
public PlannerSkillMetadataHook(SkillRegistry skillRegistry) {
|
||||||
|
this.skillRegistry = skillRegistry;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getName() {
|
||||||
|
return "planner_skill_metadata_hook";
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public int getOrder() {
|
||||||
|
return Prioritized.HIGHEST_PRECEDENCE;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||||
|
if (skillRegistry == null || skillRegistry.size() == 0 || hasCatalog(previousMessages)) {
|
||||||
|
return new AgentCommand(previousMessages);
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
List<Map<String, String>> skills = skillRegistry.listAll().stream()
|
||||||
|
.map(this::toSkillSummary)
|
||||||
|
.toList();
|
||||||
|
if (skills.isEmpty()) {
|
||||||
|
return new AgentCommand(previousMessages);
|
||||||
|
}
|
||||||
|
|
||||||
|
Map<String, Object> catalog = new LinkedHashMap<>();
|
||||||
|
catalog.put("purpose", "Planner-visible diagnosis skill metadata only.");
|
||||||
|
catalog.put("rules", List.of(
|
||||||
|
"Choose at most one primary skill.",
|
||||||
|
"Do not load full skill instructions in Planner.",
|
||||||
|
"Executor reads the selected skill before evidence collection.",
|
||||||
|
"If no skill matches, set selected_skill to null."
|
||||||
|
));
|
||||||
|
catalog.put("skills", skills);
|
||||||
|
catalog.put("required_planner_output", Map.of(
|
||||||
|
"selected_skill", "skill name or null",
|
||||||
|
"selection_reason", "short reason",
|
||||||
|
"plan", "ordered execution step list"
|
||||||
|
));
|
||||||
|
|
||||||
|
Map<String, Object> payload = Map.of("skill_catalog", catalog);
|
||||||
|
String content = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(payload);
|
||||||
|
|
||||||
|
List<Message> updatedMessages = new ArrayList<>(previousMessages.size() + 1);
|
||||||
|
updatedMessages.add(new SystemMessage(content));
|
||||||
|
updatedMessages.addAll(previousMessages);
|
||||||
|
return new AgentCommand(updatedMessages);
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.warn("Failed to inject planner skill metadata, fallback to original messages", e);
|
||||||
|
return new AgentCommand(previousMessages);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, String> toSkillSummary(SkillMetadata skill) {
|
||||||
|
Map<String, String> summary = new LinkedHashMap<>();
|
||||||
|
summary.put("name", skill.getName());
|
||||||
|
summary.put("description", skill.getDescription());
|
||||||
|
return summary;
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean hasCatalog(List<Message> messages) {
|
||||||
|
return messages.stream()
|
||||||
|
.map(Message::getText)
|
||||||
|
.anyMatch(text -> text != null && text.contains(CATALOG_MARKER));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -4,7 +4,10 @@ import org.springframework.ai.chat.model.ChatModel;
|
|||||||
import com.alibaba.cloud.ai.graph.OverAllState;
|
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook;
|
||||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||||
@@ -13,6 +16,7 @@ import com.superbiz.agent.domain.entity.AgentStep;
|
|||||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
import com.superbiz.agent.dto.AIOpsRequest;
|
import com.superbiz.agent.dto.AIOpsRequest;
|
||||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||||
|
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||||
import com.superbiz.agent.repository.AgentStepRepository;
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
@@ -32,8 +36,8 @@ import java.util.Optional;
|
|||||||
import java.util.UUID;
|
import java.util.UUID;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* AI Ops 智能运维服务
|
* AI Ops 闂備礁鎼幊妯肩磽濮樿泛绀傛俊顖滅帛娴溿倝鏌熼柇锕€鏋熸俊顖氾躬閺岋繝宕煎┑鍩裤垹鈹?
|
||||||
* 负责多 Agent 协作的告警分析流程
|
* 闂佽崵濮甸崝妤呭窗閺囥垺鍎楁俊銈勭缁?Agent 闂備礁鎲¢〃鍛崲鐎n剛绀婇柡鍐ㄧ墛閸庡秹鏌涢弴銊ヤ簼闁哥喓鍋ら幃褰掑焵椤掑嫭鏅濋柍褜鍓熷畷瑙勬償閵娿儱鍤戦梺褰掑亰閸橀箖濡堕敂鍓х<?
|
||||||
*/
|
*/
|
||||||
@Service
|
@Service
|
||||||
public class AiOpsService {
|
public class AiOpsService {
|
||||||
@@ -49,7 +53,7 @@ public class AiOpsService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private QueryMetricsTools queryMetricsTools;
|
private QueryMetricsTools queryMetricsTools;
|
||||||
|
|
||||||
@Autowired(required = false) // Mock 模式下才注册
|
@Autowired(required = false) // Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒佺箾閹寸偞灏い鎴濇閺呭爼鎮╁ù瀣亙闂侀潧顭堥崕閬嶅棘閳?
|
||||||
private QueryLogsTools queryLogsTools;
|
private QueryLogsTools queryLogsTools;
|
||||||
|
|
||||||
@Autowired
|
@Autowired
|
||||||
@@ -73,13 +77,16 @@ public class AiOpsService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private SelfEvaluationMergeService selfEvaluationMergeService;
|
private SelfEvaluationMergeService selfEvaluationMergeService;
|
||||||
|
|
||||||
|
@Autowired(required = false)
|
||||||
|
private SkillRegistry skillRegistry;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 执行 AI Ops 告警分析流程
|
* 闂備礁婀遍悷鎶藉幢閳哄倹鏉?AI Ops 闂備礁鎲$粙鎴︽晝閵娾晩鏁嗛柣鏃傚帶缁€鍡涙煕閳╁喚鐒介柍褜鍓濆Λ鍕箒婵炶揪缍€閵嗏偓闁?
|
||||||
*
|
*
|
||||||
* @param chatModel 大模型实例
|
* @param chatModel 濠电姰鍨归悥銏ゅ礋閳ь剚绗熼埀顒€鐣烽崷顓涘亾閿濆簼绨介柡澶庢閵?
|
||||||
* @param toolCallbacks 工具回调数组
|
* @param toolCallbacks 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柛褎顨呴悙濠囨煟閹邦剙顣虫繛鍫濈埣閺屸剝鎷呴悷閭︽缂?
|
||||||
* @return 分析结果状态
|
* @return 闂備礁鎲$敮鎺懳涘▎鎾村€甸柣锝呯灱绾惧ジ鏌熼幆褜鍤熷ù婊庡灦閺岋絽螣閸喚鍘梺?
|
||||||
* @throws GraphRunnerException 如果 Agent 执行失败
|
* @throws GraphRunnerException 濠电姷顣介埀顒€鍟块埀顒€缍婇幃?Agent 闂備礁婀遍悷鎶藉幢閳哄倹鏉稿┑鐘灪閸庤偐鍒掗崜褎鍠?
|
||||||
*/
|
*/
|
||||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
||||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
||||||
@@ -87,27 +94,27 @@ public class AiOpsService {
|
|||||||
|
|
||||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||||
AIOpsRequest request, String sessionId) throws GraphRunnerException {
|
AIOpsRequest request, String sessionId) throws GraphRunnerException {
|
||||||
logger.info("开始执行 AI Ops 多 Agent 协作流程");
|
logger.info("Starting AI Ops multi-agent analysis");
|
||||||
|
|
||||||
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
||||||
long startTime = System.currentTimeMillis();
|
long startTime = System.currentTimeMillis();
|
||||||
|
|
||||||
// 创建或更新诊断会话
|
// 闂備礁鎲$敮妤冪矙閹寸姷纾介柟鎹愵嚙缁狅綁鏌″鍐ㄥ缂佺虎鍨堕弻锟犲磼濞戞﹩鈧粓鏌i敂鐣屽⒌鐎殿噮鍓熼、妯衡攽閸垻宕堕梺?
|
||||||
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
||||||
diagnosisSessionRepository.save(session);
|
diagnosisSessionRepository.save(session);
|
||||||
|
|
||||||
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
// 闂佽崵濮崇粈浣规櫠娴犲鍋?ThreadLocal 濠电偞鍨堕幐鎼佹晝閿濆洨绠旈柛娑欐綑濡﹢鏌涢妷鈺婃缂佲偓閸戯箰okupKnowledgeTool 闂傚倷绶¢崑鍛┍閾忚宕查柛鎰电厛濞间即鏌曢崼婵堝缂佺媭鍨堕弻?sessionId闂?
|
||||||
SessionContextHolder.setSessionId(resolvedSessionId);
|
SessionContextHolder.setSessionId(resolvedSessionId);
|
||||||
|
|
||||||
try {
|
try {
|
||||||
// 构建 Planner 和 Executor Agent(每个 Agent 各自带 Hook)
|
// 闂備礁鎼鍛偓姘煎墰缁?Planner 闂?Executor Agent闂備焦瀵х粙鎴︽偋婵犲洦鍎婇柍鈺佸暟閳?Agent 闂備礁鎲¢懝鍓р偓姘煎灦瀹曢潧顭ㄩ崨顔芥?Hook闂?
|
||||||
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
|
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
|
||||||
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
|
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
|
||||||
|
|
||||||
// 构建 Supervisor Agent(不加 Hook)
|
// 闂備礁鎼鍛偓姘煎墰缁?Supervisor Agent闂備焦瀵х粙鎴︽偋閸涱垳绠斿鑸靛姇缁€?Hook闂?
|
||||||
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
|
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
|
||||||
.name("ai_ops_supervisor")
|
.name("ai_ops_supervisor")
|
||||||
.description("负责调度 Planner 与 Executor 的多 Agent 控制器")
|
.description("Coordinates Planner and Executor agents")
|
||||||
.model(chatModel)
|
.model(chatModel)
|
||||||
.systemPrompt(promptProperties.getSupervisor())
|
.systemPrompt(promptProperties.getSupervisor())
|
||||||
.subAgents(List.of(plannerAgent, executorAgent))
|
.subAgents(List.of(plannerAgent, executorAgent))
|
||||||
@@ -115,19 +122,19 @@ public class AiOpsService {
|
|||||||
|
|
||||||
String taskPrompt = buildTaskPrompt(request);
|
String taskPrompt = buildTaskPrompt(request);
|
||||||
|
|
||||||
logger.info("调用 Supervisor Agent 开始编排...");
|
logger.info("闂佽崵濮撮鍛村疮娴兼潙鏋?Supervisor Agent 闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢垹绱掓鏍﹂偗妤?..");
|
||||||
|
|
||||||
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
|
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
|
||||||
|
|
||||||
long duration = System.currentTimeMillis() - startTime;
|
long duration = System.currentTimeMillis() - startTime;
|
||||||
|
|
||||||
// 更新诊断会话
|
// 闂備礁鎼ú銈夋偤閵娾晛钃熷┑鐘插鐎氭艾鈹戦悩鎻掓殲闁绘帟妫勯湁闁稿繘妫挎禍銏ゆ煟?
|
||||||
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
||||||
session.setTotalDurationMs((int) duration);
|
session.setTotalDurationMs((int) duration);
|
||||||
backfillSessionMetrics(session);
|
backfillSessionMetrics(session);
|
||||||
diagnosisSessionRepository.save(session);
|
diagnosisSessionRepository.save(session);
|
||||||
|
|
||||||
// 添加调试代码
|
// 婵犵數鍎戠紞鈧い鏇嗗嫭鍙忛柣鎰仛鐎氼剟鏌涢幇闈涘箻婵¤尙顭堥湁闁绘瑥鎳愰幃濂告煟?
|
||||||
if (stateOptional.isPresent()) {
|
if (stateOptional.isPresent()) {
|
||||||
OverAllState state = stateOptional.get();
|
OverAllState state = stateOptional.get();
|
||||||
logger.debug("Final State Keys: {}", state.data().keySet());
|
logger.debug("Final State Keys: {}", state.data().keySet());
|
||||||
@@ -146,25 +153,25 @@ public class AiOpsService {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 从执行结果中提取最终报告文本
|
* 濠电偛顕慨瀵糕偓娑掓櫆閺呭爼鎮剧仦鎯т粧閻庡厜鍋撻柍褜鍓涢崚鎺楀Ω閳轰礁鍤戝┑鐘才堥崑鎾绘煠閸偄鐏存鐐存崌楠炲洭顢楅埀顒傚緤閸ф鐓涢柛顐h壘娴滃墽绱撻崒娆戭槮闁绘锕ラ幈銊╁Χ婢跺﹤绐涙繝鐢靛Т閸燁垶鎮楅鈧弻?
|
||||||
*
|
*
|
||||||
* @param state 执行状态
|
* @param state 闂備礁婀遍悷鎶藉幢閳哄倹鏉搁梻浣虹帛椤牓宕洪弽顓炵劦?
|
||||||
* @return 报告文本(如果存在)
|
* @return 闂備胶顢婄紙浼村磿闁秴绠熼柨鐔哄Т濡﹢鏌涢妷锝呭闁圭兘浜堕弻銊モ槈濡厧顣哄銈傛暘閸パ冨殤濠电姴锕ら崯浼村箺閻樼粯鐓曢柨鏃囧吹閸樻粎绱?
|
||||||
*/
|
*/
|
||||||
public Optional<String> extractFinalReport(OverAllState state) {
|
public Optional<String> extractFinalReport(OverAllState state) {
|
||||||
logger.info("开始提取最终报告...");
|
logger.info("闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢劗鎲告0浣虹獢鐎规洩缍佸浠嬪Ω閿旇法甯涚紓鍌氬€风粈渚€鎮ф繝鍐╁弿闁靛牆顦?..");
|
||||||
|
|
||||||
// 提取 Planner 最终输出(包含完整的告警分析报告)
|
// 闂備礁婀辩划顖炲礉閺嚶颁汗?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢э絿绱掓0婵嗗籍鐎规洘鐟╅幃顔锯偓闈涙憸椤︹晠姊洪崨濠勫ⅹ闁瑰啿閰i獮鍡涘醇閳垛晛浜鹃柣鐔哄濠€浼存煛閸☆厾绉柟顖氬暣瀹曠喖顢曢敐鍛畼闂佽崵濮崑鎾绘煥閺囨浜鹃梺鎼炲妼闁帮絽顕i幖浣哥疀妞ゆ挾鍊幘缁樼厱婵炴垶锕╅悡顓犵磼?
|
||||||
Optional<AssistantMessage> plannerFinalOutput = state.value("planner_plan")
|
Optional<AssistantMessage> plannerFinalOutput = state.value("planner_plan")
|
||||||
.filter(AssistantMessage.class::isInstance)
|
.filter(AssistantMessage.class::isInstance)
|
||||||
.map(AssistantMessage.class::cast);
|
.map(AssistantMessage.class::cast);
|
||||||
|
|
||||||
if (plannerFinalOutput.isPresent()) {
|
if (plannerFinalOutput.isPresent()) {
|
||||||
String reportText = plannerFinalOutput.get().getText();
|
String reportText = plannerFinalOutput.get().getText();
|
||||||
logger.info("成功提取到 Planner 最终报告,长度: {}", reportText.length());
|
logger.info("闂備胶鎳撻悺銊╁礉閺囩喐鍙忔繛鎴欏灩缁犵敻鏌熼柇锕€澧紒鎻掓健閺?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢ф稑鈹戦鍝勨偓婵嗙暦閵婏妇绡€闊洦娲滈ˇ顕€姊婚崒妤€浜鹃梺鍓茬厛閸犳牠顢? {}", reportText.length());
|
||||||
return Optional.of(reportText);
|
return Optional.of(reportText);
|
||||||
} else {
|
} else {
|
||||||
logger.warn("未能提取到 Planner 最终报告");
|
logger.warn("Unable to extract Planner final report");
|
||||||
return Optional.empty();
|
return Optional.empty();
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -197,16 +204,16 @@ public class AiOpsService {
|
|||||||
|
|
||||||
String buildQuerySummary(AIOpsRequest request) {
|
String buildQuerySummary(AIOpsRequest request) {
|
||||||
if (request == null) {
|
if (request == null) {
|
||||||
return "AI Ops 告警分析";
|
return "AI Ops alert analysis";
|
||||||
}
|
}
|
||||||
|
|
||||||
StringBuilder summary = new StringBuilder("AI Ops 告警分析");
|
StringBuilder summary = new StringBuilder("AI Ops alert analysis");
|
||||||
appendField(summary, "告警", request.getAlertName());
|
appendField(summary, "alert", request.getAlertName());
|
||||||
appendField(summary, "服务", request.getService());
|
appendField(summary, "service", request.getService());
|
||||||
appendField(summary, "等级", request.getSeverity());
|
appendField(summary, "severity", request.getSeverity());
|
||||||
appendField(summary, "时间范围", request.getTimeRange());
|
appendField(summary, "timeRange", request.getTimeRange());
|
||||||
appendField(summary, "描述", request.getDescription());
|
appendField(summary, "description", request.getDescription());
|
||||||
appendField(summary, "请求", request.getUserRequest());
|
appendField(summary, "request", request.getUserRequest());
|
||||||
return summary.toString();
|
return summary.toString();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -238,8 +245,8 @@ public class AiOpsService {
|
|||||||
|
|
||||||
String buildTaskPrompt(AIOpsRequest request) {
|
String buildTaskPrompt(AIOpsRequest request) {
|
||||||
StringBuilder prompt = new StringBuilder();
|
StringBuilder prompt = new StringBuilder();
|
||||||
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
|
prompt.append("You are an enterprise SRE handling an automated alert diagnosis task. Combine tool evidence, run a plan-execute-replan loop, and output the final alert analysis report. Do not fabricate data; if repeated queries fail, clearly state why the task cannot be completed.");
|
||||||
prompt.append("\n\n本次告警输入:\n");
|
prompt.append("\n\nAlert input:\n");
|
||||||
prompt.append(buildQuerySummary(request));
|
prompt.append(buildQuerySummary(request));
|
||||||
if (hasAlertPayload(request)) {
|
if (hasAlertPayload(request)) {
|
||||||
String knowledgeQuery = buildKnowledgeRetrievalQuery(request);
|
String knowledgeQuery = buildKnowledgeRetrievalQuery(request);
|
||||||
@@ -278,53 +285,69 @@ public class AiOpsService {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 构建 Planner Agent
|
* 闂備礁鎼鍛偓姘煎墰缁?Planner Agent
|
||||||
*/
|
*/
|
||||||
private ReactAgent buildPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
private ReactAgent buildPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||||
return ReactAgent.builder()
|
return ReactAgent.builder()
|
||||||
.name("planner_agent")
|
.name("planner_agent")
|
||||||
.description("负责拆解告警、规划与再规划步骤")
|
.description("Plans alert diagnosis steps")
|
||||||
.model(chatModel)
|
.model(chatModel)
|
||||||
.systemPrompt(promptProperties.getPlanner())
|
.systemPrompt(promptProperties.getPlanner())
|
||||||
.methodTools(buildMethodToolsArray())
|
.hooks(buildHooks("planner"))
|
||||||
.tools(toolCallbacks)
|
|
||||||
.hooks(new AgentLoggingHook(agentStepRepository, "planner"))
|
|
||||||
.outputKey("planner_plan")
|
.outputKey("planner_plan")
|
||||||
.build();
|
.build();
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 构建 Executor Agent
|
* 闂備礁鎼鍛偓姘煎墰缁?Executor Agent
|
||||||
*/
|
*/
|
||||||
private ReactAgent buildExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
private ReactAgent buildExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||||
return ReactAgent.builder()
|
return ReactAgent.builder()
|
||||||
.name("executor_agent")
|
.name("executor_agent")
|
||||||
.description("负责执行 Planner 的首个步骤并及时反馈")
|
.description("Executes the current Planner step and reports feedback")
|
||||||
.model(chatModel)
|
.model(chatModel)
|
||||||
.systemPrompt(promptProperties.getExecutor())
|
.systemPrompt(promptProperties.getExecutor())
|
||||||
.methodTools(buildMethodToolsArray())
|
.methodTools(buildMethodToolsArray())
|
||||||
.tools(toolCallbacks)
|
.tools(toolCallbacks)
|
||||||
.hooks(new AgentLoggingHook(agentStepRepository, "executor"))
|
.hooks(buildHooks("executor"))
|
||||||
.outputKey("executor_feedback")
|
.outputKey("executor_feedback")
|
||||||
.build();
|
.build();
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 动态构建方法工具数组
|
* 闂備礁鎲¢弻锝夊礉瀹ュ鐒垫い鎴f硶閸斿秹鏌f惔顔肩仩妞ゆ洘鐟╅幃婊兾熼懡銈呭箥婵犵數鍋涢ˇ鏉棵洪弽銊ヮ嚤闁圭増婢樼粈鍌炴⒑閸噮鍎愭繛鍫濆缁?
|
||||||
* 根据 cls.mock-enabled 决定是否包含 QueryLogsTools
|
* 闂備礁鎼粔鐑斤綖婢跺﹦鏆?cls.mock-enabled 闂備礁鎲¢崝鏇㈠疮閸ф鍋╁Δ锝呭暙閸欏﹥銇勯弽銊ь暡闁稿骸锕弻娑㈠冀瑜庨崳褰掓煙?QueryLogsTools
|
||||||
* 工具顺序:知识库查询优先,日志查询次之,弃用工具最后
|
* 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柟绋跨昂娴滄粓鏌涘┑鍡楊伀缁炬澘绉归弻銊モ槈濞嗘劗娈ら梺缁樻惈缁辨洟骞忛悩璇插耿婵°倕鍟惃鎴︽⒑閸濆嫯顫﹂柛搴㈡尦椤㈡艾螖娴e壊鍤ゅ┑鈽嗗灠閹碱偆鏁妷鈺傜叆婵炴垶蓱濠€鐗堜繆椤愮喐娅堢紒鐘崇☉铻栧ù锝呮惈瀵劑鏌i悩鍙夊偍闁搞劍妞介、鏇㈠礂閼测斁鏋欓柣搴到婢у海绮堟径灞稿亾濞堝灝鏋涢柛鐔跺嵆瀵偊濡堕崪浣告櫊闂侀潧顦崕鍗烆嚗閺冨牊鐓涢柛顐h壘娴滈箖姊?
|
||||||
*/
|
*/
|
||||||
private Object[] buildMethodToolsArray() {
|
private Object[] buildMethodToolsArray() {
|
||||||
if (queryLogsTools != null) {
|
if (queryLogsTools != null) {
|
||||||
// Mock 模式:包含 QueryLogsTools
|
// Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒勬⒑閹稿海鈯曢柤鐟板⒔閳ь剙鐏氶敃銏犵暦?QueryLogsTools
|
||||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools, queryLogsTools};
|
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools, queryLogsTools};
|
||||||
} else {
|
}
|
||||||
// 真实模式:不包含 QueryLogsTools(由 MCP 提供日志查询功能)
|
// Real mode excludes local QueryLogsTools because logs are provided by MCP.
|
||||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private Hook[] buildHooks(String agentName) {
|
||||||
|
AgentLoggingHook loggingHook = new AgentLoggingHook(agentStepRepository, agentName);
|
||||||
|
if (skillRegistry == null || skillRegistry.size() == 0) {
|
||||||
|
return new Hook[]{loggingHook};
|
||||||
|
}
|
||||||
|
if ("planner".equals(agentName)) {
|
||||||
|
return new Hook[]{
|
||||||
|
new PlannerSkillMetadataHook(skillRegistry),
|
||||||
|
loggingHook
|
||||||
|
};
|
||||||
|
}
|
||||||
|
return new Hook[]{
|
||||||
|
loggingHook,
|
||||||
|
SkillsAgentHook.builder()
|
||||||
|
.skillRegistry(skillRegistry)
|
||||||
|
.build()
|
||||||
|
};
|
||||||
}
|
}
|
||||||
|
|
||||||
/** 从 agent_step 和 tool_invocation 汇总指标回填 diagnosis_session */
|
/** 濠?agent_step 闂?tool_invocation 婵犳鍠氶幊鎾趁洪敃鍌氱劦妞ゆ帒鍊荤敮娑㈡倵閸偄鍝虹€殿喕绮欏畷鎯邦槼缂佲偓閳ь剚绻?diagnosis_session */
|
||||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||||
try {
|
try {
|
||||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||||
@@ -340,7 +363,7 @@ public class AiOpsService {
|
|||||||
session.setStepCount(stepCount);
|
session.setStepCount(stepCount);
|
||||||
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||||
} catch (Exception e) {
|
} catch (Exception e) {
|
||||||
logger.warn("回填会话指标失败: sessionId={}", session.getSessionId(), e);
|
logger.warn("闂備焦鎮堕崕鎶藉磻濞戙垺鏅查柣鎰綑椤曡鲸鎱ㄥΟ铏癸紞婵☆垰鐗撻弻鐔虹矙閹稿骸顦╅梺缁樼壄缁叉儳顕ラ崟顒佺秶妞ゆ劑鍎? sessionId={}", session.getSessionId(), e);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -4,7 +4,10 @@ import com.alibaba.cloud.ai.graph.OverAllState;
|
|||||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SequentialAgent;
|
import com.alibaba.cloud.ai.graph.agent.flow.agent.SequentialAgent;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.skills.SkillsAgentHook;
|
||||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
import com.fasterxml.jackson.databind.JsonNode;
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||||
@@ -13,6 +16,7 @@ import com.superbiz.agent.agent.tool.QueryLogsTools;
|
|||||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||||
|
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
|
||||||
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
||||||
import com.superbiz.agent.hook.TokenUsageHolder;
|
import com.superbiz.agent.hook.TokenUsageHolder;
|
||||||
import com.superbiz.agent.hook.VerifierInputHook;
|
import com.superbiz.agent.hook.VerifierInputHook;
|
||||||
@@ -99,6 +103,9 @@ public class ChatService {
|
|||||||
@Autowired
|
@Autowired
|
||||||
private KnowledgeDomainService knowledgeDomainService;
|
private KnowledgeDomainService knowledgeDomainService;
|
||||||
|
|
||||||
|
@Autowired(required = false)
|
||||||
|
private SkillRegistry skillRegistry;
|
||||||
|
|
||||||
@Autowired
|
@Autowired
|
||||||
private ToolTraceSummaryService toolTraceSummaryService;
|
private ToolTraceSummaryService toolTraceSummaryService;
|
||||||
|
|
||||||
@@ -160,6 +167,7 @@ public class ChatService {
|
|||||||
systemPromptBuilder.append("你是一个专业的智能助手,可以获取当前时间、查询天气信息、搜索内部文档知识库,以及查询 Prometheus 告警信息。\n");
|
systemPromptBuilder.append("你是一个专业的智能助手,可以获取当前时间、查询天气信息、搜索内部文档知识库,以及查询 Prometheus 告警信息。\n");
|
||||||
systemPromptBuilder.append("当用户询问时间相关问题时,**必须每次都调用 getCurrentDateTime 工具**,因为时间会不断变化。即使历史消息中有时间信息,也不要直接复用,必须重新查询最新时间。\n");
|
systemPromptBuilder.append("当用户询问时间相关问题时,**必须每次都调用 getCurrentDateTime 工具**,因为时间会不断变化。即使历史消息中有时间信息,也不要直接复用,必须重新查询最新时间。\n");
|
||||||
systemPromptBuilder.append("当用户需要查询公司内部文档、流程、最佳实践或技术指南时,使用 lookupKnowledgeTool 工具。\n");
|
systemPromptBuilder.append("当用户需要查询公司内部文档、流程、最佳实践或技术指南时,使用 lookupKnowledgeTool 工具。\n");
|
||||||
|
systemPromptBuilder.append("当用户的问题匹配某个诊断 Skill 时,先调用 read_skill 读取对应流程,再按流程调用证据工具。\n");
|
||||||
systemPromptBuilder.append("当用户需要查询 Prometheus 告警、监控指标或系统告警状态时,使用 queryPrometheusAlerts 工具。\n");
|
systemPromptBuilder.append("当用户需要查询 Prometheus 告警、监控指标或系统告警状态时,使用 queryPrometheusAlerts 工具。\n");
|
||||||
systemPromptBuilder.append("当用户需要查询腾讯云日志时,请调用腾讯云mcp服务查询,默认查询地域ap-guangzhou,查询时间范围为近一个月。\n\n");
|
systemPromptBuilder.append("当用户需要查询腾讯云日志时,请调用腾讯云mcp服务查询,默认查询地域ap-guangzhou,查询时间范围为近一个月。\n\n");
|
||||||
|
|
||||||
@@ -271,7 +279,7 @@ public class ChatService {
|
|||||||
.systemPrompt(systemPrompt)
|
.systemPrompt(systemPrompt)
|
||||||
.methodTools(buildMethodToolsArray())
|
.methodTools(buildMethodToolsArray())
|
||||||
.tools(getToolCallbacks())
|
.tools(getToolCallbacks())
|
||||||
.hooks(new AgentLoggingHook(agentStepRepository, "intelligent_assistant"))
|
.hooks(buildHooks("intelligent_assistant"))
|
||||||
.build();
|
.build();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -516,7 +524,7 @@ public class ChatService {
|
|||||||
.description("负责拆解问题、规划步骤")
|
.description("负责拆解问题、规划步骤")
|
||||||
.model(chatModel)
|
.model(chatModel)
|
||||||
.systemPrompt(prompt.toString())
|
.systemPrompt(prompt.toString())
|
||||||
.hooks(new AgentLoggingHook(agentStepRepository, "planner"))
|
.hooks(buildHooks("planner"))
|
||||||
.outputKey("planner_plan")
|
.outputKey("planner_plan")
|
||||||
.build();
|
.build();
|
||||||
}
|
}
|
||||||
@@ -553,11 +561,30 @@ public class ChatService {
|
|||||||
.systemPrompt(prompt.toString())
|
.systemPrompt(prompt.toString())
|
||||||
.methodTools(buildMethodToolsArray())
|
.methodTools(buildMethodToolsArray())
|
||||||
.tools(toolCallbacks)
|
.tools(toolCallbacks)
|
||||||
.hooks(new AgentLoggingHook(agentStepRepository, "executor"))
|
.hooks(buildHooks("executor"))
|
||||||
.outputKey("executor_feedback")
|
.outputKey("executor_feedback")
|
||||||
.build();
|
.build();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private Hook[] buildHooks(String agentName) {
|
||||||
|
AgentLoggingHook loggingHook = new AgentLoggingHook(agentStepRepository, agentName);
|
||||||
|
if (skillRegistry == null || skillRegistry.size() == 0) {
|
||||||
|
return new Hook[]{loggingHook};
|
||||||
|
}
|
||||||
|
if ("planner".equals(agentName)) {
|
||||||
|
return new Hook[]{
|
||||||
|
new PlannerSkillMetadataHook(skillRegistry),
|
||||||
|
loggingHook
|
||||||
|
};
|
||||||
|
}
|
||||||
|
return new Hook[]{
|
||||||
|
loggingHook,
|
||||||
|
SkillsAgentHook.builder()
|
||||||
|
.skillRegistry(skillRegistry)
|
||||||
|
.build()
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
private String resolveSessionId(String requestedSessionId) {
|
private String resolveSessionId(String requestedSessionId) {
|
||||||
if (requestedSessionId != null && !requestedSessionId.isBlank()) {
|
if (requestedSessionId != null && !requestedSessionId.isBlank()) {
|
||||||
return requestedSessionId;
|
return requestedSessionId;
|
||||||
|
|||||||
@@ -0,0 +1,38 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-aiops-alert
|
||||||
|
description: Diagnose AIOps alert payloads, active Prometheus alerts, alert scope control, HighCPUUsage, HighMemoryUsage, SlowResponse, ServiceUnavailable, and alert-driven incident reports. Use in AIOps flows or when the user asks to diagnose current alerts.
|
||||||
|
---
|
||||||
|
|
||||||
|
# AIOps Alert Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Determine scope mode.
|
||||||
|
- Payload present: treat the supplied alert as the primary diagnosis target.
|
||||||
|
- No payload: call `queryPrometheusAlerts` first and choose P0/P1 or the longest-running firing alert.
|
||||||
|
2. For payload mode, preserve alert name, service, severity, description, and time range in the `lookup_knowledge` query.
|
||||||
|
3. Confirm active alert state with `queryPrometheusAlerts` when useful, but do not diagnose unrelated alerts as the main target.
|
||||||
|
4. Query metrics/logs that match the alert type and service.
|
||||||
|
5. Produce a report that distinguishes confirmed evidence, related risks, and missing evidence.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- Alert state from payload or `queryPrometheusAlerts`.
|
||||||
|
- `lookup_knowledge` when playbook or runbook guidance is needed.
|
||||||
|
- Logs and metrics aligned to the alert type.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- If payload mode returns unrelated active alerts, mention them only as related risk.
|
||||||
|
- If three calls in the same direction fail or return no data, stop that direction and report the failure.
|
||||||
|
- Do not invent metric values, log lines, or remediation execution results.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- Use the existing alert analysis report structure.
|
||||||
|
- Keep the supplied alert as the main diagnosis target in payload mode.
|
||||||
|
- Include confidence and evidence gaps.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
RAG cases: `aiops-payment-latency-alert`, `aiops-prometheus-alert-scope`.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-jvm-memory-risk
|
||||||
|
description: Diagnose JVM memory risk, high heap usage, OOM risk, OutOfMemoryError, frequent Full GC, memory leak, pod OOMKilled, or order-service memory alerts. Use when memory, JVM, heap, GC, OOM, or OOMKilled appears.
|
||||||
|
---
|
||||||
|
|
||||||
|
# JVM Memory Risk Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Extract affected service, memory threshold, heap size, GC symptoms, pod/container events, and time window.
|
||||||
|
2. Call `query_metrics` or alert tools for heap usage, memory usage, GC count/time, and active memory alerts.
|
||||||
|
3. Call `query_logs` for Full GC warnings, OutOfMemoryError, OOMKilled, restart events, or allocation-heavy stack traces.
|
||||||
|
4. Call `lookup_knowledge` when JVM memory troubleshooting or remediation guidance is needed.
|
||||||
|
5. Decide whether the supported risk is high memory pressure, confirmed OOM, suspected leak, or insufficient evidence.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- `query_metrics` for resource pressure claims.
|
||||||
|
- `query_logs` for OOM, GC, or restart evidence.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- High memory usage alone is not proof of memory leak.
|
||||||
|
- OOM risk is stronger when high memory metrics align with Full GC, OOMKilled, or OutOfMemoryError logs.
|
||||||
|
- If evidence is incomplete, return LOW_CONFID wording and list the missing metrics/logs.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- Include immediate mitigation, heap/GC investigation, leak investigation, and monitoring recommendations.
|
||||||
|
- Do not say the issue can be ignored while memory remains above threshold.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
Fixed diagnosis case: `jvm-memory-risk`.
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-mysql-connection-pool
|
||||||
|
description: Diagnose MySQL, HikariCP, database connection pool exhaustion, connection acquisition timeout, slow SQL, connection leak, or database saturation issues. Use when the user mentions MySQL pool, HikariCP, connection pool, database timeout, order-service timeout, or connection exhaustion.
|
||||||
|
---
|
||||||
|
|
||||||
|
# MySQL Connection Pool Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Extract service, database, timeout symptom, and time window.
|
||||||
|
2. Call `lookup_knowledge` with MySQL, HikariCP, connection pool, and the affected service.
|
||||||
|
3. Call `query_logs` for connection acquisition timeout, active/max pool counts, waiting threads, leak warnings, slow query, or lock waits.
|
||||||
|
4. Call `query_metrics` when metrics are available for active connections, idle connections, wait time, DB latency, and error rate.
|
||||||
|
5. Decide whether the evidence supports pool exhaustion, slow SQL causing saturation, connection leak, or insufficient evidence.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- `lookup_knowledge` for pool configuration and diagnosis guidance.
|
||||||
|
- `query_logs` for concrete pool or SQL symptoms.
|
||||||
|
- `query_metrics` when making saturation or capacity claims.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- Confirmed pool exhaustion requires log or metric evidence such as active equals max, waiting threads, acquisition timeout, or leak warnings.
|
||||||
|
- If only request timeout is present without pool evidence, state that the pool hypothesis is unconfirmed.
|
||||||
|
- If logs show slow SQL but not pool saturation, report slow SQL as the stronger supported cause.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- Include current evidence, likely root cause, missing evidence, short-term mitigation, and long-term fix.
|
||||||
|
- Avoid saying "fully confirmed" unless at least two evidence sources align.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
Fixed diagnosis case: `mysql-pool-exhausted`.
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-payment-timeout
|
||||||
|
description: Diagnose payment API, payment gateway, ERR_TIMEOUT, gateway timeout, payment-service latency, or payment request timeout issues. Use when the user mentions payment timeout, ERR_TIMEOUT, ERR_GATEWAY_TIMEOUT, slow payment, or payment-service latency.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Payment Timeout Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Identify the affected payment service, error code, endpoint, and time window from the user request.
|
||||||
|
2. Call `read_skill` only once for this playbook, then follow the evidence order below.
|
||||||
|
3. Call `lookup_knowledge` with a narrow query containing payment, timeout, the error code if present, and the affected service.
|
||||||
|
4. Call `query_logs` for payment-service timeout, downstream dependency timeout, gateway timeout, or request duration above threshold.
|
||||||
|
5. Call `query_metrics` or alert tools for latency, error rate, saturation, and active alerts when metrics are available.
|
||||||
|
6. Compare knowledge guidance with logs and metrics before stating a root cause.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- `lookup_knowledge` for error-code or payment timeout guidance.
|
||||||
|
- `query_logs` for concrete timeout or dependency evidence.
|
||||||
|
- `query_metrics` when the question asks for impact, latency, or current alert state.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- If only knowledge is available and logs/metrics are missing, return LOW_CONFID language.
|
||||||
|
- If tools fail or return no evidence, state which evidence is missing and do not claim a confirmed root cause.
|
||||||
|
- Do not repeatedly call `lookup_knowledge` with synonym-only queries after a relevant result.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- Separate immediate mitigation from long-term remediation.
|
||||||
|
- Cite the evidence source type for each key conclusion.
|
||||||
|
- Do not claim payment provider failure unless logs or metrics support an upstream dependency issue.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
Fixed diagnosis case: `payment-timeout`.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-redis-timeout
|
||||||
|
description: Diagnose Redis timeout, Redis connection timeout, cache dependency timeout, Redis cluster unavailable, hot key, network latency, or payment-service Redis dependency failures. Use when Redis or cache timeout appears in the user request, logs, or alert payload.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Redis Timeout Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Extract affected service, Redis operation, host/cluster, timeout value, and time window.
|
||||||
|
2. Call `lookup_knowledge` for Redis timeout or cache troubleshooting guidance when knowledge evidence is needed.
|
||||||
|
3. Call `query_logs` for Redis connection timeout, retry count, host, command latency, hot key, or dependency errors.
|
||||||
|
4. Call `query_metrics` when available for Redis latency, connection count, CPU, memory, error rate, or network saturation.
|
||||||
|
5. Distinguish client timeout, Redis saturation, network issue, and missing evidence.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- `query_logs` is mandatory for a concrete Redis timeout claim.
|
||||||
|
- `lookup_knowledge` is recommended for remediation and configuration guidance.
|
||||||
|
- `query_metrics` is required before claiming Redis resource saturation.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- If only one Redis timeout log exists and no metrics are available, return LOW_CONFID wording.
|
||||||
|
- If Redis is only mentioned as a possible downstream dependency, do not make it the root cause without supporting logs.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- State whether the supported issue is client-side timeout, Redis cluster issue, network issue, or unconfirmed.
|
||||||
|
- Include retry/backoff, timeout tuning, connection pool, and monitoring recommendations only when relevant.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
Fixed diagnosis case: `redis-timeout`.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
---
|
||||||
|
name: diagnose-slow-response
|
||||||
|
description: Diagnose slow response, high P95/P99 latency, API latency regression, slow request, downstream latency, or user-service response time alerts. Use when the user mentions P99, P95, response time, slow endpoint, latency, or SlowResponse alerts.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Slow Response Diagnosis
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Extract service, endpoint, latency percentile, threshold, and time window.
|
||||||
|
2. Call `query_metrics` or alert tools to confirm latency and impact.
|
||||||
|
3. Call `query_logs` for slow request records, endpoint duration, downstream timing, cache misses, or database query timeout.
|
||||||
|
4. Call `lookup_knowledge` when process guidance, service-specific runbook, or known failure mode evidence is needed.
|
||||||
|
5. Classify the supported cause: database slow query, downstream dependency, cache miss, resource saturation, or insufficient evidence.
|
||||||
|
|
||||||
|
## Required Evidence
|
||||||
|
|
||||||
|
- `query_metrics` for latency or alert confirmation.
|
||||||
|
- `query_logs` for endpoint-level or dependency-level evidence.
|
||||||
|
|
||||||
|
## Stop Conditions
|
||||||
|
|
||||||
|
- If metrics show latency but logs do not identify a cause, say impact is confirmed but root cause is not.
|
||||||
|
- If logs identify slow SQL or dependency latency, use that as a candidate cause and mark confidence based on metric alignment.
|
||||||
|
|
||||||
|
## Report Rules
|
||||||
|
|
||||||
|
- Include impacted endpoints, observed latency, suspected bottleneck, evidence gaps, and next checks.
|
||||||
|
- Do not say there is no risk when P95/P99 remains above threshold.
|
||||||
|
|
||||||
|
## Eval Anchor
|
||||||
|
|
||||||
|
Fixed diagnosis case: `slow-response`.
|
||||||
@@ -56,17 +56,17 @@ class AiOpsServiceTest {
|
|||||||
request.setSeverity("P1");
|
request.setSeverity("P1");
|
||||||
request.setTimeRange("last_15m");
|
request.setTimeRange("last_15m");
|
||||||
request.setDescription("P95 latency is high");
|
request.setDescription("P95 latency is high");
|
||||||
request.setUserRequest("结合日志和指标排查支付超时");
|
request.setUserRequest("check logs and metrics for payment timeout");
|
||||||
|
|
||||||
String summary = service.buildQuerySummary(request);
|
String summary = service.buildQuerySummary(request);
|
||||||
|
|
||||||
assertTrue(summary.contains("AI Ops 告警分析"));
|
assertTrue(summary.contains("AI Ops alert analysis"));
|
||||||
assertTrue(summary.contains("告警: payment-service-latency-high"));
|
assertTrue(summary.contains("alert: payment-service-latency-high"));
|
||||||
assertTrue(summary.contains("服务: payment-service"));
|
assertTrue(summary.contains("service: payment-service"));
|
||||||
assertTrue(summary.contains("等级: P1"));
|
assertTrue(summary.contains("severity: P1"));
|
||||||
assertTrue(summary.contains("时间范围: last_15m"));
|
assertTrue(summary.contains("timeRange: last_15m"));
|
||||||
assertTrue(summary.contains("描述: P95 latency is high"));
|
assertTrue(summary.contains("description: P95 latency is high"));
|
||||||
assertTrue(summary.contains("请求: 结合日志和指标排查支付超时"));
|
assertTrue(summary.contains("request: "));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -99,8 +99,8 @@ class AiOpsServiceTest {
|
|||||||
assertTrue(prompt.contains("Related Risk"));
|
assertTrue(prompt.contains("Related Risk"));
|
||||||
assertTrue(prompt.contains("Recommended lookup_knowledge query: HighCPUUsage payment-service P1 CPU usage is above 80% last_15m"));
|
assertTrue(prompt.contains("Recommended lookup_knowledge query: HighCPUUsage payment-service P1 CPU usage is above 80% last_15m"));
|
||||||
assertTrue(prompt.contains("preserves alertName and service"));
|
assertTrue(prompt.contains("preserves alertName and service"));
|
||||||
assertTrue(prompt.contains("告警: HighCPUUsage"));
|
assertTrue(prompt.contains("alert: HighCPUUsage"));
|
||||||
assertTrue(prompt.contains("服务: payment-service"));
|
assertTrue(prompt.contains("service: payment-service"));
|
||||||
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -112,11 +112,11 @@ class AiOpsServiceTest {
|
|||||||
request.setSeverity(" ");
|
request.setSeverity(" ");
|
||||||
request.setDescription("P95 latency above threshold");
|
request.setDescription("P95 latency above threshold");
|
||||||
request.setTimeRange("last_10m");
|
request.setTimeRange("last_10m");
|
||||||
request.setUserRequest("结合日志和指标排查");
|
request.setUserRequest("check logs and metrics");
|
||||||
|
|
||||||
String query = service.buildKnowledgeRetrievalQuery(request);
|
String query = service.buildKnowledgeRetrievalQuery(request);
|
||||||
|
|
||||||
assertEquals("HighLatency payment-service P95 latency above threshold last_10m 结合日志和指标排查", query);
|
assertEquals("HighLatency payment-service P95 latency above threshold last_10m check logs and metrics", query);
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -142,16 +142,16 @@ class AiOpsServiceTest {
|
|||||||
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
|
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
|
||||||
DiagnosisSession session = DiagnosisSession.builder()
|
DiagnosisSession session = DiagnosisSession.builder()
|
||||||
.sessionId("aiops-session-001")
|
.sessionId("aiops-session-001")
|
||||||
.query("AI Ops 告警分析")
|
.query("AI Ops alert analysis")
|
||||||
.status("SUCCESS")
|
.status("SUCCESS")
|
||||||
.agentFlow("AI_OPS")
|
.agentFlow("AI_OPS")
|
||||||
.build();
|
.build();
|
||||||
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
|
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
|
||||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
|
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
|
||||||
|
|
||||||
service.persistFinalReport("aiops-session-001", "# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.");
|
service.persistFinalReport("aiops-session-001", "# 闁告稑锕ㄩ鐔煎礆閸℃鈧粙骞庨妷銉﹀暈\nHighCPUUsage payment-service analysis with evidence summary.");
|
||||||
|
|
||||||
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
|
assertEquals("# 闁告稑锕ㄩ鐔煎礆閸℃鈧粙骞庨妷銉﹀暈\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
|
||||||
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
|
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
|
||||||
verify(diagnosisSessionRepository).save(session);
|
verify(diagnosisSessionRepository).save(session);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,5 +1,8 @@
|
|||||||
package com.superbiz.agent.service;
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.classpath.ClasspathSkillRegistry;
|
||||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||||
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
import com.superbiz.agent.agent.tool.QueryMetricsTools;
|
||||||
@@ -194,6 +197,56 @@ class ChatServiceSequentialAgentTest {
|
|||||||
assertSame(queryMetricsTools, methodTools[3]);
|
assertSame(queryMetricsTools, methodTools[3]);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void createReactAgentInjectsSkillCatalogThroughAlibabaHook() throws Exception {
|
||||||
|
ChatService chatService = createChatService();
|
||||||
|
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||||
|
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
|
||||||
|
.classpathPath("skills")
|
||||||
|
.basePath("target/test-skills-cache")
|
||||||
|
.build();
|
||||||
|
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
|
||||||
|
|
||||||
|
ReactAgent agent = chatService.createReactAgent(chatModel, "BASE_TEST_PROMPT");
|
||||||
|
agent.call("diagnose mysql connection pool exhaustion");
|
||||||
|
|
||||||
|
assertTrue(chatModel.promptText.contains("BASE_TEST_PROMPT"));
|
||||||
|
assertTrue(chatModel.promptText.contains("## Skills System"));
|
||||||
|
assertTrue(chatModel.promptText.contains("diagnose-mysql-connection-pool"));
|
||||||
|
assertTrue(chatModel.promptText.contains("read_skill"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void plannerGetsSkillMetadataAndExecutorGetsReadSkillTool() throws Exception {
|
||||||
|
ChatService chatService = createChatService();
|
||||||
|
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||||
|
SkillRegistry skillRegistry = ClasspathSkillRegistry.builder()
|
||||||
|
.classpathPath("skills")
|
||||||
|
.basePath("target/test-skills-cache")
|
||||||
|
.build();
|
||||||
|
ReflectionTestUtils.setField(chatService, "skillRegistry", skillRegistry);
|
||||||
|
|
||||||
|
chatService.executeChatComplex(
|
||||||
|
chatModel,
|
||||||
|
new ToolCallback[0],
|
||||||
|
"diagnose mysql connection pool exhaustion",
|
||||||
|
List.of(),
|
||||||
|
"planner-skill-metadata-session"
|
||||||
|
);
|
||||||
|
|
||||||
|
assertTrue(chatModel.plannerPromptText.contains("\"skill_catalog\""));
|
||||||
|
assertTrue(chatModel.plannerPromptText.contains("diagnose-mysql-connection-pool"));
|
||||||
|
assertTrue(chatModel.plannerPromptText.contains("\"selected_skill\""));
|
||||||
|
assertFalse(chatModel.plannerPromptText.contains("## Skills System"));
|
||||||
|
assertFalse(chatModel.plannerPromptText.contains("read_skill"));
|
||||||
|
|
||||||
|
assertTrue(chatModel.executorPromptText.contains("## Skills System"));
|
||||||
|
assertTrue(chatModel.executorPromptText.contains("diagnose-mysql-connection-pool"));
|
||||||
|
assertTrue(chatModel.executorPromptText.contains("read_skill"));
|
||||||
|
assertFalse(chatModel.verifierPromptText.contains("diagnose-mysql-connection-pool"));
|
||||||
|
assertFalse(chatModel.verifierPromptText.contains("read_skill"));
|
||||||
|
}
|
||||||
|
|
||||||
private ChatService createChatService() {
|
private ChatService createChatService() {
|
||||||
ChatService chatService = new ChatService();
|
ChatService chatService = new ChatService();
|
||||||
|
|
||||||
@@ -245,6 +298,9 @@ class ChatServiceSequentialAgentTest {
|
|||||||
private static final class ScriptedChatModel implements ChatModel {
|
private static final class ScriptedChatModel implements ChatModel {
|
||||||
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
|
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
|
||||||
private String promptText = "";
|
private String promptText = "";
|
||||||
|
private String plannerPromptText = "";
|
||||||
|
private String executorPromptText = "";
|
||||||
|
private String verifierPromptText = "";
|
||||||
private boolean sawVerifierPrompt;
|
private boolean sawVerifierPrompt;
|
||||||
private final java.util.List<String> verifierOutputs;
|
private final java.util.List<String> verifierOutputs;
|
||||||
private int verifierOutputIndex;
|
private int verifierOutputIndex;
|
||||||
@@ -283,12 +339,15 @@ class ChatServiceSequentialAgentTest {
|
|||||||
String text;
|
String text;
|
||||||
if (promptText.contains("PLANNER_TEST_PROMPT")) {
|
if (promptText.contains("PLANNER_TEST_PROMPT")) {
|
||||||
agentCalls.add("chat_planner");
|
agentCalls.add("chat_planner");
|
||||||
|
plannerPromptText = promptText;
|
||||||
text = "PLANNER_PLAN";
|
text = "PLANNER_PLAN";
|
||||||
} else if (promptText.contains("EXECUTOR_TEST_PROMPT")) {
|
} else if (promptText.contains("EXECUTOR_TEST_PROMPT")) {
|
||||||
agentCalls.add("chat_executor");
|
agentCalls.add("chat_executor");
|
||||||
|
executorPromptText = promptText;
|
||||||
text = "EXECUTOR_FINAL_ANSWER";
|
text = "EXECUTOR_FINAL_ANSWER";
|
||||||
} else if (promptText.contains("VERIFIER_TEST_PROMPT")) {
|
} else if (promptText.contains("VERIFIER_TEST_PROMPT")) {
|
||||||
agentCalls.add("chat_verifier");
|
agentCalls.add("chat_verifier");
|
||||||
|
verifierPromptText = promptText;
|
||||||
sawVerifierPrompt = true;
|
sawVerifierPrompt = true;
|
||||||
int index = Math.min(verifierOutputIndex, verifierOutputs.size() - 1);
|
int index = Math.min(verifierOutputIndex, verifierOutputs.size() - 1);
|
||||||
text = verifierOutputs.get(index);
|
text = verifierOutputs.get(index);
|
||||||
|
|||||||
@@ -0,0 +1,44 @@
|
|||||||
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.skills.ReadSkillTool;
|
||||||
|
import com.alibaba.cloud.ai.graph.skills.registry.SkillRegistry;
|
||||||
|
import com.superbiz.agent.config.SkillConfig;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
class SkillCatalogServiceTest {
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void loadsDiagnosisSkillsFromClasspathRegistry() {
|
||||||
|
SkillRegistry registry = newRegistry();
|
||||||
|
|
||||||
|
assertEquals(1, registry.size());
|
||||||
|
assertTrue(registry.contains("diagnose-mysql-connection-pool"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void readSkillReturnsFullInstructionsFromOfficialTool() {
|
||||||
|
ReadSkillTool tool = new ReadSkillTool(newRegistry());
|
||||||
|
|
||||||
|
String skill = tool.apply(new ReadSkillTool.ReadSkillRequest("diagnose-mysql-connection-pool"), null);
|
||||||
|
|
||||||
|
assertTrue(skill.contains("## Workflow"));
|
||||||
|
assertTrue(skill.contains("query_logs"));
|
||||||
|
assertTrue(skill.contains("Fixed diagnosis case: `mysql-pool-exhausted`"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void readSkillToolReturnsUnknownSkillError() {
|
||||||
|
ReadSkillTool tool = new ReadSkillTool(newRegistry());
|
||||||
|
|
||||||
|
String result = tool.apply(new ReadSkillTool.ReadSkillRequest("missing-skill"), null);
|
||||||
|
|
||||||
|
assertTrue(result.contains("Skill not found: missing-skill"));
|
||||||
|
}
|
||||||
|
|
||||||
|
private SkillRegistry newRegistry() {
|
||||||
|
return new SkillConfig().skillRegistry();
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user