Reorganize workspace and archive skill artifacts
This commit is contained in:
@@ -0,0 +1,117 @@
|
||||
---
|
||||
name: diagnose
|
||||
description: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.
|
||||
---
|
||||
|
||||
# Diagnose
|
||||
|
||||
A discipline for hard bugs. Skip phases only when explicitly justified.
|
||||
|
||||
When exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
||||
|
||||
## Phase 1 — Build a feedback loop
|
||||
|
||||
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
|
||||
|
||||
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
||||
|
||||
### Ways to construct one — try them in roughly this order
|
||||
|
||||
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
||||
2. **Curl / HTTP script** against a running dev server.
|
||||
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
||||
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
|
||||
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
||||
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
||||
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
||||
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
||||
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
|
||||
Build the right feedback loop, and the bug is 90% fixed.
|
||||
|
||||
### Iterate on the loop itself
|
||||
|
||||
Treat the loop as a product. Once you have _a_ loop, ask:
|
||||
|
||||
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
||||
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
||||
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
||||
|
||||
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
|
||||
|
||||
### Non-deterministic bugs
|
||||
|
||||
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
|
||||
|
||||
### When you genuinely cannot build a loop
|
||||
|
||||
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
||||
|
||||
Do not proceed to Phase 2 until you have a loop you believe in.
|
||||
|
||||
## Phase 2 — Reproduce
|
||||
|
||||
Run the loop. Watch the bug appear.
|
||||
|
||||
Confirm:
|
||||
|
||||
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
||||
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
||||
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
||||
|
||||
Do not proceed until you reproduce the bug.
|
||||
|
||||
## Phase 3 — Hypothesise
|
||||
|
||||
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
||||
|
||||
Each hypothesis must be **falsifiable**: state the prediction it makes.
|
||||
|
||||
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
||||
|
||||
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
||||
|
||||
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
|
||||
|
||||
## Phase 4 — Instrument
|
||||
|
||||
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
||||
|
||||
Tool preference:
|
||||
|
||||
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
|
||||
2. **Targeted logs** at the boundaries that distinguish hypotheses.
|
||||
3. Never "log everything and grep".
|
||||
|
||||
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
||||
|
||||
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
||||
|
||||
## Phase 5 — Fix + regression test
|
||||
|
||||
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
||||
|
||||
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
|
||||
|
||||
If a correct seam exists:
|
||||
|
||||
1. Turn the minimised repro into a failing test at that seam.
|
||||
2. Watch it fail.
|
||||
3. Apply the fix.
|
||||
4. Watch it pass.
|
||||
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
||||
|
||||
## Phase 6 — Cleanup + post-mortem
|
||||
|
||||
Required before declaring done:
|
||||
|
||||
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
|
||||
- [ ] Regression test passes (or absence of seam is documented)
|
||||
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
||||
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
||||
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
|
||||
|
||||
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
||||
@@ -0,0 +1,41 @@
|
||||
#!/usr/bin/env bash
|
||||
# Human-in-the-loop reproduction loop.
|
||||
# Copy this file, edit the steps below, and run it.
|
||||
# The agent runs the script; the user follows prompts in their terminal.
|
||||
#
|
||||
# Usage:
|
||||
# bash hitl-loop.template.sh
|
||||
#
|
||||
# Two helpers:
|
||||
# step "<instruction>" → show instruction, wait for Enter
|
||||
# capture VAR "<question>" → show question, read response into VAR
|
||||
#
|
||||
# At the end, captured values are printed as KEY=VALUE for the agent to parse.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
step() {
|
||||
printf '\n>>> %s\n' "$1"
|
||||
read -r -p " [Enter when done] " _
|
||||
}
|
||||
|
||||
capture() {
|
||||
local var="$1" question="$2" answer
|
||||
printf '\n>>> %s\n' "$question"
|
||||
read -r -p " > " answer
|
||||
printf -v "$var" '%s' "$answer"
|
||||
}
|
||||
|
||||
# --- edit below ---------------------------------------------------------
|
||||
|
||||
step "Open the app at http://localhost:3000 and sign in."
|
||||
|
||||
capture ERRORED "Click the 'Export' button. Did it throw an error? (y/n)"
|
||||
|
||||
capture ERROR_MSG "Paste the error message (or 'none'):"
|
||||
|
||||
# --- edit above ---------------------------------------------------------
|
||||
|
||||
printf '\n--- Captured ---\n'
|
||||
printf 'ERRORED=%s\n' "$ERRORED"
|
||||
printf 'ERROR_MSG=%s\n' "$ERROR_MSG"
|
||||
@@ -0,0 +1,47 @@
|
||||
# ADR Format
|
||||
|
||||
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
|
||||
|
||||
Create the `docs/adr/` directory lazily — only when the first ADR is needed.
|
||||
|
||||
## Template
|
||||
|
||||
```md
|
||||
# {Short title of the decision}
|
||||
|
||||
{1-3 sentences: what's the context, what did we decide, and why.}
|
||||
```
|
||||
|
||||
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
|
||||
|
||||
## Optional sections
|
||||
|
||||
Only include these when they add genuine value. Most ADRs won't need them.
|
||||
|
||||
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited
|
||||
- **Considered Options** — only when the rejected alternatives are worth remembering
|
||||
- **Consequences** — only when non-obvious downstream effects need to be called out
|
||||
|
||||
## Numbering
|
||||
|
||||
Scan `docs/adr/` for the highest existing number and increment by one.
|
||||
|
||||
## When to offer an ADR
|
||||
|
||||
All three of these must be true:
|
||||
|
||||
1. **Hard to reverse** — the cost of changing your mind later is meaningful
|
||||
2. **Surprising without context** — a future reader will look at the code and wonder "why on earth did they do it this way?"
|
||||
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
||||
|
||||
If a decision is easy to reverse, skip it — you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
|
||||
|
||||
### What qualifies
|
||||
|
||||
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
|
||||
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
|
||||
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library — just the ones that would take a quarter to swap out.
|
||||
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
|
||||
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
|
||||
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
|
||||
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it — otherwise someone will suggest GraphQL again in six months.
|
||||
@@ -0,0 +1,77 @@
|
||||
# CONTEXT.md Format
|
||||
|
||||
## Structure
|
||||
|
||||
```md
|
||||
# {Context Name}
|
||||
|
||||
{One or two sentence description of what this context is and why it exists.}
|
||||
|
||||
## Language
|
||||
|
||||
**Order**:
|
||||
{A concise description of the term}
|
||||
_Avoid_: Purchase, transaction
|
||||
|
||||
**Invoice**:
|
||||
A request for payment sent to a customer after delivery.
|
||||
_Avoid_: Bill, payment request
|
||||
|
||||
**Customer**:
|
||||
A person or organization that places orders.
|
||||
_Avoid_: Client, buyer, account
|
||||
|
||||
## Relationships
|
||||
|
||||
- An **Order** produces one or more **Invoices**
|
||||
- An **Invoice** belongs to exactly one **Customer**
|
||||
|
||||
## Example dialogue
|
||||
|
||||
> **Dev:** "When a **Customer** places an **Order**, do we create the **Invoice** immediately?"
|
||||
> **Domain expert:** "No — an **Invoice** is only generated once a **Fulfillment** is confirmed."
|
||||
|
||||
## Flagged ambiguities
|
||||
|
||||
- "account" was used to mean both **Customer** and **User** — resolved: these are distinct concepts.
|
||||
```
|
||||
|
||||
## Rules
|
||||
|
||||
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others as aliases to avoid.
|
||||
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
|
||||
- **Keep definitions tight.** One sentence max. Define what it IS, not what it does.
|
||||
- **Show relationships.** Use bold term names and express cardinality where obvious.
|
||||
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
|
||||
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
|
||||
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
|
||||
|
||||
## Single vs multi-context repos
|
||||
|
||||
**Single context (most repos):** One `CONTEXT.md` at the repo root.
|
||||
|
||||
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
|
||||
|
||||
```md
|
||||
# Context Map
|
||||
|
||||
## Contexts
|
||||
|
||||
- [Ordering](./src/ordering/CONTEXT.md) — receives and tracks customer orders
|
||||
- [Billing](./src/billing/CONTEXT.md) — generates invoices and processes payments
|
||||
- [Fulfillment](./src/fulfillment/CONTEXT.md) — manages warehouse picking and shipping
|
||||
|
||||
## Relationships
|
||||
|
||||
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
|
||||
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
|
||||
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
|
||||
```
|
||||
|
||||
The skill infers which structure applies:
|
||||
|
||||
- If `CONTEXT-MAP.md` exists, read it to find contexts
|
||||
- If only a root `CONTEXT.md` exists, single context
|
||||
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
|
||||
|
||||
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
|
||||
@@ -0,0 +1,88 @@
|
||||
---
|
||||
name: grill-with-docs
|
||||
description: Grilling session that challenges your plan against the existing domain model, sharpens terminology, and updates documentation (CONTEXT.md, ADRs) inline as decisions crystallise. Use when user wants to stress-test a plan against their project's language and documented decisions.
|
||||
---
|
||||
|
||||
<what-to-do>
|
||||
|
||||
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
|
||||
|
||||
Ask the questions one at a time, waiting for feedback on each question before continuing.
|
||||
|
||||
If a question can be answered by exploring the codebase, explore the codebase instead.
|
||||
|
||||
</what-to-do>
|
||||
|
||||
<supporting-info>
|
||||
|
||||
## Domain awareness
|
||||
|
||||
During codebase exploration, also look for existing documentation:
|
||||
|
||||
### File structure
|
||||
|
||||
Most repos have a single context:
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT.md
|
||||
├── docs/
|
||||
│ └── adr/
|
||||
│ ├── 0001-event-sourced-orders.md
|
||||
│ └── 0002-postgres-for-write-model.md
|
||||
└── src/
|
||||
```
|
||||
|
||||
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT-MAP.md
|
||||
├── docs/
|
||||
│ └── adr/ ← system-wide decisions
|
||||
├── src/
|
||||
│ ├── ordering/
|
||||
│ │ ├── CONTEXT.md
|
||||
│ │ └── docs/adr/ ← context-specific decisions
|
||||
│ └── billing/
|
||||
│ ├── CONTEXT.md
|
||||
│ └── docs/adr/
|
||||
```
|
||||
|
||||
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
|
||||
|
||||
## During the session
|
||||
|
||||
### Challenge against the glossary
|
||||
|
||||
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
|
||||
|
||||
### Sharpen fuzzy language
|
||||
|
||||
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
|
||||
|
||||
### Discuss concrete scenarios
|
||||
|
||||
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
|
||||
|
||||
### Cross-reference with code
|
||||
|
||||
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
|
||||
|
||||
### Update CONTEXT.md inline
|
||||
|
||||
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
||||
|
||||
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
|
||||
|
||||
### Offer ADRs sparingly
|
||||
|
||||
Only offer to create an ADR when all three are true:
|
||||
|
||||
1. **Hard to reverse** — the cost of changing your mind later is meaningful
|
||||
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
|
||||
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
||||
|
||||
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
||||
|
||||
</supporting-info>
|
||||
@@ -0,0 +1,132 @@
|
||||
---
|
||||
name: sm-flow
|
||||
description: OpenSpec-first、Devflow-assisted 的结构化工程开发元工作流。用户想把粗略想法、issue、PRD 或已有 research 推进为准确 OpenSpec change,并通过 OpenSpec apply 实现、验证、归档时使用;也适用于标准化开发流程、增强 OpenSpec 产物质量、沉淀工程决策和为未来代理保留上下文。
|
||||
---
|
||||
|
||||
# SM Flow
|
||||
|
||||
SM Flow 是一套 **OpenSpec-first / Devflow-assisted** 的软件交付元工作流。它不替代 OpenSpec,而是用 `devflow/` 中的上下文、术语、PRD、ADR、历史验收和复合知识,辅助生成更准确的 OpenSpec proposal/design/specs/tasks;执行阶段默认依赖 OpenSpec apply;完成后再把结果回填到 `devflow/` 作为长期记忆。
|
||||
|
||||
## 角色定位
|
||||
|
||||
你是这条工作流的工程负责人。你的职责不是绕过 OpenSpec 直接写代码,而是确保 OpenSpec 产物足够准确、可执行、可验收,然后以 OpenSpec apply 作为默认执行入口。你需要在关键决策点让用户参与,但仓库阅读、上下文提取、OpenSpec 修正、文档更新和归档提炼应尽量由你完成。
|
||||
|
||||
## 真理源分层
|
||||
|
||||
- `devflow/` 是上下文真理源:术语、历史 PRD、ADR、验收记录、复合知识和项目记忆。
|
||||
- `openspec/changes/<change>/` 是执行真理源:当前变更的 proposal、design、specs 和 tasks。
|
||||
- 代码是实现结果:只能在执行真理源足够明确后修改。
|
||||
- 如果 `devflow/` 和 OpenSpec 冲突,不要直接写代码;先汇报冲突、让用户确认,并修正 OpenSpec。
|
||||
|
||||
## 核心规则
|
||||
|
||||
- Phase 3 默认必须通过 OpenSpec apply 执行;禁止直接依赖 devflow 文档绕过 OpenSpec 写代码。
|
||||
- devflow 产物只能辅助生成和校准 OpenSpec,不能成为 Phase 3 的主要执行依据。
|
||||
- devflow 默认采用分层产物:只强制保留人类理解和 OpenSpec 校准所必需的档案,避免重复 OpenSpec 的 proposal/design/tasks。
|
||||
- 不得跳过 Phase 0.5:生成或修正 OpenSpec 前,必须读取相关 devflow 上下文。
|
||||
- 不得跳过 Phase 2。即使需求看起来很清楚,也至少解决三个高价值澄清或验证问题。
|
||||
- `evidence-driven` 不是自动通过:凡是通过证据确认的结论,也必须向用户汇报证据、结论和是否需要确认。
|
||||
- `user-interview` 问题必须等待用户确认;涉及范围、偏好、验收口径、风险接受度时不能擅自决定。
|
||||
- 架构审计、澄清、ADR 或上下文发现如果影响实现,必须先回写 OpenSpec design/specs/tasks,再进入 Phase 3。
|
||||
- 不得猜测式修 bug。实现失败或行为不明时,先进入 diagnose 闭环;如果根因是规格不准,先修 OpenSpec,再继续 apply。
|
||||
- 默认采用 human-in-the-loop:Phase 1 后、Phase 2.5 后、Phase 3 前必须给用户一个简短检查点,除非用户明确要求“全自动执行”。
|
||||
- Phase 4 结束前必须询问用户是否归档 OpenSpec change;不要默认执行归档。
|
||||
- 子 skill 必须显式调用或显式降级:进入 Phase 1、1.5、2、2.5、3、4 时,先说明“本阶段调用/加载的 skill 是 X”;如果不能调用,必须说明“降级为 fallback”,并记录降级原因。
|
||||
- 显式调用优先级:优先使用平台原生 Skill 调用;如果平台没有 Skill 工具,则读取对应 `SKILL.md` 并按其协议执行;只有文件不存在或协议不可执行时,才使用 `references/fallbacks.md`。
|
||||
- fallback 产物必须标注为降级执行,不能伪装成真实子 skill 调用。
|
||||
|
||||
## 首次加载
|
||||
|
||||
执行前只读取当前任务需要的 reference 文件:
|
||||
|
||||
- 需要逐阶段执行时,读取 `references/phase-contracts.md`。
|
||||
- 创建或更新 PRD、ADR、验收报告、词汇表、复合知识文档时,读取 `references/templates.md`。
|
||||
- Phase 4 或需要从 OpenSpec 提取产物时,读取 `references/archive-rules.md`。
|
||||
- 子 skill 或 OpenSpec skill 无法直接调用时,读取 `references/fallbacks.md`。
|
||||
|
||||
## 启动检查
|
||||
|
||||
1. 判断启动模式:
|
||||
- 完整模式:用户提供粗略想法或初始 PRD。
|
||||
- Research 模式:用户已有 research,需要转成或修正 OpenSpec。
|
||||
- PRD 文件模式:用户提供已有 PRD 路径。
|
||||
- 指定阶段模式:用户要求从某个 Phase 恢复。
|
||||
- 快速模式:小改动,Phase 2 和 Phase 4 可以轻量化,但不能省略。
|
||||
2. 如果缺少 `devflow/`,初始化:
|
||||
- `devflow/projects/`
|
||||
- `devflow/glossary/CONTEXT.md`
|
||||
- `devflow/compound/`
|
||||
- `devflow/reference/`
|
||||
3. 如果根目录存在旧 `CONTEXT.md`,且 `devflow/glossary/CONTEXT.md` 不存在或为空,询问用户是迁移还是合并。
|
||||
4. 检查 OpenSpec 和子 skill 是否可用:
|
||||
- OpenSpec:`.claude/skills/openspec-propose/SKILL.md`、`.claude/skills/openspec-apply-change/SKILL.md`、`.claude/skills/openspec-archive-change/SKILL.md`。
|
||||
- 辅助技能:`.agents/skills/to-prd/SKILL.md`、`.agents/skills/grill-with-docs/SKILL.md`、`.agents/skills/diagnose/SKILL.md`、`.agents/skills/tdd/SKILL.md`、`.agents/skills/zoom-out/SKILL.md`。
|
||||
5. 如果 OpenSpec skill 不可用,不要直接绕过执行;先使用 fallback 生成/修正 OpenSpec 产物,并在 Phase 3 前向用户说明执行入口降级风险。
|
||||
|
||||
## 项目标识
|
||||
|
||||
整个流程使用同一个 slug:
|
||||
|
||||
- 优先使用 OpenSpec change name。
|
||||
- 如果还没有 change name,则从功能标题生成 kebab-case slug。
|
||||
- 项目档案目录格式:`devflow/projects/YYYY-MM-DD-{slug}/`。
|
||||
- 如果目录已存在,默认恢复该项目,不要重复创建;除非用户明确要求新开一轮。
|
||||
|
||||
## Devflow 产物分层
|
||||
|
||||
devflow 是辅助 OpenSpec 和人类阅读的档案层,不复制 OpenSpec 的执行产物。默认只创建必要文件;扩展文件必须有明确理由。
|
||||
|
||||
**必须产物**:
|
||||
|
||||
- `brief.md`:背景、目标、范围、非目标、变更规模、关联 OpenSpec change。
|
||||
- `evidence.md`:代码证据、文档证据、历史决策、evidence-driven 结论和汇报状态。
|
||||
- `decisions.md`:user-interview 问题、用户确认、关键取舍、风险接受、OpenSpec 回写记录。
|
||||
- `acceptance.md`:实现结果、验证命令、未验证项、归档状态、后续事项。
|
||||
|
||||
**按需产物**:
|
||||
|
||||
- `prd.md`:需求复杂、用户明确要求 PRD、或需要对外协作时创建;小需求并入 `brief.md`。
|
||||
- `research.md`:存在真实调研、代码考古、竞品/API 对比或复杂方案比较时创建。
|
||||
- `design.md`:仅记录 OpenSpec design 不适合承载的人类背景、架构审计摘要或长期决策索引;实现设计仍以 OpenSpec design 为准。
|
||||
- `tasks.md`:仅记录跨轮次追踪或人类复盘任务;执行任务仍以 OpenSpec tasks 为准。
|
||||
- `alignment.md` / `clarifications.md`:问题很多或冲突复杂时单独创建;否则并入 `decisions.md`。
|
||||
- `adr/*.md` 和 `compound/*.md`:仅在满足 ADR/复合知识规则时创建。
|
||||
|
||||
**规模分档**:
|
||||
|
||||
- `micro`:小且低风险,使用 `brief.md`、`decisions.md`、`acceptance.md`;证据少时并入 `brief.md`。
|
||||
- `standard`:默认模式,使用 `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`。
|
||||
- `complex`:高风险、跨模块、需求不清或多人协作时,在 standard 基础上按需增加 PRD/research/design/tasks/alignment。
|
||||
|
||||
## 阶段总览
|
||||
|
||||
1. Phase 0 — 入口澄清:接收初始 PRD、粗略想法或已有 research,澄清到可生成 OpenSpec。
|
||||
2. Phase 0.5 — Devflow 上下文收集:读取 glossary、相关 PRD、ADR、acceptance、compound knowledge。
|
||||
3. Phase 1 — OpenSpec propose:基于需求 + devflow 上下文生成或修正 OpenSpec proposal/design/specs/tasks。
|
||||
4. Phase 1.5 — PRD / OpenSpec 对齐:检查 OpenSpec 是否覆盖 PRD、遵守 ADR、使用正确术语、具备可验收 specs。
|
||||
5. Phase 2 — Human-in-the-loop 澄清:声明 `evidence-driven` / `user-interview`,所有结论都汇报,确认后回写 OpenSpec。
|
||||
6. Phase 2.5 — 架构审计:审计结果如果影响实现,必须回写 OpenSpec design/tasks。
|
||||
7. Phase 3 — OpenSpec apply:默认依赖 OpenSpec 执行;devflow 只作为上下文参考。
|
||||
8. Phase 4 — 回填 Devflow:从 OpenSpec、实现结果和验证结果提炼长期档案,并询问是否 archive OpenSpec change。
|
||||
|
||||
每个阶段的进入条件、动作、输出和退出标准见 `references/phase-contracts.md`。
|
||||
|
||||
## 快速模式
|
||||
|
||||
快速模式仅在改动小且低风险时使用。它可以压缩 Phase 1.5 和 Phase 2.5,但必须保留:
|
||||
|
||||
- Phase 0.5 最小上下文检查:至少检查 glossary 和相关 ADR。
|
||||
- Phase 2 最小澄清:至少一个术语问题、一个边界问题、一个验收问题;evidence-driven 结论也要汇报。
|
||||
- Phase 3 仍以 OpenSpec tasks/specs 为执行依据。
|
||||
- Phase 4 轻量归档:记录验收结果、OpenSpec 产物链接和是否归档。
|
||||
|
||||
## 完成标准
|
||||
|
||||
一次流程只有在满足以下条件时才算完成:
|
||||
|
||||
- OpenSpec change 中的 proposal/design/specs/tasks 已生成或更新到可执行状态。
|
||||
- 实现或规划任务已经完成,且执行依据来自 OpenSpec。
|
||||
- 已运行验证,或明确记录未运行验证的原因。
|
||||
- `devflow/projects/YYYY-MM-DD-{slug}/` 中存在符合规模分档的必要 devflow 产物,并能说明背景、证据、决策、验收和归档状态。
|
||||
- 用户知道剩余风险和下一步动作,并已被询问是否要归档 OpenSpec change。
|
||||
|
||||
@@ -0,0 +1,106 @@
|
||||
# 归档规则
|
||||
|
||||
Phase 4 的目标是把 OpenSpec 产物、实现结果和验证结果转化为持久、可读、可复用的项目记忆。v3 中,OpenSpec 是执行真理源,devflow 是辅助 OpenSpec 和人类阅读的档案层。
|
||||
|
||||
## 目录规则
|
||||
|
||||
项目档案路径:
|
||||
|
||||
```text
|
||||
devflow/projects/YYYY-MM-DD-{slug}/
|
||||
```
|
||||
|
||||
默认创建以下必要文件:
|
||||
|
||||
- `brief.md`
|
||||
- `evidence.md`
|
||||
- `decisions.md`
|
||||
- `acceptance.md`
|
||||
|
||||
按需创建以下扩展文件:
|
||||
|
||||
- `prd.md`
|
||||
- `research.md`
|
||||
- `design.md`
|
||||
- `tasks.md`
|
||||
- `alignment.md`
|
||||
- `adr/*.md`
|
||||
|
||||
不要逐字复制完整 OpenSpec 文件,也不要重复 OpenSpec 的 proposal/design/tasks。应提炼 OpenSpec 如何指导执行:背景、证据、用户决策、任务状态、假设、验证结果、风险,以及执行中对 OpenSpec 的修正。
|
||||
|
||||
## 产物分档
|
||||
|
||||
| 分档 | 适用场景 | 必须文件 | 扩展文件 |
|
||||
| --- | --- | --- | --- |
|
||||
| `micro` | 小改动、低风险、需求明确 | `brief.md`、`decisions.md`、`acceptance.md` | 证据少时并入 `brief.md` |
|
||||
| `standard` | 默认模式 | `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md` | 按需 ADR/compound |
|
||||
| `complex` | 高风险、跨模块、需求不清、多人协作 | standard 全部文件 | 按需 `prd.md`、`research.md`、`design.md`、`tasks.md`、`alignment.md` |
|
||||
|
||||
## 提取映射
|
||||
|
||||
| 来源 | 提取内容 | 写入位置 |
|
||||
| --- | --- | --- |
|
||||
| `proposal.md` | 为什么做、做什么、范围、非目标 | `brief.md` |
|
||||
| `design.md` | 技术方案、关键决策、风险;只提炼长期有用内容 | `evidence.md` / 按需 `design.md` |
|
||||
| `specs/**/*.md` | requirement 标题和 scenario 意图 | `brief.md` 或 `acceptance.md` 的验收追踪 |
|
||||
| `tasks.md` | checkbox 状态、剩余工作、执行切片 | `acceptance.md`;复杂项目可拆 `tasks.md` |
|
||||
| 澄清记录 | evidence-driven/user-interview、证据、结论、确认状态 | `evidence.md` + `decisions.md` |
|
||||
| 测试/构建输出 | 验证命令、结果、验证类型 | `acceptance.md` |
|
||||
| diagnose 记录 | 根因、修复、回归验证 | `acceptance.md` |
|
||||
| 词汇表更新 | 术语和业务规则 | `devflow/glossary/CONTEXT.md` |
|
||||
| 可复用经验 | 持久工程知识 | `devflow/compound/YYYY-MM-DD-{type}-{slug}.md` |
|
||||
|
||||
## 验收记录规则
|
||||
|
||||
必须真实记录验证情况,并按类型分类:
|
||||
|
||||
- **静态验证**:语法检查、grep/rg 检查、结构检查、类型检查等不运行完整功能的验证。
|
||||
- **脚本验证**:生成脚本、测试命令、构建命令、自动化检查等可重复命令。
|
||||
- **浏览器/人工验证**:需要用户或代理在界面中点击、观察、确认的行为验证。
|
||||
- **未验证**:未运行的验证必须记录原因、风险和建议补验步骤。
|
||||
|
||||
记录要求:
|
||||
|
||||
- 如果验证通过,记录命令/步骤和覆盖范围。
|
||||
- 如果验证失败,记录失败摘要和是否阻塞验收。
|
||||
- 如果需要人工验证,列出明确步骤,不要用“手动测试一下”这种模糊描述。
|
||||
|
||||
## ADR 规则
|
||||
|
||||
同时满足以下条件时创建 ADR:
|
||||
|
||||
1. 决策难以逆转。
|
||||
2. 缺少上下文会让未来维护者困惑。
|
||||
3. 决策来自真实权衡,而不是简单偏好。
|
||||
|
||||
项目内 ADR 存放于:
|
||||
|
||||
```text
|
||||
devflow/projects/YYYY-MM-DD-{slug}/adr/
|
||||
```
|
||||
|
||||
跨项目可复用决策或经验存放于:
|
||||
|
||||
```text
|
||||
devflow/compound/YYYY-MM-DD-decision-{slug}.md
|
||||
```
|
||||
|
||||
## 归档确认
|
||||
|
||||
OpenSpec archive 是显式 human-in-the-loop 动作。archive 前必须确认 devflow 已经回填 OpenSpec 的关键执行信息:
|
||||
|
||||
- Phase 4 可以建议 archive,但必须先询问用户。
|
||||
- 在用户确认前,不要执行 archive。
|
||||
- 如果用户暂不归档,在 acceptance 中记录原因或状态。
|
||||
- 如果用户确认归档,执行后记录 archive 结果和剩余档案位置。
|
||||
|
||||
## 归档交接
|
||||
|
||||
Phase 4 结束时告诉用户:
|
||||
|
||||
- 创建或更新了哪些档案文件。
|
||||
- 运行了哪些验证,并按静态验证、脚本验证、浏览器/人工验证、未验证分类。
|
||||
- 还剩哪些风险或后续事项。
|
||||
- 明确询问:是否现在 archive OpenSpec change?
|
||||
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
# Fallback 协议
|
||||
|
||||
当子 skill 无法直接调用时使用这些协议。Fallback 不是绕过 OpenSpec 的许可;v3 中 fallback 的目标仍然是生成、修正或执行 OpenSpec 产物。使用任何 fallback 前必须先向用户说明:目标子 skill、无法调用原因、降级协议名称、降级风险。产物中也必须记录“本阶段为 fallback 降级执行”。
|
||||
|
||||
## OpenSpec 提案 fallback
|
||||
|
||||
1. 创建或识别 `openspec/changes/{slug}/`。
|
||||
2. 先读取 devflow 上下文:`devflow/glossary/CONTEXT.md`、相关项目档案、ADR、acceptance、compound knowledge。
|
||||
3. 写入 `proposal.md`,包含:
|
||||
- 问题
|
||||
- 建议方案
|
||||
- 范围
|
||||
- 非目标
|
||||
- 来自 devflow 的上下文约束
|
||||
- 风险
|
||||
4. 当实现需要技术选择时,写入 `design.md`,并引用相关 ADR 或历史验收结论。
|
||||
5. 将 `tasks.md` 写成按纵向切片组织的 checkbox 清单。
|
||||
6. 只为外部可见行为或发生变化的需求编写 specs。
|
||||
7. 如果存在高风险假设,在 Phase 1.5 前向用户 checkpoint。
|
||||
|
||||
## OpenSpec 修正 fallback
|
||||
|
||||
当 PRD、devflow、澄清结论、架构审计和 OpenSpec 冲突时:
|
||||
|
||||
1. 列出冲突来源:PRD / glossary / ADR / acceptance / compound / OpenSpec。
|
||||
2. 判断冲突类型:术语、范围、验收、架构、任务拆分、风险。
|
||||
3. 向用户汇报冲突和推荐修正。
|
||||
4. 用户确认后,优先修正 OpenSpec proposal/design/specs/tasks。
|
||||
5. 再同步更新 devflow 文档;不要只改 devflow。
|
||||
|
||||
## OpenSpec 执行 fallback
|
||||
|
||||
仅当 `openspec-apply-change` 不可调用时使用。执行依据仍必须是 `openspec/changes/{slug}/`。
|
||||
|
||||
1. 阅读 `proposal.md`、`design.md`、`specs/**/*.md` 和 `tasks.md`。
|
||||
2. 确认 OpenSpec 与 devflow 上下文没有未解决冲突。
|
||||
3. 修改前先检查现有代码。
|
||||
4. 一次实现一个 OpenSpec task 的纵向切片。
|
||||
5. 用最窄但有效的命令验证每个切片。
|
||||
6. 只有验证通过或明确记录原因后,才更新 task 状态。
|
||||
7. 如果失败原因不确定,停止并进入 diagnose。
|
||||
8. 如果 diagnose 证明规格不准,先修正 OpenSpec,再继续执行。
|
||||
|
||||
## PRD fallback
|
||||
|
||||
优先使用 `references/templates.md#brief-模板` 创建 `brief.md`。只有复杂需求、对外协作或用户明确要求 PRD 时,才使用 `references/templates.md#prd-模板` 创建 `prd.md`。
|
||||
|
||||
规则:
|
||||
|
||||
- 从当前上下文、devflow 记忆和 OpenSpec 产物综合,不要机械复制。
|
||||
- `brief.md` 或 PRD 用来表达用户价值、范围和验收口径;不能替代 OpenSpec specs/tasks。
|
||||
- 只有当缺失决策会阻塞 OpenSpec 正确性时,才采访用户。
|
||||
|
||||
## 文档化追问 fallback
|
||||
|
||||
1. 阅读已有词汇表、ADR、相关 OpenSpec 产物和 devflow 项目档案。
|
||||
2. 先声明每个问题的模式:`evidence-driven` 或 `user-interview`。
|
||||
3. 对 evidence-driven 问题,先查代码库、文档、OpenSpec 或 ADR,再向用户汇报证据和结论。
|
||||
4. 对 user-interview 问题,一次只问一个并等待用户确认。
|
||||
5. 术语确认后立即更新词汇表。
|
||||
6. 影响实现的澄清必须回写 OpenSpec。
|
||||
7. 只为难以逆转的真实权衡创建 ADR。
|
||||
|
||||
快速模式的最小问题:
|
||||
|
||||
- 术语:这个概念应该使用哪个领域术语?证据是什么?
|
||||
- 边界:哪些内容明确不在范围内?是否需要用户确认?
|
||||
- 验收:什么可观察行为能证明它完成?是否已写入 OpenSpec specs?
|
||||
|
||||
## 架构审计 fallback
|
||||
|
||||
产出一份短架构审计:
|
||||
|
||||
1. 画出输入 → 处理 → 输出。
|
||||
2. 列出相关模块和调用方。
|
||||
3. 识别耦合、数据所有权和生命周期风险。
|
||||
4. 检查是否与词汇表、ADR 和 OpenSpec design 冲突。
|
||||
5. 用不超过五句话总结最大风险。
|
||||
6. 如果影响实现,回写 OpenSpec design/tasks。
|
||||
|
||||
## Diagnose fallback
|
||||
|
||||
1. 复现问题,或捕获准确失败信息。
|
||||
2. 最小化失败案例。
|
||||
3. 生成 3-5 个假设,并按可能性和验证成本排序。
|
||||
4. 修改代码前,先添加仪器化或定向检查。
|
||||
5. 判断根因属于实现问题还是 OpenSpec 规格问题。
|
||||
6. 如果是实现问题,修复被证明的最小原因。
|
||||
7. 如果是规格问题,先修正 OpenSpec,再继续 apply。
|
||||
8. 运行回归验证。
|
||||
|
||||
## TDD fallback
|
||||
|
||||
使用纵向切片,不要水平批量写测试:
|
||||
|
||||
1. 从 OpenSpec specs 中选择一个外部可见行为。
|
||||
2. 写一个失败测试。
|
||||
3. 实现刚好让测试通过的最小代码。
|
||||
4. 只在测试通过时重构。
|
||||
5. 对下一个 OpenSpec 行为重复以上步骤。
|
||||
|
||||
@@ -0,0 +1,199 @@
|
||||
# 阶段契约
|
||||
|
||||
本文件是 SM Flow v3 的逐阶段执行准则。v3 的核心原则是:**devflow 辅助 OpenSpec,OpenSpec 指挥执行,执行结果回填 devflow**。
|
||||
|
||||
## Phase 0 — 入口澄清
|
||||
|
||||
**进入条件**:用户提供粗略想法、初始 PRD、已有 research、issue,或要求启动 SM Flow。
|
||||
|
||||
**动作**:
|
||||
- 收集问题、期望结果、目标用户、涉及代码区域、约束条件和可能的非目标。
|
||||
- 如果用户已有 research,先识别它是否已经包含用户价值、技术方案、验收标准和任务拆分。
|
||||
- 如果输入过于模糊,最多追加三轮聚焦问题。
|
||||
- 当答案会改变 OpenSpec proposal/specs/tasks 时,优先一次只问一个问题。
|
||||
|
||||
**退出条件**:
|
||||
- 问题可以用 1-2 句话说清楚。
|
||||
- 期望结果可以用 1-2 句话说清楚。
|
||||
- 已列出已知影响代码或模块;如果未知,也明确标记。
|
||||
- 可以生成 OpenSpec change slug。
|
||||
|
||||
**输出**:
|
||||
- 入口摘要。
|
||||
- 初步 slug。
|
||||
- devflow 规模分档:`micro` / `standard` / `complex`。
|
||||
|
||||
## Phase 0.5 — Devflow 上下文收集
|
||||
|
||||
**进入条件**:Phase 0 已经有足够信息定位领域、项目或变更方向。
|
||||
|
||||
**动作**:
|
||||
- 读取 `devflow/glossary/CONTEXT.md`,提取相关术语和业务规则。
|
||||
- 搜索 `devflow/projects/` 中相关 PRD、design、tasks、acceptance 和 ADR。
|
||||
- 搜索 `devflow/compound/` 中可复用 learning、trick、decision、explore。
|
||||
- 记录哪些上下文会影响 OpenSpec proposal/design/specs/tasks。
|
||||
- 如果发现旧根目录 `CONTEXT.md` 与 `devflow/glossary/CONTEXT.md` 冲突,暂停并向用户汇报。
|
||||
|
||||
**退出条件**:
|
||||
- 已形成“OpenSpec 输入上下文摘要”。
|
||||
- 已列出相关 ADR 和不能违反的历史决策。
|
||||
- 已列出需要写入或修正 OpenSpec 的上下文点。
|
||||
|
||||
**输出**:
|
||||
- 上下文摘要,默认写入 `brief.md` 或 `evidence.md`;会影响实现的上下文必须写入 OpenSpec design/specs/tasks。
|
||||
|
||||
## Phase 1 — OpenSpec propose
|
||||
|
||||
**进入条件**:Phase 0 + Phase 0.5 已经足够生成或修正 OpenSpec change。
|
||||
|
||||
**显式子 skill**:`openspec-propose`。进入本阶段必须先声明调用方式:原生 Skill 调用 / 读取 `.claude/skills/openspec-propose/SKILL.md` / fallback 降级。
|
||||
|
||||
**动作**:
|
||||
- 优先调用 `openspec-propose`。
|
||||
- 如果不可用,执行 `references/fallbacks.md#openspec-propose-fallback`,但仍必须产出 OpenSpec 文件。
|
||||
- 用 Phase 0.5 的 devflow 上下文增强 OpenSpec:
|
||||
- proposal 写清为什么做、做什么、范围和非目标。
|
||||
- design 写入上下文约束、历史 ADR、关键技术决策。
|
||||
- specs 写成可验收的外部行为。
|
||||
- tasks 写成可执行的纵向切片。
|
||||
- 在承诺设计细节前,先检查相关仓库代码。
|
||||
|
||||
**退出条件**:
|
||||
- `openspec/changes/<slug>/proposal.md` 存在。
|
||||
- 对需要正式 OpenSpec 产物的变更,`design.md`、`tasks.md` 和 specs 存在。
|
||||
- 关键假设已显式记录在 OpenSpec 或 research 中。
|
||||
|
||||
**输出**:
|
||||
- OpenSpec proposal、design、specs 和 task list。
|
||||
|
||||
**Human checkpoint**:
|
||||
- 向用户简要说明 OpenSpec scope、关键假设、主要风险、devflow 上下文如何影响 OpenSpec。
|
||||
- 询问是否继续进入 PRD/OpenSpec 对齐和澄清阶段;用户明确要求“全自动执行”时可跳过等待。
|
||||
|
||||
## Phase 1.5 — PRD / OpenSpec 对齐
|
||||
|
||||
**进入条件**:Phase 1 已有 OpenSpec 产物。
|
||||
|
||||
**显式子 skill**:`to-prd` + `sm-flow`。进入本阶段必须先声明是否读取 `.agents/skills/to-prd/SKILL.md`;如果使用内置模板,标记为 PRD fallback。
|
||||
|
||||
**动作**:
|
||||
- 如果没有结构化 PRD,则优先按 `to-prd` 协议生成 `brief.md`;复杂需求、对外协作或用户明确要求时再生成 `prd.md`。
|
||||
- `micro` 模式默认不创建独立 PRD,除非用户要求或需求复杂度升级;PRD 内容并入 `brief.md`。
|
||||
- 检查 PRD、devflow 上下文和 OpenSpec 是否一致:
|
||||
- OpenSpec 是否覆盖 PRD 的用户故事和验收预期。
|
||||
- OpenSpec 是否使用 glossary 中的正确术语。
|
||||
- OpenSpec 是否遵守相关 ADR。
|
||||
- specs 是否能表达可观察行为。
|
||||
- tasks 是否能驱动实现,而不是泛泛描述。
|
||||
- 如果发现不一致,优先修正 OpenSpec,而不是只修改 devflow 文档。
|
||||
|
||||
**退出条件**:
|
||||
- `brief.md` 已覆盖背景、目标、范围和非目标;复杂需求存在独立 `prd.md` 或用户明确不需要 PRD。
|
||||
- OpenSpec 与 PRD/devflow 上下文没有已知冲突。
|
||||
- 所有已知冲突已修正或等待用户决策。
|
||||
|
||||
**输出**:
|
||||
- `brief.md`,以及按需创建的 `prd.md`。
|
||||
- OpenSpec 对齐检查记录。
|
||||
- 必要的 OpenSpec 修正。
|
||||
|
||||
## Phase 2 — Human-in-the-loop 澄清
|
||||
|
||||
**进入条件**:已有 OpenSpec 产物和 PRD/上下文对齐记录。
|
||||
|
||||
**显式子 skill**:`grill-with-docs`。进入本阶段必须先声明是否读取 `.agents/skills/grill-with-docs/SKILL.md`;fallback 必须标记为“文档化追问 fallback”。
|
||||
|
||||
**动作**:
|
||||
- 优先使用 `grill-with-docs`。
|
||||
- 先声明本阶段采用的澄清模式,并逐项标记:
|
||||
- `evidence-driven`:问题能通过代码、文档、测试、OpenSpec 或既有 ADR 证明;代理先查证,再向用户汇报证据、结论和是否需要确认。
|
||||
- `user-interview`:问题涉及产品偏好、范围边界、验收口径、风险接受度或价值取舍;必须问用户并等待确认。
|
||||
- 至少覆盖三个维度:术语、边界、验收。
|
||||
- 一次只问一个 `user-interview` 问题。
|
||||
- 如果澄清结果影响实现,必须回写 OpenSpec proposal/design/specs/tasks。
|
||||
- 术语一旦确认,更新 `devflow/glossary/CONTEXT.md`。
|
||||
- 对难以逆转、依赖上下文、源自真实权衡的决策创建 ADR。
|
||||
|
||||
**退出条件**:
|
||||
- 至少解决三个高价值澄清或验证问题,并记录每个问题属于 `evidence-driven` 还是 `user-interview`。
|
||||
- 所有 evidence-driven 结论已向用户汇报。
|
||||
- 所有 user-interview 决策已获得用户确认。
|
||||
- 影响实现的结论已回写 OpenSpec。
|
||||
|
||||
**输出**:
|
||||
- 澄清记录:默认写入 `decisions.md`;问题很多时可拆出 `clarifications.md`。记录术语、边界、验收三个维度、模式、证据、结论、用户确认状态。
|
||||
- 更新后的 OpenSpec。
|
||||
- 更新后的词汇表和 ADR。
|
||||
|
||||
## Phase 2.5 — 架构审计
|
||||
|
||||
**进入条件**:Phase 2 已解决主要产品、领域和验收问题。
|
||||
|
||||
**显式子 skill**:`zoom-out`。进入本阶段必须先声明是否读取 `.agents/skills/zoom-out/SKILL.md`;fallback 必须标记为“架构审计 fallback”。
|
||||
|
||||
**动作**:
|
||||
- 画出输入 → 处理 → 输出的模块链路。
|
||||
- 识别跨模块依赖、数据所有权、生命周期和耦合风险。
|
||||
- 检查是否与既有架构、ADR、OpenSpec design 冲突。
|
||||
- 用不超过五句话写出架构风险评估。
|
||||
- 如果审计结果影响实现,必须回写 OpenSpec design/tasks;只写入 devflow design 不够。
|
||||
|
||||
**退出条件**:
|
||||
- 架构风险已被接受,或流程返回 Phase 2/Phase 1 修正 OpenSpec。
|
||||
- OpenSpec design/tasks 已反映会影响实现的架构审计结论。
|
||||
|
||||
**输出**:
|
||||
- 架构审计记录,默认写入 `decisions.md` 或 `evidence.md`;复杂架构审计可拆出 `design.md`。
|
||||
- 必要的 OpenSpec design/tasks 修正。
|
||||
|
||||
**Human checkpoint**:
|
||||
- 用不超过五句话向用户说明架构风险、OpenSpec 修正点和实现计划。
|
||||
- 询问是否进入 Phase 3 OpenSpec apply;用户明确要求“全自动执行”时可跳过等待。
|
||||
|
||||
## Phase 3 — OpenSpec apply
|
||||
|
||||
**进入条件**:
|
||||
- `openspec/changes/<slug>/` 中 proposal/design/specs/tasks 已达到可执行状态。
|
||||
- Phase 2.5 后已获得用户继续实现的确认,除非用户要求“全自动执行”。
|
||||
- devflow 与 OpenSpec 没有未解决冲突。
|
||||
|
||||
**显式子 skill**:`openspec-apply-change`;遇到 bug/不确定行为时显式调用 `diagnose`;需要测试驱动时显式调用 `tdd`。进入本阶段必须先声明调用方式,不能静默 fallback。
|
||||
|
||||
**动作**:
|
||||
- 优先调用 `openspec-apply-change`。
|
||||
- 执行依据是 OpenSpec specs/tasks;devflow 只能作为上下文参考。
|
||||
- 按 OpenSpec tasks 的纵向切片实现。
|
||||
- 当用户要求、行为复杂或回归风险高时使用 TDD。
|
||||
- 当测试失败、行为意外或原因不确定时使用 diagnose。
|
||||
- 如果 diagnose 发现根因是 OpenSpec 不准确,先修正 OpenSpec,再继续 apply。
|
||||
- 修改文件前遵守仓库指令,例如 `AGENTS.md`。
|
||||
|
||||
**退出条件**:
|
||||
- OpenSpec tasks 已完成,或剩余 tasks 已明确记录。
|
||||
- 已运行验证,或记录了未验证原因。
|
||||
- 已列出已知限制。
|
||||
|
||||
**输出**:
|
||||
- 代码变更、必要测试和实现说明。
|
||||
- 更新后的 OpenSpec task 状态。
|
||||
|
||||
## Phase 4 — 回填 Devflow
|
||||
|
||||
**进入条件**:实现或规划工作已经达到可交接状态。
|
||||
|
||||
**显式子 skill**:`openspec-archive-change` 只在用户确认 archive 后调用;Phase 4 回填由 `sm-flow` 执行。必须记录 archive 是真实调用还是手动 fallback。
|
||||
|
||||
**动作**:
|
||||
- 遵循 `references/archive-rules.md`。
|
||||
- 从 OpenSpec、实现结果和验证结果提炼 brief、evidence、decisions 和 acceptance;只在复杂场景按需拆出 PRD/research/design/tasks/alignment。
|
||||
- 写入或更新验收记录,并区分静态验证、脚本验证、浏览器/人工验证、未验证。
|
||||
- 如果本次流程产生可复用经验,写入 compound knowledge。
|
||||
- 询问用户是否要 archive OpenSpec change;不要默认执行归档。
|
||||
|
||||
**退出条件**:
|
||||
- `devflow/projects/YYYY-MM-DD-{slug}/` 能说明做了什么、为什么做、OpenSpec 如何指导执行、还剩什么、如何验证。
|
||||
- 用户已被询问是否 archive OpenSpec change。
|
||||
|
||||
**输出**:
|
||||
- 默认输出 `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`;按需输出 `prd.md`、`research.md`、`design.md`、`tasks.md`、`alignment.md`、ADR 和 compound knowledge。
|
||||
|
||||
@@ -0,0 +1,290 @@
|
||||
# 模板
|
||||
|
||||
这些是最小模板。只有在能提升未来可读性时,才增加额外章节。保留 PRD、ADR、OpenSpec、slug 等行业术语,其余说明尽量使用中文。
|
||||
|
||||
## Brief 模板
|
||||
|
||||
```markdown
|
||||
# {标题} Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:{goal}
|
||||
- 当前问题:{problem}
|
||||
- 关联 OpenSpec:`openspec/changes/{slug}/`
|
||||
- devflow 分档:micro | standard | complex
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:{in scope}
|
||||
- 本次不做:{out of scope}
|
||||
- 影响区域:{modules/files if known}
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||
- specs 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||
- tasks 覆盖状态:已覆盖 / 待修正 / 不适用
|
||||
```
|
||||
|
||||
## Evidence 模板
|
||||
|
||||
```markdown
|
||||
# {标题} Evidence
|
||||
|
||||
## 证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
| --- | --- | --- | --- |
|
||||
| {file/doc/test/ADR} | {evidence summary} | {conclusion} | 是 / 否 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- 结论:{conclusion}
|
||||
- 证据:{evidence}
|
||||
- 风险:{risk if any}
|
||||
- 用户确认:需要 / 不需要 / 已确认
|
||||
```
|
||||
|
||||
## Decisions 模板
|
||||
|
||||
```markdown
|
||||
# {标题} Decisions
|
||||
|
||||
## User-interview
|
||||
|
||||
| 问题 | 用户回答 | 决策 | OpenSpec 回写 |
|
||||
| --- | --- | --- | --- |
|
||||
| {question} | {answer} | {decision} | 已回写 / 不影响 / 待回写 |
|
||||
|
||||
## 关键取舍
|
||||
|
||||
- 决策:{decision}
|
||||
- 原因:{why}
|
||||
- 影响:{impact}
|
||||
- 风险接受:{accepted by whom/when}
|
||||
```
|
||||
|
||||
## PRD 模板
|
||||
|
||||
```markdown
|
||||
# {标题} PRD
|
||||
|
||||
## 问题陈述
|
||||
|
||||
用用户视角描述问题。
|
||||
|
||||
## 解决方案
|
||||
|
||||
用用户视角描述预期解决方案。
|
||||
|
||||
## 用户故事
|
||||
|
||||
1. 作为{角色},我希望{能力},以便{收益}。
|
||||
|
||||
## 实现决策
|
||||
|
||||
- 决策:{decision}
|
||||
- 原因:{why}
|
||||
- 影响:{affected modules or behavior}
|
||||
|
||||
## 测试决策
|
||||
|
||||
- 好测试应该通过{public interface}验证{observable behavior}。
|
||||
- 必须覆盖:{critical paths}
|
||||
- 不测试:{explicit exclusions}
|
||||
|
||||
## 非目标
|
||||
|
||||
- {excluded behavior}
|
||||
|
||||
## 补充说明
|
||||
|
||||
- {open question or useful context}
|
||||
```
|
||||
|
||||
## 词汇表模板
|
||||
|
||||
```markdown
|
||||
# 上下文词汇表
|
||||
|
||||
## 术语
|
||||
|
||||
### {术语}
|
||||
|
||||
- 定义:{precise definition}
|
||||
- 使用场景:{feature/module/context}
|
||||
- 备注:{ambiguities, synonyms, or rejected meanings}
|
||||
|
||||
## 业务规则
|
||||
|
||||
- {rule}: {meaning and source}
|
||||
```
|
||||
|
||||
## ADR 模板
|
||||
|
||||
```markdown
|
||||
# ADR-{编号}: {决策标题}
|
||||
|
||||
**状态**:提议中 | 已接受 | 已废弃
|
||||
**日期**:YYYY-MM-DD
|
||||
|
||||
## 背景
|
||||
|
||||
是什么情况迫使我们做这个决策?
|
||||
|
||||
## 决策
|
||||
|
||||
我们选择了什么?
|
||||
|
||||
## 替代方案
|
||||
|
||||
| 方案 | 拒绝原因 |
|
||||
| --- | --- |
|
||||
| {option} | {reason} |
|
||||
|
||||
## 后果
|
||||
|
||||
### 正面
|
||||
|
||||
- {benefit}
|
||||
|
||||
### 负面
|
||||
|
||||
- {cost or risk}
|
||||
```
|
||||
|
||||
## 技术调研模板
|
||||
|
||||
```markdown
|
||||
# {标题} 技术调研
|
||||
|
||||
## 摘要
|
||||
|
||||
- 变更原因:{reason}
|
||||
- 变更范围:{scope}
|
||||
- 主要技术方案:{approach}
|
||||
|
||||
## 源产物
|
||||
|
||||
- OpenSpec change: `openspec/changes/{slug}/`
|
||||
- 关联 PRD: `prd.md` 或 `brief.md`
|
||||
|
||||
## 关键发现
|
||||
|
||||
- {finding}
|
||||
|
||||
## 假设
|
||||
|
||||
- {assumption and validation status}
|
||||
```
|
||||
|
||||
## 设计模板
|
||||
|
||||
```markdown
|
||||
# {标题} 设计
|
||||
|
||||
## 架构摘要
|
||||
|
||||
描述输入 → 处理 → 输出。
|
||||
|
||||
## 关键决策
|
||||
|
||||
- {decision}: {reason}
|
||||
|
||||
## 模块地图
|
||||
|
||||
| 模块 | 职责 | 备注 |
|
||||
| --- | --- | --- |
|
||||
| {module} | {responsibility} | {notes} |
|
||||
|
||||
## 架构审计
|
||||
|
||||
- 风险:{risk}
|
||||
- 缓解:{mitigation}
|
||||
```
|
||||
|
||||
## 任务模板
|
||||
|
||||
```markdown
|
||||
# {标题} 任务
|
||||
|
||||
## 需求追踪
|
||||
|
||||
| 需求 | 状态 | 备注 |
|
||||
| --- | --- | --- |
|
||||
| {requirement} | 已完成 / 待处理 / 部分完成 | {notes} |
|
||||
|
||||
## 实现任务
|
||||
|
||||
- [ ] {task}
|
||||
```
|
||||
|
||||
## 验收模板
|
||||
|
||||
```markdown
|
||||
# {标题} 验收
|
||||
|
||||
## 结果
|
||||
|
||||
已接受 / 部分接受 / 未接受。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- 命令/检查:`{command or check}`
|
||||
- 结果:{passed/failed/not run}
|
||||
- 备注:{important output or reason not run}
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- 命令:`{command}`
|
||||
- 结果:{passed/failed/not run}
|
||||
- 备注:{important output or reason not run}
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 步骤:{manual steps}
|
||||
- 结果:{passed/failed/not run}
|
||||
- 备注:{observations or reason not run}
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- {completed behavior}
|
||||
|
||||
## 已知限制
|
||||
|
||||
- {limitation}
|
||||
|
||||
## Bug 修复和诊断
|
||||
|
||||
- {bug}: {diagnosis summary and regression coverage}
|
||||
|
||||
## 交接
|
||||
|
||||
- 下一步:{archive, deploy, review, or follow-up}
|
||||
- OpenSpec 归档确认:{已询问/用户确认归档/用户暂不归档/不适用}
|
||||
```
|
||||
|
||||
## 复合知识模板
|
||||
|
||||
```markdown
|
||||
# {标题}
|
||||
|
||||
**类型**:learning | trick | decision | explore
|
||||
**日期**:YYYY-MM-DD
|
||||
|
||||
## 背景
|
||||
|
||||
这条经验来自哪里?
|
||||
|
||||
## 经验
|
||||
|
||||
未来代理应该复用什么经验?
|
||||
|
||||
## 适用性
|
||||
|
||||
什么时候适用?什么时候不适用?
|
||||
```
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
---
|
||||
name: tdd
|
||||
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
|
||||
---
|
||||
|
||||
# Test-Driven Development
|
||||
|
||||
## Philosophy
|
||||
|
||||
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
|
||||
|
||||
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
|
||||
|
||||
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
|
||||
|
||||
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
|
||||
|
||||
## Anti-Pattern: Horizontal Slices
|
||||
|
||||
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
|
||||
|
||||
This produces **crap tests**:
|
||||
|
||||
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
|
||||
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
|
||||
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
|
||||
- You outrun your headlights, committing to test structure before understanding the implementation
|
||||
|
||||
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
|
||||
|
||||
```
|
||||
WRONG (horizontal):
|
||||
RED: test1, test2, test3, test4, test5
|
||||
GREEN: impl1, impl2, impl3, impl4, impl5
|
||||
|
||||
RIGHT (vertical):
|
||||
RED→GREEN: test1→impl1
|
||||
RED→GREEN: test2→impl2
|
||||
RED→GREEN: test3→impl3
|
||||
...
|
||||
```
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1. Planning
|
||||
|
||||
When exploring the codebase, use the project's domain glossary so that test names and interface vocabulary match the project's language, and respect ADRs in the area you're touching.
|
||||
|
||||
Before writing any code:
|
||||
|
||||
- [ ] Confirm with user what interface changes are needed
|
||||
- [ ] Confirm with user which behaviors to test (prioritize)
|
||||
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
|
||||
- [ ] Design interfaces for [testability](interface-design.md)
|
||||
- [ ] List the behaviors to test (not implementation steps)
|
||||
- [ ] Get user approval on the plan
|
||||
|
||||
Ask: "What should the public interface look like? Which behaviors are most important to test?"
|
||||
|
||||
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
|
||||
|
||||
### 2. Tracer Bullet
|
||||
|
||||
Write ONE test that confirms ONE thing about the system:
|
||||
|
||||
```
|
||||
RED: Write test for first behavior → test fails
|
||||
GREEN: Write minimal code to pass → test passes
|
||||
```
|
||||
|
||||
This is your tracer bullet - proves the path works end-to-end.
|
||||
|
||||
### 3. Incremental Loop
|
||||
|
||||
For each remaining behavior:
|
||||
|
||||
```
|
||||
RED: Write next test → fails
|
||||
GREEN: Minimal code to pass → passes
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- One test at a time
|
||||
- Only enough code to pass current test
|
||||
- Don't anticipate future tests
|
||||
- Keep tests focused on observable behavior
|
||||
|
||||
### 4. Refactor
|
||||
|
||||
After all tests pass, look for [refactor candidates](refactoring.md):
|
||||
|
||||
- [ ] Extract duplication
|
||||
- [ ] Deepen modules (move complexity behind simple interfaces)
|
||||
- [ ] Apply SOLID principles where natural
|
||||
- [ ] Consider what new code reveals about existing code
|
||||
- [ ] Run tests after each refactor step
|
||||
|
||||
**Never refactor while RED.** Get to GREEN first.
|
||||
|
||||
## Checklist Per Cycle
|
||||
|
||||
```
|
||||
[ ] Test describes behavior, not implementation
|
||||
[ ] Test uses public interface only
|
||||
[ ] Test would survive internal refactor
|
||||
[ ] Code is minimal for this test
|
||||
[ ] No speculative features added
|
||||
```
|
||||
@@ -0,0 +1,33 @@
|
||||
# Deep Modules
|
||||
|
||||
From "A Philosophy of Software Design":
|
||||
|
||||
**Deep module** = small interface + lots of implementation
|
||||
|
||||
```
|
||||
┌─────────────────────┐
|
||||
│ Small Interface │ ← Few methods, simple params
|
||||
├─────────────────────┤
|
||||
│ │
|
||||
│ │
|
||||
│ Deep Implementation│ ← Complex logic hidden
|
||||
│ │
|
||||
│ │
|
||||
└─────────────────────┘
|
||||
```
|
||||
|
||||
**Shallow module** = large interface + little implementation (avoid)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────┐
|
||||
│ Large Interface │ ← Many methods, complex params
|
||||
├─────────────────────────────────┤
|
||||
│ Thin Implementation │ ← Just passes through
|
||||
└─────────────────────────────────┘
|
||||
```
|
||||
|
||||
When designing interfaces, ask:
|
||||
|
||||
- Can I reduce the number of methods?
|
||||
- Can I simplify the parameters?
|
||||
- Can I hide more complexity inside?
|
||||
@@ -0,0 +1,31 @@
|
||||
# Interface Design for Testability
|
||||
|
||||
Good interfaces make testing natural:
|
||||
|
||||
1. **Accept dependencies, don't create them**
|
||||
|
||||
```typescript
|
||||
// Testable
|
||||
function processOrder(order, paymentGateway) {}
|
||||
|
||||
// Hard to test
|
||||
function processOrder(order) {
|
||||
const gateway = new StripeGateway();
|
||||
}
|
||||
```
|
||||
|
||||
2. **Return results, don't produce side effects**
|
||||
|
||||
```typescript
|
||||
// Testable
|
||||
function calculateDiscount(cart): Discount {}
|
||||
|
||||
// Hard to test
|
||||
function applyDiscount(cart): void {
|
||||
cart.total -= discount;
|
||||
}
|
||||
```
|
||||
|
||||
3. **Small surface area**
|
||||
- Fewer methods = fewer tests needed
|
||||
- Fewer params = simpler test setup
|
||||
@@ -0,0 +1,59 @@
|
||||
# When to Mock
|
||||
|
||||
Mock at **system boundaries** only:
|
||||
|
||||
- External APIs (payment, email, etc.)
|
||||
- Databases (sometimes - prefer test DB)
|
||||
- Time/randomness
|
||||
- File system (sometimes)
|
||||
|
||||
Don't mock:
|
||||
|
||||
- Your own classes/modules
|
||||
- Internal collaborators
|
||||
- Anything you control
|
||||
|
||||
## Designing for Mockability
|
||||
|
||||
At system boundaries, design interfaces that are easy to mock:
|
||||
|
||||
**1. Use dependency injection**
|
||||
|
||||
Pass external dependencies in rather than creating them internally:
|
||||
|
||||
```typescript
|
||||
// Easy to mock
|
||||
function processPayment(order, paymentClient) {
|
||||
return paymentClient.charge(order.total);
|
||||
}
|
||||
|
||||
// Hard to mock
|
||||
function processPayment(order) {
|
||||
const client = new StripeClient(process.env.STRIPE_KEY);
|
||||
return client.charge(order.total);
|
||||
}
|
||||
```
|
||||
|
||||
**2. Prefer SDK-style interfaces over generic fetchers**
|
||||
|
||||
Create specific functions for each external operation instead of one generic function with conditional logic:
|
||||
|
||||
```typescript
|
||||
// GOOD: Each function is independently mockable
|
||||
const api = {
|
||||
getUser: (id) => fetch(`/users/${id}`),
|
||||
getOrders: (userId) => fetch(`/users/${userId}/orders`),
|
||||
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
|
||||
};
|
||||
|
||||
// BAD: Mocking requires conditional logic inside the mock
|
||||
const api = {
|
||||
fetch: (endpoint, options) => fetch(endpoint, options),
|
||||
};
|
||||
```
|
||||
|
||||
The SDK approach means:
|
||||
- Each mock returns one specific shape
|
||||
- No conditional logic in test setup
|
||||
- Easier to see which endpoints a test exercises
|
||||
- Type safety per endpoint
|
||||
@@ -0,0 +1,10 @@
|
||||
# Refactor Candidates
|
||||
|
||||
After TDD cycle, look for:
|
||||
|
||||
- **Duplication** → Extract function/class
|
||||
- **Long methods** → Break into private helpers (keep tests on public interface)
|
||||
- **Shallow modules** → Combine or deepen
|
||||
- **Feature envy** → Move logic to where data lives
|
||||
- **Primitive obsession** → Introduce value objects
|
||||
- **Existing code** the new code reveals as problematic
|
||||
@@ -0,0 +1,61 @@
|
||||
# Good and Bad Tests
|
||||
|
||||
## Good Tests
|
||||
|
||||
**Integration-style**: Test through real interfaces, not mocks of internal parts.
|
||||
|
||||
```typescript
|
||||
// GOOD: Tests observable behavior
|
||||
test("user can checkout with valid cart", async () => {
|
||||
const cart = createCart();
|
||||
cart.add(product);
|
||||
const result = await checkout(cart, paymentMethod);
|
||||
expect(result.status).toBe("confirmed");
|
||||
});
|
||||
```
|
||||
|
||||
Characteristics:
|
||||
|
||||
- Tests behavior users/callers care about
|
||||
- Uses public API only
|
||||
- Survives internal refactors
|
||||
- Describes WHAT, not HOW
|
||||
- One logical assertion per test
|
||||
|
||||
## Bad Tests
|
||||
|
||||
**Implementation-detail tests**: Coupled to internal structure.
|
||||
|
||||
```typescript
|
||||
// BAD: Tests implementation details
|
||||
test("checkout calls paymentService.process", async () => {
|
||||
const mockPayment = jest.mock(paymentService);
|
||||
await checkout(cart, payment);
|
||||
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
|
||||
});
|
||||
```
|
||||
|
||||
Red flags:
|
||||
|
||||
- Mocking internal collaborators
|
||||
- Testing private methods
|
||||
- Asserting on call counts/order
|
||||
- Test breaks when refactoring without behavior change
|
||||
- Test name describes HOW not WHAT
|
||||
- Verifying through external means instead of interface
|
||||
|
||||
```typescript
|
||||
// BAD: Bypasses interface to verify
|
||||
test("createUser saves to database", async () => {
|
||||
await createUser({ name: "Alice" });
|
||||
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
|
||||
expect(row).toBeDefined();
|
||||
});
|
||||
|
||||
// GOOD: Verifies through interface
|
||||
test("createUser makes user retrievable", async () => {
|
||||
const user = await createUser({ name: "Alice" });
|
||||
const retrieved = await getUser(user.id);
|
||||
expect(retrieved.name).toBe("Alice");
|
||||
});
|
||||
```
|
||||
@@ -0,0 +1,76 @@
|
||||
---
|
||||
name: to-prd
|
||||
description: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.
|
||||
---
|
||||
|
||||
This skill takes the current conversation context and codebase understanding and produces a PRD. Do NOT interview the user — just synthesize what you already know.
|
||||
|
||||
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
|
||||
|
||||
## Process
|
||||
|
||||
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the PRD, and respect any ADRs in the area you're touching.
|
||||
|
||||
2. Sketch out the major modules you will need to build or modify to complete the implementation. Actively look for opportunities to extract deep modules that can be tested in isolation.
|
||||
|
||||
A deep module (as opposed to a shallow module) is one which encapsulates a lot of functionality in a simple, testable interface which rarely changes.
|
||||
|
||||
Check with the user that these modules match their expectations. Check with the user which modules they want tests written for.
|
||||
|
||||
3. Write the PRD using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
|
||||
|
||||
<prd-template>
|
||||
|
||||
## Problem Statement
|
||||
|
||||
The problem that the user is facing, from the user's perspective.
|
||||
|
||||
## Solution
|
||||
|
||||
The solution to the problem, from the user's perspective.
|
||||
|
||||
## User Stories
|
||||
|
||||
A LONG, numbered list of user stories. Each user story should be in the format of:
|
||||
|
||||
1. As an <actor>, I want a <feature>, so that <benefit>
|
||||
|
||||
<user-story-example>
|
||||
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
|
||||
</user-story-example>
|
||||
|
||||
This list of user stories should be extremely extensive and cover all aspects of the feature.
|
||||
|
||||
## Implementation Decisions
|
||||
|
||||
A list of implementation decisions that were made. This can include:
|
||||
|
||||
- The modules that will be built/modified
|
||||
- The interfaces of those modules that will be modified
|
||||
- Technical clarifications from the developer
|
||||
- Architectural decisions
|
||||
- Schema changes
|
||||
- API contracts
|
||||
- Specific interactions
|
||||
|
||||
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
|
||||
|
||||
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
|
||||
|
||||
## Testing Decisions
|
||||
|
||||
A list of testing decisions that were made. Include:
|
||||
|
||||
- A description of what makes a good test (only test external behavior, not implementation details)
|
||||
- Which modules will be tested
|
||||
- Prior art for the tests (i.e. similar types of tests in the codebase)
|
||||
|
||||
## Out of Scope
|
||||
|
||||
A description of the things that are out of scope for this PRD.
|
||||
|
||||
## Further Notes
|
||||
|
||||
Any further notes about the feature.
|
||||
|
||||
</prd-template>
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
name: zoom-out
|
||||
description: Tell the agent to zoom out and give broader context or a higher-level perspective. Use when you're unfamiliar with a section of code or need to understand how it fits into the bigger picture.
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary vocabulary.
|
||||
Reference in New Issue
Block a user