SDD 14 - Framework Comparison
SDD Learning · Previous: Visual Knowledge Map · Next: Trainer Guide
Start with the failure you want to fix. If decisions disappear between sessions, improve the durable record. If the system evolves without a clear contract, manage specification changes. If the agent rushes through implementation, improve testing and review. If product, architecture and delivery drift apart, make those handoffs explicit.
Matt Pocock Skills, Spec Kit, OpenSpec, BMAD Method and Superpowers overlap, but their strongest organizing ideas differ. Choosing by that difference is more useful than choosing by repository popularity.

Colors mean the same thing in every diagram of this Learning: see the color key. The five methods are drawn in white on purpose: color never ranks a framework.
The short answer
| Your main need | Start by evaluating | Reason |
|---|---|---|
| Composable engineering workflows and project knowledge | Matt Pocock Skills | Connects clarification, domain language, specs, tickets and implementation |
| A repeatable route from a feature requirement to a plan | Spec Kit | Makes specification, planning, tasks and convergence explicit |
| Controlled evolution of an existing system contract | OpenSpec | Separates current specifications from proposed change deltas |
| Product-to-delivery coordination | BMAD Method | Provides planning, specialist perspectives, stories and build workflows |
| Better design, testing, debugging and completion habits | Superpowers | Organizes engineering routines around the agent's execution process |
These are starting hypotheses. For example, OpenSpec also works with new projects, BMAD Method has a small-change path, and Spec Kit now documents optional processes outside SDD. The categories describe emphasis, not hard limits.
Compare the durable artifacts
| Approach | Typical durable record | Main review question |
|---|---|---|
| Matt Pocock Skills | Domain docs, ADRs, tracker specs and tickets, repository agent configuration | Are the decisions and slices clear enough for another session? |
| Spec Kit | Project constitution, feature specification, technical plan and tasks | Does implementation converge on the accepted contract? |
| OpenSpec | Current specs, change proposals, deltas, design, tasks and archive | What changed in the system's promise, and was it reconciled? |
| BMAD Method | Initiative planning artifacts, requirements, architecture decisions and tickets | Can the next role or build session act without guessing? |
| Superpowers | Design and plan files for larger work, code, tests and review evidence | Did the agent follow a suitable design and verification process? |
File locations are defaults. What matters is ownership: one maintained source for each decision, a clear handoff to implementation and evidence that corresponds to the final code.
Compare the cost of the process
| Approach | Setup shape | Cost to watch |
|---|---|---|
| Matt Pocock Skills | Skills or plugin plus repository setup | Too many skills or unnecessary ticket decomposition |
| Spec Kit | Python CLI plus project and agent scaffolding | Repetitive spec and plan documents for trivial work |
| OpenSpec | Node.js CLI plus change/spec structure | Baseline specifications drifting after implementation |
| BMAD Method | Skills or plugins plus shared project setup | Handoffs and planning artifacts beyond the needs of the change |
| Superpowers | Harness-specific plugin or package | Repeated design/review interactions or more execution machinery than needed |
These are qualitative risks, not measured overhead rankings. A longer planning step can be economical if it prevents expensive rework. A short plan can be wasteful if it is unnecessary for an obvious edit.
None of these installations supplies free model inference. Agent subscriptions, API usage, tool execution and reviewer time are separate costs. Measure them in your actual setup.
A decision map

The same team may need different methods in different repositories.
Run a fair comparison
Use the starter lab. It contains a tiny working Markdown index CLI and four baseline tests. The requested change is the same on every framework track: add JSON output while preserving existing text behavior.
Walk the Matt Pocock Skills lab route first: that run is the baseline the others are compared with. Every framework has a lab route with the same levels and checks. Make a separate copy or branch from the same starting files for each trial. Keep the coding agent, model, tool permissions, starting files and feature requirements fixed. Record each framework's version or installed revision. Do not mix framework installations into the same comparison copy.
| ID | Acceptance criterion | Evidence to collect |
|---|---|---|
| AC1 | Existing text output remains identical | Baseline subprocess test |
| AC2, AC3 | JSON is an array with exactly the documented fields, in filename order | Parse stdout and compare structured values |
| AC4 | Empty input returns an empty array | A test with a temporary empty directory |
| AC5 | Bad arguments fail cleanly | Exit status, stderr and empty stdout assertions |
| AC6 | Quotes and non-ASCII titles remain correct | Round-trip JSON test with a synthetic title |
| AC7 | Top-level discovery and title fallback are unchanged | Tests with a nested file and a file without a heading |
| Exclusions | Scope stays bounded | Diff review for unrelated changes and dependencies |
The IDs are the ones the lab uses, so results from different tracks can be laid side by side.
Start with one trial to understand each workflow, then repeat promising candidates on a real feature and a real bug. A tiny CLI exercise does not reveal how a method handles a multi-repository migration or months of specification maintenance.
Use a results sheet like this, filling in observations rather than estimated scores:
| Observation | What to record |
|---|---|
| Acceptance | Which criteria passed on the final revision |
| Human steering | Clarifications, corrections and review effort |
| Rework | Wrong assumptions discovered after implementation began |
| Usage | Reported model and tool usage, where available |
| Artifact quality | Useful decisions versus duplicate or stale documents |
| Fresh-session handoff | Whether a new session understands the next task |
| Maintenance | Whether the second change updates existing knowledge correctly |
Do not treat one successful run as a reliable ranking. Agent execution has variation, and a framework can appear better because its trial received more human help.
Can I combine them?
Yes, but define which system owns the workflow and each artifact. The following are proposed composition patterns, not documented turnkey integrations:
| Primary owner | Possible complement | Boundary to define |
|---|---|---|
| OpenSpec owns change contracts | A selected TDD or debugging discipline | OpenSpec specs remain authoritative; no duplicate requirement system |
| Spec Kit owns feature planning | A focused code-review skill | Review the existing spec and plan; do not generate competing ones |
| BMAD owns product and story planning | Repository-specific testing rules | Build consumes the accepted story and normal test commands |
| Matt Pocock Skills or Superpowers owns execution | Existing ADRs and domain docs | Keep maintained knowledge in the repository's chosen locations |
Start by evaluating them separately. Combining all five means five possible answers to “what should happen next?” and several places to record the same requirement. That is a coordination problem you would have to maintain yourself.
Where the coding agent, model and MCP fit
The framework supplies instructions and artifacts. The coding agent runs the interaction and exposes tools. The model reasons over the context it receives. Files, shell commands, test runners and optional MCP servers provide capabilities or information.
MCP can be useful when a workflow needs access to a tracker or a documentation service. It is not a prerequisite for this starter lab and does not turn a prose specification into an executable test. All examples can begin with local files and the agent's normal repository tools.
Likewise, installing a method does not make every model equally reliable. Evaluate your intended agent/model combination, especially when using a smaller local model with limited context or inconsistent tool use.
Recommendations for common engineering work
A small personal utility
Begin with your existing agent and normal tests. Try Matt's bounded implementation workflow or Superpowers if clarification, debugging or verification is the recurring weakness. Add durable specifications only when they solve an actual handoff problem.
An established service with compatibility requirements
Evaluate OpenSpec first if changes repeatedly lose track of what existing consumers expect. Make the maintained specification and regression tests agree. Choose Spec Kit if the bigger gap is the route from a new feature request to a technical plan.
A new application with multiple contributors
Evaluate Spec Kit for a shared feature planning structure. Evaluate BMAD when product definition, architecture, UX and story coordination all need attention. In either case, require independently useful slices rather than a long sequence of unfinished layers.
Infrastructure templates and network tooling
Use the framework to state invariants and failure behavior: what must remain compatible, what is allowed to change, and how a rollback or failed deployment is handled. Validate with the actual platform and tooling. A generated “secure architecture” paragraph does not prove a routing policy, IAM boundary or deployment template is correct.
Use synthetic configuration and identifiers for training exercises. Real credentials and operational data are unnecessary for learning the workflow.
Read the individual guides
| Track | The tool | Levels 2 to 6 |
|---|---|---|
| Matt Pocock Skills, the baseline | Starter guide | Lab route |
| Spec Kit | Starter guide | Lab route |
| OpenSpec | Starter guide | Lab route |
| BMAD Method | Starter guide | Lab route |
| Superpowers | Starter guide | Lab route |
Checkpoint
Choose one recurring failure in your own workflow. Select one framework to evaluate and one observable improvement you would look for. Treat that choice as a hypothesis.
Comparison complete. You can now name the problem first and then choose a framework to evaluate for it.
Sources and interpretation
The tables summarize the following primary sources. The fit recommendations, comparison protocol and composition boundaries are my synthesis of those documented workflows; they are not claims made by all the maintainers.