SDD 02 - Lab
SDD Learning · Previous: Foundations · Next: Matt Pocock Skills starter guide
Levels 2 to 6 · Specify, Plan, Implement, Verify, Review. Do level 1 first if the four terms request, contract, plan and evidence are new to you.
Your task is to add JSON output to an existing Markdown indexing CLI. Your evidence must show that the new behavior works and that the old behavior remains intact.

Colors mean the same thing in every diagram of this Learning: see the color key.
This page describes what is the same for everyone: the contract, what each level asks of you and how you know you have passed it. The steps to type are on the lab route of the framework you choose.
Choose your route
| Route | Walk it when | Open |
|---|---|---|
| Matt Pocock Skills | First. This is the baseline route | Matt Pocock Skills lab route |
| Spec Kit | You want explicit specification, plan and task artifacts | Spec Kit lab route |
| OpenSpec | You want to change a maintained contract through deltas | OpenSpec lab route |
| BMAD Method | You want to match planning depth to the work | BMAD Method lab route |
| Superpowers | You want test-driven execution and fresh verification | Superpowers lab route |
Start with Matt Pocock Skills. That run becomes your baseline: when you later replay the same change on another route, you have something of your own to compare it with. Every route has the same five levels, the same contract and the same checks, so you can also begin with a different one if that is the tool you came for.
How to move through the levels
Work at your own pace. Each level ends with a check; when you pass it, go on to the next level. If a level does not come together today, write down exactly where you are and come back to it: the next attempt starts from that note.
Set up once
Download the starter lab, unpack it and work in doc-index-starter. The package contains the program, synthetic documents, four baseline tests and training/WORKSHEET.md for your observations.
$ python3 -m unittest discover -v $ python3 doc_index.py sample-docs
The four baseline tests pass and the CLI prints two tab-separated filename/title rows. Keep a clean copy of this directory for every route you walk. Set up the framework inside the lab directory as its lab route describes, and record the installed version in the worksheet. Keep credentials in the agent's normal configuration, outside the worksheet.
Level 2 · Specify: agree on the contract
Inspect the baseline code before asking the agent to modify it. The existing CLI scans top-level Markdown files, sorts by filename and uses the first non-empty H1 as the title, falling back to the filename stem.
| ID | Acceptance criterion |
|---|---|
| AC1 | Default text output stays byte-for-byte identical |
| AC2 | --format json returns a valid JSON array of records |
| AC3 | Each record has exactly the string fields file and title, sorted by filename |
| AC4 | An empty directory produces [] in JSON mode |
| AC5 | Unknown formats and missing directories fail with nonzero status, useful stderr and empty stdout |
| AC6 | Quotes and non-ASCII characters survive JSON encoding and decoding |
| AC7 | Top-level-only discovery and title fallback remain unchanged |
Exclude recursive discovery, network access, a database, a web interface and third-party Python dependencies. Whitespace and object key order in JSON are not acceptance criteria.
Every lab route carries this contract in its first prompt. Ask the agent to retain the acceptance IDs so you can trace them later. Resolve ambiguities before planning.

Level 2 check: a partner can explain the promised behavior and the exclusions from the artifact alone. Record the accepted spec path in the worksheet. If the agent generated different semantics, correct the contract before implementation.
Level 2 complete. You turned “add JSON” into a contract someone else can check. Good work: this is the step most people skip.
Level 3 · Plan: one bounded change
Produce the plan with the steps of your route. Reuse the current collection behavior. Introduce the output selection at an appropriate seam, and plan subprocess tests of the actual CLI.
Complete at least three rows of the worksheet's requirement-to-evidence table before implementation. One row must cover compatibility and one must cover an error case.
Level 3 check: the plan identifies how AC1 will be preserved and checked. A plan that only tests the new JSON path is incomplete.
Level 3 complete. Every criterion you care about now points to a task and a check. From here on you build what you have already decided.
Level 4 · Implement: build with feedback

Run the implementation step of your route. Ask for one observable behavior at a time, with a relevant failing test before the change and a passing result afterward. Inspect the resulting code diff for unrelated changes.
Do not ask for a rewrite of the whole CLI. The small existing program is intentional: the exercise is about a reviewable change and its contract.
Level 4 check: the implementation supports JSON, the original tests still pass, and the agent can point to new tests for the feature. If something is still open, describe the remaining work instead of calling it complete.
Level 4 complete. The feature exists and the old behavior is still there. Run it once more, just to see your JSON come out.
Level 5 · Verify: check the final revision
Run the commands yourself in the lab directory:
$ python3 -m unittest discover -v $ python3 doc_index.py sample-docs $ python3 doc_index.py sample-docs --format json
The JSON should parse to this value:
[{"file":"alpha.md","title":"Alpha"},{"file":"beta.md","title":"beta"}]Ask your partner which tests cover empty input, bad arguments, escaping, title fallback and top-level discovery. Record the actual commands, results and the revision or change identifier in the worksheet, so it is clear which code the evidence belongs to.
Whoever guides the session can run the independent checker from the trainer kit against your working directory, and you can run it yourself. Its checks support this exercise's criteria; they are not a complete production security or reliability assessment.
Level 5 check: the evidence is from the final code and every unmet criterion is visible. Run your route's own consistency, convergence or verification step where it has one, then inspect its findings alongside the application tests.
Level 5 complete. You can show what ran and what it returned. Enjoy the passing run.
Level 6 · Review: compare and hand off
Pair up and exchange the contract and diff. Each reviewer asks:
- Which requirement does this code change implement?
- What test would fail if the behavior were removed?
- Did any old behavior change without agreement?
- Which observations are actual executions, and which are assumptions?
- What must the next session know before continuing?
Level 6 check: your worksheet states accepted, incomplete or needs revision, and explains why. Point to the maintained specification and any remaining task. For work that is still open, a precise handoff is a good result.
Level 6 complete. You have walked the whole loop from intent to review and can point to what each step left behind. Take a moment to enjoy that before you go on.
Next: replay it on another route
Keep the worksheet from your first run: it is your baseline. Start from a fresh copy of the starter lab, reuse the accepted contract and walk another route from the table above. Keep the agent, model and task fixed where practical.
Compare the artifacts and corrections each route required. Do not rank correctness by the number of documents, skills or agents. Use the comparison protocol to record observations, and the knowledge map to see how the tracks relate.