Spec driven development with Claude Code: the workflow we use for every app
September 2026 · 6 min read
Here is how an app built with AI can go wrong. You describe an idea, the assistant starts writing code straight away, and a few days later nobody can say what the app was supposed to do. The code has drifted, tests have been loosened until they pass, and the only record of what you wanted is a long chat.
We build every app in the same phases, in the same order. Each phase ends at a gate, and nothing moves on until that gate is passed. We also wrote the workflow down as a Claude Code skill, so your coding assistant can follow it too. The link is at the end.
What is spec driven development?
Spec driven development means writing down what the app must do before writing the code for the planned work, then building and checking everything from that document. The document is the spec. It matters more when an AI assistant writes the code, because the assistant can produce a lot of code quickly, and the spec is the one place your intent is written down.
Phase 0: say what the app must do
You describe the app in plain words. Before anything gets written, the assistant asks you clarifying questions, at most five to seven, all in one message. Who uses it? Where is the data stored? What should happen when something goes wrong?
We keep the questions in one message so you can answer them together. No plan and no code are written in this phase.
Phase 1: write the spec
Your answers become a file called spec.md. It lists:
- What the app must do.
- How you will know each feature works. These are acceptance criteria: plain checks such as “an expense dated this month appears in this month’s total straight away”.
- What can go wrong, and how the app must behave when it does.
- Who is allowed to see and change what.
- What is out of scope.
Here is what part of a spec looks like for a simple expense tracker:
## F1 Add an expense
Amount (greater than 0), category from a fixed list,
date (defaults to today), optional note.
Acceptance: an expense dated this month appears in this month's total
straight away. An expense dated another month does not change this
month's total.
Failure: an amount of 0 or less is refused with a clear message,
and nothing is saved.
## Out of scope
Accounts and login, budgets, CSV export.The spec is the source of truth for the whole project, which is the core of spec driven development. The plan is written from it, the tests are written from it, and it goes with the work into every review. When something comes out wrong later, start with the spec. If a requirement is missing or wrong, fix it in the spec and regenerate the affected work. If the requirement was right and the code missed it, fix the code. What you never do is patch the code and leave the spec behind. The moment you do, the spec stops describing your app, and every check that relies on it is checking the wrong thing.
Phase 2: plan, then stop for review
The spec becomes plan.md: how the app is put together, which files it has, which database tables it needs, and which outside services it uses.
Then the workflow stops. Before a single line of code exists, the plan goes to MegaLens, which has AI models from other companies review it. The model that wrote a plan should not be the one that approves it. When a model rereads its own plan, it brings the same assumptions that shaped it. A model from another company comes with different training, so it can question assumptions the first one took for granted.
The review happens before code because a plan is the easiest place to fix a mistake. A wrong table design or a missing permission check takes a few lines to correct in a plan. Once code is built on top of it, the same fix means rewriting code that already works.
Your assistant applies the findings it agrees with. It lists the findings it rejects, each with a reason, and shows you that list. Then it waits for your go before moving on.
Phase 3: break the plan into tasks
The plan becomes tasks.md: small tasks, grouped into milestones. Each task can be finished and tested on its own, and none is bigger than one review sitting. Small tasks keep each review short enough to read properly.
Phase 4: tests first, then code
The assistant works on one task at a time:
- It writes the tests first, from the acceptance criteria in
spec.md. - It writes code until those tests pass.
- It runs lint and type checks, which are automatic checks for common coding mistakes.
Tests come before code so that what gets checked is chosen from the acceptance criteria, before any code exists. A test written after the code can end up describing what the code already does, bugs included. Passing tests tell you that those checks passed. They do not prove the app does everything in the spec.
A failing test is never weakened or deleted to get to green. Loosening a check or skipping a case does not fix the code it was meant to test. The rule leaves two options. Fix the code, or, when the test and the spec disagree, stop and bring the conflict to you. The assistant also keeps tasks.md up to date, so you can see where the build stands at any time.
Phase 5: review each milestone
When a milestone is finished, the assistant sends spec.mdand the diff (the exact lines that changed) to MegaLens for review. The chat that produced the code is left out. That chat is full of the builder’s reasoning, and a reviewer who reads it can start to see the code the way the builder did. The spec goes along as context, so the reviewers can see what the code was meant to do. Checking that the finished work meets the spec stays with you and your assistant. That is what the tests from Phase 4 are for.
Review findings are a filter, and none of them is a final ruling. Reviewers, including AI ones, sometimes report problems that are not there. We found the same in a sample of our own findings, checked by AI models with no human review: our research on AI code review findings. So the assistant checks each finding against the actual code before applying it. A finding that does not hold up is rejected, and you see which ones were rejected and why.
Phase 6: ship only on your go
Nothing deploys without your explicit go. Before you give it, the assistant hands you a short deploy report: what changed, how many tests there are, and any place where the build differs from spec.md, with the reason.
You are the one who answers for the app, so the decision to ship stays with you. The report is there so that your go is an informed one.
Rules for every phase
- Secrets, such as API keys, never go in the code or in anything the browser loads.
- Every database read or write checks who is asking and whether they are allowed to see or change that data.
- When a requirement turns out to be unclear in the middle of a task, the assistant stops and asks you. It never guesses and builds.
Use this workflow as a Claude Code skill
The workflow is a skill file in our public GitHub repository. Install it with:
npx skills add https://github.com/megalens/megalens-skillsOr copy the megalens-workflow folder into your Claude Code skills folder yourself. Then start a new project and tell your assistant:
Use the megalens-workflow skill. Here is my app idea: ...The review gates in Phase 2 and Phase 5 need MegaLens connected to your coding tool. Setup is at megalens.ai/try.