Blog
Insights & analysis
Notes on code review with several AI models, and what we learned building it.
Spec driven development with Claude Code: the workflow we use for every app
Our spec driven development workflow for Claude Code, from idea to shipped code: a written spec, a plan reviewed by AI models from other companies before any code, tests before code, and nothing ships without your go.
6 min read · install the skill from GitHub
Best AI code review tools in 2026
Eight AI code review tools compared by when they review your code, where they run, what they cost and which ones are free. Plus how to do an AI code review.
Prices checked on each vendor's site · sources listed · FAQ
OpenRouter errors: 404, 429, 401 and “provider returned error”, and how to fix each
What each error means, what causes it and the fix, with the exact messages the API sends back and code for fallbacks and retries.
Checked against OpenRouter's docs and the live API · code · FAQ
LLM speed in real code review: what public benchmarks did not tell us
We spent about $95 trying to beat our own lineup. Models that look fast timed out, the fastest setup found 4 of 10 known bugs, and a model we used disappeared.
Checked by AI models against a known answer key · method included
We ran AI code review on 30 apps. Then we fact-checked our own findings.
160 findings came back. Two AI models from different companies checked a random sample of 20. On the findings they agreed on, 44% did not hold up, and not one critical label survived both.
Validated by AI models only, no human review · method included
AI coding mistakes that pass every test: 8 patterns and how to catch them
One AI coding session passed 640 tests and was still wrong eight different ways. None of them was bad code. For each mistake: what it was, why the tests missed it, and a simple check for it.
First hand record · a check for every pattern · FAQ
My pipeline ruled a true claim false, then destroyed four correct findings with it
A claim check using web search fetched six pages, none of them the package it was judging, and filed a paraphrase of the claim as evidence against it. Then one disputed claim about two plan steps discarded all four findings, including a critical one it had nothing to do with. This post shows the stored record, all three defects, and what changed.
7 min read · stored record excerpt · fixture list
Cursor alternatives: extend Cursor before you switch
You can keep Cursor and add MCP tools that give it review from several engines. We tested Cursor alone and Cursor with MegaLens on a real project.
5 min read · comparison table · case study linked