Case Studies
Real audits. Clear outcomes.
These are working sessions on real products: what we looked at, what surfaced, and what changed afterward.
The number was right. What it counted was wrong.
One AI model analysed 32 security audits, checked its own work, and passed it. Two models from two other AI companies found the error by two different routes. What the data really said, and what shipping it would have cost.
Internal review council · not a product run · models not named
What happens when Claude Code gets independent AI reviewers?
A recorded build of an AI Email Drafter. Claude Code wrote and reviewed its own plan, then the same work went through MegaLens for independent review: 15 actionable findings, 12 of them not in Claude's own review, and 10 more during implementation. Claude stayed the decision-maker throughout.
Full screen recording · 15 findings · all 167 tests passed
Claude Code vs Cursor: both miss the same gaps.
Two real projects, two IDEs, same result. On an AI Email Drafter, 12 of MegaLens's 15 findings were not on Claude Code's own list. On a PR Description Drafter, Cursor's own review found 23 issues and MegaLens found 45. Single-model review has a ceiling.
2 projects · 2 IDEs · 60 MegaLens findings · comparison table · public repos
15 gaps caught before writing a single line of code.
Claude Code wrote the plan and the code for an AI Email Drafter. MegaLens reviewed the plan before implementation and each commit during the build. 12 of 15 findings were additions beyond our own pre-audit; 3 overlapped. Full repo included.
Plan audit + 12 per-commit reviews · 15 findings · public repo · live build video
We audited our own legal setup and found 9 risks in 5 minutes.
Before drafting a privacy policy, we audited our own product architecture. The review surfaced critical gaps in GDPR readiness, cross-border transfers and Chinese provider disclosure.
9 findings · 3 critical · 4m 55s
23 issues found before writing a single line of code.
Two reviewers checked a 7-step UI plan before any code was written. They found 23 issues across security, architecture, UX and performance. 15 changes were made to the plan before the build.
2 reviewers · 23 issues · 15 plan changes before the build
Passing tests didn't mean it was safe to ship.
74 files, 10 end-to-end tests, all passing. Two AI reviewers in our own development process still found 14 more issues: concurrency bugs, silent failures, and credential exposure risks. This review ran outside MegaLens.
2 AI reviewers · 17 total issues · 14 fixed same session
The SSRF fix that passed its tests and was still unsafe to ship.
A production SaaS had 7 security findings to patch in one session. The first-pass fixes compiled and tested. Independent review then caught two bugs hiding inside the patches themselves.
2 reviewers · 2 review rounds · 4 files · 14 tunnel-aware test cases · same-day deploy
Our audit pipeline reviewed its own expansion — from inside the editor.
We wired MegaLens into the IDE and ran the first end-to-end test on a real task: our own expansion. It caught 3 of 4 structural risks before code existed. A second pass caught 5 more that the first pass missed, including the one real high severity defect.
3 of 4 risks caught pre-code · 49/49 regression · $0.98