Case Studies

Real audits. Clear outcomes.

These are working sessions on real products: what we looked at, what surfaced, and what changed afterward.

Second Opinion39% became 6.7%Sept 2026

The number was right. What it counted was wrong.

One AI model analysed 32 security audits, checked its own work, and passed it. Two models from two other AI companies found the error by two different routes. What the data really said, and what shipping it would have cost.

Internal review council · not a product run · models not named

Claude Code167 TestsVideoLive Demo

What happens when Claude Code gets independent AI reviewers?

A recorded build of an AI Email Drafter. Claude Code wrote and reviewed its own plan, then the same work went through MegaLens for independent review: 15 actionable findings, 12 of them not in Claude's own review, and 10 more during implementation. Claude stayed the decision-maker throughout.

Full screen recording · 15 findings · all 167 tests passed

Claude CodeCursor60 FindingsIDE Comparison

Claude Code vs Cursor: both miss the same gaps.

Two real projects, two IDEs, same result. On an AI Email Drafter, 12 of MegaLens's 15 findings were not on Claude Code's own list. On a PR Description Drafter, Cursor's own review found 23 issues and MegaLens found 45. Single-model review has a ceiling.

2 projects · 2 IDEs · 60 MegaLens findings · comparison table · public repos

2 Critical7 High167 TestsPlan Audit + Build

15 gaps caught before writing a single line of code.

Claude Code wrote the plan and the code for an AI Email Drafter. MegaLens reviewed the plan before implementation and each commit during the build. 12 of 15 findings were additions beyond our own pre-audit; 3 overlapped. Full repo included.

Plan audit + 12 per-commit reviews · 15 findings · public repo · live build video

3 Critical3 HighLegal Compliance

We audited our own legal setup and found 9 risks in 5 minutes.

Before drafting a privacy policy, we audited our own product architecture. The review surfaced critical gaps in GDPR readiness, cross-border transfers and Chinese provider disclosure.

9 findings · 3 critical · 4m 55s

Plan Review23 IssuesUI Plan

23 issues found before writing a single line of code.

Two reviewers checked a 7-step UI plan before any code was written. They found 23 issues across security, architecture, UX and performance. 15 changes were made to the plan before the build.

2 reviewers · 23 issues · 15 plan changes before the build

14 Post-Test Issues50% Unique per ReviewerSecurity Review

Passing tests didn't mean it was safe to ship.

74 files, 10 end-to-end tests, all passing. Two AI reviewers in our own development process still found 14 more issues: concurrency bugs, silent failures, and credential exposure risks. This review ran outside MegaLens.

2 AI reviewers · 17 total issues · 14 fixed same session

3 Blockers4 HighsSecurity Remediation

The SSRF fix that passed its tests and was still unsafe to ship.

A production SaaS had 7 security findings to patch in one session. The first-pass fixes compiled and tested. Independent review then caught two bugs hiding inside the patches themselves.

2 reviewers · 2 review rounds · 4 files · 14 tunnel-aware test cases · same-day deploy

First IDE Test3 of 4 Risks CaughtSelf-Audit

Our audit pipeline reviewed its own expansion — from inside the editor.

We wired MegaLens into the IDE and ran the first end-to-end test on a real task: our own expansion. It caught 3 of 4 structural risks before code existed. A second pass caught 5 more that the first pass missed, including the one real high severity defect.

3 of 4 risks caught pre-code · 49/49 regression · $0.98