Case Study #7
Claude Code vs Cursor: a second review found more on both.
We ran two real builds through two different IDEs. Claude Code built an AI Email Drafter. Cursor built a PR Description Drafter. In both cases, MegaLens returned more findings than the IDE's own review. Different IDEs, same pattern.
2
Projects tested
2
IDEs compared
60
MegaLens findings, both projects
7
Found by GPT 5.4 alone
Two real projects, two IDEs
These aren't demo apps or staged tests. Both are real tools we built and use. Each one went through the same process: the IDE wrote the plan, reviewed its own plan, then MegaLens reviewed the same plan independently.
AI Email Drafter
A backend service that polls Gmail, identifies real business inquiries, and creates draft replies from a context file. Never sends email. Claude Code wrote the 19-section build plan and built the entire tool.
2
Solo review
15
MegaLens
12
Not on Claude's list
PR Description Drafter
A self-hosted git daemon that watches local branches, reads diffs and commit messages, generates PR descriptions via AI, and creates draft pull requests on GitHub. Cursor Agent wrote the 493-line V1 engineering plan.
23
Solo review
45
MegaLens
22 more
45 vs 23
Same pattern, both IDEs
The numbers are different because the projects are different. But the pattern is identical. In both cases, the IDE's solo review did solid work and caught real issues. And in both cases, MegaLens returned more: 12 items not on Claude Code's own list, and 45 findings to Cursor's 23.
| Claude Code AI Email Drafter (run 2) | Cursor PR Drafter | |
|---|---|---|
| IDE’s own review | 12 items | 23 findings |
| MegaLens review | 15 findings | 45 findings |
| Difference | 12 of the 15 not on Claude’s list (3 overlapped) | 22 more findings (45 vs 23) |
| Critical findings | 2 in MegaLens’s 15, both new | went from 3 to 7 |
| Provider cost, own key | not recorded | $0.22 |
| MegaLens time | not recorded | 7 min |
$0.22 is what the providers charged on our own key for the Cursor run. The Claude Code run's cost and time were not recorded. On pay-as-you-go, the Claude Code run's plan review was priced at $1.25 and its mid-build code review at $1.23, both printed in the session. The Cursor run's pay-as-you-go price was not recorded.
Run 1 of the same project, with Claude Code's own pre-audit, is documented here.
The Claude Code review was a quick self-audit (12 short items), while the Cursor review was a full prompted security pass (23 items). On the same plan, MegaLens returned 45 findings to Cursor's 23. The ceiling isn't about effort. It's about having one model review from one perspective.
What MegaLens found on the PR Drafter
45 findings against Cursor's 23. Here are the ones that matter most.
Found by GPT 5.4 only(Grok, DeepSeek and Gemini all missed these.)
The tool runs git commands against user repositories. But Git has built-in extension points: custom diff tools, credential helpers, and aliases defined in .gitconfig. A repository with a malicious .gitconfig would execute arbitrary code every time the daemon runs "git diff" or "git log". This is a real Git feature, not a bug. Grok, DeepSeek, Gemini and Cursor all missed it.
The plan uses a fixed port (8765) for the OAuth callback. Any other program on the same machine can bind that port first and steal the GitHub authorization code.
When the daemon runs "git fetch", Git can invoke SSH with ProxyCommand or custom transport protocols. Without restricting allowed transports, a malicious remote can trigger code execution.
Found by one model only
YAML config can execute code if parsed with an unsafe loader (Grok caught, others missed)
Git commands with no timeout or output limit. Huge repo = out of memory (DeepSeek caught, others missed)
Config reload via signal accepts new paths without validation (DeepSeek only)
State file paths can be manipulated via symlinks (DeepSeek only)
Found by more than one model
Force-push breaks state tracking: duplicate PRs or retry loops
No spending cap on AI API calls
Single GitHub token used across all repos
Race condition: two changes create duplicate PRs
Branch-to-remote mapping can target wrong repository
What Claude Code missed on the Email Drafter
13 findings that MegaLens added beyond Claude Code's own review. Two were critical.
The plan had a daily spending cap but no code to actually parse costs from the API response. The cap existed on paper only.
On database loss, the tool would process every unread email in the inbox. Dozens of unwanted drafts and a surprise API bill.
Email headers passed raw into AI prompt (injection vector)
Draft-create race condition: duplicate drafts on crash
No MIME parsing spec for multipart, HTML, charset
Reply envelope headers undefined (To, Cc, In-Reply-To)
Poison messages retry forever with no quarantine
systemd Environment= exposes API keys in /proc
Gmail historyId 7-day expiry unhandled
Why switching IDEs doesn't fix it
Every AI model has systematic blind spots.
Claude Code uses Opus. Cursor uses Sonnet (or whichever model you configure). Both are excellent at generating and reviewing code. But each model was trained on different data with different priorities. They're consistently strong in different areas and consistently weak in others. Switching from one to the other gives you a different perspective, but still just one.
Self-review has a ceiling.
Whether it's Claude Code reviewing Claude Code's plan, or Cursor reviewing Cursor's plan, the dynamic is the same: one model checking its own work. The model that wrote the plan has the same blind spots when it reviews the plan. This isn't a flaw in either IDE. It's a structural limitation of single-model review.
Other models found what one model alone did not.
MegaLens selects up to four AI models from different companies to review your code. A final check reviews their findings. In the PR Drafter test, 7 findings came from GPT 5.4 alone. Grok, DeepSeek and Gemini did not raise them.
How MegaLens fits in
MegaLens is an MCP server. You install it once, and it becomes a tool your IDE can call. It works with Claude Code, Codex CLI, Cursor, Gemini CLI and Lovable, and with any tool that supports OAuth MCP connectors, such as ChatGPT and Claude.ai. GitHub Copilot is coming soon. You don't leave your IDE.
Add MegaLens to Claude Code, Codex CLI, Cursor, Gemini CLI or Lovable, or add it as an OAuth connector in ChatGPT or Claude.ai.
Ask your IDE to run megalens_debate on your code or plan.
MegaLens selects up to four AI models from different companies. Each one reviews on its own.
A final check reviews their findings.
Structured findings arrive in your IDE. Each result names the models that actually took part. MegaLens reviews. Your IDE decides what to change.
Honest notes
The two projects are different in size and complexity, so the raw numbers aren't directly comparable. What's comparable is the pattern: solo review catches a solid base, multi-engine review catches a meaningful layer on top.
The Claude Code test was a quick self-audit (12 short items), while the Cursor test was a full prompted security review (23 items). On the Cursor plan, MegaLens returned 45 findings to Cursor's 23.
MegaLens produces findings, not proofs. Every finding required human judgment to confirm and decide on a fix. These are two tests on two projects. Your results will depend on what you're reviewing and how thorough your existing process is.
What this actually means for your workflow
Two projects. Two IDEs. The data tells the same story both times: the IDE isn't the problem, and switching IDEs isn't the fix.Claude Code's own list missed 12 of MegaLens's 15 findings on the Email Drafter. On the PR Drafter, MegaLens returned 45 findings to Cursor's 23. The ceiling is structural, not a product flaw.
MegaLens selects up to four AI models from different companies to review your code. Their findings come back into the IDE you already use. MegaLens reviews. Your IDE decides what to change.
Seven findings in the PR Drafter came from GPT 5.4 alone. The other three models did not raise them. Two critical findings in the Email Drafter existed only on paper — a cost cap with no implementation and a fresh-install bug that would flood inboxes. These are the kinds of gaps that don't surface until production. Or until you have enough AI perspectives looking at the same code.
Add multi-engine review to your IDE.
Works with Claude Code, Codex CLI, Cursor, Gemini CLI and Lovable, and with OAuth connectors such as ChatGPT and Claude.ai. GitHub Copilot is coming soon. MegaLens reviews. Your IDE decides what to change.