Case Study #5
Our first IDE test: the audit pipeline reviewed its own expansion.
We wired MegaLens into the editor and ran it on a real task: our own expansion from three audit styles to five. It reviewed the design before a line of code was written, then audited the finished diff. This was the first end to end run on work headed for production.
5
Findings the second pass added
~$2.07
Pay as you go price at the April rate
3 of 4
Risks caught pre-code
49/49
Regression pass rate
How we choose the models and combine their work is proprietary to MegaLens.
The setup
The task was narrow but real. We were expanding MegaLens from three audit styles to five, adding one style shaped for research and one shaped for diffs.
The founder was working inside his IDE. Instead of pasting context into a separate chat, the coding assistant in the editor called MegaLens through the connector and got back a structured consultation. We used that path twice: once before any code existed, once after the implementation was done.
We ran it this way to measure three things: context savings for the host assistant, coverage of blind spots, and raw execution metrics.
Before code was written
Design consultation from inside the editor. The goal was to flag structural risks before they ship.
The pre-code pass surfaced 3 of the 4 load-bearing risks in the design. These were not cosmetic notes. They were the kind of mistakes that send implementation down the wrong path and make the later review more expensive. One of them would have led us to apply an obvious-looking fix that was actually wrong.
We changed the implementation plan before touching a file. Our own separate review process caught one more risk in the next round.
What the pre-code pass actually caught
A drift risk between the public product surface and what the IDE path exposed.
Two declared capability sets had to stay in sync. The pass before coding flagged the drift pattern before the new styles were routed, which forced a single source of truth for both.
A classification mistake that would have routed design requests the wrong way.
Without this catch, the obvious fix would have been to relabel the style type. The pass before coding pointed at the routing layer instead, which was the actual defect.
A guardrail gap where design-from-scratch work could slip past the intended boundary.
The new style could have been used for unbounded design work. We tightened the guardrail before the capability was live, not after.
After the diff was written
Audit consultation on the finished implementation. A first pass of reviewing models, then a second pass over their findings.
The first pass surfaced a set of issues. Then the second pass went over the same diff and added 5 findings the first pass had missed, including the one real high severity defect in the review set.
The defect was a rubber stamp pattern, meaning approval without a real check. Under a specific input shape, a diff with no meaningful changes could have been approved without the layer checking that the target file was present. A false approval like this only shows up later, when something downstream depends on it being correct.
Diff-review cross-check
5
Second pass findings the first pass missed
1
Real high-severity defects caught
3
Rounds of our separate review before ship
49 / 49
Regression cases passing
In this session the second pass, the read after the first read, did the most important work. It is the reason the high severity defect did not ship.
Efficiency in the IDE
What the IDE connector sent and received in this session.
Host token savings were not measured in this session.
The IDE called MegaLens for two reviews and received structured findings.
108,613
Consultation exchange (input + output)
The numbers
108,613
Consultation tokens (input + output)
$0.214
Raw provider cost
~$2.07
Pay as you go price, April rate
3 of 4
Pre-code risks caught
5
Extra catches in the second pass
49 / 49
Regression pass rate
Two latency numbers are worth naming. The pre-code consultation took 153 seconds. The diff audit took 327 seconds. Staff-engineer review speeds — designed for the moments that matter: before a branch, before a ship, at the inflection points of a change.
What this changed our mind about
In this session, the IDE did not have to write the review itself.
The IDE called MegaLens for two reviews and received structured findings. The consultation used 108,613 tokens. Host token savings were not measured.
A second pass over the first one matters.
The second pass caught 5 issues the first pass missed, including the one real high severity defect. This is the cleanest signal in the whole session. The blind spots are not theoretical. They surface on our own code too.
Two separate reviews caught all four risks.
MegaLens caught 3 of 4 design risks before code. Our own separate review process caught the fourth in the next round.
Limitations
One session, one product change. Ratios from a single run do not generalize to every task. We are publishing this because it was our first end to end IDE run, not because it is a universal benchmark.
Host token savings were not measured in this session.
The product reviewed its own expansion. In this session, the review improved the tool itself. Structural risks came up before code existed. After the code existed, a high severity defect was caught that the first reviewer pass missed. Council review added the final cross check on top, which is how we want every serious ship to go.
Implementation details, tool names, and architecture specifics are left out on purpose. We share the process and the numbers.
Try MegaLens Free