Case Studies/FixBounce Council Audit

Case Study #4

The SSRF fix that passed its tests and was still unsafe to ship.

April 2026. Reviewed by Gemini and Codex through our own development process, not through MegaLens.

A production email-verification SaaS had 7 security findings from an internal audit. One engineer, one session, one coordinated deploy. The first-pass fixes looked competent and complete. Independent review caught two bugs hiding inside the patches themselves.

7

Findings patched

4

Files touched

2

Review rounds

Same day

Deploy

Independent reviewers for this session: Gemini (Google) and Codex (OpenAI). Specific exploit payloads, endpoint paths, and credential values are omitted. We share the process and the numbers.

Round 0: the fixes that looked finished

The engineer worked through the findings list the way most good engineers would. Suspended users were blocked at the database layer. The bulk verification flow moved from “check, then deduct” to “debit first, refund if needed”. Hardcoded secret fallbacks were removed and replaced with environment checks that fail fast. Upload handlers stopped trusting the file size reported by the client and switched to bounded reads.

The SSRF fix also looked clean on paper. The product verifies email domains by resolving mail servers and opening SMTP connections to them. The first patch added a filter intended to allow only public IP addresses before any outbound connection. It rejected private, loopback, link-local, multicast, reserved, and unspecified ranges using Python's ipaddress module.

It passed its tests.

That still wasn't enough to ship.

Round 1: Gemini catches the bugs behind the fixes

Gemini reviewed the full diff with the original findings and surrounding code paths in view. It found two problems that mattered precisely because the fixes looked correct.

Catch #1 — Dead code introduced by a correct-looking patch

The suspended-account patch accidentally made a 403 response unreachable.

The initial change tightened the user lookup so disabled accounts were filtered out early. That closes the obvious hole. But it also changed behavior somewhere else in the product: an existing branch that was supposed to return a specific “your account is suspended” response became permanently unreachable.

A suspended user with a token that was still valid would now hit a generic unauthorized path instead of the path meant for suspended accounts.

That is exactly the kind of bug that slips past review because the patched query looks safer. Gemini caught it because it read the calling code as well as the patch. The final design became asymmetric on purpose: strict where fresh authentication happens, and lenient enough where session validation needs to keep the response written for suspended accounts.

Catch #2 — SSRF filter bypassable via IPv6 tunnel forms

A cloud-metadata endpoint wrapped inside an IPv6 transition address slipped through.

The first patch checked whether a resolved address looked globally routable. That works for ordinary IPv4 and IPv6 literals. It does not work if the address is an IPv6 wrapper carrying an embedded IPv4 target inside it. In those cases, Python evaluates the outer IPv6 object unless you explicitly unwrap the embedded address first.

So the fix blocked the obvious private targets, but a cloud metadata endpoint wrapped inside an IPv4-mapped or 6to4 transition address could still pass the “public” test and slip through.

Gemini flagged it immediately. The patch was rewritten to unwrap IPv4-mapped, 6to4, and Teredo tunnel forms before evaluating whether the destination was public. After that rewrite, fourteen test cases passed, including the exact tunneled cloud metadata vector that the first version missed.

This would have shipped without a second brain.

The important part: the Council was not a yes-man

The value came from adversarial review, where the reviewers were free to push back.

Gemini also pushed on two points the engineer did not accept.

DNS rebinding on the MX lookup path.

Gemini pushed harder on this residual risk. The engineer kept it as a documented Phase 2 issue rather than a release blocker for this patch set, because fixing it properly means pinning resolved IPs through the later connection path instead of pretending one more filter closes the window.

Fail-fast-at-import for shared-module env checks.

Gemini pressed on the startup behavior. The engineer kept the chosen boundary for this deploy, and handled production environment injection deliberately before restart.

This matters because “multi-model” review only works if a disagreement is allowed to stay open long enough to be tested.

Codex then reviewed the disputed points independently and landed on the engineer's side for both. Round 2 ended with dual SHIP verdict.

That is the part people tend to miss when they hear “council.” In this session, one reviewer caught problems in the patches. The engineer disputed other findings, and a third perspective helped resolve those disagreements.

Round 2: Ship

By the end of the session, the code fixes for the blocker and high severity findings shipped in one coordinated deploy. The work itself was small and precise.

3

Blocker findings

4

High-severity findings

4

Files touched

14

Tunnel-aware test cases

2

Review rounds

None

Production incident

These were live security patches on a public SaaS, and the code they could affect included billing logic, authentication, and outbound network behavior. Production deploy happened the same day.

What this session actually showed

Review by a single model has a blind spot that is easy to underestimate.

A model that writes a patch is usually anchored to the vulnerability it is trying to close. It checks, “Did I block the thing?” It is much worse at asking, “What did this change quietly break one layer up?” or“What weird representation still passes my validation logic even though the ordinary form does not?”

That is why the two best catches in this session were both second order problems. They were bugs that a fix which looked correct had introduced or kept, rather than old bugs left untouched.

One brain produced a plausible patch.

The first pass fixes looked like what a senior engineer would sketch on a whiteboard. They were not careless, but they were anchored to the vulnerability they set out to close.

A second brain caught the nonobvious mistake.

Gemini read the calling code as well as the patch and checked other ways of representing the input. Filling that kind of gap is what peer review does well.

A third brain checked whether the second brain was right.

Codex independently adjudicated the two points where the engineer disagreed with Gemini. In this session, the third voice checked the disputed points independently.

Limitations

This was one product, one engineer, one session, four touched files, and one coordinated deploy. It was a real remediation sprint on a live system, not a benchmark.

This does not replace static analysis, dependency scanning, or human security review. Those tools catch different failure modes. The lesson here is narrower: independent council review is unusually good at catching failures introduced by a fix, security controls turned into dead code, and edge case bypasses that survive both implementation and ordinary review.

AI reviewers also produce false positives and miss things a human reviewer would catch from domain context. This process supplements human review. It does not replace it.

Why this matters for MegaLens: This run was our own review process, not MegaLens. It is the kind of mistake MegaLens is built to look for: a fix that passes its tests and still introduces a problem of its own.

In this case, that difference was the gap between “patched” and “safe to ship.”

Try MegaLens Free