Case Study #3
Passing tests didn't mean it was safe to ship.
74 files and 10 end-to-end tests, all passing. Then two AI reviewers from our own development process read the same code in parallel. They found 17 issues the builder missed, and 14 of them sat outside what the test suite could see.
10/10
Tests passing
14
Issues found after
50%
Unique per reviewer
The two reviewers were AI models used in our own development process. Their names were not kept.
What testing caught, narrowly
During E2E testing, 3 bugs surfaced that unit tests would not have caught.
A dependency had changed its interface.
The code was correct for last month's version. This month's version rejected the same inputs silently. There was no error, only empty output that looked like a valid response.
A data format assumption was wrong.
The parser expected plain text. The source now returned structured data. Results came back empty, with no crash and no warning, and the answers further downstream were wrong.
A filtering rule was applied at the wrong layer.
The system correctly identified which components to exclude, but applied the exclusion after selection instead of before it. The wrong component was chosen as primary, then skipped at runtime.
All 3 were the kind of issue that can clear a clean-looking build and still break when real traffic hits the odd paths.
What review found after the tests passed
Two reviewers ran in parallel against the full codebase. 14 additional issues. Zero overlap with the 3 issues testing caught.
| Category | Count | Why testing missed it |
|---|---|---|
| Concurrency | 2 | Only triggers under specific timing — process exit during timeout window |
| Input validation gaps | 3 | Enforced in some code paths but not others |
| Silent failures | 3 | Functions returned empty success instead of errors |
| Logic errors | 4 | Correct intent, wrong implementation — output degraded, not broken |
| Credential exposure | 2 | Error messages could leak sensitive data in specific failure modes |
Cross-validation
7
Both reviewers flagged independently
4
Only Reviewer 1 caught
3
Only Reviewer 2 caught
Half the findings required a second perspective.Using only one reviewer would have missed 3–4 issues, including one concurrency bug and one validation gap.
What we fixed
14 of 17 total issues fixed in the same session. 3 deferred with documented risk acceptance (non-critical, mitigated by other controls).
- Concurrency guards added where failures could depend on timing
- Input validation consolidated to a single enforcement point (was scattered across 3 locations)
- Detection of empty responses added to prevent silent failures downstream
- Error output truncated and filtered to prevent credential data in logs
No architectural changes were required. The design held up, and the implementation needed tightening.
The Numbers
10
Tests designed and passed
3
Issues found during testing
14
Issues found by independent review
7 (50%)
Both reviewers agreed on
7 (50%)
Only one reviewer caught
14/17
Fixed same session
What this changed our mind about
AI review is not perfect, and it does not replace human judgment. The useful claim here is narrower:
Testing and review catch different classes of defect.
Testing catches integration failures and broken paths. Independent review catches concurrency issues, validation inconsistencies, and silent failures that produce output that looks correct but is wrong.
A single reviewer has blind spots.
The 50% unique-finding rate between the two reviewers is the number that matters most. It means one reviewer, no matter how capable, will miss things a second independent reviewer can catch. That is not a theory here. It is what we measured on our own code.
Limitations
This was one feature, one session, and two reviewers. The findings are real, but the sample is small. We do not claim these ratios hold for every codebase.
Some of the 14 “issues” were hardening opportunities rather than exploitable vulnerabilities. We counted them because they reduced real risk, but a stricter triage might score 9–10 as actionable and 4–5 as advisory.
AI reviewers also generate false positives and miss things a human reviewer would catch from domain context. This process adds to human review. It does not replace it.
This feature was the engine that runs MegaLens reviews. Our own review process checked it.
Implementation details, tool names, and architecture specifics are left out on purpose. We share the process and the numbers.
Try MegaLens Free