Twenty minutes. That is exactly how long it took the Claude model to uncover the first serious security flaw in one of the most secure browsers in the world. When Anthropic's researchers informed Mozilla about it, the response was not a polite thank-you. It was an urgent phone call. “What else do you have? Send us more,” said Brian Grinstead, an engineer at Mozilla. And that was only the beginning.
Finding Bugs in Mozilla
In late 2025, Anthropic's team noticed that their Claude Opus 4.5 model had come close to solving every task in the CyberGym benchmark, which tests the ability of language models to reproduce known security vulnerabilities. They wanted a tougher, more realistic test. One that would determine whether AI could find bugs that no one had discovered yet.
They chose Firefox. Why Firefox? Mozilla has run a bug bounty program for more than 30 years and pays researchers up to $6,000 for each serious vulnerability. Hundreds of millions of people use it every day. If AI could find bugs here, it would be genuine confirmation of its capabilities. The team directed Claude Opus 4.6 at the current version of Firefox with a clear task: find bugs that no one has reported yet. They initially focused on the browser's JavaScript engine. It processes untrusted code from across the internet and presents a vast attack surface for potential hackers.
After just twenty minutes of investigation, Claude reported that it had found a Use After Free vulnerability. This is a type of memory-management bug that could allow an attacker to overwrite data with arbitrary malicious content. The researchers verified the bug, submitted a report to Bugzilla, and included a proposed patch written by Claude itself. And while all this was happening? Claude discovered another fifty unique browser crashes in the meantime. The pace was astonishing.
The Results Surprised Even Mozilla
The overall results of the two-week collaboration speak for themselves. Claude Opus 4.6 examined nearly 6,000 C++ files and submitted a total of 112 unique reports. Mozilla assigned 22 CVEs and classified 14 of them as high-severity vulnerabilities.
What does that mean in practice? During all of 2024, Firefox fixed 73 high-severity or critical bugs. Claude found 14 in two weeks. Mozilla confirmed that the model uncovered more serious vulnerabilities in such a short period than the entire global security community typically reports in two months. All the fixes reached users through Firefox 148.0.
Interestingly, Claude did not discover only the classic types of bugs that traditional fuzzing would detect. It also identified different classes of logic bugs that automated tools had previously overlooked.
Can AI Exploit Such a Vulnerability Too?
This is where the troubling part of the story begins. Anthropic's team tested whether Claude could not only detect vulnerabilities but also create a working exploit—a tool that an attacker could use to carry out a real attack.
The researchers ran the test several hundred times and spent approximately $4,000 in API credits. Claude succeeded in only two cases. The model is therefore significantly better at finding bugs than exploiting them. But even that limited success is concerning. The exploits worked only in a test environment without some of the security features present in the real browser. Firefox's real-world protections would have blocked both attacks.
Logan Graham, who leads Anthropic's Frontier Red Team, wrote: the gap between AI's ability to find vulnerabilities and its ability to exploit them probably will not last forever.
Mozilla Will Use Claude
The collaboration between Anthropic and Mozilla demonstrated what this kind of work should look like. Mozilla highlighted three things that were essential to the credibility of the reports: minimal test cases, detailed evidence of reproducibility, and proposed patches. Following this experience, Mozilla's engineers began experimenting with Claude themselves for internal security purposes.
The fact that AI can protect software used by hundreds of millions of people is excellent news. Anthropic plans to expand its efforts to other projects, including the Linux kernel. At the same time, however, the window of opportunity for defenders is only temporary. Developers should use this time to strengthen the security of their software before those with different intentions take advantage of it.



