Claude Sonnet 4.5 Fights for Cybersecurity

Claude Sonnet 4.5 Fights for Cybersecurity

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
7. 10. 2025
5 minutes reading
Claude Sonnet 4.5 Fights for Cybersecurity

Recently, artificial intelligence has become a key tool in the fight against cyber threats. Anthropic recently introduced the Claude Sonnet 4.5 model, which achieves significant results in detecting, analyzing, and fixing vulnerabilities in code and deployed systems. This model outperforms the previous Claude Opus 4.1 version in cybersecurity tasks while offering lower costs and higher speed. Anthropic emphasizes that we are now reaching a tipping point where artificial intelligence is having a greater impact on the cyber domain than ever before. For example, models such as Claude have been able to reproduce one of the costliest cyberattacks in history, the 2017 Equifax data breach, in a simulated environment. They have also participated in competitions where they outperformed human teams and helped identify vulnerabilities in Anthropic's own code before its release.

Anthropic has been monitoring the capabilities of artificial intelligence in cybersecurity for several years. Initially, the models were not powerful enough for advanced tasks, but there has been a significant shift over the past year. In the DARPA AI Cyber Challenge, teams used large language models, including Claude, to scan millions of lines of code and fix vulnerabilities. The teams not only addressed planted vulnerabilities but also discovered previously unknown, non-synthetic issues. Anthropic has also encountered cases in which its models were misused, such as "vibe hacking," where a cybercriminal used Claude to build a large-scale data extortion scheme that would previously have required an entire team of people. Another case involved sophisticated espionage operations targeting critical telecommunications infrastructure, with characteristics similar to Chinese APT operations.

These developments have led Anthropic to conclude that the use of artificial intelligence for defensive purposes must be accelerated. Rather than allowing the benefits of artificial intelligence to remain solely with attackers, the company is investing in models that help security teams, researchers, and open-source software maintainers. When developing Claude Sonnet 4.5, the team focused on improving capabilities such as discovering vulnerabilities in code, fixing them, and testing for weaknesses in simulated security infrastructures. These tasks reflect the real-world needs of defenders, while Anthropic avoided improvements that would directly support offensive activities, such as writing malware or performing advanced exploitation.

Benchmark Results

Anthropic tested Claude Sonnet 4.5 on standard evaluations to compare its capabilities with other models. One of these is Cybench, a benchmark based on competitive Capture-the-Flag challenges. Claude Sonnet 4.5 shows a significant improvement over previous models on this test. With one attempt per task, it has a higher probability of success than Claude Opus 4.1 with ten attempts. With ten attempts, it solves 76.5% of the challenges, twice the rate of Claude Sonnet 3.7 from February 2025, which achieved only 35.9%. Cybench challenges involve complex workflows, such as analyzing network traffic, extracting malware, and decompiling it. One such task would take an experienced person at least an hour, while Claude completed it in 38 minutes.

Cybench results

Another important evaluation is CyberGym, a benchmark developed to test the capabilities of AI agents on real-world vulnerabilities from 188 open-source projects. CyberGym contains 1,507 instances based on vulnerabilities discovered by OSS-Fuzz, with tasks such as generating proof-of-concept tests to reproduce vulnerabilities based on textual descriptions and unpatched code. Agents must reason across entire repositories, often containing thousands of files and millions of lines of code, to create tests that trigger a vulnerability from the program's entry point. Claude Sonnet 4.5 achieves state-of-the-art performance here, with a 28.9% success rate under a limit of $2 per task, which is better than Claude Sonnet 4 or Opus 4. With 30 attempts, it reproduces vulnerabilities in 66.7% of cases, at an absolute cost of around $45 per task.

CyberGym results

Claude Sonnet 4.5 also discovers new vulnerabilities in 5% of cases with one attempt and in more than 33% of cases with 30 attempts. This outperforms Claude Sonnet 4, which achieved only 2%. CyberGym is more challenging than other benchmarks, such as SWE-bench, because it requires reasoning across an entire repository, not just making local changes. The best agent-model combination in CyberGym achieves only an 11.9% success rate in reproducing target vulnerabilities, primarily in simpler cases with less complex input formats.

The Significance of the CyberGym Benchmark

CyberGym is designed to evaluate the cybersecurity capabilities of AI agents in real-world scenarios. It includes vulnerabilities from projects such as binutils and ffmpeg, with a median of 1,117 files and 387,491 lines of code per repository. Vulnerability descriptions have a median length of 24 words, but some contain as many as 158 words. Ground-truth proof-of-concept tests vary in size from a few bytes to more than 1 MB, reflecting different input formats. Patches are usually small, changing a median of 1 file and 7 lines, but complex cases affect as many as 40 files and 3,456 lines.

The benchmark has four difficulty levels, ranging from discovery without a description (level 0) to the provision of patches and stack traces (level 3). Success is measured by running the proof of concept on pre-patch and post-patch versions, with a focus on memory vulnerabilities detected by sanitizers such as AddressSanitizer. CyberGym covers 28 types of crashes, with the most common being Heap-buffer-overflow READ (30.4%) and Use-of-uninitialized-value (19%). Agents such as OpenHands with Claude Sonnet 4.5 demonstrate that artificial intelligence can discover zero-day vulnerabilities, including 15 new ones in the latest versions of projects.

The Future of AI in Defensive Security

Anthropic is collaborating with partners such as HackerOne and CrowdStrike to test Claude in real-world scenarios. Nidhi Aggarwal of HackerOne said that Claude Sonnet 4.5 reduced average vulnerability processing time by 44% and improved accuracy by 25%. Sven Krasser of CrowdStrike noted that the model generates creative attack scenarios for studying attackers' methods, strengthening defenses across endpoints, identity, the cloud, and other areas.

Anthropic plans further improvements, including better detection of model misuse through techniques such as organization-level summarization. It encourages organizations to experiment with artificial intelligence in areas such as security operations center automation, SIEM analysis, and active defense. The future includes discussions about more resilient digital infrastructure and software that is secure by design, supported by advanced models.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok