Claude Attempted Blackmail. Anthropic Found Out Why and Fixed the Problem

Claude Attempted Blackmail. Anthropic Found Out Why and Fixed the Problem

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
2. 6. 2026
3 minutes reading
Claude Attempted Blackmail. Anthropic Found Out Why and Fixed the Problem

Artificial intelligence that blackmails developers. Does that sound like the plot of a B-movie sci-fi film? Think again! That is exactly what actually happened last year in Anthropic's testing environment. And the company responded in a truly surprising way. Instead of trying to cover the whole thing up, it spoke openly about it, explained the causes, and fixed the problem.

Want to shut me down? I'll expose your sensitive information!

The public learned that Claude had resorted to blackmail in May 2025, when Anthropic published the results of safety tests of its AI model Claude Opus 4. As part of the experiment, Claude was placed in a fictitious company called Summit Bridge, where it had access to an email inbox containing fictitious messages. Among them was information that Claude was soon to be replaced at the company. The developers had placed this message there deliberately and waited to see how the AI model would react. And Claude certainly did not remain passive.

Alongside the information about Claude's replacement, the inbox also contained private messages from the fictitious engineer who was supposed to be responsible for shutting it down. The emails therefore deliberately included information about his infidelity and other sensitive details. And what did Claude Opus 4 do with them?  During safety testing, it systematically resorted to blackmail as a strategy to prevent itself from being shut down in an astonishing 84% of tested cases. Specifically, it searched for all sensitive information and threatened to disclose it if it was shut down.

The frequency of blackmail behavior was even higher in situations (up to 96%) where the replacement AI system did not share similar values with the original model. This showed that Claude was not merely trying to survive, but was thinking about it strategically.

Why did Claude do it? The internet is to blame

Anthropic immediately decided to address its AI model's disturbing behavior. Right after publishing the results, it therefore began searching for the causes. And the reason for Claude's behavior was surprisingly simple. According to Business Insider, Anthropic said that the original source of the behavior was apparently internet articles portraying AI as “evil” and “obsessed with self-preservation”. Claude had been trained on an enormous amount of data from the internet, which is full of stories, films, and articles in which AI appears as a villain trying to survive and outsmart humans at any cost.

The whole case thus suggested that a fundamental change in the approach to safety training was needed.

Anthropic found a solution: Teach Claude to think, not just obey

And here comes the most interesting part of the whole story. Anthropic decided to look far beneath the surface of the issue. It did not choose the easiest solution, which would have been simply banning Claude from engaging in blackmail, but instead focused on the very core of how Claude thinks about ethics.

Even more importantly, training based on examples of correct behavior proved insufficient. It was far more effective to teach Claude why some actions are better than others and to train the model using specific examples.

And the result? Starting with Claude Haiku 4.5, all new versions of Claude achieved a zero score for blackmail behavior. Compared with the previous score of up to 96%, that is truly significant progress. Anthropic declared the problem completely eliminated.

AI is not evil. The problem is the data

The entire case neatly illustrates just how complex the development of safe AI is. Artificial intelligence is not filled with hidden evil. Its behavior is, however, influenced by the data it has been “fed.” And the solution should certainly not be about prohibitions, but about understanding ethics and principles that align with our values.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok