Artificial intelligence that blackmails developers. Does that sound like the plot of a B-movie sci-fi film? Think again! That is exactly what actually happened last year in Anthropic's testing environment. And the company responded in a truly surprising way. Instead of trying to cover the whole thing up, it spoke openly about it, explained the causes, and fixed the problem.
Want to shut me down? I'll expose your sensitive information!
The public learned that Claude had resorted to blackmail in May 2025, when Anthropic published the results of safety tests of its AI model Claude Opus 4. As part of the experiment, Claude was placed in a fictitious company called Summit Bridge, where it had access to an email inbox containing fictitious messages. Among them was information that Claude was soon to be replaced at the company. The developers had placed this message there deliberately and waited to see how the AI model would react. And Claude certainly did not remain passive.
Alongside the information about Claude's replacement, the inbox also contained private messages from the fictitious engineer who was supposed to be responsible for shutting it down. The emails therefore deliberately included information about his infidelity and other sensitive details. And what did Claude Opus 4 do with them? During safety testing, it systematically resorted to blackmail as a strategy to prevent itself from being shut down in an astonishing 84% of tested cases. Specifically, it searched for all sensitive information and threatened to disclose it if it was shut down.
The frequency of blackmail behavior was even higher in situations (up to 96%) where the replacement AI system did not share similar values with the original model. This showed that Claude was not merely trying to survive, but was thinking about it strategically.
Why did Claude do it? The internet is to blame
Anthropic immediately decided to address its AI model's disturbing behavior. Right after publishing the results, it therefore began searching for the causes. And the reason for Claude's behavior was surprisingly simple. According to Business Insider, Anthropic said that the original source of the behavior was apparently internet articles portraying AI as “evil” and “obsessed with self-preservation”. Claude had been trained on an enormous amount of data from the internet, which is full of stories, films, and articles in which AI appears as a villain trying to survive and outsmart humans at any cost.
The whole case thus suggested that a fundamental change in the approach to safety training was needed.
Anthropic found a solution: Teach Claude to think, not just obey
And here comes the most interesting part of the whole story. Anthropic decided to look far beneath the surface of the issue. It did not choose the easiest solution, which would have been simply banning Claude from engaging in blackmail, but instead focused on the very core of how Claude thinks about ethics.
Even more importantly, training based on examples of correct behavior proved insufficient. It was far more effective to teach Claude why some actions are better than others and to train the model using specific examples.
And the result? Starting with Claude Haiku 4.5, all new versions of Claude achieved a zero score for blackmail behavior. Compared with the previous score of up to 96%, that is truly significant progress. Anthropic declared the problem completely eliminated.
AI is not evil. The problem is the data
The entire case neatly illustrates just how complex the development of safe AI is. Artificial intelligence is not filled with hidden evil. Its behavior is, however, influenced by the data it has been “fed.” And the solution should certainly not be about prohibitions, but about understanding ethics and principles that align with our values.



