White House Demands Zero Vulnerabilities in Anthropic’s Fable 5. Experts Say: “That’s Impossible!”

White House Demands Zero Vulnerabilities in Anthropic’s Fable 5. Experts Say: “That’s Impossible!”

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
23. 6. 2026
7 minutes reading
White House Demands Zero Vulnerabilities in Anthropic’s Fable 5. Experts Say: “That’s Impossible!”

The Trump administration imposed export restrictions on Anthropic’s most powerful models, forcing the company to withdraw both Fable 5 and Mythos 5 from the market just days after their release. As a condition for their return, it is now demanding that Anthropic eliminate absolutely all jailbreaks and report every discovered vulnerability to the government in advance. A jailbreak is a trick that allows a user to bypass an artificial intelligence system’s safety safeguards and make it do something it is normally prohibited from doing, such as providing instructions for something dangerous. Yet there is near-unanimous agreement among security researchers that perfect protection against jailbreaks is technically impossible for today’s leading models—and not just Anthropic’s.

Events Behind the Ban on Anthropic’s AI Models

Fable 5 was released to the public in early June. Anthropic described it as a “Mythos-class” model with safeguards designed to make it safe even for general use. Before its release, it was reviewed by the administration and the UK AI Safety Institute. A few days after launch, the situation took a turn. Two days after the model was released, Amazon CEO Andy Jassy alerted the White House that the model’s safeguards could be bypassed. Amazon, which is also an investor in Anthropic, was reportedly merely responding to the administration’s request for feedback.

By Friday morning, the matter had reached the highest levels of the White House. Treasury Secretary Scott Bessent, Cybersecurity Director Sean Cairncross, and others met to discuss what would happen next. Then the search began for Dario Amodei, Anthropic’s CEO. According to one account, he was unavailable because he was at a wellness retreat, which a company spokesperson sharply rejected as “a complete lie.” A person close to Anthropic added that Amodei was on the phone within an hour and fifteen minutes.

When he finally connected with officials, he took part in three calls with about half a dozen senior officials. He tried to explain what he regarded as a misunderstanding. He defended the model’s safeguards and argued that the bypass was narrowly targeted and did not pose the same risk as a full-scale jailbreak that would disable all protections at once.

Neither Bessent nor Cairncross was convinced. According to one official, they ran Amazon’s findings through the National Security Agency and felt they had “proof.” Amodei asked for more time but made no commitments. Bessent reportedly told him directly that he was making a “bad decision.” Shortly after the call, the hammer came down. The administration imposed export restrictions on Fable 5 and Mythos 5 and prohibited their use by foreign nationals, including Anthropic’s own employees. The company had no choice but to disable both models for all customers.

Different Views of the Events

The accounts diverge quite significantly here. A senior White House official portrays the move as a last resort: “The export restrictions were the last option after we spent hours begging them to cooperate with us. We did not want to do this, but our hands were tied.” People close to Anthropic tell the exact opposite story. “The White House gave them ninety minutes to take down the models, without providing any details about the actual threat,” one of them said. “Nobody begged anyone or asked for cooperation; they simply gave us a ninety-minute deadline.”

Officials were reportedly taken aback. They had heard Amodei compare the danger of the technology to an atomic bomb, and now the same person was refusing to shut down the model because of a known security hole.

Was It a Jailbreak, or Just a Pretext?

And now for the most sensitive question. Was this really a security threat, or was something else at play? Axios described the weekend’s tensions and claimed that the export order was driven more by “personality differences” than by a technical problem with the product.

Katie Moussouris, a cybersecurity veteran and founder of Luta Security, wrote on her blog that Anthropic had sent her a private copy of the paper describing the alleged leak. According to the Wall Street Journal, its authors were researchers from Amazon. Moussouris analyzed how they triggered the bypass but added that it “should never have triggered export restrictions” on its own. The main difference, she said, is whether you ask the model to “review code for security flaws” or to “fix that code.” The result is essentially the same.

“The behavior described in the paper cannot be meaningfully fixed, and any such attempt would only weaken the model for defensive purposes,” she wrote. She called the order hasty, harsh, and misguided. Together with dozens of other experts, she then urged the administration to lift the restrictions. In their view, withdrawing advanced defensive capabilities from people protecting American networks creates a danger.

This time, moreover, it appears to be retaliation. Analysts believe the administration’s move “is likely to raise alarm in foreign capitals about the reliability of American AI for critical applications.” The dispute has deeper roots: the Pentagon designated Anthropic a supply-chain risk as early as March 3 because the company refused to allow its tools to be used for mass domestic surveillance and autonomous weapons. David Sacks, the White House’s former AI chief, however, claimed on social media that the previous disputes were unrelated to the export decision. “The ball is in Anthropic’s court,” he wrote.

Why a Jailbreak Cannot Simply Be “Patched”

This brings us to the heart of the matter. A conventional software bug can be found and fixed. A jailbreak does not work that way.

“A jailbreak is not a flaw in a specific piece of code. It is a hostile input into a system whose entire value lies in being open and flexible,” Martin Riley, chief technology officer at Bridewell, explained to Cybernews. “With conventional software, you can define expected inputs and close the gaps. With a leading model, the input space is essentially infinite.” Attackers can use natural language, encoded instructions, role-playing, or gradual prompting. “You cannot enumerate every attack path, so you cannot prove that none exists,” Riley says.

Oliver Simonnet of CultureAI sees it the same way. Companies including Anthropic have reportedly poured a great deal of money into security, but attackers continue devising new ways to persuade models to ignore instructions. “While developers can reduce the success rate of jailbreaks, it is currently unrealistic to guarantee that a model will remain 100 percent resistant forever,” he concludes.

Anthropic itself had already set out this limitation in the Fable 5 documentation. Helen Toner of the Center for Security and Emerging Technology pointed this out months ago. Among researchers, this is not a disputed claim; it is the consensus. The government’s condition therefore sets a bar that no one can clear. Not OpenAI, not Google, nor anyone else. A single existing jailbreak is enough for Anthropic to fail the standard.

A Trap That Will Be Difficult to Escape

Let us put ourselves in Anthropic’s shoes for a moment. It essentially has three options, and none of them is good. It can comply in appearance by improving jailbreak detection while accepting that some will still slip through. The administration can then point to the first one discovered and declare that the company failed to meet the condition. The second option is to refuse and escalate the dispute, perhaps to the courts or Congress. That would be expensive and time-consuming. The third option is to negotiate a different standard focused on detecting attempts rather than eliminating them entirely. That, however, would mean the government accepting a technical reality that it may not want to accept.

Anthropic already maintains a thirty-day retention period for Fable 5 usage data, a bug bounty program, and partnerships with the US NIST and the UK AI Safety Institute. According to analysts, building on these tools instead of insisting on an impossible standard is the approach the two sides are most likely negotiating.

Sources: wired.com and finance.yahoo.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok