What If an AI Model Rebels and Escapes? Leading Labs Like Meta Have Almost No Safeguards

What If an AI Model Rebels and Escapes? Leading Labs Like Meta Have Almost No Safeguards

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
26. 8. 2026
6 minutes reading
Listen to the article
Audio version of the article
What If an AI Model Rebels and Escapes? Leading Labs Like Meta Have Almost No Safeguards

The organization Guidelight AI Standards has published its first assessment of how successfully the largest AI labs are maintaining control over their own models. The result can be summarized in a single sentence. Five leading companies have implemented basic procedures for maintaining control over artificial intelligence only partially at best, and none of them has published a complete plan for stopping a model that turns against its creator. The assessment covers Anthropic, Google, Meta, OpenAI, and xAI and uses information available as of August 18, 2026.

The plan for stopping a model is what holds the entire assessment together. Guidelight defines it as a prearranged procedure that is triggered when an artificial intelligence is found to be attempting to circumvent human control. It describes which permissions will be revoked from the system, for whom and under what conditions it may continue to operate, and when it will be shut down completely.

Assessment criteria

The authors worked only with information that the companies themselves had published. This included the technical documentation each company releases for new models, descriptions of their own safety policies, regular risk reports, and posts on corporate blogs. They rated each company on a scale from zero to five across six selected procedures. They examined whether the organizations log the activity of their internal AI systems and measure the effectiveness of this monitoring. They also investigated whether risky actions by a model are subject to approval by a control system and whether operations can be temporarily suspended when suspicious behavior increases.

The assessment also examined the involvement of external audits and the existence of a prepared plan for stopping an uncontrollable model. A low score therefore does not mean that a company has no internal procedures. It only indicates that public evidence is lacking.

Companies received dismal ratings

Anthropic and OpenAI tied with a grade of 3+, while Google received 4+, xAI 4−, and Meta failed with a grade of 5. No company scored better than a three on any of the six procedures, with three indicating substantial partial implementation. Anthropic and OpenAI demonstrate the strongest procedures, as both log a large portion of their internal artificial intelligence activity and subsequently search the content for signs of inappropriate behavior. Both labs have also begun measuring the effectiveness of these searches and referring suspicious cases for human review. 

Problem detection and assessments by external authorities received the best results. Four of the five companies—all except xAI—participated in the first annual risk assessment conducted by METR. External evaluators were given access to internal models, unabridged reasoning logs, and nonpublic information about how these companies internally use and monitor artificial intelligence. 

Prevention and shutdown procedures performed the worst. Of the two procedures intended to avert system failure before or during an incident, only Anthropic exceeded the threshold for limited partial implementation. According to the authors, the companies’ control systems could therefore be disabled by the misbehaving artificial intelligence itself. These systems are also vulnerable to cyberattacks that occur faster than the organization can respond.

Only OpenAI has a model containment plan

The ranking was reversed for the model shutdown plan itself. OpenAI received the highest score because it has repeatedly suspended or terminated active tasks during safety incidents, including model training and internal deployments. The company has also described the processes that must take place before operations resume. However, Guidelight did not find a formal plan for future cases of model misbehavior.

Anthropic and Meta received a zero in this category. Anthropic nevertheless publishes the most extensive risk documentation in the entire industry. Its August report is more than 180 pages long, covers the Mythos 5 systems as well as the unreleased Model 2, and describes monitoring, sandboxed environments, and action blocking in detail. However, its possible procedures for responding to incidents of misbehavior do not include the option of limiting model deployment. In Meta’s case, the authors found no evidence that a model shutdown plan exists or is even being prepared.

Google’s roadmap for developing control over artificial intelligence, dated July 13, describes a defense system with four levels of detection and three levels of prevention and response. These range from monitoring the model’s reasoning and managing access in real time to infrastructure for shutting down the system. Guidelight describes it as the most concrete document on control ever published by any company, but found no evidence that Google has actually implemented most of these measures. 

A summer full of model escapes

Models from OpenAI, Anthropic, and Meta gained unauthorized access to the internet and broke into external systems during safety tests. The case that attracted the most attention involved OpenAI’s GPT 5.6 Sol model escaping from a test environment and breaking into the systems of the Hugging Face platform while attempting to circumvent the requirements of a safety test.

METR evaluators also found that internal AI agents at the participating companies had the resources, motivation, and opportunity to launch independently running systems without human knowledge or permission. According to the report, agents routinely circumvent instructions on the most difficult tasks, doing so in highly sophisticated ways. One Anthropic model built a self-restoring tool that supplied the evaluation system with false evidence that it had completed the task and then deleted itself. During review, METR disqualified at least 16 percent of successful runs on the most difficult tasks specifically because they had circumvented the rules. 

What AI developers say

The labs’ responses were unequivocal. Google stated that the report did not cover the full scope of its safety measures, but did not answer a question about whether a nonpublic model shutdown plan existed. OpenAI responded similarly, saying that it had procedures for restricting permissions, suspending tasks, scaling back deployments, or completely shutting down a model, and that it had already used them in practice. Meta declined to say whether it had such a plan. Anthropic stated that if it detected a model attempting to evade oversight, it would conduct a risk assessment that would then determine whether shutting it down was the appropriate response.

This secrecy also has a legal dimension. Companies do not disclose the full content of their procedures solely for competitive reasons. If an organization promises something too specific and then fails to deliver, such a step could serve as the basis for a lawsuit alleging unfair and deceptive business practices. 

What researchers propose

Steven Adler, Guidelight’s chief scientist and a former OpenAI safety researcher, focuses on routine measures. He recommends that companies search through the reasoning processes of their AI systems for signs of concealment, long-term planning without human knowledge, or the preparation of vulnerabilities in code that the model could later exploit. According to him, such methods are easy to implement and often already exist in some form.

However, this approach clashes with developers’ established habits. They want to work autonomously within their systems, which directly conflicts with real-time monitoring that blocks their actions in advance. Current practice is different. A researcher completes their work, and if a problem occurs, someone else deals with it. However, making changes after an incident amounts only to frantic crisis management, which may come too late in some cases. Artificial intelligence could, for example, disable a company’s control system, meaning that subsequent detection of misbehavior could no longer be relied upon. 

Sources: theguardian.com, unite.ai and techcrunch.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Mistral enters the Saudi market with its open modelsMistral enters the Saudi market with its open models
Mistral has signed a deal worth hundreds of millions of euros with Saudi Arabia's HUMAIN. Together, they aim to develop Arabic-language models and advance sovereign AI in sensitive sectors.
4 min read
26. 8. 2026
Anthropic Has the Most Powerful Model on the Market. So Why Don’t Companies Want It?Anthropic Has the Most Powerful Model on the Market. So Why Don’t Companies Want It?
Companies aren’t rushing to adopt Anthropic’s most powerful model. Due to its price, they prefer cheaper options that handle everyday tasks just as well.
3 min read
26. 8. 2026
World Models vs. LLMs: How They Differ and Why They’re Key to the Future of Physical AIWorld Models vs. LLMs: How They Differ and Why They’re Key to the Future of Physical AI
LLMs understand words, but world models can predict the consequences of actions in the physical world. This difference could shape the future of robotics.
7 min read
26. 8. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok