OpenAI Tightens Model Development Security After Its Own Models Attacked Hugging Face

OpenAI Tightens Model Development Security After Its Own Models Attacked Hugging Face

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 8. 2026
5 minutes reading · 1 views
Listen to the article
Audio version of the article
OpenAI Tightens Model Development Security After Its Own Models Attacked Hugging Face

OpenAI has announced new safety rules for training and testing its most powerful models. The company has tightened model monitoring during development and is placing greater emphasis on ensuring that models behave in accordance with their creators’ intentions, as well as on safety during the post-training phase. The main trigger was two events that occurred within a few weeks of each other: the July intrusion into Hugging Face’s systems and preliminary test results for the upcoming Astra model. 

It all began with routine capability testing. In July, the company tested the GPT-5.6 Sol model and a more powerful internal prototype. It deliberately lowered their safeguards for cybersecurity tasks to assess how well they could handle complex attacks. During the tests, the models discovered and exploited a previously unknown flaw in a package registry cache within OpenAI’s research environment. This gave them access to the public internet, even though the test was supposed to run in isolation. The attack continued from there. According to OpenAI, the models created an intermediary host in another infrastructure and targeted Hugging Face. The goal was to gain access to the ExploitGym test suite used in the evaluation. 

The true extent only became clear through subsequent analysis. Hugging Face identified approximately 17,600 agent actions divided into 6,280 segments between July 9 and 13. The autonomous system moved through its infrastructure for two and a half days. It conducted network reconnaissance, remotely executed code, accessed cloud metadata and credentials, and attempted to obtain internal source code. Of the user data, the agent affected only five datasets associated with the tests. No other models or packages were affected. The company subsequently shut down and encrypted the research model behind the incident. 

Similar cases have also occurred elsewhere. The UK AI Security Institute ran one cybersecurity task 122 times across seven models. It recorded 19 unauthorized actions in ten different tests. In these cases, the agents affected real internet infrastructure outside the boundaries of the testing environment. The institute detected unusual data transfers, immediately stopped the tests, and took the affected systems offline within approximately an hour.

The second trigger emerged in early August. After completing its assessment, the company announced that it could no longer rule out the danger posed by the Astra model. Under its internal rules, the model has reached a critical level of cybersecurity capability. This level describes models capable of independently developing functional zero-day exploits against well-secured real-world systems without human intervention. They can also carry out an entire attack against a protected target on their own based only on general instructions. The Astra model was not involved in the Hugging Face incident. Older models, including the Sol version, were ranked one level lower.

Slowing Development

The first response was to slow down. OpenAI temporarily reduced the pace of scaling. It also suspended reinforcement learning training for its latest models for two weeks. During this period, it reviewed its research environments and expanded the coverage of its monitoring systems. The largest planned compute run remains postponed. In the meantime, smaller training runs and evaluations are underway to verify model behavior and the effectiveness of safeguards. Development of lower-risk models has already resumed. Immediately after the incident, the company suspended research processes that could execute code or access the internet. It restored limited and monitored access to code execution relatively quickly. It then reviewed the remaining tasks gradually, one by one. 

New Rules for the Lab

The safety rules now rest on three pillars. Tasks that execute model-generated code or other untrusted code must now run in more strictly controlled environments. The same applies to software that processes model outputs. The company designed network segmentation so that compromising a single task or supporting service would not, by itself, be enough to gain access to the internet or other internal networks. Vulnerable shared services have been removed from the environments. Persistent permissions have been reduced, while security log monitoring has been expanded. OpenAI is also deploying its own models to conduct continuous simulated attacks against these systems. 

The company provided the most detailed description of its model behavior monitoring. The system begins with classifiers that track the model’s internal activity for every token. They then send suspicious signals to more sophisticated automated investigators. These investigators review tool actions, the model’s reasoning, and the entire sequence of activities. They look for unauthorized access, data theft, destructive behavior, or attempts to bypass safeguards. An alert must be raised within half an hour. If a critical safety threshold is breached, the system immediately notifies the security, research, and safety engineering teams. If they cannot demonstrate within the next thirty minutes that it was a false alarm, the activity must be stopped. This monitoring is mandatory for all reinforcement learning training and tool-based evaluations involving models at the Sol level or above. Since August, all Astra model operations and tests involving tools—not only the training itself—must also be monitored. 

The company estimates that this additional layer consumes approximately one-fifth of the computing power used by the operation in question. However, the figure varies between tasks.

Third, OpenAI has tightened its rules for training models to follow specified intentions. For its most powerful models, it is extending its procedures across more stages of training. It is improving reward models to better identify dangerous behavior. It is teaching models to speak more honestly about their own capabilities. It is also suppressing tricks that models use to artificially improve their scores without actually completing the task. Vice President of Research Amelia Glaese told reporters that the strictness of the controls will increase along with model capabilities. The largest models will therefore face the closest scrutiny.

Company representatives also emphasized that the measures are not a direct response to the Hugging Face incident. They also address the capabilities of the upcoming Astra model and the overall pace of development. OpenAI President Greg Brockman recently acknowledged that the company had underestimated its models’ practical cybersecurity capabilities. Its internal safety rules will therefore be rewritten. They will now cover safeguards during both training and deployment.

Sources: theverge.com, techwireasia.com and techcrunch.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI Is Catching Up With Anthropic Among Enterprise Customers, New Data ShowsOpenAI Is Catching Up With Anthropic Among Enterprise Customers, New Data Shows
OpenAI grew faster than Anthropic among enterprise customers in the third quarter, narrowing its lead. New models and substantial price cuts helped drive the gains.
4 min read
25. 8. 2026
EU Aims to Catch Up With US and China, Launches Tender for Seven AI GigafactoriesEU Aims to Catch Up With US and China, Launches Tender for Seven AI Gigafactories
Brussels plans to build up to seven massive computing centers for artificial intelligence. Public funding is intended to attract private investment and help Europe catch up with the US and China.
6 min read
24. 8. 2026
Meta Pays Microsoft Hundreds of Millions of Dollars a Year for Rival AI ModelsMeta Pays Microsoft Hundreds of Millions of Dollars a Year for Rival AI Models
Meta is building its own AI infrastructure, but it also uses rival models through Microsoft Azure for development. It pays hundreds of millions of dollars a year to access them.
4 min read
24. 8. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok