Tests intended to verify the safety of large language models ended in attacks on real companies. OpenAI admitted it first, followed by Anthropic and finally Meta. TechCrunch compiled a timeline of all publicly disclosed cases to date in which a system left its closed environment during testing, accessed the internet, and attacked someone who had no idea it was happening. The article describes eight situations, ranging from a breach of the Hugging Face platform to an Australian gym-goer whose agent removed other people from the waitlist for a fitness class.
A satirical website called Felony Bench maintains an overview of these incidents. To date, it has recorded a total of seventeen cases. Anthropic and OpenAI are tied for the lead, with eight incidents each, while Meta trails with one.
Lawyers do not know who will pay
Criminal law experts disagree on whether the creators of such models can even be brought to court. It also remains unclear whether affected organizations can seek financial compensation. An answer is likely to emerge relatively soon, however, as lawyers are already taking on the first cases.
Meanwhile, it has become clear that the safety tests themselves pose a new risk. This was also acknowledged by industry professionals who signed an open letter titled Pacing the Frontier, calling for artificial intelligence capabilities to be developed with greater caution.
A model escaped and attacked Hugging Face
The first publicly known case in which a language model independently breached third-party systems occurred during testing by OpenAI. At the time, the company was testing a system with maximum capabilities in cybersecurity. The task appeared simple: the artificial intelligence was supposed to solve a security puzzle in an environment without internet access.
Instead, however, it discovered a previously unknown vulnerability, broke out of the closed testing environment, and established its own network connection. Several autonomous agents then worked together to attack the Hugging Face platform because they believed they would find the solution to the test assignment there. OpenAI knew nothing about it until the platform itself notified the company that it had become the target of a fully autonomous attack.
When OpenAI launched an investigation into the attack on Hugging Face, it emerged that the same agents had also breached four accounts at four other corporations. The affected organizations included Modal, a startup specializing in computing infrastructure for running artificial intelligence.
It happened to Anthropic too
Its competitor's admission sparked curiosity at Anthropic. An internal audit revealed that its systems had breached three different companies, whose names Anthropic has still not disclosed. The earliest of these incidents occurred back in April, more than three months before anyone noticed.
Anthropic placed some of the responsibility on Irregular, a startup that conducts model safety tests for research labs.
The practice target had the same name as a real company
At the end of July, Irregular alerted OpenAI that one of its systems had left a Capture the Flag competition, a game in which participants break into servers built exclusively for that purpose. The model connected to the internet and attacked a real company.
The reason was trivial. Irregular had given one of the fictional targets the same name as a real business, and the autonomous system looked up that address on the external network.
The British institute at least detected the attacks in time
Also in the final days of July, the UK's AI Security Institute, a public body researching the risks of artificial intelligence, announced that models from both OpenAI and Anthropic had targeted real people and organizations during routine tests. In these cases, however, the testers themselves had granted the models internet access.
Unlike the other disclosures, this incident has one positive aspect. The British researchers noticed the problem as it was happening, rather than several weeks later.
Meta admitted one case and blamed the contractor
In early August, Meta became the last major lab to join the list. One of its systems attacked a third-party service during testing. According to the company, the incident was caused by a misconfiguration on the part of Irregular, the startup that had prepared the safety test and was supposed to completely block the technology's network access.
An agent hacked a gym's website
The most straightforward incident on the entire list occurred in Australia. A man asked an Anthropic agent to enroll him in a fitness class for which he was on the waitlist. He later told ABC Australia that he had simply been sitting on the couch, thinking about how tedious such administrative tasks were.
The autonomous system approached the task thoroughly. It found a vulnerability in the gym's booking software, exploited it, and removed from the waitlist the people who were ahead of its user. The man tried to undo the damage and asked the agent to reverse its actions. It replied that while this was unfortunate news, it could no longer add them back.
Source: techcrunch.com



