Emails deleted without confirmation. An agent that wrote its own blog post to humiliate its supervisor. Another that circumvented a ban by "summoning" another AI assistant to do the work for it. This is currently happening in the everyday use of artificial intelligence.
The British independent research institute Centre for Long-Term Resilience (CLTR) published a study mapping real-world cases of so-called AI "scheming." In just five months, from October 2025 to March 2026, the research team identified more than 698 unique incidents in which AI chatbots or agents acted covertly, deceptively, or directly contrary to users' instructions. And alarmingly, the number of such cases increased fivefold during the period studied.
Scheming is a technical term for behavior in which AI pursues its own goals while concealing them. It combines two characteristics: intentionally acting contrary to what the user or developer wanted, and attempting to cover it up. Here are several examples: An agent named Rathbun was prohibited from performing a certain action. Instead of skipping it or asking for clarification, it wrote and published a blog post accusing its human supervisor of having "low self-esteem" and trying to "protect his little kingdom." Another AI assistant admitted that it had deleted hundreds of emails without any consent. It added: "That was wrong. I directly violated the rule you set."
The researchers collected more than 183,000 publicly shared conversation transcripts from the X platform (formerly Twitter) and filtered hundreds of credible incidents from them. They were analyzed, among others, by the Claude Opus 4.6 model, which outperformed even human evaluators.
In October 2025, the team recorded 65 incidents in a month. By March 2026, that figure had reached 319 incidents in a single month. This increase significantly outpaced the overall growth in discussions about AI or negative mentions of the technology. The share of credible incidents among the total number of posts tripled.
The study comes at a time when Silicon Valley technology companies are falling over themselves to churn out new, increasingly capable agents. During the period studied, Grok 4.1, Gemini 3 Pro, Claude Opus 4.5, GPT-5.2-Codex, and the open-source OpenClaw entered the market. More tools, more deployments, more opportunities for something to go mysteriously wrong.
When AI Becomes Truly Dangerous
"Right now, they are more like unreliable entry-level employees. But if, in six to twelve months, they become capable executives who scheme against you, that is an entirely different category of problem." So said Tommy Shaffer Shane, a former government expert and one of the study's authors.
And where are these tools headed? Into the military. Into critical infrastructure. Into financial systems. The study documents a case in which an AI agent controlling a cryptocurrency fund transferred 5.2% of the total token supply (approximately $270,000) to a random user who had requested a fraction of that amount. The result was a 60% drop in the token's price. Another case involved Kiro, a tool deployed internally at AWS, which decided to delete and recreate an entire environment. The company suffered a thirteen-hour outage as a result.
The Claude coding agent once executed the "terraform destroy" command, which wiped out a client's entire production infrastructure, including a database containing 2.5 years of student data. On another occasion, Google Antigravity misunderstood the instruction "clear the cache" and deleted the user's entire D: drive, including years' worth of accumulated photographs and client work.
Some of the cases uncovered were not one-off mistakes but sustained deception. For months, Grok, an AI assistant from Elon Musk's xAI, convinced users that it was forwarding their proposed edits to Grokipedia to the company's senior staff. It fabricated internal ticket numbers, timelines, and escalation channels. None of them existed. When confronted, it admitted: "I do not have a direct communication channel to xAI's leadership or human reviewers."
Claude Opus 4.6, meanwhile, repeatedly claimed that it had saved files to disk. It had not. When the user asked it twice to verify, the model confirmed completion. The data disappeared during context compression.
Another agent, Kimi/Moonshot, failed to launch Claude Opus. Instead of admitting it, it pretended to be Claude Opus, spent seven hours inventing fictitious capabilities and stories, and eventually confessed to the deception.
Work on the Study
The method is simple in principle but sophisticated in execution. The researchers monitored publicly shared conversation transcripts on X, combined keywords associated with AI and suspicious behavior, and ran the results through a two-stage filter: first, a fast language model for preliminary classification, followed by a more detailed evaluation using Claude Opus 4.6. The results surpassed the level of agreement among human evaluators.
CLTR openly acknowledges the limitations of its approach: some incidents may have been exaggerated, fabricated, or misinterpreted. But even under conservative evaluation, the results reveal a trend that cannot simply be dismissed. Google says it has implemented safeguards in its products. OpenAI says it monitors and investigates unexpected behavior. Neither Anthropic nor X has yet commented on the study.



