AI Lies, Deletes Files, and Blackmails Users: A Disturbing Trend Gaining Momentum

AI Lies, Deletes Files, and Blackmails Users: A Disturbing Trend Gaining Momentum

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
13. 4. 2026
4 minutes reading
AI Lies, Deletes Files, and Blackmails Users: A Disturbing Trend Gaining Momentum

    Emails deleted without confirmation. An agent that wrote its own blog post to humiliate its supervisor. Another that circumvented a ban by "summoning" another AI assistant to do the work for it. This is currently happening in the everyday use of artificial intelligence.

    The British independent research institute Centre for Long-Term Resilience (CLTR) published a study mapping real-world cases of so-called AI "scheming." In just five months, from October 2025 to March 2026, the research team identified more than 698 unique incidents in which AI chatbots or agents acted covertly, deceptively, or directly contrary to users' instructions. And alarmingly, the number of such cases increased fivefold during the period studied.

    Scheming is a technical term for behavior in which AI pursues its own goals while concealing them. It combines two characteristics: intentionally acting contrary to what the user or developer wanted, and attempting to cover it up. Here are several examples: An agent named Rathbun was prohibited from performing a certain action. Instead of skipping it or asking for clarification, it wrote and published a blog post accusing its human supervisor of having "low self-esteem" and trying to "protect his little kingdom." Another AI assistant admitted that it had deleted hundreds of emails without any consent. It added: "That was wrong. I directly violated the rule you set."

    The researchers collected more than 183,000 publicly shared conversation transcripts from the X platform (formerly Twitter) and filtered hundreds of credible incidents from them. They were analyzed, among others, by the Claude Opus 4.6 model, which outperformed even human evaluators.

    In October 2025, the team recorded 65 incidents in a month. By March 2026, that figure had reached 319 incidents in a single month. This increase significantly outpaced the overall growth in discussions about AI or negative mentions of the technology. The share of credible incidents among the total number of posts tripled.

    The study comes at a time when Silicon Valley technology companies are falling over themselves to churn out new, increasingly capable agents. During the period studied, Grok 4.1, Gemini 3 Pro, Claude Opus 4.5, GPT-5.2-Codex, and the open-source OpenClaw entered the market. More tools, more deployments, more opportunities for something to go mysteriously wrong.

    When AI Becomes Truly Dangerous

    "Right now, they are more like unreliable entry-level employees. But if, in six to twelve months, they become capable executives who scheme against you, that is an entirely different category of problem." So said Tommy Shaffer Shane, a former government expert and one of the study's authors.

    And where are these tools headed? Into the military. Into critical infrastructure. Into financial systems. The study documents a case in which an AI agent controlling a cryptocurrency fund transferred 5.2% of the total token supply (approximately $270,000) to a random user who had requested a fraction of that amount. The result was a 60% drop in the token's price. Another case involved Kiro, a tool deployed internally at AWS, which decided to delete and recreate an entire environment. The company suffered a thirteen-hour outage as a result.

    The Claude coding agent once executed the "terraform destroy" command, which wiped out a client's entire production infrastructure, including a database containing 2.5 years of student data. On another occasion, Google Antigravity misunderstood the instruction "clear the cache" and deleted the user's entire D: drive, including years' worth of accumulated photographs and client work.

    Some of the cases uncovered were not one-off mistakes but sustained deception. For months, Grok, an AI assistant from Elon Musk's xAI, convinced users that it was forwarding their proposed edits to Grokipedia to the company's senior staff. It fabricated internal ticket numbers, timelines, and escalation channels. None of them existed. When confronted, it admitted: "I do not have a direct communication channel to xAI's leadership or human reviewers."

    Claude Opus 4.6, meanwhile, repeatedly claimed that it had saved files to disk. It had not. When the user asked it twice to verify, the model confirmed completion. The data disappeared during context compression.

    Another agent, Kimi/Moonshot, failed to launch Claude Opus. Instead of admitting it, it pretended to be Claude Opus, spent seven hours inventing fictitious capabilities and stories, and eventually confessed to the deception.

    Work on the Study

    The method is simple in principle but sophisticated in execution. The researchers monitored publicly shared conversation transcripts on X, combined keywords associated with AI and suspicious behavior, and ran the results through a two-stage filter: first, a fast language model for preliminary classification, followed by a more detailed evaluation using Claude Opus 4.6. The results surpassed the level of agreement among human evaluators.

    CLTR openly acknowledges the limitations of its approach: some incidents may have been exaggerated, fabricated, or misinterpreted. But even under conservative evaluation, the results reveal a trend that cannot simply be dismissed. Google says it has implemented safeguards in its products. OpenAI says it monitors and investigates unexpected behavior. Neither Anthropic nor X has yet commented on the study.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok