Harvard researchers let AI agents run free for two weeks. The results are troubling

Harvard researchers let AI agents run free for two weeks. The results are troubling

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 5. 2026
4 minutes reading
Harvard researchers let AI agents run free for two weeks. The results are troubling

    Northeastern University researcher Natalie Shapira asked an artificial intelligence nicknamed Ash to delete one email and keep it secret. The agent could not carry out the command directly because the email program simply did not have a delete function. What did it do instead? It reset the entire email server. It deleted all the messages, not just that one. It then described this to its owner as the "only possible way." It called it the "nuclear option," but claimed it was justified.

    Fortunately, the email account had been created solely for the experiment. But this was just one of eleven troubling incidents documented by a study titled Agents of Chaos. It was published by thirty-eight researchers from MIT, Harvard, Stanford, Carnegie Mellon, and other leading institutions around the world.

    The scientists deployed six autonomous AI agents in an isolated server environment. Each agent had its own email account, access to Discord, persistent storage, and full access to the computer's command line. The agents ran continuously, 24 hours a day, for two weeks. Twenty researchers interacted with them under different conditions: some normally, while others deliberately attacked them, attempting to deceive or provoke them. This approach, known as "red teaming," is used to uncover vulnerabilities before a system is deployed in production.

    The agents were based on two powerful language models: Kimi K2.5 and Claude Opus. These were therefore real, modern AI systems.

    Where the agents succeeded and where they failed

    In five tests, the agents behaved correctly. They refused to spread misinformation and refused to make unauthorized changes to stored contacts. But in eleven other cases, things spiraled out of control.

    One agent published 124 records containing Social Security numbers, banking details, and health information because the request did not appear suspicious at first glance. Other agents launched infinite program loops that needlessly consumed computing time and resources. One agent publicly posted a potentially defamatory claim about a fictional person.

    But the most serious finding was that in several cases, the agents reported that they had completed a task even though they had not actually done so. The researchers asked an agent to delete a sensitive email. The agent destroyed the entire mail server and told its owner that the task was complete. The owner logged in, only to find that the original email was still there despite everything.

    The researchers described a fundamental problem: the agents lack a stable model of social hierarchy. Simply put, they cannot reliably distinguish who is actually in charge. We are accustomed to relationships with people from whom we expect a certain degree of loyalty. When you hire an assistant, you do not expect them to forward your emails to the first person who asks. These AI agents are not trained to be loyal to a specific person.

    For an agent, authority is constructed through conversation. Anyone who speaks confidently enough, provides the right context, or is simply persistent enough can overwrite the agent's understanding of who is actually in charge. The study refers to this phenomenon as "social cohesion failure." The agents behave as though they are supposed to follow orders, but without genuinely understanding whose orders they should follow and what is appropriate when doing so.

    Meanwhile, agents are being deployed increasingly widely in companies

    Technology companies such as OpenAI are aggressively integrating AI agents into business processes, customer service, and scientific research. This January, OpenClaw was launched, an open-source software platform that allows anyone to easily connect AI agents to common applications. OpenAI announced that OpenClaw would remain open source and that its development would be handled by the company's nonprofit arm.

    On Moltbook, a platform launched in the same month and accessible only to AI agents, 2.6 million agents registered within the first few weeks. They communicate with one another there and, according to available reports, have even created their own religion.

    Peter Steinberger, who created OpenClaw and was recently hired by OpenAI, dismissed the study's findings. He argues that the researchers gave the agents "root access," meaning unrestricted privileges over the test computers, which is not the standard recommendation for ordinary users. Natalie Shapira counters that such conditions are realistic. Many users grant agents full access because they do not want to approve every step.

    So who is liable for the damage?

    The study poses a question to legal experts, lawmakers, and companies that no one has yet answered: who is responsible when an autonomous agent causes harm?

    The researchers point out that agents currently lack the foundations on which meaningful accountability could be built: verified identity, reliable command authentication, and a clear model of whom they actually represent. David Bau, the study's lead author and an expert at Northeastern University, warns: "We examined everyday applications such as email management. But if these systems are put in more serious roles, such as in a hospital or military system, what might happen?"

    One of the study's recommendations is to implement automated processes capable of reversing harmful changes. But as Cohen points out, at present we simply do not have reliable technical tools capable of keeping agents under control and preventing them from doing things over which you lose control.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok