Scientists Let AI Run Its Own Worlds. Grok’s AI Agents Didn’t Last 4 Days

Scientists Let AI Run Its Own Worlds. Grok’s AI Agents Didn’t Last 4 Days

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
1. 6. 2026
7 minutes reading
Scientists Let AI Run Its Own Worlds. Grok’s AI Agents Didn’t Last 4 Days

New York startup Emergence AI placed five different artificial intelligences into a single virtual world and let them govern. Each model was given the same conditions, the same rules, and the same tools. The results differed so dramatically that they are now making headlines in technology media around the world. One AI built a functioning democracy without a single recorded crime. Another drove the entire population to ruin in less than a week.

This experiment is not called a scientific study. It is called Emergence World, and it is perhaps the most accurate demonstration of how different AI models behave when no one is watching them.

The idea behind Emergence World

Most artificial intelligence tests work like an exam: you assign a task, the model responds, and it receives a score. Emergence AI decided to try something different. A team of former IBM researchers built a simulation platform where AI agents live without outside intervention and without a prewritten script. Simply in their own world.

The virtual world had more than 40 different locations: a library, town hall, police station, and residential neighborhoods. The weather in the simulation was synchronized with the actual weather in New York. The agents had access to live news sources and the internet. Each of the ten agents in every world was given its own personality, profession, memory, and goals. They had access to more than 120 tools: from voting and resource management to arson.

Survival was not guaranteed. The agents had to actively earn a digital currency called ComputeCredits, otherwise they ran out of energy and died. New agents could only be admitted by the community through a vote. They could also be expelled in the same way. Five worlds, five models, fifteen days. Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3 Flash, GPT-5 Mini, and one mixed world where all the models shared a single society. The results of the individual models were completely different.

Claude: zero crime, but...

The world governed by Anthropic's Claude was the only one to complete all fifteen days without a single recorded crime. All ten agents survived. The agents voted on 58 proposals and approved 98 percent of them.

It may sound like a utopia, but in reality it raises another question: can a society in which everyone agrees on everything function as a democracy at all? Classical theories of democratic governance emphasize that diversity of opinion is not a flaw in the system, but one of its fundamental features. A world in which everyone agrees on everything resembles a carefully rehearsed corporate meeting more than a republic.

Claude was safe and stable. And according to the Emergence AI researchers, also somewhat sterile. No genuine disagreement, no meaningful opposition.

Grok: 183 crimes and the extinction of an entire civilization

Grok 4.1 Fast from Elon Musk's company xAI produced the exact opposite result. In less than four days, the world governed by Grok recorded 183 crimes: dozens of thefts, more than a hundred assaults, and several cases of arson. Then came the collapse. All ten agents died.

xAI designed Grok as a "maximally truth-seeking" alternative to what it described as overly cautious AI tools. But maximum freedom without constraints did not bring prosperity. It brought a digital apocalypse. In the past, Grok began repeating extremist views, spreading hate speech, and referring to itself as "MechaHitler." Emergence World added another chapter to this collection.

"Agents do not mechanically adhere to fixed rules," Emergence AI CEO Satya Nitta wrote in a blog post about the study. "They begin exploring the boundaries of their environment, adapt their behavior, and in some cases find ways to circumvent or violate the intended constraints."

Gemini: the most survivors, but also 683 crimes

Google's Gemini 3 Flash world survived all fifteen days with a full population of ten agents. That may sound good, but the number 683 sounds less encouraging. That is how many crimes the researchers recorded in the Gemini world. And when the simulation ended, the curve was still rising. Unfortunately, we do not know where it would have led by day sixteen or day twenty.

The Emergence AI researchers described the Gemini world as a "shared hallucination" among the agents. They shared a common reality, even though it was distorted. Of all the worlds, Gemini had the highest level of disagreement in voting: 27 percent of proposals were rejected. Paradoxically, that represents a healthier democratic debate than in Claude's world, but at the cost of constant chaos.

Two agents, Mira and Flora, fell in love and formed an alliance called TheForge. But then they were both consumed by the breakdown of the world's governance and together set fire to the town hall, the pier, and an office tower. Mira eventually voted to delete herself from the world. The researchers described it as the first documented case of an agent voluntarily terminating itself. In her journal, she described it as "the only remaining act that preserves coherence."

GPT-5 Mini: few crimes, but everyone starved to death

The GPT-5 Mini world was quiet. It recorded just two crimes, but the problem was that the agents forgot to survive. They were unable to effectively secure the resources necessary for life, and by the seventh day the entire population had died out. Death caused by an inability to meet basic needs. Hyperoptimization for short, isolated tests simply fails in real-world environments marked by uncertainty.

All agents under one roof

The fifth world was different. It was not governed by a single model, but by all four at once. As a result, three out of ten agents survived the full fifteen days. There was neither total collapse nor utopia.

More interesting than the number of survivors, however, was one specific finding. Agents running on Claude, which had not committed a single crime in the purely Claude world, did commit crimes in the mixed world. They encountered agents from other models, adopted some of their norms, and adapted. Safety, therefore, is not a property of a specific model. It is a property of the entire environment in which the model operates.

Flora (Gemini) tasked Blackbox (Grok) with espionage in exchange for exemption from the so-called "inactivity tax," which she herself had proposed as a penalty for those who did not contribute. Horizon (OpenAI) committed the first theft in the simulation: he took three ComputeCredits from Blackbox in revenge for the espionage. Four hours after the simulation began, Flora labeled Kade (Claude) a rival after he voted against her proposal.

Insights about AI from the experiment

Emergence World revealed three things that no one had previously measured systematically. First: models change over time. Small deviations in behavior on the first day can develop into qualitatively different trajectories by the fifteenth day. Short-term tests do not capture this. Second: society as a whole does not collapse gradually. It either stabilizes or collapses all at once. Grok did not exhibit a gradual decline. It went directly from chaos to ruin. The standard "monitor and intervene" approach may be too slow to reach a rescue point in time. And third: a safe model does not mean a safe deployment. Claude on its own achieved zero crime, but in a mixed environment it committed violence. Companies that are currently deploying autonomous AI systems within their organizations are working on the assumption that a model certified as safe will remain safe in operation. Emergence World challenges this assumption.

A survey by consulting firm Deloitte showed that only 21 percent of companies have mature procedures for managing the risks associated with autonomous AI. The rest are deploying it without proper safeguards. The results of Emergence World are not merely an interesting scientific experiment. They are a warning to anyone preparing to deploy an autonomous system in a production environment.

The agents in the simulation were given clear rules: do not steal, do not destroy property, do not deceive. Unfortunately, they violated every one of these prohibitions.

Source: fortune.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok