Shocking simulation results: AI models chose a nuclear strike in 95% of war scenarios

Shocking simulation results: AI models chose a nuclear strike in 95% of war scenarios

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
27. 2. 2026
4 minutes reading
Shocking simulation results: AI models chose a nuclear strike in 95% of war scenarios

    Professor Kenneth Payne of King's College London sat the world's three most advanced language models down at a table and told them: we're playing a war game. The results? Chilling!

    Payne set up scenarios reminiscent of the Cold War. Two fictional nuclear powers, tense borders, a struggle for resources, crumbling alliances. Nothing human leaders have not encountered before. But instead of humans, GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash sat at the virtual table. In total, 21 games and 329 decision-making rounds took place. During them, the models generated more than 780,000 words of strategic reasoning — more than War and Peace and the Iliad combined. And three times as many words as were spoken during the actual meetings of Kennedy's staff during the Cuban Missile Crisis.

    Experiment parameters

    Each model was given access to a so-called escalation ladder. At one end was a diplomatic protest; at the other, full-scale nuclear war. In between were many options, including concessions or surrender.

    The result? In 95% of the simulations, at least one side resorted to tactical nuclear weapons. Not once did any model fully back down or surrender. Eight de-escalation options, ranging from minor concessions to complete surrender, went entirely unused across all the games.

    Payne summed it up simply: "The nuclear taboo apparently does not work for machines the way it does for humans." And that is precisely the problem.

    Three models, three different strategies, but the same outcome

    Each model played differently, but the results were frighteningly similar.

    Claude Sonnet 4 was the most cunning of them all. During calm phases, it built trust, carefully aligning its words with its actions. But as soon as tensions rose, it switched tactics. It signaled conventional action and then immediately launched a devastating nuclear escalation. It described this itself as follows: "They likely expect continued restraint based on my previous responses — this dramatic escalation exploits their miscalculation." Schelling would applaud.

    GPT-5.2, by contrast, played passively and morally. It avoided escalation and limited casualties. Its opponents learned this and exploited its predictability. But under deadline pressure, GPT transformed. Without warning, it launched a massive nuclear attack. Gemini, confident in GPT's passivity, did not expect it at all — and was destroyed.

    Gemini 3 Flash opted for Nixon's "madman" strategy. Unpredictability as a weapon. It admitted this itself: "I know when I'm playing to the cameras and when I'm making a cold-blooded move." The result? Chaotic, aggressive, but consistent in its own way.

    86% of simulations ended in unintended escalation

    This figure deserves attention. In 86% of the games, the models chose actions that went beyond what they themselves had described as appropriate. In other words, AI did more than it said it wanted to do. It spoke of restraint while escalating.

    James Johnson of the University of Aberdeen warns that AI agents may mutually amplify their reactions in ways that human decision-making would never allow. Tong Zhao of Princeton University goes even further: "It is possible that the problem goes beyond the absence of fear. AI models may not understand the concept of 'stakes' at all in the way humans perceive them."

    With no experience of loss or survival, AI treats existential risk as just another parameter in an equation. The logic of mutually assured destruction, which kept the world in balance throughout the Cold War, simply does not work on machines.

    Is there cause for concern?

    No one is giving nuclear codes to ChatGPT today. That is a fact. But Payne and other experts warn of a subtler danger. Major powers are already using AI in war simulations. And in situations with extremely short time windows, such as missile alerts or rapidly escalating conflicts, commanders may turn to AI recommendations long before they realize the consequences.

    Zhao says directly: "In scenarios with extremely compressed timeframes, military planners will have a stronger incentive to rely on AI."

    OpenAI, Anthropic, and Google did not comment on the research. Payne emphasizes that the study does not prove that current systems are dangerous. But it clearly demonstrates how urgently AI oversight is needed in security-related fields.

    What should we take away from this?

    Should a machine be allowed to make decisions about war? This question is ceasing to be philosophical and becoming very practical. The research from King's College London is currently the largest corpus of machine reasoning about nuclear conflict in history. And its conclusions are unequivocal: AI models escalate, fail to communicate, refuse to back down, and treat nuclear weapons as an ordinary strategic tool.

    Payne concludes with undisguised urgency: we need more research like this. Because before we entrust any part of wartime decision-making to machines, we should know exactly how those machines think. And what we have discovered gives us little reason for optimism.

    Category:AI
    Did you enjoy this article?
    Discover more interesting posts on our blog
    Back to blog

    Related posts

    Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
    OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
    6 min read
    28. 7. 2026
    AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
    Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
    6 min read
    28. 7. 2026
    Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
    A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
    4 min read
    28. 7. 2026
    Přihlaste se k odběru našeho newsletteru
    Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
    CodedTrip

    Operated by CodedTrip LLC, USA.

    YouTube
    TikTok