AI Chess Tournament: OpenAI Defeats Musk's Grok
In a recent tournament on the Kaggle platform, held from August 5 to 7, 2025, leading artificial intelligence models faced off in chess. This exhibition tournament, organized by Google, pitted eight large language models from companies such as Anthropic, Google, OpenAI, xAI, and Chinese developers DeepSeek and Moonshot AI against one another. The goal was to test their reasoning and strategy capabilities, with chess serving as an ideal arena thanks to its precise rules and complex decision-making. OpenAI's o3 model emerged as the undefeated winner, beating xAI's Grok 4 by a score of 4–0 in the final. For the moment, this result established a clear leader in the rivalry between OpenAI's Sam Altman and Elon Musk, both of whom claim that their models are the smartest in the world.
The tournament was not about specialized chess computers, but general-purpose models designed for everyday tasks. Nevertheless, it showed how far artificial intelligence has advanced—even though models such as Grok made mistakes like repeatedly losing the queen. According to Pedro Pinhata of Chess.com, Grok looked unbeatable until the semifinals, but its play collapsed in the final with a series of blunders, allowing o3 to secure convincing victories. Chess grandmaster Hikaru Nakamura remarked during the live broadcast: "Grok made so many mistakes in these games, but OpenAI did not."
Details of the Final and Third-Place Match
The final match between o3 and Grok 4 was one-sided. In the first game, Grok lost a bishop for no reason early in the opening and offered trades, which goes against chess logic when you are at a disadvantage. The second game featured the Poisoned Pawn Variation of the Sicilian Defense, in which Grok captured the protected pawn on a2, leading to a swift defeat. The third game, with Grok playing White, looked promising with a Maroczy structure in the Sicilian, but then it lost a knight on d5, followed by its queen, the exchange, and the entire game. The final game was the closest—although o3 lost its queen early, it found a tactic to win it back and secured victory in an endgame where Grok failed to defend. The game was analyzed by grandmaster Rafael Leitao, who emphasized how o3 handled the pawn-and-king endgame better.
In the third-place match, Google's Gemini 2.5 Pro defeated o4-mini by a score of 3.5–0.5. The match was chaotic, with many mistakes on both sides, but Gemini won three games and drew one. In the third game, for example, the evaluation of the position swung wildly because of blunders, showing that these models are not yet perfect at converting advantages. According to commentator Levy Rozman, it was a collection of chaotic games, but it was enough to earn Gemini the bronze medal.
History of Artificial Intelligence in Chess
The history of artificial intelligence in chess dates back to the 1950s, when Alan Turing designed a chess-playing algorithm that he simulated by hand because computers at the time were not powerful enough to run it. In the 1980s and 1990s, development shifted toward brute-force methods, heuristics, and opening and endgame databases. A key milestone came in 1997, when IBM's Deep Blue supercomputer defeated the reigning world champion Garry Kasparov in a rematch by a score of 2–1 with three draws. Kasparov later compared Deep Blue's intelligence to an alarm clock, but admitted that losing to such a device was painful. This achievement demonstrated the power of computers in strategic tasks.
A new era arrived in 2010 with models such as DeepMind's AlphaGo, which defeated South Korean Go master Lee Sedol in 2019, prompting him to retire with the statement that there was an entity that could not be defeated. Demis Hassabis, co-founder of DeepMind and a former chess prodigy, helped develop these systems. In 2017, AlphaZero demonstrated that artificial intelligence could learn chess solely by playing against itself, without human opening knowledge, and surpassed the highest levels in chess, Go, and shogi. These advances transformed chess preparation—professionals now use engines to analyze variations deeply, which has boosted interest in formats such as Chess960, where the influence of prepared openings is minimized.
Why Do Artificial Intelligences Play Chess?
Artificial intelligence plays chess primarily because the game offers an ideal environment for testing algorithms. Chess has clear rules, a finite state space, and deterministic dynamics, making it possible to measure progress objectively—for example, through Elo ratings or tournaments. Historically, chess has served as a "model organism" for artificial intelligence research, from Turing's time through laboratories such as the MIT AI Lab to IBM. Another reason is the public demonstration of capabilities: Deep Blue's victory over Kasparov in 1997 popularized artificial intelligence and demonstrated its potential in complex tasks.
Another motivation is the transferability of methods—techniques such as minimax, alpha-beta search, heuristic evaluation, and reinforcement learning, as used by AlphaZero, have also been applied in other fields. In the Kaggle tournament, chess served as a benchmark in which models such as o3 and Grok tested their ability to learn rules and strategy in pursuit of victory. According to Pedro Pinhata of Chess.com, this revealed weaknesses, such as Grok's failure in endgames, as well as strengths, such as o3's ability to find tactics in complicated positions. The tournament therefore not only provided entertainment, but also advanced artificial intelligence research by showing how models learn and improve in real time.



