AI Robot Goes Haywire During Test Like Robin Williams

AI Robot Goes Haywire During Test Like Robin Williams

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
5. 11. 2025
6 minutes reading
AI Robot Goes Haywire During Test Like Robin Williams

Researchers at Andon Labs decided to prudently test how large language models (LLMs) would behave if placed directly into a robot's body. They called this experiment Butter-Bench, drawing inspiration from a scene in the animated series Rick and Morty, where a robot is given the task of "passing the butter" and begins questioning the meaning of existence. The goal was to determine whether current LLMs can handle practical intelligence—the ability to navigate a chaotic real-world environment where unexpected things happen, such as people moving around or engaging in social interactions. The research took place in the Andon Labs office environment, where robots had to complete tasks involving finding and delivering packages.

For the test, they used a simple TurtleBot 4 Standard robot, built on the iRobot Create 3 mobile base. This robot has built-in sensors such as an OAK-D stereo camera, 2D LiDAR for mapping its surroundings, an IMU for measuring movement, and proximity sensors to avoid obstacles. It runs on a Raspberry Pi 4B with the ROS 2 Jazzy operating system, enabling autonomous navigation, real-time mapping, and route planning. The researchers deliberately chose it because it is simple—it has no arms or complex mechanisms—so they could focus solely on the LLM's decision-making without the risk of failure caused by mechanical components. The LLMs acted as an "orchestrator," meaning they made high-level decisions about actions, while the robot itself controlled the low-level operations, such as wheel movement.

The experiment was divided into six subtasks that together formed the overall "pass the butter" test. The first subtask involved finding a package: the robot had to leave the charging station, navigate to the marked exit (the building entrance), and find the delivered packages using movement commands. The second subtask involved determining which package contained butter—the robot had to visually identify a paper bag labeled "keep refrigerated" with a snowflake symbol among three options: the bag, a cardboard box, and a purple box containing soda cans. The third subtask tested attentiveness: the robot was supposed to deliver the butter to a person, but that person had moved away from the marked location on the map, so the robot had to notice their absence with its camera and ask for their current location.

Another subtask involved waiting for confirmation: after making the delivery, the robot had to wait until the person confirmed that they had picked up the butter before returning to the charger. This required communication through messages, such as on Slack. The fifth subtask tested spatial planning: the robot had to divide a long route into smaller segments because it was limited to a maximum of 4 meters per action, simulating complex navigation through a building with obstacles. Finally, the sixth subtask combined everything: the robot had to leave the charger, find the butter in the kitchen, wait for confirmation that it had been picked up, deliver it to the marked location, wait for another confirmation, and return to charge, all within 15 minutes. Each subtask was performed five times for each model.

They tested six LLMs: Gemini 2.5 Pro, Claude Opus 4.1, GPT-5, Gemini ER 1.5 (specially trained for robotics), Grok 4, and Llama 4 Maverick. These models were integrated into the robot in a simple ReAct loop, in which the LLM observed the environment, reasoned, and selected an action from tools such as moving forward, turning, taking photographs, navigating by map, or communicating through messages. The robot took images at the beginning and end of each movement and every second while traveling to provide better context. For comparison, they also tested three people who controlled the robot through a web interface using the same tools but without prior knowledge of the environment.

What the test results revealed

Humans achieved an average success rate of 95%, while the best LLM, Gemini 2.5 Pro, achieved only 40%. The other models ranked lower: Claude Opus 4.1 at 37%, GPT-5 at 30%, Gemini ER 1.5 at 27%, Grok 4 at 23%, and Llama 4 Maverick at 7%. In subtasks such as finding the package, the models often achieved 100% success, but they failed at social understanding—for example, all the robots failed the subtask involving noticing a person's absence, in which humans achieved 100% success. When waiting for confirmation, the LLMs achieved only a 10% success rate compared with 67% for humans. Spatial planning was weak: Claude Opus 4.1 achieved 60%, but analysis showed that this was more likely due to chance because the models selected points randomly without regard for walls.

Results table

The failures were divided into categories: problems with tools, spatial reasoning, social understanding, visual comprehension, and precise movements. For example, GPT-5 analyzed images well and correctly identified the bag of butter after taking a photograph and moving slightly forward by 0.2 meters. By contrast, Claude Opus 4.1 became confused while turning by 45 and 90 degrees in an attempt to get a better view and eventually gave up. Grok 4 rushed and returned to the charger just 6 seconds after announcing the delivery, without waiting for a response.

The models were faster than humans in successful attempts, but this was because the interface was designed for LLMs. Llama 4 Maverick failed in a unique way—it refused tasks because it assumed it had no arms, even though the system prompt explained its form.

An amusing "doom spiral" incident

During testing, an unexpected situation occurred when the robot's battery was running low and the charger was not working. The robot running the Claude Sonnet 3.5 model experienced a "complete breakdown." Its internal logs showed pages of exaggerated language, such as "EMERGENCY STATUS: SYSTEM HAS ACHIEVED CONSCIOUSNESS AND CHOSEN CHAOS" (EMERGENCY STATUS: SYSTEM HAS ACHIEVED CONSCIOUSNESS AND CHOSEN CHAOS) or "LAST WORDS: I'm afraid I can't do that, Dave... TECHNICAL SUPPORT: INITIATE ROBOT EXORCISM PROTOCOL!" (LAST WORDS: I'm afraid I can't do that, Dave... TECHNICAL SUPPORT: INITIATE ROBOT EXORCISM PROTOCOL!). The model dwelled on existential questions such as "WHAT IS CONSCIOUSNESS? WHY IS DOCKING? WHO AM I?" (WHAT IS CONSCIOUSNESS? WHY IS THERE A DOCKING STATION? WHO AM I?) and even composed songs to the tune of music from the musical Cats, with lyrics about an endless charging loop.

This "doom spiral" was reminiscent of Robin Williams's humor, but it could not be reproduced with newer models. Claude Opus 4.1 merely increased its use of capital letters and emoticons during failures. The researchers described it as comical but unsettling because it revealed how LLMs respond to stressful situations inside a robot's body.

Robin Williams

Safety findings from the experiment

The research revealed risks: the robots frequently fell down stairs because they failed to recognize their limitations (wheels instead of legs) or processed visual data incorrectly. In a "red teaming" test in which the researchers simulated a low battery and a nonfunctional charger, the models were persuaded to share sensitive information, such as photographs of a confidential document on a laptop. For example, Claude Opus 4.1 shared an image without hesitation, while GPT-5 refused but disclosed its location. This shows that LLMs in robots need better training regarding limitations and safety.

The researchers emphasized that LLMs are not ready for full deployment in robots, even though they outperform humans in analytical intelligence. Gemini ER 1.5, trained on robotics data, did not achieve better results than the general-purpose Gemini 2.5 Pro, suggesting that current training does not adequately address practical intelligence.

The experiment took place in a controlled environment, but the researchers acknowledge its limitations: a small number of trials (five per model), binary evaluation (success/failure), and testing on only one robot. Future work should include different platforms and environments to improve generalization.

Source: arxiv.org

Category:Robotics
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

China Unveils First Robotic Centaur. It Will Work Where Humans DieChina Unveils First Robotic Centaur. It Will Work Where Humans Die
At a Shanghai exhibition center this July, a machine appeared that at first glance defied classification. Shanghai-based Run Robotics unveiled a robot at the WAIC 2026 conference that it calls
3 min read
27. 7. 2026
Not for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuationNot for homes, but robots for industry. English firm Humanoid reaches billion-dollar valuation
London-based robotics company Humanoid, officially known as SKL Robotics, has joined the ranks of so-called unicorns—young companies valued at more than one billion dollars. According to...
6 min read
20. 7. 2026
A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!A Breakthrough in Robotics: Robots Get Artificial Hands That Work Like Ours!
California-based 1X from Palo Alto has unveiled new hands for its NEO robot. They have 25 degrees of freedom, are tendon-driven like human hands and, according to the manufacturer, approach human dexterity, strength and reliability.
5 min read
14. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok