Large language models pass the Turing test for the first time

Large language models pass the Turing test for the first time

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 4. 2025
3 minutes reading
Large language models pass the Turing test for the first time

Revolution in AI: Large Language Models Pass the Turing Test for the First Time

A recently published study on arXiv presents a breakthrough discovery in the field of artificial intelligence: large language models have finally passed the standard Turing test. This result could mark a fundamental turning point in our understanding of the capabilities of artificial intelligence and its potential impact on society.

 

What Is the Turing Test and Why Is It Important?

The Turing test was proposed by mathematician Alan Turing in 1950 as a way to assess whether a machine can exhibit intelligent behavior indistinguishable from that of a human. In the classic version of the test, a human evaluator has a conversation with an unknown counterpart and subsequently decides whether they communicated with a human or a machine. If the evaluator cannot reliably distinguish between the machine and the human, the machine has "passed" the test. Although this test has become an iconic benchmark in the field of artificial intelligence, no system had yet successfully passed it in a controlled scientific environment – until now.

Study Methodology

Researchers Cameron R. Jones and Benjamin K. Bergen conducted two randomized, controlled, preregistered Turing tests with independent groups of participants. The study included four different systems:

  1. ELIZA (a historical chatbot from the 1960s)
  2. GPT-4o
  3. LLaMa-3.1-405B
  4. GPT-4.5

Participants held five-minute conversations simultaneously with another human participant and one of these systems. They were then asked to decide which of their conversational partners was human. And what were the results?

  • GPT-4.5, when instructed to adopt a human persona, was identified as human in 73% of cases – significantly more often than actual human participants! This model therefore clearly passed the Turing test.
  • LLaMa-3.1-405B, with the same instruction, was identified as human in 56% of cases – statistically, it did not differ from actual humans, which means that it also passed the test.
  • The baseline models ELIZA and GPT-4o achieved significantly lower results (23% and 21%), which were substantially below the chance level.

The First Empirical Evidence in History

These results represent the first empirical evidence that an artificial system has passed the standard three-party Turing test. It is a historic moment that has long been predicted by technology visionaries and researchers in the field of artificial intelligence.

This study has far-reaching implications:

  1. Reconsidering AI Intelligence: The results raise questions about the nature and quality of the intelligence demonstrated by large language models. The fact that a machine can convince people that it is human challenges some existing notions about the limits of machine intelligence.
  2. Social Impacts: AI's ability to behave in a manner indistinguishable from humans could dramatically affect a range of areas, from customer service and education to social interactions online.
  3. Economic Consequences: The potential of these models could lead to the transformation of jobs and industries dependent on human communication.
  4. Ethical Questions: The research raises important questions about transparency, consent, and truthfulness in technology-mediated communication.

What Comes Next?

Passing the Turing test represents the beginning rather than the end of the research journey. Questions that now arise include:

  • How will these models continue to evolve?
  • What standards and regulations will be needed for systems that can be mistaken for humans?
  • How can we ensure that these capabilities are used ethically and responsibly?
  • How will our understanding of intelligence, communication, and even humanity change?

The study "Large Language Models Pass the Turing Test" represents a historic milestone in the development of artificial intelligence. For the first time, we have solid scientific evidence that machines can communicate in a way that is indistinguishable to human evaluators from communicating with other humans – and in the case of GPT-4.5, even more convincing than actual humans. This moment is a turning point that forces us to reconsider our assumptions about the limits of artificial intelligence and to begin seriously addressing the social, economic, and philosophical consequences of a world in which machines can communicate like humans – or even better than humans.

 
Note: This article is based on the research study "Large Language Models Pass the Turing Test" by authors Cameron R. Jones and Benjamin K. Bergen, published on arXiv on March 31, 2025.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok