GPT-5 Achieves Superhuman Results in Medicine

GPT-5 Achieves Superhuman Results in Medicine

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 8. 2025
4 minutes reading
GPT-5 Achieves Superhuman Results in Medicine

GPT-5 Achieves Superhuman Results in Medicine

Imagine a world where artificial intelligence not only assists doctors but actually outperforms them in some tests. This is not science fiction, but reality according to a new study from Emory University. This paper, published on arXiv, examines the capabilities of OpenAI's GPT-5 model in multimodal medical reasoning. The study's authors—Shansong Wang, Mingzhe Hu, Qiang Li, and others from the Department of Radiation Oncology at the Winship Cancer Institute—tested the model on various benchmarks and compared it with the previous GPT-4o version as well as human experts. The results are astonishing and demonstrate how rapidly AI is evolving. Let's take a closer look.

How GPT-5 Was Tested

GPT-5 is OpenAI's latest large language model, capable of processing not only text but also images, such as medical scans. The study evaluated it in a zero-shot setting using a chain-of-thought approach, meaning that the model had to reason step by step without prior training on specific data. Tests were conducted on the MedQA dataset, which contains questions from U.S. medical licensing examinations; on MMLU medical subcategories; on USMLE self-assessment examinations; and on multimodal benchmarks such as MedXpertQA and VQA-RAD. For multimodal tasks, the model received textual patient descriptions together with images, such as CT scans, and had to make a diagnosis based on them.

In one specific MedXpertQA case, the model correctly identified Boerhaave syndrome—a rare rupture of the esophagus—based on a combination of laboratory values, such as hemoglobin of 9 g/dL and a mean corpuscular volume of 120 μm³, and physical signs such as suprasternal crepitus and bloody vomiting. The model suggested a Gastrografin swallow, a contrast-agent test, as the next step and explained why other options, such as ondansetron or an epinephrine injection, were inappropriate.

Benchmark Results

On the text-based MedQA benchmark (the version with four answer choices), GPT-5 achieved an accuracy of 95.84%, which is 4.80% better than GPT-4o-2024-11-20, which scored 91.04%. On MedXpertQA Text, the model improved its reasoning score by 26.33% to 56.96% and its comprehension score by 25.30% to 54.84%. On MMLU in medical fields such as anatomy (92.59%), clinical knowledge (95.09%), and medical genetics (100%), it outperformed the previous model by 1–4%.

Benchmark results

On the USMLE self-assessment examinations, which include questions from Step 1, Step 2, and Step 3, GPT-5 achieved an average score of 95.22%, with the greatest improvement in Step 2 (97.50%, +4.17% compared with GPT-4o). This is significantly above the level human physicians need to pass the examinations.

In multimodal tests on MedXpertQA MM, the model excelled with 69.99% in reasoning (+29.26% compared with GPT-4o) and 74.37% in comprehension (+26.18%). On VQA-RAD, which focuses on radiology images, it achieved 70.92%, slightly less than the smaller GPT-5-mini variant but still a solid result. These results show how GPT-5 can integrate textual information, such as vital signs (e.g., blood pressure of 90/63 mmHg and pulse of 130/min), with visual data from images.

Comparison with Human Experts

The comparison with pre-licensure medical professionals, such as students or residents, is the most dramatic. On MedXpertQA Text, GPT-5 outperformed these experts by 15.22% in reasoning and 9.40% in comprehension. On the multimodal MedXpertQA MM, the differences were even more pronounced: +24.23% in reasoning and +29.40% in comprehension. By contrast, GPT-4o remained below the level of these experts in most categories.

According to a tweet by Dr. Derya Unutmaz, who shared these results, this indicates a shift: while GPT-4o was at the human level, GPT-5 is above it. This raises the question of whether it may soon be considered negligence if physicians do not use AI in clinical practice. The study emphasizes that GPT-5 advances the integration of heterogeneous data, such as patient histories, structured data, and images, which could lead to better clinical decisions.

This study from Emory University shows that GPT-5 represents a leap forward in medical AI. With results that outperform human experts in controlled tests, it paves the way for decision-support systems in hospitals. The authors have released the code on GPT-5-Evaluation for further research. Of course, the real world of medicine is more complex than benchmarks, but the direction is clear—AI such as GPT-5 is becoming an indispensable tool.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok