GPT-5 Achieves Superhuman Results in Medicine
Imagine a world where artificial intelligence not only assists doctors but actually outperforms them in some tests. This is not science fiction, but reality according to a new study from Emory University. This paper, published on arXiv, examines the capabilities of OpenAI's GPT-5 model in multimodal medical reasoning. The study's authors—Shansong Wang, Mingzhe Hu, Qiang Li, and others from the Department of Radiation Oncology at the Winship Cancer Institute—tested the model on various benchmarks and compared it with the previous GPT-4o version as well as human experts. The results are astonishing and demonstrate how rapidly AI is evolving. Let's take a closer look.
How GPT-5 Was Tested
GPT-5 is OpenAI's latest large language model, capable of processing not only text but also images, such as medical scans. The study evaluated it in a zero-shot setting using a chain-of-thought approach, meaning that the model had to reason step by step without prior training on specific data. Tests were conducted on the MedQA dataset, which contains questions from U.S. medical licensing examinations; on MMLU medical subcategories; on USMLE self-assessment examinations; and on multimodal benchmarks such as MedXpertQA and VQA-RAD. For multimodal tasks, the model received textual patient descriptions together with images, such as CT scans, and had to make a diagnosis based on them.
In one specific MedXpertQA case, the model correctly identified Boerhaave syndrome—a rare rupture of the esophagus—based on a combination of laboratory values, such as hemoglobin of 9 g/dL and a mean corpuscular volume of 120 μm³, and physical signs such as suprasternal crepitus and bloody vomiting. The model suggested a Gastrografin swallow, a contrast-agent test, as the next step and explained why other options, such as ondansetron or an epinephrine injection, were inappropriate.
Benchmark Results
On the text-based MedQA benchmark (the version with four answer choices), GPT-5 achieved an accuracy of 95.84%, which is 4.80% better than GPT-4o-2024-11-20, which scored 91.04%. On MedXpertQA Text, the model improved its reasoning score by 26.33% to 56.96% and its comprehension score by 25.30% to 54.84%. On MMLU in medical fields such as anatomy (92.59%), clinical knowledge (95.09%), and medical genetics (100%), it outperformed the previous model by 1–4%.

On the USMLE self-assessment examinations, which include questions from Step 1, Step 2, and Step 3, GPT-5 achieved an average score of 95.22%, with the greatest improvement in Step 2 (97.50%, +4.17% compared with GPT-4o). This is significantly above the level human physicians need to pass the examinations.
In multimodal tests on MedXpertQA MM, the model excelled with 69.99% in reasoning (+29.26% compared with GPT-4o) and 74.37% in comprehension (+26.18%). On VQA-RAD, which focuses on radiology images, it achieved 70.92%, slightly less than the smaller GPT-5-mini variant but still a solid result. These results show how GPT-5 can integrate textual information, such as vital signs (e.g., blood pressure of 90/63 mmHg and pulse of 130/min), with visual data from images.
Comparison with Human Experts
The comparison with pre-licensure medical professionals, such as students or residents, is the most dramatic. On MedXpertQA Text, GPT-5 outperformed these experts by 15.22% in reasoning and 9.40% in comprehension. On the multimodal MedXpertQA MM, the differences were even more pronounced: +24.23% in reasoning and +29.40% in comprehension. By contrast, GPT-4o remained below the level of these experts in most categories.
According to a tweet by Dr. Derya Unutmaz, who shared these results, this indicates a shift: while GPT-4o was at the human level, GPT-5 is above it. This raises the question of whether it may soon be considered negligence if physicians do not use AI in clinical practice. The study emphasizes that GPT-5 advances the integration of heterogeneous data, such as patient histories, structured data, and images, which could lead to better clinical decisions.
Also, GPT-5 Pro is even better and performs at the level of leading clinical specialists.
— Derya Unutmaz, MD (@DeryaTR_) August 13, 2025
At this point failing to use these AI models in diagnosis and treatment, when they could clearly improve patient outcomes, may soon be regarded as a form of medical malpractice. https://t.co/yffWk68Z4d
This study from Emory University shows that GPT-5 represents a leap forward in medical AI. With results that outperform human experts in controlled tests, it paves the way for decision-support systems in hospitals. The authors have released the code on GPT-5-Evaluation for further research. Of course, the real world of medicine is more complex than benchmarks, but the direction is clear—AI such as GPT-5 is becoming an indispensable tool.



