GPT-5 Achieves Superhuman Results in Medicine

GPT-5 Achieves Superhuman Results in Medicine

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 8. 2025
4 minutes reading · 8 views
GPT-5 Achieves Superhuman Results in Medicine

GPT-5 Achieves Superhuman Results in Medicine

Imagine a world where artificial intelligence not only assists doctors but actually outperforms them in some tests. This is not science fiction, but reality according to a new study from Emory University. This paper, published on arXiv, examines the capabilities of OpenAI's GPT-5 model in multimodal medical reasoning. The study's authors—Shansong Wang, Mingzhe Hu, Qiang Li, and others from the Department of Radiation Oncology at the Winship Cancer Institute—tested the model on various benchmarks and compared it with the previous GPT-4o version as well as human experts. The results are astonishing and demonstrate how rapidly AI is evolving. Let's take a closer look.

How GPT-5 Was Tested

GPT-5 is OpenAI's latest large language model, capable of processing not only text but also images, such as medical scans. The study evaluated it in a zero-shot setting using a chain-of-thought approach, meaning that the model had to reason step by step without prior training on specific data. Tests were conducted on the MedQA dataset, which contains questions from U.S. medical licensing examinations; on MMLU medical subcategories; on USMLE self-assessment examinations; and on multimodal benchmarks such as MedXpertQA and VQA-RAD. For multimodal tasks, the model received textual patient descriptions together with images, such as CT scans, and had to make a diagnosis based on them.

In one specific MedXpertQA case, the model correctly identified Boerhaave syndrome—a rare rupture of the esophagus—based on a combination of laboratory values, such as hemoglobin of 9 g/dL and a mean corpuscular volume of 120 μm³, and physical signs such as suprasternal crepitus and bloody vomiting. The model suggested a Gastrografin swallow, a contrast-agent test, as the next step and explained why other options, such as ondansetron or an epinephrine injection, were inappropriate.

Benchmark Results

On the text-based MedQA benchmark (the version with four answer choices), GPT-5 achieved an accuracy of 95.84%, which is 4.80% better than GPT-4o-2024-11-20, which scored 91.04%. On MedXpertQA Text, the model improved its reasoning score by 26.33% to 56.96% and its comprehension score by 25.30% to 54.84%. On MMLU in medical fields such as anatomy (92.59%), clinical knowledge (95.09%), and medical genetics (100%), it outperformed the previous model by 1–4%.

Benchmark results

On the USMLE self-assessment examinations, which include questions from Step 1, Step 2, and Step 3, GPT-5 achieved an average score of 95.22%, with the greatest improvement in Step 2 (97.50%, +4.17% compared with GPT-4o). This is significantly above the level human physicians need to pass the examinations.

In multimodal tests on MedXpertQA MM, the model excelled with 69.99% in reasoning (+29.26% compared with GPT-4o) and 74.37% in comprehension (+26.18%). On VQA-RAD, which focuses on radiology images, it achieved 70.92%, slightly less than the smaller GPT-5-mini variant but still a solid result. These results show how GPT-5 can integrate textual information, such as vital signs (e.g., blood pressure of 90/63 mmHg and pulse of 130/min), with visual data from images.

Comparison with Human Experts

The comparison with pre-licensure medical professionals, such as students or residents, is the most dramatic. On MedXpertQA Text, GPT-5 outperformed these experts by 15.22% in reasoning and 9.40% in comprehension. On the multimodal MedXpertQA MM, the differences were even more pronounced: +24.23% in reasoning and +29.40% in comprehension. By contrast, GPT-4o remained below the level of these experts in most categories.

According to a tweet by Dr. Derya Unutmaz, who shared these results, this indicates a shift: while GPT-4o was at the human level, GPT-5 is above it. This raises the question of whether it may soon be considered negligence if physicians do not use AI in clinical practice. The study emphasizes that GPT-5 advances the integration of heterogeneous data, such as patient histories, structured data, and images, which could lead to better clinical decisions.

This study from Emory University shows that GPT-5 represents a leap forward in medical AI. With results that outperform human experts in controlled tests, it paves the way for decision-support systems in hospitals. The authors have released the code on GPT-5-Evaluation for further research. Of course, the real world of medicine is more complex than benchmarks, but the direction is clear—AI such as GPT-5 is becoming an indispensable tool.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok