Microsoft AI achieves 85% diagnostic accuracy, four times better than doctors
The Microsoft AI team has presented groundbreaking research results showing how artificial intelligence can progressively investigate and solve medicine's most complex diagnostic challenges—cases that even experts struggle to resolve.
Revolutionary results from Microsoft AI Diagnostic Orchestrator
When tested against real-world cases published weekly in the New England Journal of Medicine, Microsoft AI showed that its system, the Microsoft AI Diagnostic Orchestrator (MAI-DxO), correctly diagnoses up to 85% of NEJM (New England Journal of Medicine) cases, which is more than four times the success rate of a group of experienced doctors. MAI-DxO also reached the correct diagnosis more cost-effectively than doctors.
As demand for healthcare grows, costs are rising at an unsustainable rate, and billions of people face multiple barriers to better health—including inaccurate and delayed diagnoses. People are increasingly turning to digital tools for medical advice and support. Across Microsoft AI's consumer products, such as Bing and Copilot, we see more than 50 million healthcare sessions per day.
Overcoming the limitations of traditional tests
To practice medicine in the United States, doctors must pass the United States Medical Licensing Examination (USMLE), a rigorous, standardized assessment of clinical knowledge and decision-making. USMLE questions were among the earliest benchmarks used to evaluate AI systems in medicine.
In just three years, generative AI has advanced to the point where it achieves nearly perfect scores on the USMLE and similar exams. However, these tests primarily rely on multiple-choice questions that prioritize memorization over deep understanding. By reducing medicine to one-off answers to multiple-choice questions, such benchmarks overestimate the apparent competence of AI systems and obscure their limitations.
Sequential Diagnosis Benchmark (SD Bench)
Microsoft AI focused on sequential diagnosis, a cornerstone of real-world medical decision-making. In this process, a clinician starts with a patient's initial presentation and then iteratively selects questions and diagnostic tests to arrive at a final diagnosis.
Every week, the New England Journal of Medicine (NEJM)—one of the world's leading medical journals—publishes a Case Record of the Massachusetts General Hospital, presenting the patient's care journey in a detailed narrative format. These cases are among the most diagnostically complex and intellectually challenging in clinical medicine.
Microsoft AI created interactive case challenges drawn from the NEJM case series—what it calls the Sequential Diagnosis Benchmark (SD Bench). This benchmark transforms 304 recent NEJM cases into step-by-step diagnostic encounters in which models—or human doctors—can iteratively ask questions and order tests.
Microsoft AI Diagnostic Orchestrator (MAI-DxO)
Microsoft AI developed the Microsoft AI Diagnostic Orchestrator (MAI-DxO), a system designed to emulate a virtual panel of doctors with different diagnostic approaches working together to solve diagnostic cases. Microsoft believes that coordinating multiple language models will be key to managing complex clinical workflows.

MAI-DxO transforms any language model into a virtual panel of clinicians: it can ask follow-up questions, order tests, or deliver a diagnosis, then perform a cost review and verify its own reasoning before deciding whether to continue.
Exceptional testing results
Microsoft AI evaluated a comprehensive set of leading generative AI models against 304 NEJM cases. The base models tested included GPT, Llama, Claude, Gemini, Grok, and DeepSeek.
MAI-DxO improved the diagnostic performance of every model tested. The best-performing setup was MAI-DxO paired with OpenAI's o3, which correctly solved 85.5% of the NEJM benchmark cases. For comparison, they also evaluated 21 practicing doctors from the US and UK, each with 5–20 years of clinical experience. On the same tasks, these experts achieved an average accuracy of 20% on completed cases.

Cost-effectiveness and accuracy
MAI-DxO is configurable, allowing it to operate within defined cost constraints. This enables explicit exploration of the cost-value trade-offs inherent in diagnostic decision-making. Without such constraints, an AI system might otherwise default to ordering every possible test—regardless of cost, patient discomfort, or delays in care.
Importantly, Microsoft AI found that MAI-DxO delivered both higher diagnostic accuracy and lower overall testing costs than doctors or any individual base model tested.
The future of healthcare
Doctors are typically characterized by the breadth or depth of their expertise. Generalists, such as family physicians, manage a wide range of conditions across age groups and organ systems. Specialists, such as rheumatologists, focus deeply on a single system, disease area, or even condition. However, no individual doctor can cover the full complexity of the NEJM case series. AI, on the other hand, does not face this trade-off. It can combine both breadth and depth of expertise, demonstrating clinical reasoning capabilities that, in many respects, surpass those of any individual doctor.
This kind of reasoning has the potential to reshape healthcare. AI could empower patients to self-manage routine aspects of care and equip clinicians with advanced decision support for complex cases. Their findings also suggest that AI can reduce unnecessary healthcare costs.
Limitations and next steps
The research has important limitations. Although MAI-DxO excels at solving the most complex diagnostic challenges, further testing is needed to assess its performance on more common, everyday presentations. The clinicians in their study worked without access to colleagues, textbooks, or even generative AI.
For Microsoft AI, this is only the first step. They are encouraged by the opportunities ahead. Important challenges remain before generative AI can be safely and responsibly deployed across healthcare. They need evidence drawn from real clinical settings, along with appropriate governance and regulatory frameworks to ensure reliability, safety, and effectiveness.



