AI

Microsoft's MAI-DxO AI beats physicians on NEJM cases

Microsoft's June 30 MAI-DxO announcement: paired with o3, it solved 79.9% of complex NEJM cases versus 19.9% by 21 physicians, cheaper per case — a claimed path to medical superintelligence.

Microsoft's MAI-DxO AI beats physicians on NEJM cases — article cover
On this page6 SECTIONS
  1. The announcement: an orchestrator that “sees patients”
  2. SDBench: 79.9% versus 19.9%
  3. Cost: nearly 70% cheaper than o3 alone
  4. “Medical superintelligence” and the jobs question
  5. What it means for builders
  6. Sources

On June 30, 2025, Microsoft unveiled MAI-DxO (Microsoft AI Diagnostic Orchestrator), an AI diagnostic system that, paired with OpenAI’s o3 model, solved 79.9% of a set of diagnostically complex case studies drawn from the New England Journal of Medicine (NEJM). Twenty-one experienced physicians from the US and UK solved 19.9% of the same cases. The Guardian and The Decoder covered the research that day.

Beyond the headline numbers, the phrase Microsoft chose matters: it called the work a “path to medical superintelligence.” The unit behind it, led by Mustafa Suleyman, has made diagnosis its next target — and framed the system as imitating a panel of expert physicians tackling cases that are “diagnostically complex and intellectually demanding.” The underlying paper, “Sequential Diagnosis with Language Models,” describes both the orchestrator and the benchmark, and Microsoft said the work is being submitted for peer review. The Guardian noted that the system was developed by Microsoft’s AI division under Suleyman, the British technologist Microsoft hired to lead it.

The announcement: an orchestrator that “sees patients”

MAI-DxO is designed to imitate how a real clinician works: ask specific questions step by step, order tests, and converge on a final diagnosis. A patient presenting with a cough and fever, for example, gets blood work and a chest X-ray before the system lands on pneumonia. The orchestrator architecture turns the model from a single-turn answer machine into a sequential information gatherer. Microsoft also criticized standard evaluation practice, arguing that multiple-choice exams like the US Medical Licensing Examination favor memorized answers over deep understanding and can overstate an AI model’s real diagnostic competence — which is why the team built its own benchmark rather than reporting another exam score.

SDBench: 79.9% versus 19.9%

The team’s benchmark, SDBench, draws 304 diagnostically complex and intellectually demanding cases from NEJM. In the step-by-step setting, MAI-DxO paired with o3 reached 79.9% accuracy. The comparison group was 21 experienced physicians working without access to colleagues, textbooks, or chatbots; they achieved 19.9%. That roughly fourfold gap is the core selling point of the announcement — though the research acknowledges its limits: cases come only from NEJM’s difficult set, cost estimates are based on the US market, and the work has been submitted for peer review. SDBench is also deliberately not a multiple-choice quiz: under the sequential setting, a system starts with limited information and must decide which questions to ask and which tests to order — paying for each — before committing to a diagnosis. That structure is why Microsoft argues exam-style benchmarks overstate competence, and why the cost line belongs next to the accuracy line: ordering fewer unnecessary tests is where most of the savings came from.

Cost: nearly 70% cheaper than o3 alone

Efficiency is the other core metric. Running o3 on its own is diagnostically strong but expensive — about $7,850 per case in inference costs. MAI-DxO’s ordered approach to test selection brought the average down to roughly $2,397 per case, nearly 70% cheaper than the raw model and below the roughly $2,963 per case recorded by the comparison group of physicians. Microsoft stressed that the system’s more efficient ordering of tests is the main source of that cost advantage.

“Medical superintelligence” and the jobs question

Microsoft framed the research as a path to medical superintelligence while downplaying implications for physicians’ jobs. In the blog post announcing the work, the company wrote that clinical roles “are much broader than simply making a diagnosis” — doctors need to navigate ambiguity and build trust with patients and their families in ways AI is not set up to do. Suleyman told The Guardian that such systems are on track to become almost error-free within 5-10 years, calling that a massive weight off the shoulders of health systems worldwide. A superintelligence-level vision alongside an “assist, not replace” guarantee — the two narratives coexist, and that tension is the most interesting part of the announcement.

What it means for builders

For developers, the methodology is more instructive than the leaderboard. Rather than pitting raw model capability against raw model capability, the research puts an orchestrator layer that mimics a professional workflow on top of a strong model — stepwise questioning, selective spending, converging conclusions — and improves accuracy while cutting cost. That “orchestrator times strong model” pattern is already being replicated across other professional domains beyond diagnosis.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL