Analysis · Benchmarks
The scoreboard that can’t be drawn
Fifteen studies say AI matches or beats clinicians in global health. Grouped honestly, the claim fractures: sometimes there is no fair test, sometimes AI beats the weakest human but trails the best, sometimes it genuinely ties the experts — and sometimes it loses. On a dozen scales that don’t compare, and almost never measured on a patient.
Every few weeks a study says an AI has matched or beaten doctors, and the obvious next move is a bar chart: AI against human, tallest bar wins. We tried to build that chart for fifteen “AI-versus-clinician” results in global health. It can’t be built — the studies don’t share a scale, and “beats doctors” turns out to mean at least five different things. So instead of a false ladder, the table below groups them by what the comparison actually was. Every cell is traced to the study’s primary source; nothing is averaged or ranked across rows.
| Study & claim | What the score actually graded | A comparable human number? | The gap, on the study’s own scale | A patient outcome? |
|---|---|---|---|---|
| 1 · No fair test — the human is the answer key, handicapped, or on a different exam | ||||
| Rwanda: LLMs “outperformed” GPsNature Health 2026 · PMC12880909 | Answer quality on an 11-metric, 1–5 rubric (506 written Q&A pairs, 6 physician raters). | No single clinician score exists. The paper reports per-metric gaps, not one human bar to place beside the models. | Top model averaged +0.83 of a point over local GPs (range 0.38–1.10 across the 11 metrics); models ≈ 4.5/5. | No. |
| Microsoft MAI-DxO: “four times higher” than physiciansSDBench · arXiv:2506.22405 | Diagnostic accuracy (%) on 304 deliberately rare NEJM text cases, answer already known. | Not on the same test. Physicians scored ≈ 20% on a 56-case held-out split, barred from references; the AI figures are on the full 304-case set. | o3 alone 78.6%, MAI-DxO+o3 80%, max ensemble 85.5% (≈ US$7,184/case) vs physicians ≈ 20%. | No. |
| Google AMIE: “superior” to PCPsNature 2025 · s41586-025-08866-7 | Consultation-quality rubric in a text-chat OSCE with simulated patient-actors. | Yes — but handicapped. The PCPs were made to work through an unfamiliar text-chat the authors call not representative of usual practice. | AMIE rated higher on 30 of 32 axes by specialists and 25 of 26 by patient-actors (published Nature figures). | No. Simulated patients. |
| OpenAI HealthBencharXiv:2505.08775 | A weighted, normalized score over 48,562 physician-written rubric criteria (some worth negative points), graded by an AI (GPT-4.1). | A physician-written response baseline exists — but it is a writing task, not clinical practice, and the grader is conflicted: OpenAI builds it, an OpenAI model marks it (judge–physician F1 ≈ 0.71), an OpenAI model tops it. | Best model (o3) scored 0.60 on the weighted rubric; the human reference is physicians’ own written responses, not a clinician taking a diagnostic test. | No. |
| TB computer-aided detection (12 products)External validation · PMC11339183 | Chest-X-ray TB detection, sensitivity/specificity on a South-African cohort. | None — no radiologist arm. The tools are validated against a microbiological reference, not against the humans they would replace. | ≈ 64% sensitivity at a borrowed literature threshold; one product’s specificity fell ≈ 87% → 37% in previously-treated patients. | No. |
| Rheumatic-heart-disease echo, Ugandan childrenJAHA 2024 · PMID 38226515 | RHD detection from 511 children’s echocardiograms (GOAL trial). | None scored. Expert cardiologists are the answer key the AI is graded against, not a measured arm — so the authors’ conclusion that AI has “the potential to detect RHD as accurately as expert cardiologists” is a hope, not a result. | AI detection AUC 0.84 (recall 0.98); a mitral-regurgitation model AUC 0.93. No clinician score on the same test to compare. | No. |
| 2 · AI beats the weakest human — and trails the best | ||||
| Cervical AI (AVE): beats VIALancet Global Health · PubMed 41386246 | Sensitivity (and specificity) for precancer (CIN2+) against biopsy, 18,086 women, five countries. | Yes, matched — VIA (a health-worker visual screen) at 36.6%. But the recommended screen, HPV, beats the AI. | AVE 60.1% sensitivity vs VIA 36.6% (+23.5 pts), and −30 pts vs HPV’s 90.4%. On specificity the trade reverses: AVE 81.9% vs VIA 94.2%. | No. No cancer-incidence, mortality or over-treatment measured. |
| Oral-cancer smartphone AI, IndiaSci Rep 2022 · PMC9395355 | Sensitivity for oral (pre)cancer from smartphone images — the AI on a 1,416-image test set; frontline workers across 4,728 people screened. | Matched vs frontline health workers — but both are scored against a remote specialist, so it measures agreement with a human, not truth. | AI 82–87% sensitivity vs frontline workers 60%. Onsite specialists reach 94% — but on a different, biopsy-checked subset (n=102). | No. |
| ENT / otoscopy LLM, rural KenyaOtolaryngol HNS 2025 · PMID 40525718 | Diagnostic concordance with the on-site ENT specialist on 63 real patients (a Claude-3.5 model). | Two humans, same cases: beats primary-care practitioners (79.4% vs 50.8%), and matches the specialist — but the specialist is the reference, so that’s agreement, not accuracy. | 79.4% concordance with the specialist; 96.8% management alignment. n=63, single site, single author. | No. |
| 3 · A genuine match with experts — on narrow perception tasks with a clear ground truth | ||||
| Diabetic-retinopathy screening, ThailandLancet Digital Health 2022 · PMID 35272972 | Vision-threatening DR on fundus images, ~7,651 patients, 9 sites, national programme. | Matched vs retina specialists on the same images — and here the AI genuinely wins. | DL sensitivity 91.4% vs retina specialists 84.8% (p=0.024); specificity tied (~95%). A real head-to-head win. | No — accuracy only; no vision or treatment outcome. |
| AI ultrasound gestational age, ZambiaJAMA 2024 · PMID 39088200 | Gestational-age error (days) from blind-sweep ultrasound, 399 pregnancies, 14–27 wk. | Matched vs credentialed sonographers — and equivalent. The twist: the AI arm was run by untrained novices. | AI 3.2 vs sonographers 3.0 days mean error (diff 0.2 d, 95% CI −0.1 to 0.5). Cohort partly US; dating pre-curated. | No. |
| 4 · AI underperforms the experts | ||||
| Malaria microscopy AI (EasyScan GO)Malaria Journal 2022 · PMC9004086 | Parasite detection, species ID and density on 2,250 slides across 11 countries, vs expert microscopy. | Matched — and the AI loses. Graded on the same WHO scale experts are certified on, it reached Level 2, below the Level 1 “Expert” bar. | Detection sensitivity 89% / species ID 84% (Expert needs >90%); density within ±25% in only 23% of slides; specificity drops to 76% on field-quality slides. | No. |
| 5 · Not “AI versus doctor” at all — AI assisting a clinician, or graded only on paper | ||||
| OpenAI–Penda Health: “16% fewer errors”arXiv:2507.16947 | Physician-rated documentation errors on de-identified visit notes. | N/A — clinician-with-AI vs clinician-without (help-the-doctor, not beat-the-doctor). | 16% fewer diagnostic and 13% fewer treatment errors in the documented decisions. | Attempted — and null: 8-day “feeling better” 3.8% vs 4.3% (n.s.; ~60% unreachable; authors call it exploratory). |
| AI-assisted physicians RCT, PakistanNature Health 2026 · s44360-025-00007-8 | Diagnostic-reasoning score on simulated vignettes; 58 physicians, single-blind RCT. | N/A — GPT-4o assisting physicians vs physicians alone (not AI vs clinician). | Physicians +GPT-4o 71.4% vs 42.6% alone (+27.5 pts). But only after a 20-hour AI-literacy course, and on vignettes. | No — simulated cases. |
| Locally-authored vignettes, KenyamedRxiv 2025 · preprint | Blinded Kenyan physicians rating 5 LLMs vs local clinicians on 507 local vignettes (safety, accuracy). | Matched on the rubric — LLMs rated safer (harm score 4.3–4.8 vs clinicians’ 3.2). But a writing task, not care. | LLMs > clinicians on most safety/accuracy domains, at ~$0.01 vs $3.35 per vignette. Preprint. | No — and in a separate Nigeria field trial of LLM support, the on-site physicians saw little to no improvement in health workers’ decisions. |
A red cell flags what is missing or mismatched — no comparable human number, a handicapped human arm, a conflicted grader, or no patient outcome. Every value is traced to the primary source in the last section.
1 · No common scale. The fifteen report on a dozen non-comparable measures — rubric points, diagnostic accuracy, screening sensitivity, gestational-age error in days, a WHO competence tier, cost per vignette. There is no ladder to rank them on.
2 · “AI beats doctors” is at least five different claims — the five groups above. And the genuinely convincing wins (group 3) are all narrow perception tasks with an objective ground truth; on open-ended clinical reasoning, there is no clean win.
3 · Almost nobody measured a patient. Fourteen of the fifteen stop at a proxy — a graded answer, a paper diagnosis, a screen’s sensitivity. The one study that asked patients directly (Penda) found no effect; and a separate field trial in Nigeria that put LLM support into real clinics found little to no improvement in health workers’ decisions, judged by the on-site physicians who treated the same patients.
These are the questions from our field guide on “AI beats doctors” claims, turned back on the studies that prompted it — the table is that guide applied at scale. Two patterns are worth naming.
Where AI wins, it wins narrow
Strip out the studies with no fair comparison and the ones that only beat the weakest human, and a real signal survives: on diabetic-retinopathy screening in Thailand the AI genuinely out-read retina specialists (91.4% vs 84.8% sensitivity), and novice-acquired ultrasound in Zambia matched credentialed sonographers. Both are narrow perception tasks with an objective ground truth — a graded image, a dated scan — which is exactly where machine vision has long been strong, and a genuinely useful result. It is also the opposite of the open-ended clinical reasoning the LLM headlines claim: where the task is a whole consultation, the “win” keeps dissolving back into an unfair or absent comparison. And on malaria microscopy, graded on the very scale experts are certified against, the AI simply lost.
The number nobody reports
Across all fifteen, the closest anyone came to the thing that matters was the Penda Health trial’s eight-day “feeling better?” call — and it was non-significant. Push a little further and a separate field trial in Nigeria that deployed LLM support in real clinics found little to no improvement in health workers’ decisions, judged by the on-site physicians who treated the same patients. Every other row stops at a graded answer, a paper diagnosis, or a screen’s sensitivity. “AI beats clinicians” has been measured on almost everything except the thing that matters.
How this was built
This is an original synthesis: no prior source places these fifteen studies on one honest table and flags which “wins” have no comparable human number. The risk here is not a mistyped figure but a coding error — a study filed in the wrong group, a missing human arm quietly treated as a low one. So every row was built twice: two different frontier models (Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5, via the Codex CLI) coded each study independently from its primary source, and every disagreement was reconciled back to the source, never split.
Fifteen studies · two independent codings each · the second model raised seven corrections across two rounds · every one held · each reconciled to its primary source.
The corrections are recorded here rather than hidden, because on this publication the reconciliation is part of the evidence. Among them: AMIE’s axis count (the preprint says 28 of 32; the peer-reviewed Nature paper we cite says 30 of 32); that HealthBench does keep a physician-written baseline, which a first pass had called absent; the cervical study’s specificity figures, which are in fact reported (AVE 81.9%, VIA 94.2%); that the oral-cancer AI’s accuracy is on a 1,416-image test set, not the full screened cohort; and that a Nigeria field trial we had leaned on measures clinicians’ decisions, not patient outcomes. Corrections to any cell will continue to be made in public.
What to take from the table
None of this says the systems are bad or the studies dishonest — most are careful and state their own limits. It says that “AI beats clinicians,” as a genre of headline, is at least five different claims measured on a dozen scales, and it has yet to show a patient is better off. Before repeating one of these numbers, find its row. The portable version of the questions is the field guide.
Cite this analysis
Ground Truth (2026). The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured. Ground Truth. https://groundtruth.health/ai-vs-clinician-scoreboard/
The underlying scoreboard is released as an open dataset under CC BY 4.0 — reuse it with attribution. The primary source for every cell is listed below.
Sources
- Rwanda / LLMs vs GPs — “Large language models for frontline healthcare support in low-resource settings,” Nature Health (2026); six physician raters, 11-metric rubric, 506 question–response pairs, four districts; top model +0.83 pt over local GPs (range 0.38–1.10); funded by the Gates Foundation (INV-068056). pmc.ncbi.nlm.nih.gov/articles/PMC12880909
- Microsoft MAI-DxO / SDBench — Nori et al., “Sequential Diagnosis with Language Models,” arXiv:2506.22405 (2026); physicians ≈ 20% on the 56-case held-out split (references barred), o3 78.6% / MAI-DxO+o3 80% / max ensemble 85.5% (≈ US$7,184/case) on the full 304-case benchmark. arxiv.org/abs/2506.22405
- Google AMIE — Tu et al., “Towards Conversational Diagnostic AI,” Nature (2025); randomised text-chat OSCE (159 scenarios, 20 PCPs) with trained patient-actors; AMIE rated superior on 30/32 specialist axes and 25/26 patient-actor axes; interface “not representative of usual clinical practice.” Published figures: nature.com/articles/s41586-025-08866-7 (preprint arXiv:2401.05654 reports the earlier 28/32, 24/26).
- OpenAI HealthBench — arXiv:2505.08775 (2025); best model (o3) scored 0.60 on a weighted, normalized rubric over 48,562 physician-written criteria (some worth negative points), marked by a GPT-4.1 judge that agrees with physicians at macro F1 ≈ 0.71; a physician-written response baseline exists. arxiv.org/abs/2505.08775
- OpenAI–Penda Health — arXiv:2507.16947 (2025); clinician-randomized quality-improvement study; 16% / 13% relative reductions in physician-rated diagnostic / treatment documentation errors; 8-day patient-reported outcome non-significant (3.8% vs 4.3%, ~60% unreachable, “exploratory”). See our full investigation. arxiv.org/abs/2507.16947
- Cervical AI (AVE) — five-country prospective study (Malawi, Rwanda, Senegal, Zambia, Zimbabwe), Lancet Global Health; sensitivity for CIN2+ AVE 60.1% / VIA 36.6% / HPV 90.4%, specificity AVE 81.9% / VIA 94.2% / HPV 80.1%, among 18,086 women with confirmed status (526 CIN2+). pubmed.ncbi.nlm.nih.gov/41386246
- TB computer-aided detection — external validation of 12 commercial CAD products on a South-African cohort with no radiologist arm; ≈ 64% sensitivity at a borrowed threshold; qXR specificity ≈ 87% → 37% in previously-treated patients at a fixed threshold. pmc.ncbi.nlm.nih.gov/articles/PMC11339183
- Rheumatic-heart-disease echo (Uganda) — J Am Heart Assoc 2024;13(2):e031257; AI RHD-detection AUC 0.84 (mitral-regurgitation model 0.93), 511 pediatric echoes from the GOAL trial; expert cardiologists provided the reference labels, not a separately-scored comparator arm. pubmed.ncbi.nlm.nih.gov/38226515
- Oral-cancer smartphone AI (India) — Birur et al., Sci Rep 2022; AI 82–87% sensitivity (1,416-image test set) vs frontline health workers 60% (both against remote-specialist telediagnosis, 4,728 screened); onsite specialists 94% but against histology on n=102. pmc.ncbi.nlm.nih.gov/articles/PMC9395355
- ENT / otoscopy LLM (rural Kenya) — Lechien, Otolaryngol Head Neck Surg 2025;173(4):1024-1027; a Claude-3.5 model on 63 real patients: 79.4% concordance with the on-site otolaryngologist, and 79.4% vs 50.8% against primary-care practitioners (P=.001). Single site, single author. pubmed.ncbi.nlm.nih.gov/40525718
- Diabetic-retinopathy screening (Thailand) — Lancet Digital Health 2022; in the national programme, DL sensitivity for vision-threatening DR 91.4% vs retina-specialist over-readers 84.8% (p=0.024) at matched specificity (~95%), ~7,651 patients across nine sites. pubmed.ncbi.nlm.nih.gov/35272972
- AI ultrasound gestational age (Zambia + US) — JAMA 2024;332(8):649-657; blind-sweep AI acquired by novices MAE 3.2 days vs credentialed sonographers 3.0 days (difference 0.2 d, 95% CI −0.1 to 0.5; ±2-day equivalence), 399 pregnancies, 14–27 weeks. pubmed.ncbi.nlm.nih.gov/39088200
- Malaria microscopy AI (EasyScan GO) — Malaria Journal 2022; 11-country field evaluation, 2,250 slides vs expert microscopy; reached WHO-TDR competence Level 2 (below expert Level 1): detection sensitivity 89% / species ID 84% (Expert requires >90%), parasite density within ±25% of reference in only 23% of slides. The AI did not match experts. pmc.ncbi.nlm.nih.gov/articles/PMC9004086
- AI-assisted physicians RCT (Pakistan) — Nature Health 2026; single-blind RCT, 58 physicians; those with GPT-4o access scored 71.4% vs 42.6% with conventional resources (adjusted +27.5 pts, 95% CI 22.8–32.2) on simulated vignettes, after a 20-hour AI-literacy course. nature.com/articles/s44360-025-00007-8
- Locally-authored vignettes (Kenya) — preprint, medRxiv 2025.10.25.25338798; a blinded panel of six Kenyan family physicians rated five LLMs vs local clinicians on 507 local vignettes (11-domain rubric); LLMs rated safer (harm scores 4.3–4.8 vs clinicians’ 3.2) at ~$0.01 vs $3.35 per vignette. Not yet peer-reviewed. medrxiv.org/content/10.1101/2025.10.25.25338798v1
- Nigeria field trial (referenced above; not one of the fifteen) — Abaluck, Pless, Ravi, Sautmann & Schwartz, “Does LLM Assistance Improve Healthcare Delivery? An Evaluation Using On-site Physicians and Laboratory Tests,” medRxiv 2025.10.31.25339278; LLM decision support for health workers at two outpatient clinics in Nigeria. Health workers changed prescribing for more than half of patients, and blinded retrospective academic reviewers rated the assisted plans more favorably — but the on-site physicians who evaluated and treated the same patients “observed little to no improvement in diagnostic alignment or treatment decisions,” and laboratory testing showed mixed effects with no significant increase in detection rates. Not peer-reviewed. medrxiv.org/content/10.1101/2025.10.31.25339278v2 (also NBER Working Paper 34660).
Disclosures & provenance
- Published
- 5 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Corrections
- None to date. Corrections log → · Challenge this analysis