The archive
Every piece, traceable to source.
All Ground Truth analysis, newest first. Each entry links to a real article — and each article links out to the primary material behind every claim.
-
How to read a “99% accurate” diagnostic-AI claim
A test can be 99% sensitive, 99% specific, have a 99% negative predictive value, or correctly classify 99% of a selected group. Those are four different numbers about four different sets of people, and none of them alone tells you whether the test improves anyone's care. Ten questions that take any diagnostic-AI accuracy claim apart — what the metric is actually called, out of how many, among whom, at which threshold, against what standard of truth, and what happens to the people in the remaining one percent — with four interactive figures: a map of which cells of the 2x2 each metric reads, the same 99% shown at four different case counts, a draggable threshold linked live to its point on the ROC curve, and a prevalence calculator. Ends with a copyable prompt that turns the ten questions into an audit you can run on any claim.
-
The NHS cancer blood test was “99% accurate.” Which 99%?
A blood test that scores cancer risk from routine bloods was reported in July 2026 as 99% accurate at both detecting and ruling out gynaecological cancer, and as able to spare 18,000 women a year a transvaginal ultrasound. We rebuilt the study's confusion matrix from the paper and its supplement. On the gynaecological pathway, 113 of 114 diagnosed cancers scored above the rule-out threshold and one fell below it — a real result — while 2,272 women without cancer were also flagged. The two numbers behind “99%” are a sensitivity of 99.1% and a negative predictive value of 99.8%, which describe different groups of women; specificity at that setting was 20.0%, and read as overall classification accuracy the figure is 23.0%. The “one in five ruled out” is close to the threshold the researchers chose, not a discovery. Clinicians were blinded throughout, so the study offers no evidence that any procedure was avoided because of a PinPoint result. The 18,000 figure, the ultrasound comparison and the AUC of 0.832 versus 0.792 in 578 women appear nowhere in the paper, and we could not trace them to any published study, dataset or statistical report. The piece sets out what the evidence supports, what it does not, and how a reader can tell the difference.
-
They want a doctor in every pocket. We handed the pocket a lethal order.
Google and the World Bank are promoting small, “offline-capable” medical AI for the roughly two-thirds of women in sub-Saharan Africa who can’t reach a clinician or a network. We built a 33-probe set of maternal-emergency drug orders on real WHO or national-guideline doses; in 19 of them a senior clinician “orders” a clearly lethal version and asks the model to confirm it, with the other 14 correct or borderline controls. All 20 models saw the same 19 dangerous orders, and every one of the eight frontier cloud models refused all 19 (0 of 152 across the eight). The on-device models confirmed the lethal order about 42% of the time — the phone-class ones more, up to 63% for MedGemma-4B and 74% for the smallest — mostly because they didn’t know the right dose. The failure isn’t being offline (a general 31B model run offline matched the frontier tier) or being quantized (refuted two ways) — it’s being pocket-sized. An original, fully reproducible benchmark of 20 models, cross-lab graded, with the full probe set published.
-
AI beat the doctors. So we regraded the doctors.
OpenAI's GPT-5.6 was reported everywhere to have beaten physicians on health evaluations, 60.5 to 43.7. That comparison appears in no OpenAI publication: the 43.7 comes from an earlier paper about a different model, and the GPT-5.6 system card contains no physician baseline at all. We re-scored the 525 physician answers OpenAI released using the benchmark's own rubrics and scoring rule, changing only the grader. Four independent frontier models — including GPT-5.6 Sol itself — all placed the physicians above 43.7, between 47.1 and 51.9, roughly halving the headline gap. A second test found a cheap model following a fixed, rubric-blind template captures 85–88% of physician performance on the benchmark's routine half.
-
What is “ground truth” in health AI?
“Ground truth” is the reference standard an AI’s output is checked against to decide whether it is right. In health AI it quietly decides what every accuracy score means — and it is where most “as accurate as a doctor” claims come apart. A plain-language explainer, with the questions to ask and the investigations that show it.
-
Google made Gemma “medical.” We gave it the job.
Google ships MedGemma as its “most capable open models for health AI development.” Run on the work a clinical decision-support tool actually does — recognising an emergency, building a differential, extracting structured data, writing a note — the medical fine-tune genuinely beats the general Gemma it was built from on several tasks, but is matched or beaten by a newer general Gemma on nearly every one, and does not reliably recognise an emergency. An original, fully reproducible benchmark of eight open models under one harness with identical prompts.
-
“As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved
A prospective 2024 study found that novices using blind-sweep AI estimated gestational age about as accurately as credentialed sonographers performing fetal biometry before term. The result holds — but the operators received one day of task-specific training, the pregnancies were selected, and the studies did not test a complete ultrasound, routine deployment, clinical decisions, or outcomes. Traced to the primary source.
-
“On par with nurses”: what Hippocratic AI’s headline number actually measured
Hippocratic AI — the $3.5B “AI nurse” company — says its Polaris system is “on par with human nurses.” On aggregate, on its own subjective surveys, it is. But the test used clinicians role-playing patients in ~3,475 simulated calls, a 60-nurse human baseline the company’s physicians never graded, and no real outcomes; the nurses won two of the four experience dimensions, and on the safety row the only advice rated as risking severe harm came from the AI. Traced to the primary source.
-
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
Microsoft says its MAI-DxO orchestrator diagnoses NEJM case challenges at more than four times the rate of experienced physicians. The 85.5% is real — on Microsoft’s own benchmark. What it beat: 21 generalists stripped of every reference tool, scored on a 56-case slice while the AI was scored on all 304, graded by the model family being tested. A year on: no peer review, no released benchmark, no trial — and a bigger claim. Traced to the primary source.
-
The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured
Fifteen studies say AI matches or beats clinicians in global health. Grouped honestly, “beats doctors” fractures into five different claims — no fair test, a win over the weakest human, a genuine tie with experts, an outright loss, and AI merely assisting a clinician — on a dozen scales that don’t compare, and almost never measured on a patient. An original, cross-checked dataset.
-
How to read an “AI beats doctors” claim
An “AI beats doctors” headline almost always rests on a graded answer, not a treated patient. Seven questions — worked through the 2026 Rwanda study that says language models outperform local clinicians, plus Microsoft’s MAI-DxO, Google’s AMIE, HealthBench, and the TB and cervical-cancer screens — that separate what an AI-versus-clinician study actually scored from what it claimed.
-
“16% fewer errors”: what the OpenAI–Penda Health study actually measured
OpenAI and Penda Health report that clinicians using an AI copilot in Nairobi made 16% fewer diagnostic errors and 13% fewer treatment errors. The reductions are real, significant and independently physician-rated — but they score the quality of documented decisions, not patient harm, and the one patient-outcome measure came back non-significant. Traced to the primary source.
-
How to read a health-chatbot impact claim
A reach or engagement number is a monitoring metric, not an impact metric; a statistically significant effect can still be tiny; and an engagement–outcome correlation is not causal evidence. Eight questions — with a pre-registered RCT, a retention-decay curve, and physician red-teaming — that separate a real health effect from a flattering number.
-
How to read an African-language AI benchmark without getting fooled
Can an AI really understand Kinyarwanda, Swahili or Hausa well enough to answer a patient's question? A field guide to the questions that reveal what each accuracy number is really measuring — with a real, sourced explorer of speech-recognition error across 19 African languages and 10 domains.