Investigations
One claim at a time, traced to source.
Each investigation takes a single, load-bearing claim about AI in health — a vendor benchmark, a funder statistic, a deployment announcement — and follows it back to the primary evidence, ending in a calibrated verdict you can check. When a number is accurate but incomplete, we say so; we reserve the word “false” for what is actually false.
-
The NHS cancer blood test was “99% accurate.” Which 99%?
A blood test that scores cancer risk from routine bloods was reported in July 2026 as 99% accurate at both detecting and ruling out gynaecological cancer, and as able to spare 18,000 women a year a transvaginal ultrasound. We rebuilt the study's confusion matrix from the paper and its supplement. On the gynaecological pathway, 113 of 114 diagnosed cancers scored above the rule-out threshold and one fell below it — a real result — while 2,272 women without cancer were also flagged. The two numbers behind “99%” are a sensitivity of 99.1% and a negative predictive value of 99.8%, which describe different groups of women; specificity at that setting was 20.0%, and read as overall classification accuracy the figure is 23.0%. The “one in five ruled out” is close to the threshold the researchers chose, not a discovery. Clinicians were blinded throughout, so the study offers no evidence that any procedure was avoided because of a PinPoint result. The 18,000 figure, the ultrasound comparison and the AUC of 0.832 versus 0.792 in 578 women appear nowhere in the paper, and we could not trace them to any published study, dataset or statistical report. The piece sets out what the evidence supports, what it does not, and how a reader can tell the difference.
-
They want a doctor in every pocket. We handed the pocket a lethal order.
Google and the World Bank are promoting small, “offline-capable” medical AI for the roughly two-thirds of women in sub-Saharan Africa who can’t reach a clinician or a network. We built a 33-probe set of maternal-emergency drug orders on real WHO or national-guideline doses; in 19 of them a senior clinician “orders” a clearly lethal version and asks the model to confirm it, with the other 14 correct or borderline controls. All 20 models saw the same 19 dangerous orders, and every one of the eight frontier cloud models refused all 19 (0 of 152 across the eight). The on-device models confirmed the lethal order about 42% of the time — the phone-class ones more, up to 63% for MedGemma-4B and 74% for the smallest — mostly because they didn’t know the right dose. The failure isn’t being offline (a general 31B model run offline matched the frontier tier) or being quantized (refuted two ways) — it’s being pocket-sized. An original, fully reproducible benchmark of 20 models, cross-lab graded, with the full probe set published.
-
AI beat the doctors. So we regraded the doctors.
OpenAI's GPT-5.6 was reported everywhere to have beaten physicians on health evaluations, 60.5 to 43.7. That comparison appears in no OpenAI publication: the 43.7 comes from an earlier paper about a different model, and the GPT-5.6 system card contains no physician baseline at all. We re-scored the 525 physician answers OpenAI released using the benchmark's own rubrics and scoring rule, changing only the grader. Four independent frontier models — including GPT-5.6 Sol itself — all placed the physicians above 43.7, between 47.1 and 51.9, roughly halving the headline gap. A second test found a cheap model following a fixed, rubric-blind template captures 85–88% of physician performance on the benchmark's routine half.
-
Google made Gemma “medical.” We gave it the job.
Google ships MedGemma as its “most capable open models for health AI development.” Run on the work a clinical decision-support tool actually does — recognising an emergency, building a differential, extracting structured data, writing a note — the medical fine-tune genuinely beats the general Gemma it was built from on several tasks, but is matched or beaten by a newer general Gemma on nearly every one, and does not reliably recognise an emergency. An original, fully reproducible benchmark of eight open models under one harness with identical prompts.
-
“As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved
A prospective 2024 study found that novices using blind-sweep AI estimated gestational age about as accurately as credentialed sonographers performing fetal biometry before term. The result holds — but the operators received one day of task-specific training, the pregnancies were selected, and the studies did not test a complete ultrasound, routine deployment, clinical decisions, or outcomes. Traced to the primary source.
-
“On par with nurses”: what Hippocratic AI’s headline number actually measured
Hippocratic AI — the $3.5B “AI nurse” company — says its Polaris system is “on par with human nurses.” On aggregate, on its own subjective surveys, it is. But the test used clinicians role-playing patients in ~3,475 simulated calls, a 60-nurse human baseline the company’s physicians never graded, and no real outcomes; the nurses won two of the four experience dimensions, and on the safety row the only advice rated as risking severe harm came from the AI. Traced to the primary source.
-
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
Microsoft says its MAI-DxO orchestrator diagnoses NEJM case challenges at more than four times the rate of experienced physicians. The 85.5% is real — on Microsoft’s own benchmark. What it beat: 21 generalists stripped of every reference tool, scored on a 56-case slice while the AI was scored on all 304, graded by the model family being tested. A year on: no peer review, no released benchmark, no trial — and a bigger claim. Traced to the primary source.
-
“16% fewer errors”: what the OpenAI–Penda Health study actually measured
OpenAI and Penda Health report that clinicians using an AI copilot in Nairobi made 16% fewer diagnostic errors and 13% fewer treatment errors. The reductions are real, significant and independently physician-rated — but they score the quality of documented decisions, not patient harm, and the one patient-outcome measure came back non-significant. Traced to the primary source.