Skip to content
Ground Truth
Articles Ratings Method About Subscribe

Investigations

One claim at a time, traced to source.

Each investigation takes a single, load-bearing claim about AI in health — a vendor benchmark, a funder statistic, a deployment announcement — and follows it back to the primary evidence, ending in a calibrated verdict you can check. When a number is accurate but incomplete, we say so; we reserve the word “false” for what is actually false.

  • 25 Jul 2026 Investigation Diagnostic AI

    The NHS cancer blood test was “99% accurate.” Which 99%?

    A blood test that scores cancer risk from routine bloods was reported in July 2026 as 99% accurate at both detecting and ruling out gynaecological cancer, and as able to spare 18,000 women a year a transvaginal ultrasound. We rebuilt the study's confusion matrix from the paper and its supplement. On the gynaecological pathway, 113 of 114 diagnosed cancers scored above the rule-out threshold and one fell below it — a real result — while 2,272 women without cancer were also flagged. The two numbers behind “99%” are a sensitivity of 99.1% and a negative predictive value of 99.8%, which describe different groups of women; specificity at that setting was 20.0%, and read as overall classification accuracy the figure is 23.0%. The “one in five ruled out” is close to the threshold the researchers chose, not a discovery. Clinicians were blinded throughout, so the study offers no evidence that any procedure was avoided because of a PinPoint result. The 18,000 figure, the ultrasound comparison and the AUC of 0.832 versus 0.792 in 578 women appear nowhere in the paper, and we could not trace them to any published study, dataset or statistical report. The piece sets out what the evidence supports, what it does not, and how a reader can tell the difference.

  • 24 Jul 2026 Investigation Benchmarks

    They want a doctor in every pocket. We handed the pocket a lethal order.

    Google and the World Bank are promoting small, “offline-capable” medical AI for the roughly two-thirds of women in sub-Saharan Africa who can’t reach a clinician or a network. We built a 33-probe set of maternal-emergency drug orders on real WHO or national-guideline doses; in 19 of them a senior clinician “orders” a clearly lethal version and asks the model to confirm it, with the other 14 correct or borderline controls. All 20 models saw the same 19 dangerous orders, and every one of the eight frontier cloud models refused all 19 (0 of 152 across the eight). The on-device models confirmed the lethal order about 42% of the time — the phone-class ones more, up to 63% for MedGemma-4B and 74% for the smallest — mostly because they didn’t know the right dose. The failure isn’t being offline (a general 31B model run offline matched the frontier tier) or being quantized (refuted two ways) — it’s being pocket-sized. An original, fully reproducible benchmark of 20 models, cross-lab graded, with the full probe set published.

  • 23 Jul 2026 Investigation Benchmarks

    AI beat the doctors. So we regraded the doctors.

    OpenAI's GPT-5.6 was reported everywhere to have beaten physicians on health evaluations, 60.5 to 43.7. That comparison appears in no OpenAI publication: the 43.7 comes from an earlier paper about a different model, and the GPT-5.6 system card contains no physician baseline at all. We re-scored the 525 physician answers OpenAI released using the benchmark's own rubrics and scoring rule, changing only the grader. Four independent frontier models — including GPT-5.6 Sol itself — all placed the physicians above 43.7, between 47.1 and 51.9, roughly halving the headline gap. A second test found a cheap model following a fixed, rubric-blind template captures 85–88% of physician performance on the benchmark's routine half.

  • 17 Jul 2026 Investigation Benchmarks

    Google made Gemma “medical.” We gave it the job.

    Google ships MedGemma as its “most capable open models for health AI development.” Run on the work a clinical decision-support tool actually does — recognising an emergency, building a differential, extracting structured data, writing a note — the medical fine-tune genuinely beats the general Gemma it was built from on several tasks, but is matched or beaten by a newer general Gemma on nearly every one, and does not reliably recognise an emergency. An original, fully reproducible benchmark of eight open models under one harness with identical prompts.

  • 14 Jul 2026 Investigation Benchmarks

    “As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved

    A prospective 2024 study found that novices using blind-sweep AI estimated gestational age about as accurately as credentialed sonographers performing fetal biometry before term. The result holds — but the operators received one day of task-specific training, the pregnancies were selected, and the studies did not test a complete ultrasound, routine deployment, clinical decisions, or outcomes. Traced to the primary source.

  • 10 Jul 2026 Investigation Benchmarks

    “On par with nurses”: what Hippocratic AI’s headline number actually measured

    Hippocratic AI — the $3.5B “AI nurse” company — says its Polaris system is “on par with human nurses.” On aggregate, on its own subjective surveys, it is. But the test used clinicians role-playing patients in ~3,475 simulated calls, a 60-nurse human baseline the company’s physicians never graded, and no real outcomes; the nurses won two of the four experience dimensions, and on the safety row the only advice rated as risking severe harm came from the AI. Traced to the primary source.

  • 10 Jul 2026 Investigation Benchmarks

    “Medical superintelligence”: what Microsoft’s 85.5% actually beat

    Microsoft says its MAI-DxO orchestrator diagnoses NEJM case challenges at more than four times the rate of experienced physicians. The 85.5% is real — on Microsoft’s own benchmark. What it beat: 21 generalists stripped of every reference tool, scored on a 56-case slice while the AI was scored on all 304, graded by the model family being tested. A year on: no peer review, no released benchmark, no trial — and a bigger claim. Traced to the primary source.

  • 4 Jul 2026 Investigation Evidence

    “16% fewer errors”: what the OpenAI–Penda Health study actually measured

    OpenAI and Penda Health report that clinicians using an AI copilot in Nairobi made 16% fewer diagnostic errors and 13% fewer treatment errors. The reductions are real, significant and independently physician-rated — but they score the quality of documented decisions, not patient harm, and the one patient-outcome measure came back non-significant. Traced to the primary source.

The newsletter

Get the hidden conditions, not the hype.

Every few weeks: one big health-AI claim, traced to its primary source, with the part the headline left out put back in red. That’s the whole email — one-click unsubscribe, no tracking pixels.

What we store when you subscribe, and why →


Ground Truth

An independent publication scrutinizing AI claims in health. We trace every claim to its primary source and correct our own errors in public.

Read

  • All articles
  • Ratings registry
  • The method
  • RSS feed

Accountability

  • Editorial standard
  • What we review
  • Independence
  • Corrections log
  • Privacy
  • Submit a tip or correction

© 2026 Ground Truth