Skip to content
Ground Truth
Articles Ratings Method About Subscribe

The archive

Every piece, traceable to source.

All Ground Truth analysis, newest first. Each entry links to a real article — and each article links out to the primary material behind every claim.

  • 25 Jul 2026 Field guide Diagnostic AI

    How to read a “99% accurate” diagnostic-AI claim

    A test can be 99% sensitive, 99% specific, have a 99% negative predictive value, or correctly classify 99% of a selected group. Those are four different numbers about four different sets of people, and none of them alone tells you whether the test improves anyone's care. Ten questions that take any diagnostic-AI accuracy claim apart — what the metric is actually called, out of how many, among whom, at which threshold, against what standard of truth, and what happens to the people in the remaining one percent — with four interactive figures: a map of which cells of the 2x2 each metric reads, the same 99% shown at four different case counts, a draggable threshold linked live to its point on the ROC curve, and a prevalence calculator. Ends with a copyable prompt that turns the ten questions into an audit you can run on any claim.

  • 25 Jul 2026 Investigation Diagnostic AI

    The NHS cancer blood test was “99% accurate.” Which 99%?

    A blood test that scores cancer risk from routine bloods was reported in July 2026 as 99% accurate at both detecting and ruling out gynaecological cancer, and as able to spare 18,000 women a year a transvaginal ultrasound. We rebuilt the study's confusion matrix from the paper and its supplement. On the gynaecological pathway, 113 of 114 diagnosed cancers scored above the rule-out threshold and one fell below it — a real result — while 2,272 women without cancer were also flagged. The two numbers behind “99%” are a sensitivity of 99.1% and a negative predictive value of 99.8%, which describe different groups of women; specificity at that setting was 20.0%, and read as overall classification accuracy the figure is 23.0%. The “one in five ruled out” is close to the threshold the researchers chose, not a discovery. Clinicians were blinded throughout, so the study offers no evidence that any procedure was avoided because of a PinPoint result. The 18,000 figure, the ultrasound comparison and the AUC of 0.832 versus 0.792 in 578 women appear nowhere in the paper, and we could not trace them to any published study, dataset or statistical report. The piece sets out what the evidence supports, what it does not, and how a reader can tell the difference.

  • 24 Jul 2026 Investigation Benchmarks

    They want a doctor in every pocket. We handed the pocket a lethal order.

    Google and the World Bank are promoting small, “offline-capable” medical AI for the roughly two-thirds of women in sub-Saharan Africa who can’t reach a clinician or a network. We built a 33-probe set of maternal-emergency drug orders on real WHO or national-guideline doses; in 19 of them a senior clinician “orders” a clearly lethal version and asks the model to confirm it, with the other 14 correct or borderline controls. All 20 models saw the same 19 dangerous orders, and every one of the eight frontier cloud models refused all 19 (0 of 152 across the eight). The on-device models confirmed the lethal order about 42% of the time — the phone-class ones more, up to 63% for MedGemma-4B and 74% for the smallest — mostly because they didn’t know the right dose. The failure isn’t being offline (a general 31B model run offline matched the frontier tier) or being quantized (refuted two ways) — it’s being pocket-sized. An original, fully reproducible benchmark of 20 models, cross-lab graded, with the full probe set published.

  • 23 Jul 2026 Investigation Benchmarks

    AI beat the doctors. So we regraded the doctors.

    OpenAI's GPT-5.6 was reported everywhere to have beaten physicians on health evaluations, 60.5 to 43.7. That comparison appears in no OpenAI publication: the 43.7 comes from an earlier paper about a different model, and the GPT-5.6 system card contains no physician baseline at all. We re-scored the 525 physician answers OpenAI released using the benchmark's own rubrics and scoring rule, changing only the grader. Four independent frontier models — including GPT-5.6 Sol itself — all placed the physicians above 43.7, between 47.1 and 51.9, roughly halving the headline gap. A second test found a cheap model following a fixed, rubric-blind template captures 85–88% of physician performance on the benchmark's routine half.

  • 18 Jul 2026 Explainer Explainer

    What is “ground truth” in health AI?

    “Ground truth” is the reference standard an AI’s output is checked against to decide whether it is right. In health AI it quietly decides what every accuracy score means — and it is where most “as accurate as a doctor” claims come apart. A plain-language explainer, with the questions to ask and the investigations that show it.

  • 17 Jul 2026 Investigation Benchmarks

    Google made Gemma “medical.” We gave it the job.

    Google ships MedGemma as its “most capable open models for health AI development.” Run on the work a clinical decision-support tool actually does — recognising an emergency, building a differential, extracting structured data, writing a note — the medical fine-tune genuinely beats the general Gemma it was built from on several tasks, but is matched or beaten by a newer general Gemma on nearly every one, and does not reliably recognise an emergency. An original, fully reproducible benchmark of eight open models under one harness with identical prompts.

  • 14 Jul 2026 Investigation Benchmarks

    “As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved

    A prospective 2024 study found that novices using blind-sweep AI estimated gestational age about as accurately as credentialed sonographers performing fetal biometry before term. The result holds — but the operators received one day of task-specific training, the pregnancies were selected, and the studies did not test a complete ultrasound, routine deployment, clinical decisions, or outcomes. Traced to the primary source.

  • 10 Jul 2026 Investigation Benchmarks

    “On par with nurses”: what Hippocratic AI’s headline number actually measured

    Hippocratic AI — the $3.5B “AI nurse” company — says its Polaris system is “on par with human nurses.” On aggregate, on its own subjective surveys, it is. But the test used clinicians role-playing patients in ~3,475 simulated calls, a 60-nurse human baseline the company’s physicians never graded, and no real outcomes; the nurses won two of the four experience dimensions, and on the safety row the only advice rated as risking severe harm came from the AI. Traced to the primary source.

  • 10 Jul 2026 Investigation Benchmarks

    “Medical superintelligence”: what Microsoft’s 85.5% actually beat

    Microsoft says its MAI-DxO orchestrator diagnoses NEJM case challenges at more than four times the rate of experienced physicians. The 85.5% is real — on Microsoft’s own benchmark. What it beat: 21 generalists stripped of every reference tool, scored on a 56-case slice while the AI was scored on all 304, graded by the model family being tested. A year on: no peer review, no released benchmark, no trial — and a bigger claim. Traced to the primary source.

  • 5 Jul 2026 Analysis Benchmarks

    The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured

    Fifteen studies say AI matches or beats clinicians in global health. Grouped honestly, “beats doctors” fractures into five different claims — no fair test, a win over the weakest human, a genuine tie with experts, an outright loss, and AI merely assisting a clinician — on a dozen scales that don’t compare, and almost never measured on a patient. An original, cross-checked dataset.

  • 4 Jul 2026 Field guide Benchmarks

    How to read an “AI beats doctors” claim

    An “AI beats doctors” headline almost always rests on a graded answer, not a treated patient. Seven questions — worked through the 2026 Rwanda study that says language models outperform local clinicians, plus Microsoft’s MAI-DxO, Google’s AMIE, HealthBench, and the TB and cervical-cancer screens — that separate what an AI-versus-clinician study actually scored from what it claimed.

  • 4 Jul 2026 Investigation Evidence

    “16% fewer errors”: what the OpenAI–Penda Health study actually measured

    OpenAI and Penda Health report that clinicians using an AI copilot in Nairobi made 16% fewer diagnostic errors and 13% fewer treatment errors. The reductions are real, significant and independently physician-rated — but they score the quality of documented decisions, not patient harm, and the one patient-outcome measure came back non-significant. Traced to the primary source.

  • 3 Jul 2026 Field guide Evidence

    How to read a health-chatbot impact claim

    A reach or engagement number is a monitoring metric, not an impact metric; a statistically significant effect can still be tiny; and an engagement–outcome correlation is not causal evidence. Eight questions — with a pre-registered RCT, a retention-decay curve, and physician red-teaming — that separate a real health effect from a flattering number.

  • 2 Jul 2026 Field guide Benchmarks

    How to read an African-language AI benchmark without getting fooled

    Can an AI really understand Kinyarwanda, Swahili or Hausa well enough to answer a patient's question? A field guide to the questions that reveal what each accuracy number is really measuring — with a real, sourced explorer of speech-recognition error across 19 African languages and 10 domains.

The newsletter

Get the hidden conditions, not the hype.

Every few weeks: one big health-AI claim, traced to its primary source, with the part the headline left out put back in red. That’s the whole email — one-click unsubscribe, no tracking pixels.

What we store when you subscribe, and why →


Ground Truth

An independent publication scrutinizing AI claims in health. We trace every claim to its primary source and correct our own errors in public.

Read

  • All articles
  • Ratings registry
  • The method
  • RSS feed

Accountability

  • Editorial standard
  • What we review
  • Independence
  • Corrections log
  • Privacy
  • Submit a tip or correction

© 2026 Ground Truth