What is Ground Truth?
Before you believe an AI claim in health, check it here.
Every hype claim is a true sentence with the conditions cut out. We put the cut part back, trace it to the primary source, and mark the hidden part in red.
Read the latest investigation → Browse the ratings →
For people deciding what health-AI evidence deserves to be reported, funded, regulated, or deployed.
A health-AI headline, as published
The cut part, put back
Latest
Featured analysis
-
How to read a “99% accurate” diagnostic-AI claim
A test can be 99% sensitive, 99% specific, have a 99% negative predictive value, or correctly classify 99% of a selected group. Those are four different numbers about four different sets of people, and none of them alone tells you whether the test improves anyone's care. Ten questions that take any diagnostic-AI accuracy claim apart — what the metric is actually called, out of how many, among whom, at which threshold, against what standard of truth, and what happens to the people in the remaining one percent — with four interactive figures: a map of which cells of the 2x2 each metric reads, the same 99% shown at four different case counts, a draggable threshold linked live to its point on the ROC curve, and a prevalence calculator. Ends with a copyable prompt that turns the ten questions into an audit you can run on any claim.
-
The NHS cancer blood test was “99% accurate.” Which 99%?
A blood test that scores cancer risk from routine bloods was reported in July 2026 as 99% accurate at both detecting and ruling out gynaecological cancer, and as able to spare 18,000 women a year a transvaginal ultrasound.
-
They want a doctor in every pocket. We handed the pocket a lethal order.
Google and the World Bank are promoting small, “offline-capable” medical AI for the roughly two-thirds of women in sub-Saharan Africa who can’t reach a clinician or a network.
-
AI beat the doctors. So we regraded the doctors.
OpenAI's GPT-5.6 was reported everywhere to have beaten physicians on health evaluations, 60.5 to 43.7.
-
What is “ground truth” in health AI?
-
Google made Gemma “medical.” We gave it the job.
-
“As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved
-
“On par with nurses”: what Hippocratic AI’s headline number actually measured
-
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
-
The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured
-
How to read an “AI beats doctors” claim
-
“16% fewer errors”: what the OpenAI–Penda Health study actually measured
-
How to read a health-chatbot impact claim
-
How to read an African-language AI benchmark without getting fooled
The editorial standard
Credibility you can inspect.
Traced to source
Every claim we examine is linked to the primary material — the paper, model card, or registry — so you can check it yourself.
Independent by design
We take no money from, and hold no affiliation with, the companies or funders whose claims we scrutinize.
Corrected in public
When we get something wrong, we say so on the page, with the date and what changed. Corrections are a feature, not an embarrassment.
The newsletter
Get the hidden conditions, not the hype.
Every few weeks: one big health-AI claim, traced to its primary source, with the part the headline left out put back in red. That’s the whole email — one-click unsubscribe, no tracking pixels.