Reading a diagnostic-accuracy claim

The ten questions that turn “99% accurate” into something you can actually judge. Four interactive figures, a worked example from a real NHS claim, and a copy-paste evaluation prompt live in the article: How to read a “99% accurate” diagnostic-AI claim →

  1. Ninety-nine percent of what? Replace the word "accurate" with the metric's exact technical name — sensitivity, specificity, positive or negative predictive value, overall classification accuracy, AUC, agreement, or positive/negative percent agreement (PPA/NPA, which resembles sensitivity and specificity but concedes there is no reference standard). If the source never names it, say so: that is the finding.
  2. Out of how many? Give the numerator and denominator behind the headline percentage, quoting the source. If the source reports a confidence interval, quote it. If it does not, write "no interval reported" — and only then, if and only if you have the exact numerator and denominator, compute one, show the arithmetic and name the method. Never state an interval you cannot derive from numbers printed in the source. A percentage resting on a handful of cases is compatible with performance nobody would accept.
  3. Among whom, and how common was the disease? Predictive values move with prevalence. Quote the disease rate in the study population. Then say whether the source states where the tool is intended to be used and at what disease rate there — if it does not, write "not stated" and do not supply a figure from your own knowledge. If you have relevant background knowledge, put it on a separate line marked "outside the source, unverified", give a range rather than a point, and do not use it in any calculation. Finally, give the negative predictive value that clearing people at random would have scored in the study population (1 − prevalence), and set the reported NPV against it.
  4. At which threshold, and who chose it? Was the operating point fixed before the analysis or picked after? Check whether any headline proportion — "rules out one in five" — is simply the threshold restated.
  5. What happens to the people it gets wrong? Take each direction separately, and judge it against what would actually be done with the result. Note anything the replaced test was also catching incidentally.
  6. How was "truth" decided? Biopsy, pathology, imaging, a coding record, or just no diagnosis recorded within a window? A model cannot be more reliable than the labels it was graded against.
  7. Who was left out? Compare the number enrolled with the number analysed, and check which of the two is the denominator of the headline. Separately, ask how many samples or images returned no usable result at all — invalid, insufficient, indeterminate, QC failure — and whether those people appear in that denominator.
  8. Compared with what? Current care, an existing marker, clinician judgement — or nothing? Beware AUC standing in for performance at the threshold clinicians will actually use, and check whether any head-to-head comparison is in the paper at all.
  9. Was it measured or modelled? Distinguish internal validation, external validation, prospective silent validation, live implementation and impact evaluation. For any projected benefit — procedures saved, lives saved, money saved — separate the steps that were measured from the steps that were assumed. Then answer plainly: did the study show the tool changing a single real clinical decision, yes or no?
  10. Who ran it, who funded it, and has anyone unconflicted repeated it? Treat disclosed conflicts as a signal about what verification is still needed, not as proof of bad faith. Check what "regulated" or "CE/UKCA marked" actually means for this device class.

Run these on a claim →

Reading an “AI beats doctors” claim

The seven questions that separate what an AI-versus-clinician study scored from what it claimed. Full worked examples — built around the Rwanda answer-quality study, with MAI-DxO, AMIE, HealthBench and the imaging screens — a rubric-score chart, primary sources, and a copy-paste prompt live in the article: How to read an “AI beats doctors” claim →

  1. Did the human get to practice medicine — or just answer a paragraph? Real medicine is information-gathering. If the clinician couldn't take a history, examine, or order a test, the written case handed the model the workup for free.
  2. What did the score actually grade — answer quality, or a patient outcome? Completeness, empathy, and fluent prose are what models are optimized to produce. A right diagnosis, the right treatment, a patient who got better — that's a different, harder measurement.
  3. Were the conditions matched? Same inputs, language, format, time, and grading rubric for human and AI? Or did the clinicians work in an unfamiliar interface, a second language, or through a translation step?
  4. A clean vignette, or an undifferentiated patient? A curated case with a known answer, or a real queue of mislabeled, comorbid, mid-workup people? Benchmarks are pre-cleaned; clinics are not.
  5. Does it beat the doctor, or help the doctor? Model-alone versus the same clinician with the tool are different studies with opposite deployment implications. Ask which was measured.
  6. Does an off-the-shelf model already score this high? If a plain model without the special scaffold nearly matches the ‘specialized’ system — and the test cases could sit in its training data — the reported gain may be contamination, not capability.
  7. Realistic prevalence — and what does a false positive cost? On a balanced benchmark, high accuracy is cheap. At true clinic prevalence, ask the false-positive rate and who pays for it: the confirmatory test, the referral, the anxiety, the alert fatigue.

Run these on a claim →

Reading a health-chatbot or digital-health impact claim

The eight questions that separate a real health effect from a flattering number. Full worked examples — built around the PROMPTS/Jacaranda randomized trial — sources, effect-size charts, and a copy-paste prompt live in the article: How to read a health-chatbot impact claim →

  1. Is this a monitoring number or an impact number? Reach and engagement are monitoring; impact needs an outcome measured against a counterfactual.
  2. Significant — but how big, in units I can picture? Ask for the effect size and its confidence interval.
  3. Is this a measured health outcome, or a proxy for one? Separate self-report from a measured outcome; ask if the study was powered for it.
  4. Has this been tested against a control — and what happened when it was? Ask for the randomized evidence.
  5. What does the retention curve look like — not the install count? Reach is a moment; use is a curve.
  6. Is this a randomized comparison, or a correlation among the already-engaged? Correlation is selection, not effect.
  7. Validated on what — and what happens when it's wrong? A benchmark is not a bedside; ask where the human fallback sits.
  8. What's the cost per outcome — not per user? With the effect size in the denominator and the full economic cost in the numerator.

Run these on a claim →

Reading a language or benchmark claim

The nine questions that separate a real capability from a headline number. Full worked examples, a per-language error explorer, primary sources, and a copy-paste evaluation prompt live in the article: How to read an African-language AI benchmark without getting fooled →

  1. On what was it tested? Read or spontaneous speech, clean or noisy audio, whose accents, which subject matter.
  2. What does the metric count as a mistake — and what does it ignore? For morphologically rich languages, ask for character error rate beside word error rate.
  3. How was "right" decided, and by whom? A human native speaker, or an AI judge — and against an answer in which language?
  4. Over how many items — and would the gap survive noise? A few dozen questions can't separate two close models.
  5. Which languages, exactly — and on what base model? A leaderboard entry is not a shipped, working model.
  6. Is this shipping or a preview — and is my language in the tested set? A language count is not a coverage guarantee.
  7. Who produced the benchmark, and are they a player in it? A self-graded result needs an independent second opinion.
  8. Is this a continental average hiding wide variance? Performance tracks transcribed-data volume, not difficulty.
  9. Is a low score a verdict, or a specification? Which layer is failing — real misunderstanding, or an accent the recognizer mis-transcribes?

Run these on a claim →

How we rate a claim

The questions above are the tools; a rating is the record. Selected claims get a structured, versioned rating: the claim quoted verbatim, seven dimensions scored against the primary sources, every hidden condition marked in red, and a composite band. Every rating lives in the registry. Two rules govern all of them. A rating measures the credibility of the claim as stated — never the performance of the system, and never a ranking of systems against each other. And no rating ships without its primary sources linked: no source, no rating. Ratings are never for sale; no rated entity pays us, sponsors us, or sees a rating before it publishes.

The composite bands

These five labels are the only rating vocabulary used anywhere on the site: every machine-readable rating — a standalone claim marked inside a field guide, a verdict on an investigation, or a rated subject in the registry — sits on this same 1–5 scale with these same names.

The seven dimensions

  1. Fair test. Was the comparison designed so the claim could lose — randomized, controlled, blinded where it matters?
  2. What was measured. The thing in the claim, or a proxy for it — answers instead of outcomes, notes instead of patients?
  3. Real-world validation. A deployed setting or an external dataset — or a curated lab test?
  4. Effect size. Big enough to matter clinically, not just statistically?
  5. Generalizability. Whom and where does the result actually cover, versus whom the claim implies?
  6. Independence & conflicts. Who ran, funded, and graded the study?
  7. Traceability. Is the primary source public, complete, and linked?

Each dimension gets a mark — holds, conditional, or does not hold — a one-line rationale, and, where one exists, the hidden condition in red. The composite band follows the weakest load-bearing dimension: a claim that fails on a rigged comparison cannot rate well merely because its other dimensions look good. Nothing is averaged.

Rubric versions

The rubric is versioned like any methodology. Every rating names the version it was made under, and a rating can be revised — with a dated public note and a version bump, under the same corrections policy as everything else we publish.

The question underneath all of them

Is this number measuring the world the tool will be used in — the clean lab, or the messy point of care? Every rubric here is a way of asking that one question in a form you can check. A score is a property of a measurement, never of a model; keeping those two apart is the whole job, and it is one anyone can do.