What Ground Truth does

We examine the benchmarks, accuracy scores, and capability announcements that shape what gets funded, deployed, and believed in health — a frontier model’s medical benchmark, a clinical copilot’s error rate, a screening algorithm’s sensitivity. Ground Truth began, and stays rooted, in global health, where evidence is hardest to verify and a wrong number costs the most; the method proven there applies unchanged to any health-AI claim. For each claim, we find the primary source, read what it actually measured, and mark the part that was left out. The goal is not to be against AI. It is to keep two things apart: what a system has been shown to do, and what it is being sold as able to do.

The method

Every claim traced to a primary source

We do not report a number without linking to the material that produced it — the paper, the model card, the registry entry. If a claim can only be traced to a press release, we say so, and we treat it as unverified. Our readers should never have to take our word for a fact they could check themselves.

Self-contained, plainly written

Each piece leads with the answer, then shows the working. We write so that a single passage still makes sense when it is lifted out and quoted — by a reader, a journalist, or an answer engine — without distorting what we meant.

The claim, un-redacted

Our recurring analytical unit takes a claim, states what you are meant to conclude, marks what is hidden in red, and gives you the question to ask. It is the same move every time, because the move is the point: hype is a true sentence with the conditions cut out, and our job is to put the conditions back.

Ratings

Selected claims get more than an article: a structured rating — the claim quoted verbatim, seven dimensions scored against the primary sources, every hidden condition marked in red, and a composite band on a five-step scale from Holds up to Misleading. A rating measures the credibility of the claim as stated, never the performance of the system; we do not rank systems against each other. The rubric is public and versioned — how we rate — and every rating lives in the registry. No rating ships without its primary sources linked: no source, no rating.

What we review — and how to answer back

Everything we examine is already public: the paper, the preprint, the model card, the registry entry, the press release, the marketing page, the executive's post. We do not publish confidential documents, leaked material, or private communications, and we do not report on what was said behind a closed door. If a fact cannot be cited to something a reader can open, it does not appear. Where we searched for a source and came up empty, we write that we could not find it — not that it does not exist.

We rate the claim, not the company. A band assesses a published sentence against the evidence published beneath it; it is never a finding about anyone's motives, competence, or honesty. Most claims we rate are literally true. What makes them overstated is the conditions that fell away between the study and the headline, and our job is to put those conditions back and show the arithmetic while we do it.

Because of that, we publish without seeking advance comment. There is no undisclosed allegation here for a subject to answer: the claim is quoted verbatim, the evidence is linked, the method is written down, and any reader — including the party being examined — can check every step without our help. Pre-publication comment is how you handle private conduct that only the subject can confirm. This is the opposite of that: a public sentence measured against a public source, in the open, with the working shown.

The right of reply is permanent and open to anyone. If you are the subject of an article or a rating and believe we got something wrong, write to corrections@groundtruth.health. Challenges are read by a human and answered against the primary source, not filed. If we are wrong, we fix it on the page with a dated note and move the band if the band was wrong — publicly, in the corrections log and the registry. If we are right, we will show you why in the same detail we would show a reader. Evidence we did not have is the fastest way to move a rating, and a rating can move up as readily as down.

When we run our own tests, we publish the working: what we ran, how many times, the prompts, and the raw outputs, so a result can be reproduced or falsified by someone who doubts it. We test only what is publicly available to test — open weights, public APIs, published datasets — and we do not circumvent access controls, paywalls, or terms of use to obtain a number. Where a system cannot be independently tested at all, that is itself a finding, and we report it as one rather than guess.

Independence

Ground Truth takes no funding from, and holds no affiliation with, the companies or funders whose work it examines. We accept no sponsorship, advertising, or paid placement from AI vendors, model developers, or the philanthropies that finance them. If that ever changes, it will be disclosed on this page before it appears anywhere else. Independence is not a marketing line here; it is the reason the analysis can be trusted.

The same wall holds for ratings: they are never for sale. No rated entity pays us, sponsors us, or sees a rating before it publishes, and nothing a company could offer changes a band. If we are ever paid for private analysis, it will be disclosed here first — and it will buy no influence over the public record.

Corrections policy

We correct our own errors in public. When we get something wrong, we fix it on the page, add a dated correction note explaining what changed, and update the article's modified date. We do not quietly edit and move on. A correction is not an embarrassment to hide; it is the standard working as intended, and a signal that the record can be trusted precisely because it is kept honest.

Ratings are held to the same rule. A rating can be revised — up or down — when the evidence changes or we got it wrong: the revision is dated on the rating card and marked in red in the registry, and if the rubric itself changed, the version number says so. A rating with a correction on it is not a weaker rating; it is the policy working.

Submit a claim, tip, or correction

If you have found a claim worth scrutinizing, spotted an error in our work, or can point us to a primary source we missed, tell us. We read everything and correct what needs correcting.

The most useful submissions include: the exact claim, word for word; where it appeared (a link or citation); the primary source, if you know it; why it matters; and whether you want to be credited if we take it up. A one-line tip is welcome too.