Life

FDA Cleared 1,357 AI Medical Devices, and Only Three Were Outcome-Tested

A systematic review published in PLOS Digital Health finds that of 1,357 AI and machine-learning medical devices cleared or approved by the FDA through December 2025, only 34 — 2.5 percent — have ever been linked to a registered clinical trial. Just 12 posted results, 12 reached peer-reviewed publication, and only three were ever evaluated for outcomes that actually matter to a patient: death, major complications, or hospital readmission. [1] That is 0.2 percent of every device now cleared to guide a clinician's decision in a hospital.

The gap traces back to how these devices reach the market in the first place. The overwhelming majority, the study's authors write, are cleared through the FDA's 510(k) pathway, which requires only that a device demonstrate "substantial equivalence" to something already on the market — not prospective evidence that it improves patient outcomes. [1] Radiology accounts for 78 percent of all cleared devices, and fewer than 1 percent of those have a registered prospective trial; cardiovascular and neurology devices do modestly better, at roughly 9.5 and 9.7 percent respectively, while anesthesiology has 22 cleared devices and not a single registered trial among them. [1]

Even the small number of trials that do exist are not built to detect much. Nearly three-quarters enrolled fewer than 500 participants, and a quarter enrolled fewer than 100 — sample sizes the authors say are too small to reliably power subgroup analysis. Sixty-eight percent were conducted entirely within the United States, limiting what they can say about how a device performs elsewhere. Only nine of the 34 trials — 27 percent — reported any subgroup analysis at all; sex was examined in five, age in four, race or ethnicity in three, and language in none. Pregnant patients were excluded from 42 percent of cardiovascular trials and a third of radiology trials, and pediatric patients were almost universally excluded across specialties. Ninety-four percent of the 34 trials that exist were industry-led. [1]

The study argues the consequences are not hypothetical. It cites the Epic Sepsis Model, deployed across more than half of U.S. hospitals, which in independent testing failed to identify 67 percent of patients who actually had sepsis while generating alerts on 18 percent of all hospitalized patients — a model that reached national scale without the kind of outcome testing this study finds is nearly universal across the category. It also cites IBM Watson for Oncology, cleared in the U.S. and deployed internationally, which achieved only 33 percent concordance with expert treatment recommendations in Denmark and just 12 percent for gastric cancer cases in China. [1] A separate analysis the authors cite as a benchmark found that 6 percent of a sample of 950 FDA-cleared AI devices were associated with recall events, nearly half of those within the first year of clearance. [1]

Cardiologist and researcher Eric Topol amplified the study's central number directly on X, writing: "Out of >1,350 medical AI's cleared or arrived by FDA, a total of 3 (0.2%) were tested for patient outcomes only 34 (2.5%) linked to prospective, registered clinical trials." The post drew replies from clinicians framing the finding not as a surprise but as confirmation of what practicing physicians already suspected — that the 1,357-device figure, when it circulates as a marker of how far medical AI has advanced, obscures how little of that adoption has been checked against whether it actually helps anyone.

The study's authors are not arguing AI devices should be pulled from use, and they stop short of claiming the devices are unsafe — the FDA maintains non-public post-market monitoring through manufacturer reporting and its MAUDE adverse-event database that this analysis could not access. Their argument is narrower and, in its way, more damning: a system that cleared 1,357 devices in roughly a decade has never built an evidentiary requirement for outcomes to keep pace with the speed of that clearance in the first place. They propose a three-phase framework modeled on drug development — retrospective validation before clearance, mid-size prospective studies at clearance, and large multi-center outcome trials afterward — precisely because none of those stages currently exist as a requirement. Until one does, the number that matters is not 1,357. It is three.

-- NORA WHITFIELD, Chicago

Get the New Grok Times in your inbox

A weekly digest of the stories shaping the timeline — delivered every edition.

No spam. Unsubscribe anytime.