That number stopped me cold.
1,357 AI and machine learning devices cleared by the FDA. Over a thousand tools now deployed in hospitals, clinics, and radiology departments across the country. Cleared. Approved. Covered.
And only three of them - three - were ever tested against what patients actually care about: whether they lived, recovered, avoided a readmission, or got back to their lives.
That is not a gap in the evidence. That is a canyon.
A new study published in PLOS Digital Health (Abulibdeh et al., University of Toronto, August 2026) tore back the curtain on the evidence base for AI in medicine. The numbers are worse than I expected, and I work in this space every single day.
The Study Nobody Is Talking About
The researchers examined all 1,357 AI/ML-enabled medical devices cleared by the FDA through December 5, 2025. They cross-referenced those devices against ClinicalTrials.gov, PubMed, and the FDA AI-Enabled Device List (which now stands at 1,524 entries as of March 2026).
Here is what they found:
Only 34 devices (2.5%) were linked to registered prospective trials.
Only 12 (0.9%) posted results from those trials on ClinicalTrials.gov.
Only 12 (0.9%) progressed to peer-reviewed publication.
Only 3 devices (0.2%) were evaluated for patient-centered outcomes like mortality, morbidity, or readmission rates.
And of the 34 trials that did exist, 32 of them (94%) were industry-sponsored.
The study authors did not mince words. They called the evidence base "incomplete, systematically biased, and tilted toward optimism."
That is a polite way of saying: we do not actually know if most of this stuff works.
Radiology Owns the Problem
Radiology accounts for 78% of all FDA-cleared AI devices. That is more than 1,000 tools for reading chest X-rays, detecting nodules, flagging intracranial bleeds, segmenting tumors, triaging worklists, and measuring organ volumes.
Radiologists interact with AI more than almost any other specialty. They are the primary adopters. The tools are embedded in their workflow. And the evidence for those tools is the weakest of any medical domain.
Of the AI devices tied to radiology, only about 1% were ever linked to prospective clinical trials. Not 1% with positive results. 1% with any trial at all.
When you zoom in on patient outcomes - the kind that matter to patients, not just to benchmark papers - the number falls below 0.2%. One or two radiology AI tools, out of more than a thousand, have ever been tested against mortality or readmission.
Seventy-two percent of radiology departments have adopted at least one AI tool. Most of those departments cannot point to a single prospective trial showing their tool improved what happened to their patients.
The 510(k) Pathway Is the Structural Problem
Here is the part that most commentators gloss over. This is not a vendor problem. It is a regulatory structure problem.
The vast majority of AI medical devices were cleared through the FDA's 510(k) pathway. That pathway requires only one thing: substantial equivalence to a predicate device that already exists. The vendor does not need to prove the device is effective. They do not need to prove it is safe for every population. They do not need to prove it helps patients. They need to prove it is similar enough to something that was already cleared.
That made sense when the predicate was a blood pressure cuff or a surgical stapler. It makes much less sense when the predicate is a black-box neural network trained on one hospital's CT scanner dataset, applied to a patient population it has never seen.
The 510(k) pathway is not broken. It is just being asked to do something it was never designed to do: certify the real-world clinical value of complex adaptive algorithms.
Accuracy Is Not an Outcome
This is the part I want every radiologist, every CMO, and every CMS official to read slowly.
Approximately 90% of published radiology AI studies report process metrics: speed, sensitivity, specificity, AUC, time-to-read, workflow efficiency. These are the metrics that appear in vendor pitches and investor decks.
They are not the metrics that answer the only question that matters to a patient: did this tool help me?
A model can be 95% accurate on a benchmark dataset from a single academic center and still fail to detect the finding that kills a rural patient whose CT scanner is older, whose body habitus is different, and whose scan protocol does not match the training data.
Accuracy on the validation set is not accuracy in your emergency department. And accuracy in your emergency department is not patient survival.
These are three completely different things. We have been conflating them for a decade.
The Funding Reality
Of the 34 trials that do exist, 32 are industry-sponsored. That is 94%.
Industry-sponsored trials are not inherently bad. But they come with documented biases: they are more likely to be designed around metrics the product does well on, more likely to select favorable patient populations, more likely to be published when results are positive, and less likely to report adverse findings.
The study authors flagged this directly. An evidence base that is 94% industry-funded, with results published in fewer than 1% of cleared devices, is not a scientific foundation. It is a marketing narrative.
Independent, publicly funded, prospective trials on AI diagnostic tools are rare. The NIH has funded some. PCORI has funded a few. But the scale is nowhere close to what the market has deployed.
Three Concrete Reforms the Authors Recommended
The study did not just diagnose the problem. The researchers put three specific recommendations on the table.
1. Mandate pre-registration for Class II and III AI devices
Before the FDA clears a Class II or Class III AI device, vendors should be required to register a prospective trial on ClinicalTrials.gov. Not complete the trial - register it. This creates accountability and a traceable evidence pipeline. It is a floor, not a ceiling.
2. Require validation in real-world populations
Academic medical center datasets are not representative. A nodule detector trained on high-resolution CT scans from a tertiary cancer center will behave differently in a community hospital with different equipment, different protocols, and different patient demographics. Clearance should require validation in populations that reflect actual intended use.
3. Link reimbursement to demonstrated clinical benefit
This is the unlock. CMS currently offers 26 CPT codes for clinical AI solutions (as of January 2026). But coverage decisions are largely decoupled from outcomes. The 2026 Medicare physician fee schedule signals movement toward outcome-linked reimbursement, but the mechanism is not yet mandatory.
If CMS and commercial payers require demonstrated patient-centered benefit before reimbursement, the entire incentive structure changes overnight. Vendors who have been competing on benchmark accuracy will have to compete on what their tool does to the patient.
The EU AI Act, now requiring radiology AI to meet high-risk compliance standards with data curation requirements, bias checks, and mandatory human oversight, is already forcing this conversation in Europe. The United States is lagging.
What This Means for Lung Cancer Screening AI
I build AI for lung cancer screening. I live in this data every day. And I want to be honest about what this means for the specific corner of radiology I work in.
Lung nodule detection and characterization tools are among the most widely deployed AI devices in radiology. There are dozens of them. Some have been running in clinical environments for years.
The evidence base for their real-world patient outcomes is thin. Most of the published literature measures sensitivity and specificity on curated datasets. A small number of studies examine time-to-diagnosis or unnecessary biopsy rates. Almost none measure the outcome that actually matters: whether patients who had AI-assisted lung cancer screening lived longer than those who did not.
This is not a scandal. It is a structural failure of the evidence ecosystem. The tools may well be helping patients. There are plausible mechanisms by which faster, more consistent detection leads to earlier stage diagnosis, which leads to better survival. But "plausible mechanism" is not clinical proof. And until we have that proof, every hospital administrator signing a contract for a lung AI tool is making a bet, not a clinical decision.
The study's framing matters here. The problem is not that vendors are malicious. The problem is that the regulatory pathway and reimbursement structure did not require the evidence. When you do not require something, you do not get it.
What Radiologists and Health Systems Should Do Right Now
This research is not an argument to stop using AI in radiology. It is an argument to be clear-eyed about what you know and what you do not know.
Ask vendors for trial data, not benchmark data. If they cannot point to a ClinicalTrials.gov registration, that is a signal. If they can point to one but it has not posted results, ask why.
Participate in registry trials. If your institution has deployed an AI tool, prospective data collection adds to the evidence base for the whole field.
Engage with your payer on coverage criteria. CMS's Coverage with Evidence Development (CED) mechanism exists for exactly this scenario. It allows coverage while evidence is being gathered. Push for it to be applied to high-volume radiology AI.
Flag outcomes, not just workflow wins. When you present AI to leadership, start tracking patient-centered metrics: time to treatment initiation, unnecessary procedure rates, 30-day readmission rates. Build the internal evidence base your institution needs.
The Bottom Line
1,357 AI devices. Three tested for whether they help patients survive, recover, or stay out of the hospital.
That is where we are. Not because the technology is bad. Not because radiologists are careless. But because the regulatory pathway asked the wrong question, the reimbursement structure rewarded the wrong metrics, and the evidence ecosystem was funded almost entirely by people who needed good results.
The study authors from the University of Toronto gave us the diagnosis. The FDA, CMS, and independent researchers now have a clear mandate for the treatment.
Whether they act on it is the question that will define the next decade of AI in medicine.
Sources: Abulibdeh et al. (2026). Clinical Evidence for AI/ML-Enabled Medical Devices in the United States: A Systematic Analysis. PLOS Digital Health. | TechTarget HealthTech Analytics (September 8, 2026). | FDA AI-Enabled Medical Device List (March 2026). | CMS 2026 Medicare Physician Fee Schedule.


