1. Among 1,357 FDA-authorized artificial intelligence and machine learning-enabled medical devices, only 34 were linked to registered prospective trials and only three had been evaluated using patient-centered outcomes such as mortality, morbidity, or readmission.
2. FDA authorization and clinical validation answer different questions, making prospective workflow and outcome evidence increasingly important when physicians and health systems decide whether an authorized algorithm should actually be adopted into practice.
Artificial intelligence-enabled medical devices have accumulated regulatory authorizations much faster than evidence showing whether using them makes patients healthier. A systematic analysis published August 19, 2026 in PLOS Digital Health examined 1,357 artificial intelligence and machine learning-enabled devices authorized by the U.S. Food and Drug Administration (FDA) through December 5, 2025 and linked them with trial registrations and peer-reviewed publications. Only 34 devices, or 2.5%, were associated with registered prospective clinical trials, while only 12 had posted trial results and 12 had corresponding peer-reviewed publications. The funnel narrowed further at the endpoint physicians ultimately care about, with only three of 1,357 devices evaluated using patient-centered outcomes such as mortality, morbidity, or hospital readmission. Even the available prospective evidence was generally limited in scale and diversity, with nearly three quarters of registered trials enrolling fewer than 500 participants and roughly one quarter enrolling fewer than 100. Only nine of the 34 trials reported any subgroup analysis, while race or ethnicity appeared in only three, leaving substantial uncertainty about performance across the populations in which these tools may ultimately be deployed. A highly accurate algorithm can still fail clinically if clinicians ignore its output, false positives generate unnecessary testing, workflow changes undermine its value, or performance deteriorates when the deployment population differs from the development dataset.
Regulatory clearance therefore should not be interpreted as evidence that an AI device has demonstrated improved clinical outcomes, just as a technically impressive validation study does not establish that implementing the technology will benefit patients. The policy challenge becomes more complicated as generative AI enters medical devices, because outputs may be less deterministic and some systems may evolve in ways conventional static validation does not capture well. That creates a stronger argument for postdeployment monitoring, version tracking, prospective evaluation, and predefined processes for detecting changes in performance rather than treating authorization as the end of evidence generation. For physicians, the immediate implication is practical: when a hospital introduces an FDA-authorized algorithm, asking what prospective evidence supports its use is not redundant with asking whether the FDA cleared it. Health systems should also determine whether evidence comes from populations and workflows resembling their own, what endpoint was actually measured, whether subgroup performance is known, and what happens when the model makes an incorrect recommendation. The extraordinarily small proportion of devices with patient outcome data does not establish that the remaining technologies are ineffective, but it does mean their clinical benefit often remains unproven. As medical AI moves from novelty to infrastructure, success should increasingly be measured by what happens to patients after deployment rather than by the number of algorithms that reach the market.
Image: PD
©2026 2 Minute Medicine, Inc. All rights reserved. No works may be reproduced without expressed written consent from 2 Minute Medicine, Inc. Inquire about licensing here. No article should be construed as medical advice and is not intended as such by the authors or by 2 Minute Medicine, Inc.




