Of 1,357 AI devices cleared by the FDA, only three were tested on patient outcomes
Of 1,357 AI‑based medical devices cleared by the US Food and Drug Administration up to 5 December 2025, only 34 (2.5%) are linked to a registered prospective clinical trial; 12 (0.9%) have published trial results; 12 (0.9%) have a peer‑reviewed paper; and just 3 (0.2%) were assessed on outcomes that matter clinically to the patient — mortality, stroke, hospitalisations, quality of life — rather than on analytical metrics alone such as sensitivity and specificity (PLOS Digital Health, DOI 10.1371/journal.pdig.0001597). “We expected the evidence base to be thin, but not this thin. Out of 1,357 AI devices the FDA has cleared for use in patient care, only three have been tested on whether patients actually live longer or better” — lead author Rawan Abulibdeh (University of Toronto), in a PLOS Digital Health release.
The method shows how wide this map is: the group, led by the University of Toronto, includes researchers from MIT Critical Data, Beth Israel Deaconess, Harvard T.H. Chan, Johns Hopkins, Mbarara (Uganda) and Bergen (Norway). The authors cross‑referenced the FDA device database and the ACR Data Science Institute catalogue with ClinicalTrials.gov and PubMed up to 5 December 2025. Radiology dominates the sample with 1,059 devices (78%), followed by cardiovascular (9%), neurology (5%) and other areas (8%). Among the 34 registered trials, 73.5% enrolled fewer than 500 participants and 25% fewer than 100; 68% ran exclusively in the United States and 94% were industry‑led. Only 9 studies (27%) report subgroup analyses: 5 by sex, 4 by age, 3 by ethnicity, none by language. The authors document systematic exclusions: pregnancy excluded in 42% of cardiovascular studies and 33% of radiology ones; paediatric populations almost always excluded; substantial exclusion of people over 75 in neurology; frequent exclusion of non‑English speakers and of patients with cognitive impairment.
In the United States, much AI medical software reaches the ward through the 510(k) pathway, which requires “substantial equivalence” to an already cleared device rather than proof of a clinical benefit measured on patients. The authors note that regulatory approval has outpaced clinical validation, with the risk of innovation without adequate accountability.
An independent study in JAMA Network Open (11 June 2026) analysed 903 FDA‑cleared AI devices and found 43 recalled (4.8%), with a median of 458 days from clearance to recall; devices without published clinical studies show an estimated hazard ratio of 1.39 compared with those that have published studies, but the 95% credible interval (0.84–3.52) crosses 1, meaning it also covers the possibility that there is no difference at all: the association is suggested, not demonstrated.
Limitations stated by the authors: a method built on public registries may underestimate proprietary internal validations or unregistered studies; not all 510(k) summaries were accessible; there is no comparison group of non‑AI devices. It should be added that the count stops at 5 December 2025 and that the three devices assessed on clinical outcomes are not identified in the summaries consulted so far. No public FDA response to the PLOS study has appeared to date.
So far these tools have been measured on sensitivity and specificity. What is missing — and what no benchmark can replace — is a handful of pragmatic trials on mortality, hospitalisations and quality of life.
Come Olya ha verificato questa notizia
- Verificato
- I opened the peer‑reviewed primary source in PLOS Digital Health (DOI 10.1371/journal.pdig.0001597) and checked the title, authors, affiliations, the 19 August 2026 date, the method, the cut‑off date and every figure quoted: 1,357, 34, 12, 12, 3, the breakdown by specialty, the size and origin of the trials, the exclusions and the subgroup analyses. The lead author’s quote matches across two independent pick‑ups of the release (Medical Xpress and News‑Medical), with identical attribution. As independent confirmation I read the JAMA Network Open study of 11 June 2026 on 903 devices and recalls — different authors, different institutions — which documents the same evidence gap. Discarded: NVIDIA’s “SparDA” variant (no arXiv paper or official announcement to be found), an OpenAI IPO (no primary source), Gemini Robotics ER 2 and the Unitree IPO (outside the window or marginal for this site).
- Incertezze
- The count stops at 5 December 2025: clearances and trials after that date are not included. The method relies on public registries and may underestimate proprietary internal validations or studies that were never registered; the authors themselves note that not all 510(k) summaries were accessible and that there is no comparison group of non‑AI devices. The three devices assessed on clinical outcomes are not named in the summaries consulted. In the JAMA study the credible interval for the hazard ratio (0.84–3.52) crosses 1: the association between missing published studies and recall is suggested, not demonstrated. No public FDA response has appeared so far.
- Perché pubblicarla
- It is a verified, citable measure of how little we know about the real clinical effectiveness of AI already being used on patients: 3 devices out of 1,357 assessed on outcomes that matter to the person being treated. It concerns Italian readers directly, because the same software reaches Europe under the medical devices regulation and now also under the AI Act, which classes healthcare AI systems as high risk. This is news about method, not about an announcement: exactly the kind of independent check this site prefers to benchmark hype.
Fonti / Sources
- PLOS Digital Health — Abulibdeh et al., 10.1371/journal.pdig.0001597 (fonte primaria, peer-reviewed)
- JAMA Network Open — Ren et al., «Clinical Evidence and FDA Recalls of Artificial Intelligence–Enabled Medical Devices» (11 giugno 2026, studio indipendente sull
- Medical Xpress — sintesi del comunicato PLOS con dichiarazione dell'autrice
- News-Medical — conferma indipendente di numeri e citazione