Evidence explainer

Imaging and radiology

How to Judge an FDA Cleared Radiology AI Tool

An FDA clearance means a radiology AI tool reached the market, usually by matching an existing device through the 510(k) pathway, not by proving clinical benefit. Ask what was tested, and on whom.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. How to Judge an FDA Cleared Radiology AI Tool

How to Judge an FDA Cleared Radiology AI Tool#

An FDA clearance tells you a radiology AI tool is legal to sell. It does not tell you the tool will help your patients. That single distinction is where most confusion begins, because the same word, "cleared," gets read as if it meant "proven." To judge a tool for what it actually is, treat the clearance as a starting point and then ask three plain questions: what was tested, who was in the test, and how were the reported accuracy numbers produced.

Start with what the clearance certifies, and what it does not#

The FDA keeps a public list of AI-enabled medical devices, and radiology is by a wide margin the largest category on it, roughly three quarters of the entries. Almost all of these products enter the market through premarket notification, the pathway known as 510(k). Under it, a manufacturer shows that its device is "substantially equivalent" to a legally marketed predicate: same intended use, and either the same technological characteristics or differences that raise no new questions of safety and effectiveness. That phrasing is the FDA's own, and the word to sit with is equivalent. The bar is a resemblance to something that already exists, not fresh evidence that radiologists read more accurately or that patients do better.

An audit by Khunte and colleagues, released first as a medRxiv preprint and later in Clinical Radiology, put numbers on this. Of 151 imaging AI algorithms cleared through late 2021, 148 went through 510(k), and only three used the De Novo route meant for genuinely new devices. Hardly any faced the rigorous premarket approval studies expected of high-risk devices. So a clearance letter certifies that a regulatory comparison was reviewed and accepted. Think of it as the floor a product must clear to enter the room, not the ceiling of what it can do once inside.

Three accuracy numbers, and what each one hides#

Most imaging AI summaries lean on three figures, and each answers only a narrow question.

Sensitivity is the fraction of truly abnormal cases the tool catches. A very sensitive tool misses few cancers but tends to sound more false alarms. Specificity is the fraction of truly normal cases it correctly leaves alone. A very specific tool spares patients false alarms but can overlook subtle disease. These two move in opposite directions, and here is the catch: a vendor can slide the operating threshold up or down and quote whichever number flatters the product.

Area under the curve, or AUC, was designed to close that loophole. It summarizes the sensitivity to specificity trade-off across every possible threshold, running from 0.5 (a coin flip) to 1.0 (perfect). A high AUC looks reassuring, and it does describe good discrimination. But it stays silent on two things you cannot ignore. First, it says nothing about the single threshold your hospital will actually run the tool at. Second, it takes no account of how common the disease is in your patients.

That second gap is the one that surprises clinicians. Consider a model validated on an enriched dataset where half the scans are abnormal, then deployed in a screening clinic where fewer than one in a hundred are. Sensitivity and specificity can hold steady while positive predictive value, the number a radiologist lives with case by case, falls through the floor, because most positives in a low-frequency setting are false. One more question is worth asking of any figure: did it come from the algorithm reading alone, or from radiologists reading with and without it? Only the second design tells you whether the tool changes what a human decides.

Where a tool built elsewhere can slip#

A model learns the world it was trained on. Change the scanner, the reconstruction settings, the field strength, or the patient mix, and performance can drift without anyone noticing, because the dashboard still shows the old AUC. The published record suggests this risk is often invisible at the point of purchase, because the clearance summaries leave out exactly the details you would need to gauge it.

Comparison bar chartShare of the 151 cleared imaging AI summaries reporting each validation feature. Values: Reported clinical validation data, 64%; Validation data described as multicenter, 34%; Disclosed study population demographics, 4%Reported clinical validation data64%Validation data described as multicenter34%Disclosed study population demographics4%Scale maximum: 100%
Share of the 151 cleared imaging AI summaries reporting each validation feature.
View the constructed data table
Chart values
MeasureValue
Reported clinical validation data64%
Validation data described as multicenter34%
Disclosed study population demographics4%

In the Khunte audit, 97 of the 151 algorithms, about 64 percent, reported using clinical data to validate the device at all, meaning roughly a third described no clinical validation in their public summary. Only 51, about 34 percent, called their validation data multicenter; the rest largely did not say. Much of the testing was retrospective, run on stored images rather than in a live reading room. The generalizability gaps are wider still: only six summaries, about 4 percent, disclosed the demographic makeup of the study population, and only around 5 percent named the scanner models used. When those facts are missing, a published AUC is a claim you cannot check against your own department. The pattern is a fair reason to treat vendor numbers as hypotheses to test locally.

Questions to run through before you trust the dashboard#

When a tool crosses your desk, a short interrogation usually separates the strong products from the well-marketed ones:

None of this makes FDA-cleared imaging AI untrustworthy. It means clearance and clinical proof sit at different heights, and that the FDA's public list plus the peer-reviewed audits of it give you enough to place a given product between them. Reading diagnostic-accuracy statistics carefully is a general clinical-epidemiology skill, and it applies to a vendor's slide deck the same way it applies to a journal abstract. The best tools will welcome these questions. The weakest will answer with a clearance number and little else.

Sources and further reading

  1. FDA AI-Enabled Medical Devices List
  2. FDA Premarket Notification 510(k)
  3. Khunte et al., Trends in Clinical Validation of FDA-Cleared Imaging AI (medRxiv preprint)
  4. Khunte et al., Clinical Radiology 2023 (PubMed)

Questions and answers

Does FDA clearance mean a radiology AI tool was proven to help patients?

No. Clearance through the 510(k) pathway means the device was judged substantially equivalent to a product already on the market. It confirms a regulatory comparison, not that the tool improves diagnosis or patient outcomes. Those require separate clinical evidence, which the public summary may or may not contain.

Why can a tool with a high AUC still perform poorly in my clinic?

AUC measures how well a model separates abnormal from normal across all thresholds, but it ignores how common the disease is in your patients. In a low-frequency setting such as screening, most flagged cases can be false positives even when sensitivity and specificity look excellent, so positive predictive value drops sharply.

What single question is most useful when evaluating imaging AI?

Ask who was in the validation study and how similar they are to your patients and equipment. If the summary does not disclose demographics, scanner models, or whether testing was multicenter and prospective, treat the reported accuracy as a hypothesis to confirm locally rather than a guarantee.