When an ACR Appropriateness Criteria document tells you a CT for a given complaint is "Usually Appropriate," that phrase is the compressed output of two separate machines running in sequence: a graded review of the published literature, and a structured vote by a multispecialty physician panel on a 1 to 9 scale. Scores of 7 to 9 read as "Usually Appropriate," 4 to 6 as "May Be Appropriate," and 1 to 3 as "Usually Not Appropriate." The number is a transparent summary of expert judgment tied to evidence for a typical patient in that scenario. It is not a rule that a particular test must or must not be done for the person in front of you.
That distinction is the whole point of reading these documents well. The American College of Radiology maintains the library at scale, and its current release spans hundreds of diagnostic imaging and interventional topics, thousands of clinical scenarios, and more than a thousand clinical variants, refreshed annually by panels that draw on hundreds of volunteer physicians. A clinician who understands how one rating is assembled can treat it as graded evidence plus disciplined consensus, rather than as a verdict handed down from somewhere unseen.
Key points#
- An appropriateness rating is built in two stages: a graded literature review, then a panel vote on a 1 to 9 scale using the RAND/UCLA method.
- The final category for a scenario is set by the panel's median score, and the underlying evidence table is open for inspection.
- "May Be Appropriate" is the most information-rich band, and a specific label flags scenarios where the panel disagreed.
- A separate Relative Radiation Level sits beside each rating so dose stays visible as its own fact, not folded into appropriateness.
What "appropriate" is actually claiming#
Before anyone votes, the method pins down what the word means. Following the RAND/UCLA Appropriateness Method User's Manual, the ACR treats an imaging procedure as appropriate when its expected health benefit exceeds its expected harms by a margin wide enough to justify performing it, judged separately from cost. That single definition tells you what a high score asserts and, just as usefully, what it does not. A "Usually Appropriate" rating is a claim about the benefit-to-harm balance for a typical patient who matches the scenario. It is not a statement about cost-effectiveness, insurance coverage, or any individual's odds.
Holding that definition in mind keeps a reader from over-interpreting the number. A high rating is not a guarantee of a useful scan, and a low rating is not a prohibition. Each is a summary of a balance struck for a described situation.
Stage one: the evidence table underneath the score#
Every topic opens with a systematic literature search, and the retrieved studies are condensed into an evidence table that lists each included study and rates its quality by design. The ACR aligns this appraisal with the GRADE framework (Grading of Recommendations Assessment, Development and Evaluation) and with the Institute of Medicine's standards for trustworthy guidelines.
GRADE is useful here because it keeps two ideas apart that people routinely fuse: how certain we are about the evidence, and how strong the resulting recommendation is. High-certainty evidence can still support a cautious recommendation, and a strong recommendation can rest, when honestly labeled, on lower-certainty evidence if the benefit-to-harm balance is clear. The payoff is that the appropriateness number sits on top of an evidence table you can open and read. Where the studies are thin or contradictory, the panel says so, and that uncertainty is meant to ride along with the rating instead of disappearing behind it.
Stage two: the panel vote on a 1 to 9 scale#
With the evidence assembled, the panel rates each scenario through a modified Delphi process. Members score every procedure privately on the 1 to 9 scale, so no single voice steers the room. The anonymous scores are tabulated and sent back to the group, and the panel votes again over successive rounds. The scenario's published category is set by the median of the final scores, which means the number reflects where collective judgment settled after members had seen both the evidence and one another's ratings.
Two design choices give this its credibility. Private, iterative voting is built to blunt the pull of the most senior or most vocal person present. And because each panel is multispecialty, seating both the clinicians who order a study and the radiologists who perform it, the resulting rating carries more than one professional vantage point rather than a single specialty's preference.
Why the middle of the scale carries the most information#
The most misread part of the scale is the middle band. "May Be Appropriate" does not reliably mean "a coin flip on the merits." Under the RAND/UCLA logic the ACR uses, a scenario can land there for genuinely different reasons: the benefits and harms really are balanced, the evidence is sparse or conflicting, specific subpopulations complicate the picture, or the panel simply could not converge. When scores scatter too widely around the median, the ACR marks the scenario as "May Be Appropriate (Disagreement)" and assigns it a 5.
That label is a feature, not a hedge. It signals that the number reflects unresolved professional conflict rather than a settled tie, and those are very different messages for a clinician to act on. A 5 is not one thing, so reading it well means opening the narrative to see which of those situations produced it.
The radiation symbols sitting beside the score#
Alongside each appropriateness rating, the criteria carry a Relative Radiation Level, a symbolic scale where more symbols signal a higher range of expected effective dose. Dose is placed next to appropriateness rather than mixed into it, so a reader can weigh a test's expected yield against its radiation burden as two separate facts. The dose bands are set by an ACR radiation dose subcommittee and have been validated against real-world dose measurements in the ACR Dose Index Registry, which anchors the symbols in observed practice rather than estimate alone.
Keeping dose visible and distinct matters most when appropriateness is high but the radiation level is also high, or when a lower-dose study earns a similar rating. Those are exactly the trade-offs a clinician wants surfaced rather than buried inside a single figure.
Reading a rating as three open layers#
Put together, an ACR appropriateness rating is best read as three layers stacked in plain view: a graded body of evidence, a structured expert vote summarized as a median, and a companion radiation estimate. A high number says the benefit-to-harm case is strong for the typical patient in that scenario. A middle number is an invitation to read why. The radiation symbols keep dose in the frame throughout. None of it substitutes for the clinical judgment that adapts a category to a real person with a real history.
The strength of the framework is that it shows its work. When a recommendation is this legible, from literature search to graded evidence to a documented vote, you are equipped to appraise it as evidence rather than accept or dismiss it as an edict.
Sources and further reading
Questions and answers
Does a "Usually Not Appropriate" rating mean the test is forbidden?
No. It means that, for a typical patient matching the scenario, the expected harms and low yield outweigh the expected benefit. An individual patient's history can still justify a study the general rating discourages. The rating informs the decision; it does not make it.
Why is the final category based on the median rather than an average?
The median is more stable against a few extreme scores and fits the RAND/UCLA method's aim of describing where the group settled. Pairing the median with a formal disagreement check also lets the ACR flag scenarios where consensus was not reached instead of hiding that spread inside a single averaged number.
Is the appropriateness rating the same as a cost or coverage decision?
No. The method defines appropriateness by the benefit-to-harm balance considered apart from cost. Insurance coverage and cost-effectiveness are separate judgments made by other bodies, so a high appropriateness rating does not by itself settle whether a payer will authorize a study.