Evidence explainer

Prevention, nutrition, and travel health

How a USPSTF Letter Grade Is Actually Decided

A USPSTF letter grade is two judgments crossed: how sure the Task Force is of its net-benefit estimate, and how large that benefit is. An I statement flags missing evidence, not a service to avoid.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. The two dials behind every letter
  3. Building the certainty dial
  4. Sizing the magnitude dial
  5. Where the two dials cross
  6. An I is not a soft D
  7. Why the letter follows the patient out the door

Key points#

Reach for a familiar habit when you see a preventive-care recommendation graded A, B, C, D, or I: read it as one verdict, thumbs up or thumbs down. That habit is the source of most confusion about these grades. Each letter is not a single score. It is the meeting point of two independent questions the U.S. Preventive Services Task Force asks separately and then crosses together: how sure are we, and how much good does it do. Understanding the grade means keeping those two questions apart.

The two dials behind every letter#

Think of the Task Force turning two dials before it prints a letter.

The first dial is magnitude, the size of the net benefit, defined as the benefit of a preventive service minus its harm when delivered to an ordinary primary care population. Magnitude lands in one of four settings: substantial, moderate, small, or zero to negative.

The second dial is certainty, defined as the likelihood that the Task Force's estimate of that net benefit is actually correct. Certainty has three settings: high, moderate, or low.

The reason the two must stay separate is that they answer genuinely different questions. You can be very sure that a benefit is small. You can be quite unsure about something that could be large. Folding both dials into a single good-or-bad impression is exactly where misreadings start.

Building the certainty dial#

Certainty is not read off a single trial. The Task Force first draws an analytic framework, a map that runs from the population being screened to the health outcomes patients actually care about, passing through every intermediate step and possible harm on the way. Each connection on that map is a key question, and each key question needs its own evidence for the full route to hold.

For every key question, reviewers weigh six things: whether the study design suits the question, the internal quality of the studies, how well the results carry over to routine U.S. primary care, the number and size and precision of the studies, how consistent the findings are, and supporting signals such as a dose-response pattern or biological plausibility. The evidence behind each link is then graded convincing, adequate, or inadequate.

Here the map metaphor earns its keep. A route is only as reliable as its weakest verified connection, so one inadequate link can cap the certainty of an entire recommendation no matter how solid the other steps are.

That grading produces the three settings. High certainty rests on consistent results from sound studies in representative populations, where further research is unlikely to overturn the conclusion. Moderate certainty means the evidence is enough to reach a conclusion but is held back by inconsistency, shakier generalizability, or gaps in the chain, so new findings could still move the estimate. Low certainty means the evidence is simply insufficient: too few studies, flawed ones, contradictory ones, or a chain that never quite reaches real health outcomes.

Sizing the magnitude dial#

To set the magnitude dial, the Task Force builds outcome tables that project concrete health results for a hypothetical population, one copy of it receiving the service and one not. Benefits and harms are both expressed in those same tangible terms, cases prevented, harms incurred, then set against each other and sorted into one of the four magnitude bins. The whole point of the table is to force benefit and harm onto a shared scale before anyone reaches for a letter.

Where the two dials cross#

Line certainty up against magnitude and the published grade definitions fall straight out of the grid.

Low certainty breaks the pattern. Whenever the certainty dial sits at low, the grid does not return a letter at all, no matter how large the benefit might eventually prove to be. It returns an I statement.

An I is not a soft D#

This is the distinction most worth carrying away. A D grade is a firm recommendation against a service, grounded in reasonable confidence that it does not help or that it does net harm. An I statement says the Task Force cannot yet weigh benefits against harms because the evidence is missing, poor, or conflicting.

Put simply, a D reflects knowledge and an I reflects its absence. Reading an I as though it were a D turns a genuine "we do not yet know" into a false "stop," and that mistake can unfairly discredit services whose only fault is that they have not been studied well enough yet.

Why the letter follows the patient out the door#

Some of the weight these grades carry is legal rather than merely clinical. Under the Task Force's congressional mandate and the coverage rules attached to it, services rated A or B are the ones private health plans, and through a separate provision many Medicaid programs, must cover with no patient cost-sharing.

That is why the boundary between a B and a C, and between a B and an I, matters so much in practice, and why the two-dial logic behind each grade rewards being understood rather than flattened into a thumbs up or thumbs down.

Sources and further reading

  1. USPSTF Grade Definitions
  2. USPSTF Methods for Estimating Certainty and Magnitude of Net Benefit
  3. USPSTF Congressional Mandate (Procedure Manual, Appendix I)

Questions and answers

Does an I statement mean the service does not work?

No. An I statement means the U.S. Preventive Services Task Force found the evidence too limited, poor, or conflicting to weigh benefits against harms. It reports missing knowledge, not a decision against the service. A D grade is the recommendation against, and it rests on reasonable confidence that the service does not help or causes net harm.

What is the difference between an A and a B grade?

Both recommend the service and both require at least moderate net benefit backed by adequate certainty. An A reflects high certainty of a substantial benefit; a B reflects high certainty of a moderate benefit or moderate certainty of a moderate to substantial one. For insurance purposes the two are usually treated the same.

Why does the letter grade affect what I pay?

Under the Task Force's congressional mandate, most private health plans must cover services rated A or B with no patient cost-sharing, and a separate provision extends similar coverage to many Medicaid programs. That is why the line between a B and a C or an I can change the out-of-pocket cost of a preventive service.