Evidence explainer

Evidence and research methods

High, Moderate, Low, Very Low: How GRADE Rates the Certainty of Evidence

GRADE rates evidence as high, moderate, low, or very low certainty. The label says how much to trust that a study's number reflects reality.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Certainty is not the size of the benefit
  3. Every body of evidence has a starting line
  4. Five reasons the rating drops
  5. When weaker evidence gets promoted
  6. The label you should double-check

Look at the bottom of almost any modern clinical guideline and you will find one of four words attached to each recommendation: high, moderate, low, or very low. That word is a GRADE certainty rating, and it answers a narrow but essential question: how confident can we be that the effect measured in the studies is close to the effect that actually exists? It is not a rating of how big the effect is, how new the drug is, or how strongly the authors feel. When a recommendation is tagged "low certainty," the guideline is warning you that the real answer may sit some distance from the number in the summary table.

Key points#

Certainty is not the size of the benefit#

The most common misreading of GRADE is to treat "high" as "this works really well." It does not mean that. Certainty describes our confidence in the estimate, not the magnitude of what it estimates. A treatment that shrinks risk by a trivial amount can be rated high certainty if the studies are large and clean. A treatment that halves mortality can be rated low certainty if the only evidence comes from a handful of small, flawed studies. Keeping these two axes separate is the whole skill: one axis is "how good is the evidence," the other is "how big is the effect," and GRADE only speaks to the first.

The four tiers map onto plain statements of trust. High certainty means we are very confident the true effect is close to what the studies found. Moderate means it is probably close, but a meaningfully different reality is possible. Low means our confidence is limited and the truth could differ substantially. Very low means the estimate is barely more than an educated starting point.

Every body of evidence has a starting line#

GRADE does not treat all studies as equal at the outset. It assigns a starting rating based on study design, then moves the rating up or down from there.

Randomized controlled trials begin at high certainty. The reason is randomization itself: when done properly, allocating people to groups by chance balances out both the confounders we know about and the ones we have never thought of, so a difference between groups is more plausibly caused by the treatment. Non-randomized and observational studies begin at low certainty, because without randomization, confounding and selection effects can masquerade as treatment effects. From those two starting lines, the certainty rating gets adjusted.

Five reasons the rating drops#

Reviewers can downgrade a body of trial evidence for any of five problems. Each one, if serious, can knock the rating down one or even two levels.

Risk of bias concerns how the studies were actually conducted. Unconcealed allocation, missing blinding, large numbers of participants lost to follow-up, or selective reporting of only the favourable outcomes all let bias seep into the result. A very large trial run with sloppy methods does not buy back certainty; size cannot launder bias.

Inconsistency shows up when studies asking the same question return answers that scatter far more than chance alone would explain. Unexplained heterogeneity like this suggests something hidden, a difference in the patients, the dose, or how an outcome was defined, is pulling the results apart, which makes any single pooled figure less trustworthy. Consistent findings from independent teams, on the other hand, are reassuring.

Indirectness is a mismatch between the evidence and the question. A trial in frail hospitalized elders is only indirect evidence for healthy middle-aged adults. A trial that tracked a lab marker rather than an outcome patients feel, such as a heart attack avoided or a year of life gained, is indirect for what actually matters. Comparisons stitched together across separate trials, rather than tested head to head, are indirect too.

Imprecision is the problem of wide confidence intervals. Too few participants or too few events produce a range of plausible truths so broad that it spans both a worthwhile benefit and no benefit at all, sometimes even harm. When the interval covers outcomes that would drive opposite decisions, the estimate is imprecise and the rating falls. This is the domain that most often pulls a tidy set of small trials down to low certainty.

Publication bias reflects what never made it into print. Studies with underwhelming results are less likely to be published, so the surviving literature can flatter a treatment. Signs such as an asymmetric funnel plot, a field crowded with small industry-sponsored trials, or a conspicuous absence of null results prompt a downgrade for the studies that likely exist but were never reported.

When weaker evidence gets promoted#

Observational evidence is not stuck at the bottom. GRADE allows an upgrade in three situations. A very large effect, roughly a doubling or halving of risk with no credible confounder to explain it away, raises confidence; this is how the connection between smoking and lung cancer earned firm standing without a randomized trial. A dose-response gradient, where more of the suspected cause tracks with more of the effect, adds weight. And when every plausible bias would have pushed toward a smaller effect, yet an effect still appears, the true effect is probably at least as large as observed. These promotions rarely lift observational data all the way to high, but they can carry it to moderate.

The label you should double-check#

Certainty is only half of what GRADE produces. It deliberately separates the certainty of the evidence from the strength of the recommendation, and the two do not have to agree.

A strong recommendation means the benefits clearly outweigh the harms for nearly everyone, so most well-informed people would choose the option. A conditional (or weak) recommendation means the balance is close, or the evidence is shaky, so the right choice hinges on a person's own values and circumstances. Higher certainty naturally tends to support stronger recommendations, and GRADE guidance cautions against issuing strong recommendations on low or very low certainty evidence.

Yet the mismatch happens. A 2023 cross-sectional analysis in BMC Medical Research Methodology examined a set of national guidelines and found strong recommendations frequently attached to low or very low certainty evidence, sometimes justified by GRADE's recognized exceptions, such as immediately life-threatening situations, and sometimes not. That pairing, a strong recommendation built on thin evidence, is the single most useful thing to check when you read a guideline, because the effect underneath it could still move.

Sources and further reading

  1. Cochrane Handbook Chapter 14 (GRADE and Summary of Findings)
  2. GRADE: an emerging consensus (Guyatt 2008, BMJ)
  3. Strong recommendations from low-certainty evidence (Chong 2023, BMC Med Res Methodol)

Questions and answers

Does high certainty mean a treatment works well?

No. It means we can trust that the measured effect is close to the true one. The effect itself might be small. Certainty is about confidence in the number, not the size of the benefit.

Why do randomized trials start higher than observational studies?

Because randomization balances confounders, including unknown ones, across groups by chance. Observational studies cannot guarantee that, so a measured difference is more easily explained by something other than the treatment. GRADE reflects this by starting trials at high and observational evidence at low.

Can a strong recommendation rest on low-certainty evidence?

It can, and it sometimes should, for example in urgent, life-threatening scenarios where waiting for perfect evidence is not an option. But because GRADE generally advises against it, that combination is worth reading carefully rather than taking at face value.