Evidence explainer

Evidence and research methods

How to Read a Cardiology Guideline: Class of Recommendation and Level of Evidence

Class of Recommendation says how strongly to act. Level of Evidence says how certain the data are. Because they are rated separately, a firm 'should do' can rest on expert consensus.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Two labels, two questions
  3. Level of Evidence: how good is the data
  4. Class of Recommendation: how hard should you act
  5. Why the two labels do not move together
  6. What the numbers actually look like
  7. A quick way to read any recommendation

Open almost any recommendation in a modern American College of Cardiology and American Heart Association (ACC/AHA) guideline and you will see two small labels sitting side by side, something like "Class I (LOE A)." They answer two different questions. The Class of Recommendation (COR) answers "how firmly should I act on this?" The Level of Evidence (LOE) answers "how sure is anyone about the data underneath?" The single most useful habit a reader can build is to stop treating those as one verdict. A confident instruction and a solid evidence base are not the same thing, and a guideline is written so you can see exactly where they line up and where they part ways.

Key points#

Two labels, two questions#

Think of the pair as a "what to do" tag and a "how well we know" tag. The writing committee sets the first from its overall read of benefit versus risk, feasibility, and patient values. It sets the second purely from the trials, registries, and studies that exist. Because the two are built from different inputs, they do not have to agree, and the guideline deliberately keeps them apart so a reader can inspect each one.

It helps to read the evidence tag first. Knowing the quality of the data frames how much weight the instruction can really bear.

Level of Evidence: how good is the data#

LOE describes the amount, quality, and consistency of the underlying research, with no reference to how forceful the recommendation is. The current framework uses five tiers.

LOE A is the top rung: high-quality data from more than one randomized controlled trial, or meta-analyses of such trials. LOE B-R is moderate-quality evidence from one or more randomized trials, where the R flags a randomized design. LOE B-NR is moderate-quality evidence from well-conducted nonrandomized, observational, or registry work, with NR for nonrandomized. LOE C-LD signals limited data: small studies, registries with methodological gaps, or reasoning from physiology. LOE C-EO is expert opinion, the consensus of the writing group where direct evidence is thin or absent.

The 2015 revision, published in Circulation in 2016, split the old "Level B" and "Level C" into randomized and nonrandomized components. That change lets a reader tell at a glance whether "moderate evidence" came from a trial that assigned treatment at random or from watching what happened in practice, a difference that changes how much a single new study should shift your view.

Class of Recommendation: how hard should you act#

COR translates the committee's judgment about net benefit into an instruction, and the standardized wording is the tell. The 2015 framework fixed the verbs for each class so the phrasing itself signals strength.

Class I is strong: benefit clearly outweighs risk, and the text reads is recommended, is indicated, or should be performed.

Class IIa is moderate: benefit still outweighs risk, but less decisively, phrased as is reasonable or can be useful.

Class IIb is weak: benefit is only marginally greater than or roughly equal to risk, softened to may be reasonable or may be considered.

Class III is split into two distinct messages. Class III: No Benefit means the intervention simply does not help (is not recommended, is not indicated). Class III: Harm means it is expected to cause injury (should not be performed, is potentially harmful). Flattening those into one "don't" throws away real information: one is telling you a thing is useless, the other is telling you it is dangerous.

The verbs are the recommendation. When a committee writes may be reasonable instead of is recommended, that downgrade in language is the whole point.

Why the two labels do not move together#

Here is the crux most readers slide past. A high LOE does not force a high COR, and a low LOE does not rule out a strong one.

You can find a Class I recommendation supported only by LOE C-EO. The usual reason is an action so plainly beneficial, or so impossible to randomize ethically, that no trial exists yet withholding it would be indefensible. The committee is confident in the action even though the evidence tier is low. Run it the other way and a large randomized trial (LOE A) can yield only a Class IIb recommendation when the measured benefit is small, uncertain in a subgroup, or cancelled out by cost or side effects. Firm data, tentative advice.

That independence is a feature, not a defect. COR blends the evidence with judgment about benefit, risk, feasibility, and values; LOE reports only how far the data reach. Squash them into a single score and you lose the difference between "we are confident this helps a great deal" and "we are confident the trial was well run."

What the numbers actually look like#

This distinction matters more once you see how uncommon top-tier evidence is. In a 2019 analysis in JAMA, Fanaroff and colleagues reviewed 26 current ACC/AHA guidelines and found that only about 8.5 percent of recommendations rested on LOE A, while roughly half were LOE B and the rest LOE C. Even among the strongest Class I recommendations, only about one in seven carried LOE A, so the large majority of "should do" statements were not backed by multiple randomized trials. The share of LOE A recommendations had not meaningfully climbed over the prior decade.

That is not a reason to distrust guidelines. It is a sign the labels are doing their job, marking which advice is anchored in randomized data and which is careful consensus filling a genuine gap. A Class I, LOE A line and a Class I, LOE C-EO line ask the same action of you, but they invite very different scrutiny when a fresh trial reports or when a specific patient does not fit the usual mold.

A quick way to read any recommendation#

Build a two-step reflex. First read the COR to learn what the committee wants done and how strongly. Then read the LOE to learn how firm the ground is. A Class IIb, LOE C-LD line is a soft suggestion on thin data, fair to weigh against patient preference and to revisit as evidence matures. A Class I, LOE A line is about as settled as cardiology gets. And when a guideline is revised, the movement of the labels is often the real news: an LOE climbing as trials report, or a COR shifting as the benefit-and-risk picture changes.

Sources and further reading

  1. ACC/AHA Further Evolution of the Recommendation Classification System (Circulation 2016)
  2. Fanaroff et al., Levels of Evidence Supporting ACC/AHA and ESC Guidelines 2008-2018 (JAMA 2019)
  3. Fanaroff 2019, full text (PMC open access)

Questions and answers

Can a strong recommendation have weak evidence?

Yes. Class I means the committee is confident in the action, not that randomized trials prove it. Some Class I recommendations rest on expert opinion (LOE C-EO), often for interventions that are obviously helpful or cannot be ethically randomized.

What is the difference between LOE B-R and LOE B-NR?

Both are moderate-quality, but B-R comes from randomized trials and B-NR comes from nonrandomized, observational, or registry studies. The tag tells you at a glance which study design produced the "moderate" rating.

Why split Class III into No Benefit and Harm?

Because the two carry different clinical meanings. "No Benefit" says an intervention does not help; "Harm" says it is expected to cause injury. Keeping them separate preserves information that a single "do not" would erase.