Evidence explainer

Evidence and research methods

The Hierarchy of Evidence, Explained

A randomized trial is usually strong for treatment benefit, a cohort better for prognosis or rare harm, a cross-sectional study right for diagnostic accuracy. Evidence has an order only once you fix the question.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Why a hierarchy exists
  2. Start with the exact question
  3. The Oxford levels are question-specific
  4. A randomized label is not a quality certificate
  5. Observational evidence is not one category
  6. A systematic review is a lens, not a throne
  7. One study versus a body of evidence
  8. GRADE separates design from certainty
  9. Risk of bias, indirectness, and imprecision are different
  10. Certainty is not the same as effect size
  11. Evidence strength is not recommendation strength
  12. When randomized trials are unavailable
  13. Match evidence to the decision, not just the claim
  14. A compact way to appraise any evidence claim
  15. The question-first conclusion
  16. References

The familiar evidence pyramid places laboratory ideas and case reports near the bottom, observational studies in the middle, randomized trials above them, and systematic reviews at the top. It teaches a useful first lesson: some designs protect causal comparisons from predictable biases better than others.

Taken literally, the picture becomes wrong.

A systematic review of biased trials does not become trustworthy by being a review. A randomized trial is not the right design for estimating disease prevalence. A large cohort may reveal a rare delayed harm that a short trial could never detect. A high-quality diagnostic-accuracy study uses a different architecture from a treatment trial.

Treat the hierarchy as a question-specific starting point, then do the appraisal yourself. Look at the conduct, the directness, and the precision. Look at the consistency and the reporting.

Why a hierarchy exists#

Research designs differ in the counterfactual they can support, and a treatment question asks what would happen to the same type of people if they received one option rather than another. Because you cannot send one person down both paths at once, a study needs a credible comparison group.

Random allocation makes treatment choice unrelated to measured and unmeasured baseline prognosis in expectation. That property gives a well-conducted randomized trial a strong starting position for estimating treatment benefit.

In an observational comparison, people and clinicians choose treatments for reasons related to prognosis. Age, disease severity, access, contraindications, and preference can influence both the treatment and the outcome. Adjustment can address measured factors, but unmeasured and poorly measured confounding may remain. The hierarchy therefore encodes typical vulnerability to bias. It does not pronounce a verdict on a particular paper.

Start with the exact question#

“What is the best evidence?” is incomplete until “for what?” is answered.

Does an intervention improve outcomes?#

A randomized trial is often the strongest primary design. A systematic review of comparable, low-bias randomized trials can estimate the total evidence and consistency.

Pragmatic randomized trials can test effectiveness in routine settings. Explanatory trials can test efficacy in more controlled circumstances. Both randomize, but they answer different versions of the treatment question.

What causes harm?#

Trials can measure common, near-term harms, but they may be too small, short, or selective for rare and delayed events. Large cohorts, registries, spontaneous safety reports, and self-controlled designs are usually where those signals first appear and where they get estimated. Observational safety evidence needs careful work on confounding, outcome validation, time-related bias, and selective prescribing. Its lower causal protection does not make it optional.

Is a diagnostic test accurate?#

The core design is usually a representative or consecutive sample of people in whom the diagnosis is uncertain, and all participants receive the index test and an appropriate reference standard, with interpretation protected from knowledge of the other result.

A case-control study using obvious advanced cases and very healthy controls may exaggerate accuracy. Randomization is not needed merely to estimate sensitivity and specificity. It may be needed later to ask whether using the test improves patient outcomes.

What is the prognosis?#

An inception cohort follows people from a similar, clearly defined point in disease. Complete follow-up, valid outcome measurement, and appropriate handling of competing events matter. A treatment trial's control arm can contribute, but restrictive eligibility may limit generalizability.

How common is a condition?#

A current, representative population sample or census is more useful than a randomized trial. Sampling frame, participation, measurement validity, and weighting determine credibility.

What is a patient's experience?#

Interviews, focus groups, ethnography, and other qualitative methods can examine meaning and barriers. They can examine communication and lived experience. Their rigor rests on appropriate sampling, data collection, and reflexivity. It rests on analysis and transparency, not random allocation.

How might a mechanism work?#

Cell, tissue, animal, genetic, and physiological studies can establish pathways and generate targets. They can show that a mechanism is possible. They do not by themselves establish net clinical benefit in humans.

The Oxford levels are question-specific#

The Oxford Centre for Evidence-Based Medicine's 2011 Levels of Evidence use separate columns for prevalence, diagnosis, prognosis, treatment benefit, treatment harm, and screening; that structure is more informative than one universal pyramid.

For treatment benefit, systematic reviews of randomized trials and individual randomized trials sit high. For prognosis, systematic reviews of inception cohorts and individual inception cohorts lead. For diagnosis, cross-sectional studies with a consistently applied reference standard are central. The table also includes prompts to downgrade evidence because of poor quality, imprecision, indirectness, or inconsistency, and the level is not meant to be read apart from its introductory and background documents.

A randomized label is not a quality certificate#

Randomization protects a study only if implemented credibly and followed by an analysis aligned with the randomized comparison.

Important threats include:

Cochrane's RoB 2 framework assesses bias for a specific result, not a whole paper in the abstract, and one trial can have a low-bias mortality result and a more concerning subjective symptom result.

Randomization also does not guarantee applicability. A trial may exclude older adults, people with multiple conditions, pregnant participants, or communities with limited access. Internal validity and relevance are separate questions.

Observational evidence is not one category#

“Observational study” covers very different designs and data quality.

A prospective cohort with prespecified variables, validated outcomes, near-complete follow-up, an active comparator, and careful causal analysis is not equivalent to an uncontrolled chart review; a case-control study nested in a defined cohort is not equivalent to asking a selected group to remember past behavior.

Credibility depends on whether the design emulates the comparison of interest:

FDA's real-world evidence framework recognizes that routine-care data can support randomized or observational designs. Data source and design are distinct. Electronic health records do not automatically create observational evidence, and a registry can host a randomized trial.

A systematic review is a lens, not a throne#

Systematic reviews use explicit methods to define a question, search for studies, select eligible evidence, appraise bias, and synthesize results; done well, they reduce cherry-picking and show the full pattern of evidence.

Their conclusions remain constrained by:

A meta-analysis is the statistical combination inside some reviews. It can produce a precise pooled estimate from consistently biased inputs. Precision does not cancel systematic error.

The “new evidence pyramid” proposed by Murad and colleagues reframes systematic reviews as a lens through which studies are viewed, rather than a design floating above them. That image captures an important idea: synthesis quality depends on both the lens and the underlying evidence.

One study versus a body of evidence#

An individual study may be compelling, but health decisions usually depend on a body of evidence.

Replication across different investigators and settings helps distinguish a stable effect from chance, local practice, analytic choice, or hidden bias; consistency is not mandatory if real effect modification explains differences, but unexplained conflict lowers confidence.

A body can include complementary designs. Randomized trials estimate benefit. Large cohorts extend safety follow-up. Diagnostic studies define who has the condition. Qualitative research identifies burdens and acceptability. Mechanistic work explains plausibility. The designs do not need to compete for one trophy. They answer connected parts of a decision.

GRADE separates design from certainty#

GRADE assesses certainty for a specific outcome across a body of evidence. The four levels are high, moderate, low, and very low.

For intervention effects, randomized evidence generally starts at high certainty and can be rated down for:

Nonrandomized intervention evidence traditionally starts at low certainty and may be rated up for factors such as a large effect, a dose-response pattern, or plausible confounding that would reduce the observed effect. When ROBINS-I is used, reviewers can conceptually start at high and then account explicitly for confounding and selection concerns. In practice, strong justification is needed to retain high certainty. The rating belongs to an outcome, not a study or intervention forever, and evidence for short-term symptom relief may be high certainty while evidence for a rare serious harm is low certainty.

Risk of bias, indirectness, and imprecision are different#

These domains are often collapsed into “study quality,” but they ask distinct questions.

Risk of bias asks whether methods may systematically distort the estimate.

Indirectness asks whether the evidence matches the target population, intervention, comparator, and outcome. A flawless trial of a related dose or a surrogate endpoint may be indirect for the decision.

Imprecision asks whether the estimate is stable enough to distinguish among decisions. A low-bias trial with ten events can remain highly uncertain.

Inconsistency concerns unexplained variation across studies. Publication bias concerns missing results related to their direction or size. Keeping the domains separate makes the reason for uncertainty actionable. Better concealment addresses bias; a larger sample addresses imprecision; a trial in the target population addresses indirectness.

Certainty is not the same as effect size#

High-certainty evidence can show little or no benefit. Low-certainty evidence can suggest a dramatic effect. Certainty describes confidence that the effect lies near the estimated range, not whether the effect is favorable.

Look at both:

A precise change in a laboratory surrogate may be less decision-relevant than an imprecise estimate of mortality. Outcome importance is not encoded by study design alone.

Evidence strength is not recommendation strength#

A guideline recommendation combines evidence with judgments.

Panels consider the balance of benefits and harms, patient values, and resource use. They consider feasibility, acceptability, equity, and health-system context. A small certain benefit may not justify major burden; a strong recommendation may sometimes rest on lower-certainty evidence when harms of inaction are grave and alternatives are limited, but the rationale should be transparent.

USPSTF methods, for example, build analytic chains for preventive services: screening evidence can require test accuracy, benefits and harms of downstream evaluation, treatment effects in screen-detected disease, and the link from intermediate outcomes to health outcomes. No single study design answers the whole chain.

When randomized trials are unavailable#

Randomization may be infeasible or unethical when an effect is enormous, an event is rare, a harmful condition cannot be assigned, or a long latency exceeds practical follow-up.

Anesthesia for surgery did not require a placebo trial against restraint to establish a dramatic immediate effect, and the relationship between smoking and lung cancer emerged from converging observational, pathological, dose-response, temporal, and mechanistic evidence.

These examples do not mean observational evidence is always sufficient. They show that causal confidence can grow from a coherent body whose alternative explanations become implausible; natural experiments, instrumental variables, regression discontinuity, interrupted time series, and difference-in-differences designs can strengthen causal inference in certain settings. Each depends on assumptions that must be stated and tested where possible.

Match evidence to the decision, not just the claim#

A clinician asking whether to start a medicine needs treatment benefit and harm evidence. But they also need baseline risk, comorbidities, and competing treatments. They need preferences and feasibility.

A regulator may ask whether evidence supports a product indication. A public-health agency may ask about population impact and equity. A patient may care most about function, symptom burden, or treatment burden. The same body of studies can support different decisions because thresholds and consequences differ. Evidence appraisal should make those decision rules visible.

A compact way to appraise any evidence claim#

Ask these questions in order:

  1. What exact question is being asked?
  2. Which design best answers that question in principle?
  3. Was this study conducted and reported well?
  4. Does its population, intervention, comparator, and outcome match the decision?
  5. How large and precise is the effect?
  6. Do other studies agree, and are missing studies likely?
  7. What is the certainty for the outcome that matters?
  8. How do benefits, harms, burdens, values, and context change the action?

That sequence keeps the useful discipline of the hierarchy and stops a design label from ending your appraisal for you.

The question-first conclusion#

The evidence hierarchy remains useful when it reminds readers that design shapes bias, and it fails when it becomes a universal ladder that ranks papers without asking what each one was built to answer.

Define your question first. Choose the design that best addresses it. Appraise the conduct that actually occurred. Then judge the body of evidence for bias, consistency, directness, precision, and missing results.

Good evidence-based reasoning is not obedience to a pyramid. It is a transparent match among question, method, result, certainty, and decision.

References#

Questions and answers

Are systematic reviews always the highest level of evidence?

No. A rigorous review is a strong synthesis method, but its conclusion depends on the included studies, search completeness, bias appraisal, comparability, and analysis.

Are randomized trials always better than observational studies?

They usually offer stronger causal protection for treatment benefits. Observational designs can be more appropriate for prognosis, prevalence, rare harms, long latency, and some policy questions.

Can a case report ever be important?

Yes. A case report can identify a new disease, unexpected adverse event, or striking mechanism. It usually cannot estimate frequency or prove a treatment effect.

What does high-certainty evidence mean?

It means reviewers have high confidence that the true effect for a specified outcome is close to the estimated range, but it does not mean the effect is large or beneficial.

Why can a guideline make a strong recommendation with low-certainty evidence?

Sometimes the consequences of inaction, plausible benefit, limited alternatives, or values make one course preferable despite uncertainty, and the panel should explain that judgment and avoid implying stronger evidence than exists.