Evidence explainer

Evidence and research methods

Prediction Intervals in Meta-Analysis: What They Add Beyond a Confidence Interval

A confidence interval summarizes uncertainty around an average effect. A prediction interval asks where the true effect of a similar new study might lie, given estimated heterogeneity.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Three quantities readers should keep separate
  2. The logic of the calculation
  3. A simple interpretation example
  4. What “a new similar study” really means
  5. Why I-squared is not a substitute
  6. Few studies create fragile intervals
  7. A wide interval is information, not a failed analysis
  8. Risk of bias and publication bias remain outside the formula
  9. Prediction for an effect is not prediction for a patient
  10. When a prediction interval is most useful
  11. A reading checklist

A random-effects meta-analysis often ends with one pooled estimate and one confidence interval. Those numbers answer a useful question about the mean effect across a distribution of study effects. They do not say that every setting shares that effect.

A prediction interval adds a different question: assuming the model is appropriate, what range might contain the underlying true effect in a new study sufficiently similar to those synthesized? When effects vary substantially, the interval can cross no effect or clinically important thresholds even while the pooled confidence interval looks convincing. That makes prediction intervals valuable without making them a universal forecast: they inherit every weakness in the studies and the model, and they can be unstable when the meta-analysis contains few studies.

Three quantities readers should keep separate#

First is the observed effect in each study. It contains sampling error because the study enrolled a finite number of participants.

Second is the average underlying effect across the random-effects distribution. The pooled estimate targets this mean. Its confidence interval represents uncertainty about where that mean lies.

Third is the underlying effect in another study drawn from the modeled distribution. This can differ from the average because populations, implementation, comparison care, outcome measurement, follow-up, and other effect modifiers vary. A prediction interval tries to represent that additional spread. The distinction explains why a meta-analysis can estimate the mean fairly precisely while remaining uncertain about a new setting: adding studies can narrow the uncertainty around the mean, but genuine heterogeneity does not disappear merely because its average is well estimated.

The logic of the calculation#

In a conventional random-effects model, each study has its own underlying effect. Those study effects are assumed to vary around an overall mean with between-study variance tau squared.

A common prediction interval has the general form:

pooled effect plus or minus a critical value multiplied by the square root of the uncertainty in the pooled mean plus the estimated between-study variance.

The exact formula and critical value depend on method. On a log risk-ratio or log odds-ratio scale, the interval is calculated on that scale and then transformed back, and this matters because the interval you finally see is asymmetric on the ratio scale. The tau-squared term usually makes a prediction interval wider than the confidence interval for the mean, and if estimated heterogeneity is zero the two may be similar, although uncertainty in the heterogeneity estimate and the choice of method still matter.

A simple interpretation example#

Suppose a meta-analysis reports a pooled risk ratio of 0.80 with a 95 percent confidence interval from 0.74 to 0.87. That supports an average relative reduction under the model. Now suppose its 95 percent prediction interval runs from 0.58 to 1.11.

The prediction interval says that the underlying effect in a new similar study could plausibly range from substantial benefit to little benefit or possible harm, but it does not cancel the average finding. It changes the claim from “the treatment works everywhere” to “the mean favors treatment, but effects may differ meaningfully across settings.”

The clinical interpretation should also consider absolute risk. The same relative effect can produce different absolute benefits at different baseline risks. A prediction interval for a relative effect does not directly state the range of absolute outcomes.

What “a new similar study” really means#

The phrase does substantial work. A prediction interval does not cover every country, age group, dose, outcome definition, or future era. It assumes the new study belongs to the same conceptual population of studies represented by the synthesis.

If the meta-analysis mixes fundamentally different interventions or comparisons, the interval may summarize a collection that should not have been pooled, and if it includes only tightly selected trials, the interval does not automatically transport to routine practice. If standards of care change, a study five years later may not be exchangeable with the older evidence; authors should define the intended universe of settings and explain why the random-effects distribution is scientifically meaningful. Statistical heterogeneity cannot repair clinical incoherence.

Why I-squared is not a substitute#

I-squared describes the proportion of observed variation associated with between-study heterogeneity rather than sampling error under the meta-analytic framework, but it does not give the magnitude or clinical consequence of that heterogeneity.

I-squared can be high when studies are very precise even if effects differ by a modest amount, and lower when studies are imprecise despite important variation, and it also depends on within-study standard errors. A prediction interval works on the effect scale, which lets you compare its bounds with no effect and with clinical thresholds. Tau squared supplies an absolute estimate of between-study variance on the analysis scale, and even then you still have to translate the range into quantities that matter clinically.

Few studies create fragile intervals#

Estimating between-study variance is difficult when only a handful of studies are available: tau squared may be estimated as zero despite real heterogeneity, or may swing widely based on one study. Methods that treat the estimate as known can understate uncertainty.

The assumed normal distribution of true effects is also hard to examine with few studies. A conventional interval may have poor coverage. Alternative estimators and small-sample adjustments can produce different results, sometimes dramatically.

A report should name the estimator, the interval method, the number of studies, and the sensitivity to reasonable alternatives. With very few studies, a numerical prediction interval may imply more knowledge than the evidence supports. A transparent narrative about plausible effect modifiers may be more honest.

A wide interval is information, not a failed analysis#

You may be tempted to read a wide prediction interval as an inconvenient result best left out. In fact, it may reveal the most decision-relevant feature of the synthesis: uncertainty about transfer across settings.

The next question is why effects differ. Prespecified subgroup analysis or meta-regression may examine dose, baseline risk, duration, intervention fidelity, risk of bias, or outcome definition. Such analyses need caution because there are usually few studies per characteristic, study-level associations can be confounded, and multiple exploratory analyses create chance findings. A narrower interval after subgrouping is credible only if the grouping has clinical rationale, adequate information, and validation. Dividing studies until heterogeneity disappears is not an evidence-based solution.

Risk of bias and publication bias remain outside the formula#

A prediction interval quantifies modeled heterogeneity and sampling uncertainty. It does not automatically account for biased randomization, missing outcomes, selective analysis, or publication bias. If all available studies overestimate benefit, the interval can be centered on the wrong value.

Small-study effects may also distort both the pooled mean and heterogeneity; funnel plots and statistical tests have limited power with few studies, the same situation in which prediction intervals are most fragile. Risk-of-bias judgments and sensitivity analyses remain essential.

Prediction for an effect is not prediction for a patient#

The interval describes an underlying study effect, usually a relative or mean difference for a population under a study design. It does not tell an individual patient the range of personal treatment response. Variation among people within each study is not the same as variation in average effects between studies.

Nor does the interval predict the observed estimate that a future study will report. A future observed estimate has additional sampling variation. Authors should state which predictive target their method uses.

When a prediction interval is most useful#

It is particularly informative when a random-effects synthesis contains enough studies, clinical diversity is expected, and a decision depends on whether benefit persists across settings. It can help guideline panels avoid treating an average as universal and can guide research toward effect modifiers.

It may be less useful when studies are too few, pooling is not clinically sensible, heterogeneity cannot be estimated with any stability, or the analysis uses a fixed-effect model because the target is limited to the included studies. Omitting it should be explained rather than hidden.

A reading checklist#

When a meta-analysis presents a prediction interval, check:

Then read the pooled confidence interval and prediction interval together. One summarizes the average; the other tests how far that average can be generalized within the modeled study population.

Sources and further reading

  1. Cochrane Handbook, Chapter 10, Analysing Data and Undertaking Meta-Analyses
  2. IntHout and colleagues, Plea for Routinely Presenting Prediction Intervals in Meta-Analysis, BMJ Open (2016)
  3. Riley and colleagues, Interpretation of Random Effects Meta-Analyses, BMJ (2011)
  4. Cochrane Handbook, Chapter 15, Interpreting Results and Drawing Conclusions

Questions and answers

Is a prediction interval always wider than a confidence interval?

Usually in a random-effects meta-analysis because it adds between-study variation. With an estimated tau squared of zero or unusual methods, they may be very similar.

Does crossing the null mean the treatment does not work?

No. It means the model allows underlying effects on both sides of no effect in similar future studies. The mean effect, evidence quality, absolute outcomes, and reasons for heterogeneity still matter.

Can a prediction interval be trusted with three studies?

It can be calculated, but estimation of heterogeneity and coverage are often unstable. The number should be presented with strong caveats and sensitivity analysis. A prediction interval makes heterogeneity visible on the outcome scale. Its greatest contribution is intellectual discipline: it prevents a precise average from being mistaken for a universal effect.