Evidence explainer

Evidence and research methods

When Not to Pool a Meta-Analysis

Statistical software can average almost any set of numbers. The harder judgment is whether the studies estimate a clinically meaningful common quantity and whether an average helps rather than hides.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What a pooled estimate claims
  2. Start with the target question
  3. Clinical diversity can make an average meaningless
  4. Outcomes that share a label may not share a meaning
  5. Methodological diversity can generate false agreement
  6. Risk of bias is not another source of random variation
  7. Statistical heterogeneity comes after clinical judgment
  8. Why I-squared is not a traffic light
  9. Random effects do not repair incompatibility
  10. Directional conflict deserves special attention
  11. Incompatible effect measures and denominators
  12. Duplicate populations can create artificial precision
  13. Publication and reporting bias can shape the pool
  14. What to do instead of pooling
  15. Transparency makes a no-pool decision credible
  16. Certainty is outcome-specific
  17. A practical decision sequence
  18. References

A meta-analysis is a weighted average of study estimates. Done well, it improves precision, tests consistency, and shows the range of evidence. Done mechanically, it can compress different patients, treatments, and outcomes into a number that describes no real clinical question. It can do the same with designs and biases.

The software rarely knows the difference. It will calculate a pooled effect if the analyst supplies compatible-looking numbers. The scientific decision comes first: are these studies similar enough in what they estimate that an average has a coherent interpretation?

Sometimes the correct systematic-review result is a set of carefully organized estimates without a pooled diamond. Choosing not to pool is not a failure when the alternative is more truthful.

What a pooled estimate claims#

A fixed-effect meta-analysis commonly assumes that included studies estimate one common effect and that observed differences arise from sampling error, and its summary gives larger weight to more precise studies under that model.

A random-effects meta-analysis assumes studies estimate different but related effects drawn from a distribution. Its center is an average effect, and between-study variance contributes to uncertainty. Smaller studies usually receive relatively more weight than under a fixed-effect model.

Both approaches make a scientific claim about relatedness. If one study tests prevention and another treatment after disease, their effects are not merely variable versions of one quantity. No variance estimator turns them into the same question.

Start with the target question#

Define the population, intervention, and comparator. Define the outcome and study design, often summarized as PICO. Add timing, setting, effect measure, and estimand. “Does exercise help arthritis?” is too broad for pooling. The dose, type of arthritis, comparator, follow-up, and outcome definition can change the effect.

An estimand states the quantity being estimated, including how intercurrent events such as discontinuation, rescue treatment, death, or crossover are handled; a trial's treatment-policy effect can differ from an effect among adherent participants. Combining them without reconciliation obscures meaning. Grouping rules should be written in the protocol before results are known. Post hoc grouping can be influenced by which combination produces the preferred answer.

Clinical diversity can make an average meaningless#

Patients may differ in disease stage, baseline risk, or age. They may differ in comorbidity, prior therapy, diagnostic criteria, or setting. Interventions may differ in drug, dose, or route. They may differ in intensity, duration, co-intervention, or implementation quality. Comparators may range from placebo to active usual care.

Some diversity is acceptable and can make a summary more generalizable. The question is whether the interventions and effects remain meaningfully related and whether an average would inform a decision. For example, pooling a brief low-intensity counseling program with bariatric surgery under “weight loss intervention” would yield a category average with little clinical use. Separate syntheses or a structured comparison is more defensible.

Outcomes that share a label may not share a meaning#

“Response” can mean a two-point scale change, a 50 percent symptom reduction, clinician judgment, or a composite. “Major cardiovascular event” can include different combinations of death, infarction, stroke, hospitalization, or revascularization.

Different scales can sometimes be combined using a standardized mean difference when they measure the same construct. Standardization does not make depression, anxiety, and quality of life one outcome. It also changes the unit to standard deviations, which can be hard to interpret when variability differs across populations.

Time points matter. A pain effect at two weeks and one at one year may have different mechanisms and value; pooling the nearest reported time from each study can mix acute and durable effects while appearing uniform.

Methodological diversity can generate false agreement#

Randomized trials, uncontrolled before-after studies, and confounded observational comparisons have different bias structures. A similar numerical effect does not mean the same evidentiary strength.

The Cochrane chapter on non-randomized intervention studies emphasizes confounding, selection, and classification. It emphasizes deviations, missing data, outcome measurement, and selective reporting. These concerns vary across designs and analyses. Randomized and observational evidence can inform the same review, but separate synthesis is usually clearer; a combined hierarchy or Bayesian model requires a prespecified rationale and assumptions about bias, not a simple pooled row.

Risk of bias is not another source of random variation#

A random-effects model treats effect differences as a distribution; it does not identify which differences result from bias. If small unblinded studies show large benefits and larger rigorous studies show none, averaging can give the biased studies more relative weight.

Excluding a study solely because its result is inconvenient also introduces bias. Risk-of-bias criteria should be applied without knowing the preferred direction, followed by prespecified sensitivity analysis. When all available studies have critical flaws, a precise pooled number can exaggerate confidence. The synthesis should foreground why the evidence may not identify the causal effect.

Statistical heterogeneity comes after clinical judgment#

The Cochrane Handbook chapter 10 distinguishes clinical diversity, methodological diversity, and statistical heterogeneity. Statistical heterogeneity means observed effects vary more than expected from sampling error alone.

The chi-squared test has low power when studies are few or small and excess power when studies are numerous. A nonsignificant result is not proof of consistency.

I-squared estimates the proportion of observed variability attributed to between-study heterogeneity rather than sampling error under a model, though it does not measure the size or clinical importance of effect differences. Its uncertainty can be substantial, especially with few studies.

Why I-squared is not a traffic light#

An I-squared of 0 percent can occur when imprecise studies lack power to reveal real heterogeneity, and an I-squared of 80 percent can occur when large precise studies differ modestly in magnitude while all favor the same treatment.

Read the point estimates, the confidence intervals, and the directions. Read the between-study variance and the prediction interval. Then ask what plausible clinical cause could produce that spread. What it costs you depends on the decision: variation between a small benefit and a moderate one may matter far less than variation between benefit and harm. A universal cutoff such as “pool below 50 percent” replaces scientific judgment with a noisy statistic. Cochrane explicitly advises against simple thresholds as the decision rule.

Random effects do not repair incompatibility#

The phrase “we used random effects because heterogeneity was high” is incomplete. Random effects estimate an average across a modeled distribution. They do not show why effects vary or make the average relevant to every setting.

With few studies, between-study variance is poorly estimated. Conventional confidence intervals can be too narrow, depending on the method; a prediction interval can show the range where a future true effect might lie, but it too can be unstable. If the distribution plausibly includes important benefit and harm, reporting only its mean is hazardous. If studies estimate different constructs, even a perfectly estimated distribution is the wrong model.

Directional conflict deserves special attention#

Suppose half of studies show benefit and half show harm. An average near zero may be mathematically correct but clinically false for both groups. The pattern may reflect dose, baseline risk, or population. It may reflect implementation, bias, or chance.

Check extraction and coding first, including whether a scale was reversed. Then examine prespecified effect modifiers and design differences. Subgroup analysis should use a direct interaction test rather than comparing whether each subgroup is separately statistically significant. Exploratory explanations generated after seeing the pattern should be labeled hypotheses. When variation cannot be resolved, not pooling may communicate the evidence more accurately.

Incompatible effect measures and denominators#

Odds ratios, risk ratios, risk differences, hazard ratios, and incidence-rate ratios answer related but distinct questions. Some can be converted under assumptions. Careless conversion can distort effects when baseline risk or follow-up varies.

Cluster-randomized and crossover trials require unit-of-analysis adjustments. Multiple intervention arms can double-count a shared control if entered as separate independent comparisons. Repeated outcomes from one cohort are not separate studies.

Rare-event studies with zero events in both groups contribute no information to common ratio measures, but they still describe event scarcity, and the method should follow the event process and estimand, not whichever option produces a result.

Duplicate populations can create artificial precision#

One trial may generate a conference abstract, primary paper, subgroup report, long-term follow-up, and registry analysis. Pooling them as separate cohorts counts participants multiple times.

Overlapping claims or registry datasets can also share patients without stating exact identifiers. Map the recruitment sites, the dates, and the investigators. Map the sample sizes, the baseline characteristics, and the trial registration numbers before you treat two reports as two studies. Choose the most complete report for each outcome and time point, link companion papers, and avoid duplicate weighting. More publications do not necessarily mean more evidence.

Publication and reporting bias can shape the pool#

Meta-analysis summarizes available results, which may exclude unpublished studies, unreported outcomes, or unfavorable time points, and funnel plots and statistical tests have limited power with few studies and can reflect mechanisms other than publication bias.

Search trial registries, protocols, regulatory reviews, dissertations, and conference records as appropriate. Compare reported outcomes with prespecified plans. Missing evidence should affect interpretation and certainty even when the visible studies are statistically consistent. Pooling cannot recover an unknown set of suppressed estimates. Sensitivity analyses can show how conclusions depend on assumptions but do not create the missing data.

What to do instead of pooling#

A systematic review can show forest plots without a summary diamond, grouped by clinically coherent categories. Tables can present effect estimates, confidence intervals, and sample sizes. They can present design, risk of bias, and direction.

The Cochrane chapter on other synthesis methods describes structured approaches when meta-analysis is not possible. Vote counting based only on statistical significance should be avoided because significance depends on sample size and discards effect magnitude.

The SWiM guideline provides nine reporting items for synthesis without meta-analysis, including grouping, standardized metrics, and synthesis method. The items also cover prioritization, heterogeneity, and certainty. They cover data presentation and limitations. It improves transparency but does not prescribe one universal method.

Transparency makes a no-pool decision credible#

State eligibility and synthesis groups in the protocol. Explain exactly why a pooled estimate would be misleading. Show each study's result and uncertainty where possible. Describe which studies contributed to each conclusion.

PRISMA 2020 is a reporting guideline, not a risk-of-bias tool or guarantee of correct methods. Its checklist and flow diagram help readers see search, selection, synthesis, and reporting decisions. Report deviations from the protocol and the reasons for them. A principled no-pool decision reads differently from one made after an unfavorable aggregate appeared on somebody's screen.

Certainty is outcome-specific#

Whether pooled or not, assess certainty for each important outcome. Cochrane chapter 14 describes GRADE considerations: risk of bias, inconsistency, indirectness, imprecision, and publication bias.

Not pooling does not automatically mean very-low certainty. Several consistent, rigorous studies may resist pooling only because they report incompatible effect measures, and a tidy low-heterogeneity meta-analysis, meanwhile, can still carry low certainty because every study in it is biased or indirect. So the conclusion has to give you the range and the certainty, and it must never imply that no pooled estimate means no information.

A practical decision sequence#

First, specify the estimand and synthesis groups. Second, compare PICO, timing, design, and outcome definitions. Third, audit extraction, overlap, and effect-measure compatibility. Fourth, assess bias and missing evidence.

Only then inspect statistical heterogeneity and choose a model, and ask yourself whether the average has a clinical interpretation, whether it is hiding variation that matters, and whether the prediction interval would change what you do.

If the answer is no, use a structured alternative and say why. The purpose of evidence synthesis is not to manufacture one number. It is to represent the evidence faithfully enough to support a judgment.

References#

  1. Cochrane Handbook chapter 10
  2. Cochrane Handbook chapter 12
  3. Cochrane Handbook chapter 14
  4. SWiM reporting guideline
  5. PRISMA 2020 statement
  6. Cochrane Handbook chapter 24

Questions and answers

Does a high I-squared value automatically forbid meta-analysis?

No. I-squared has uncertainty and depends on effect size, precision, direction, and context; the clinical and methodological reasons for variation matter more than a universal cutoff.

Does a random-effects model solve heterogeneity?

No. It models a distribution of effects and estimates an average, but it does not make incompatible studies comparable or remove bias.

Can a systematic review be complete without a pooled number?

Yes. A transparent review can use structured tables, forest plots without a summary, grouped effect estimates, and other prespecified synthesis methods.

Should randomized and observational studies be pooled together?

Usually they should be synthesized separately because their estimands and bias structures differ, unless a clear protocol-based rationale supports a combined model.

What should authors report when they decide not to pool?

State the decision rule and reasons, show study-level effects and uncertainty, group studies logically, assess bias and certainty, and explain the alternative synthesis method.