Evidence explainer

Evidence and research methods

Quantifying Heterogeneity: What I-Squared and Tau-Squared Each Mean

I-squared gives the share of observed variation attributed to differences between studies. Tau-squared gives the between-study variance on the effect scale. Neither can explain why studies differ.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Start with the forest plot and the question
  2. Cochran's Q is a test, not a measure of importance
  3. What I-squared actually estimates
  4. What tau-squared adds
  5. Why I-squared and tau-squared can appear to disagree
  6. Prediction intervals are often more decision-friendly
  7. Heterogeneity may be clinical, methodological, or statistical
  8. A practical appraisal sequence

Meta-analysis gives an average estimate, but the average is only part of the answer. Trials may estimate genuinely different effects because their participants, interventions, comparators, outcome definitions, follow-up, or methods differ. The name for this between-study variation is heterogeneity. I-squared, tau-squared, and a prediction interval illuminate different aspects of it; none is a stand-alone verdict on whether pooling was sensible.

Start with the forest plot and the question#

Before you interpret a statistic, identify the outcome, effect measure, direction of benefit, and population. A tau-squared of 0.04 has no universal clinical meaning. It means something different for a log risk ratio than for a mean difference measured in points. I-squared also cannot tell whether estimates fall on both sides of a clinically important boundary.

Read the forest plot. Are confidence intervals overlapping? Are estimates consistently beneficial but different in size, or do some suggest benefit and others harm? Is one study unlike the rest? Does the apparent pattern align with dose, baseline risk, setting, follow-up, or risk of bias? A single summary number loses all of that structure.

Heterogeneity is also outcome-specific. The same trials can be homogeneous for mortality and heterogeneous for a symptom scale. It is result-specific, not a permanent property of a collection of studies.

Cochran's Q is a test, not a measure of importance#

Cochran's Q compares each study estimate with a common-effect estimate, weighting more precise studies more heavily, and under the null hypothesis, all studies share one underlying effect and their observed differences arise from sampling error. A large Q relative to its degrees of freedom produces a small p-value and evidence against that hypothesis.

The test is sensitive to information size. With only a few imprecise studies, Q may miss substantial real differences. With many highly precise studies, it can detect variation too small to matter clinically. The conventional 0.05 threshold therefore does not divide coherent evidence from incoherent evidence. A nonsignificant Q is not proof of no heterogeneity, and a significant Q does not quantify its magnitude.

Q remains useful because I-squared is calculated from it and some models use it, but resist the familiar mistake of treating its p-value as your main appraisal.

What I-squared actually estimates#

I-squared is commonly described as the percentage of total observed variation across study estimates attributable to heterogeneity rather than chance; in its familiar form, it is derived from Q and the degrees of freedom and constrained not to fall below zero.

That wording needs care. Sampling error is not literally partitioned from truth with certainty. I-squared is itself an estimate, sometimes highly uncertain. With a small meta-analysis, its confidence interval can span from little inconsistency to very substantial inconsistency.

I-squared depends on within-study precision. Imagine that the underlying effects differ by the same modest amount in two meta-analyses. If one contains small noisy studies, sampling error may dominate and I-squared may be low, and if the other contains enormous precise studies, even the modest real spread may dominate and I-squared may be high. The clinical variation is similar; the ratio of heterogeneity to total observed variation is not.

Thresholds such as 25%, 50%, and 75%, or labels such as low, moderate, and high, are rough descriptions rather than biological cut points. Cochrane advises interpreting I-squared with the size and direction of effects, the strength of evidence for heterogeneity, and uncertainty in the estimate. Reporting “I-squared was 68%, therefore heterogeneity was high” is incomplete.

What tau-squared adds#

Random-effects meta-analysis assumes that included studies estimate a distribution of underlying effects rather than one identical effect. Tau-squared is the estimated variance of that distribution. Tau, its square root, is the corresponding standard deviation.

This is closer to the amount of underlying variation, but it is tied to the effect scale. For a log risk ratio, tau is in log-risk-ratio units. For an unstandardized mean difference, it is in the outcome's units. Tau-squared values therefore should not be compared casually across different outcomes or effect measures.

Estimating tau-squared is difficult when few studies are available. DerSimonian-Laird, restricted maximum likelihood, Paule-Mandel, and other estimators can give different answers, especially with sparse or unbalanced evidence. A point estimate of zero does not prove identical effects; it may reflect limited information. A careful report names the estimator and, where possible, gives uncertainty around tau-squared or tests sensitivity to reasonable alternatives.

The random-effects weighting method matters too. Merely adding tau-squared to each study's variance can yield confidence intervals that are too optimistic in small meta-analyses; methods such as Hartung-Knapp adjustments may better reflect uncertainty in some settings, although their behavior also requires judgment. “Random effects” is not one automatic procedure.

Why I-squared and tau-squared can appear to disagree#

Suppose five small trials have effect estimates spread meaningfully apart, but each confidence interval is wide. The observed scatter contains a large share of sampling error. Tau may indicate appreciable underlying spread while I-squared remains modest.

Now suppose five very large trials produce tightly estimated effects that differ only slightly. Tau may be small on the clinical scale, yet I-squared can be large because sampling error is tiny. Both statistics can be internally consistent: one describes an absolute variance and the other a proportion relative to study precision.

This is why comparing I-squared across reviews can mislead. A higher I-squared is not necessarily a more clinically variable evidence base. The outcome scale, study sizes, and tau estimate must travel with it.

Prediction intervals are often more decision-friendly#

A confidence interval around the pooled random-effects mean answers how precisely the average underlying effect has been estimated under the model. It does not show the full spread of effects across settings.

A prediction interval combines uncertainty in the average with estimated between-study variation to indicate where the underlying effect of a future study might lie, assuming that future setting is exchangeable with those analyzed. If the pooled estimate favors treatment but the prediction interval includes no effect or harm, that is important for transportability and for whatever decision you are trying to make.

Prediction intervals are not promises. They are unstable with few studies, depend on the tau estimator and distributional assumptions, and can be unjustified when studies are not meaningfully comparable. A wide interval can honestly display uncertainty; it cannot diagnose its source.

Heterogeneity may be clinical, methodological, or statistical#

Clinical heterogeneity includes differences in participant risk, disease definition, dose, co-interventions, comparator, adherence, and follow-up. Methodological heterogeneity includes allocation problems, missing data, measurement differences, and selective reporting. Statistical heterogeneity is the observed manifestation after those studies are placed on a common effect scale.

Subgroup analysis or meta-regression may explore prespecified effect modifiers, but the number of studies, not the number of participants, largely limits this work. Testing many explanations invites false discoveries. A difference in statistical significance between subgroups is not evidence that subgroup effects differ; an interaction test is required. Even then, study-level relationships may not represent participant-level effect modification.

Sometimes pooling should be reconsidered. An average of fundamentally different interventions or incompatible outcomes can be precise and meaningless. In other cases, a pooled mean plus a prediction interval is useful precisely because effects vary, and the decision depends on the estimand and on whether an average across the included settings answers the question you actually have.

A practical appraisal sequence#

First, confirm that the studies address a coherent question and that outcome definitions and effect directions match. Second, inspect the forest plot and look for design or population differences that could explain patterns. Third, read Q and I-squared with their uncertainty, not as pass/fail tests. Fourth, locate tau-squared, its estimator, and the effect scale. Fifth, examine a prediction interval if enough comparable studies exist.

Then ask whether sensitivity analyses changed the conclusion: excluding high-risk-of-bias studies, using another tau estimator, changing the effect measure, or addressing an outlier. Finally, translate statistical spread into the clinical scale. The important question is not “Is I-squared high?” but “Could the effect differ enough across plausible settings to change the decision you are making?”

Sources and further reading

  1. Cochrane Handbook, Chapter 10, Analysing Data and Undertaking Meta-Analyses
  2. Higgins and Thompson, Quantifying Heterogeneity in a Meta-Analysis, Statistics in Medicine (2002)
  3. Higgins and colleagues, Measuring Inconsistency in Meta-Analyses, BMJ (2003)
  4. IntHout and colleagues, Plea for Routinely Presenting Prediction Intervals in Meta-Analysis, BMJ Open (2016)
  5. Riley and colleagues, Interpretation of Random Effects Meta-Analyses, BMJ (2011)

Questions and answers

Does an I-squared of 0% prove that every study has the same true effect?

No. It means the method estimated no inconsistency beyond sampling error, often with considerable uncertainty. A small or imprecise meta-analysis can have 0% I-squared even when important heterogeneity cannot be excluded.

Is a random-effects model the fix for high heterogeneity?

No. It models a distribution of effects and changes the weighting and uncertainty. It does not make incompatible studies comparable, remove bias, or explain why effects differ.

Can tau-squared be compared between two meta-analyses?

Only cautiously when the effect measure and outcome scale are the same and the analyses are otherwise comparable. Tau-squared on a log odds-ratio scale cannot be directly compared with tau-squared for a mean difference.

What should a useful meta-analysis report?

At minimum, the forest plot, model and estimator, pooled effect with confidence interval, tau-squared, I-squared with appropriate context, and, when justified, a prediction interval. Clinical and methodological sources of variation should be discussed separately.