Evidence explainer

Evidence and research methods

Meta-Regression Explains Differences Between Studies, Not Between People

Meta-regression asks whether effects vary with study-level characteristics. What it finds is observational, and it rarely tells you anything about an individual patient.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. What the regression is actually fitting
  3. A randomized evidence base can yield an observational slope
  4. Ecological bias changes the level of the claim
  5. Heterogeneity is not a license to search
  6. Too few studies create fragile results
  7. Baseline risk needs special handling
  8. Why individual participant data can help
  9. A practical appraisal sequence

Meta-regression can test whether studies with one characteristic report different treatment effects from studies without it. The unit of evidence is the study, however, so a study-level slope should not be translated automatically into a claim about which individual patients benefit.

Key points#

What the regression is actually fitting#

A conventional meta-analysis estimates an average effect and describes variation around it. Meta-regression adds a study-level predictor, sometimes called a moderator, and the predictor might be mean age, average baseline risk, treatment intensity, follow-up length, year of publication, or a feature of study methods.

Each study contributes an effect estimate, such as a log risk ratio, and a value for the predictor; the model estimates a slope: how much the reported effect changes as the study characteristic changes. More precise studies generally receive greater weight. In a random-effects model, residual between-study variation also enters the weighting and uncertainty calculation, and this resembles ordinary regression, but the data set may contain only a dozen rows, because each study supplies one row. Pooling thousands of participants does not give you thousands of observations.

A randomized evidence base can yield an observational slope#

Within a well-run trial, treatment assignment is randomized. Across trials, mean age, dose, setting, follow-up, and risk of bias are not, because a trial with an older average population might also use a different treatment schedule, recruit in a different era, and define outcomes differently.

If trials with older participants show larger effects, age may be the reason. It may also be acting as a marker for one of those other differences. The meta-regression coefficient is observational because the moderator was not allocated across studies.

This is the core distinction:

Calling the component studies randomized controlled trials does not remove confounding from the between-study comparison.

Ecological bias changes the level of the claim#

Ecological bias, also called aggregation bias, arises when group summaries hide or distort individual relationships. Mean age is not the same as each participant's age: a study with a mean age of 70 can contain a wide mixture of ages, while a study with a mean of 60 can overlap substantially.

Consider four trials. The two trials with older mean ages used a more intensive protocol and reported larger effects. A study-level regression may show larger benefit as mean age rises. Within every trial, however, older and younger participants may respond similarly. The apparent age relationship came from protocol differences between trials, not age differences between people.

The direction can even reverse. That is why a meta-regression based on trial means cannot establish that an older person will respond differently from a younger person; what it can establish is that older-skewing studies reported different estimates, and then leave you with the question of why.

Binary summaries have the same problem. The percentage of participants with a condition is a study property; a relationship between that percentage and effect size is not the same as a treatment-by-condition interaction measured within each trial.

Reviewers often start meta-regression after observing a large heterogeneity statistic, and that sequence can encourage a broad search for an explanation. With enough candidate moderators, one will appear associated by chance.

Heterogeneity statistics also need care. The I² statistic describes the proportion of observed variation attributed to between-study inconsistency rather than sampling error under the model, and it does not tell you how clinically important the inconsistency is, and it can behave differently as study precision changes. Tau-squared estimates the amount of between-study variance on the effect scale, but is uncertain when few studies are available. A moderator that reduces tau-squared may be useful, yet an apparent percentage of heterogeneity “explained” can be unstable, so do not read it as a definitive causal R-squared from a large, well-specified data set.

Too few studies create fragile results#

Meta-regression commonly has low power to detect real effect modification and poor control of false positives when models are crowded. Cochrane guidance advises that meta-regression generally should not be considered with fewer than 10 studies, and even 10 may be too few when characteristics are unevenly distributed.

The often repeated idea of roughly 10 studies per examined characteristic is a caution, not a guarantee. Effective information can be much smaller when most studies sit on one side of a binary moderator, predictor values cover a narrow range, or small studies receive little weight.

Model choices matter. Standard errors can be too optimistic with sparse studies. Methods such as Knapp-Hartung adjustments may better reflect uncertainty in some random-effects settings, but no adjustment creates information that is absent. Influential-study checks are essential because one outlying study can determine the slope.

Baseline risk needs special handling#

Relating treatment effect to the control-group event rate looks attractive because baseline risk is clinically meaningful. It is also technically difficult. The control event rate is estimated with error and contributes to some effect measures, creating mathematical coupling. Regression to the mean can then generate a relationship even when true effect modification is absent.

Baseline risk may also summarize many case-mix and care differences at once. Specialized models can address parts of the problem, but a simple plot of effect against observed control risk should be treated as exploratory.

The effect scale changes interpretation too. A constant relative effect naturally produces larger absolute benefit when baseline risk is higher. That does not necessarily mean the relative treatment effect is modified. Check whether the analysis concerns relative or absolute effects before making anything of a risk gradient.

Why individual participant data can help#

An individual participant data meta-analysis obtains records for participants in the component studies. It can estimate a treatment-by-characteristic interaction within each trial and then combine those interaction estimates, which preserves the randomized comparison and separates within-trial information from differences between trials.

Patient-level data are not a universal cure. Availability may be selective, variables may be defined differently, and flexible modeling can still produce chance findings. Yet for a claim that an individual characteristic modifies treatment response, within-trial interaction evidence is generally closer to the question than a regression on study averages.

A practical appraisal sequence#

When a review hands you a meta-regression, ask:

  1. Was the moderator and its direction specified in the review protocol?
  2. Is there a plausible reason it would modify the effect?
  3. How many studies, rather than participants, informed the model?
  4. How many moderators and alternative cut points were examined?
  5. Are moderator values distributed across enough studies?
  6. Does the result survive removal of influential studies?
  7. Is residual heterogeneity still substantial?
  8. Is the conclusion stated at the study level or promoted to the patient level?
  9. Does patient-level interaction evidence support the same pattern?

Meta-regression is valuable when it narrows uncertainty and guides a future test. Its most defensible conclusion is often, “studies with this feature reported different effects.” A claim about who should receive treatment requires evidence at the level of the person.

Sources and further reading

  1. Statistics in Medicine, guidance on conducting and interpreting meta-regression
  2. Statistics in Medicine, quantifying heterogeneity in meta-analysis
  3. Cochrane Handbook, Chapter 10 on meta-analysis and meta-regression

Questions and answers

Is meta-regression the same as subgroup analysis?

They address similar questions. Subgroup analysis places studies into categories, while meta-regression can model a continuous characteristic or several predictors. Both remain vulnerable to between-study confounding when they use study summaries.

Does a significant slope explain heterogeneity?

It identifies an association under the fitted model. It does not prove the moderator caused the heterogeneity, and uncertainty may be understated when there are few studies or many tested predictors.

Can a null meta-regression show there is no effect modification?

No. Sparse meta-regressions often have low power. A wide uncertainty interval may be compatible with clinically important modification in either direction.