Evidence explainer

Evidence and research methods

What a Subgroup Analysis Shows

The central subgroup question is whether treatment effects differ between groups, not whether one group's p value crossed a threshold and another group's did not.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key takeaways
  2. Start with the question the trial was built to answer
  3. The common error: comparing two p values
  4. What an interaction test does
  5. Prespecification makes the claim easier to trust
  6. Multiple analyses generate patterns by chance
  7. Power for the main effect is not power for interaction
  8. Randomization helps, but smaller groups can still be imbalanced
  9. Continuous variables should usually remain continuous
  10. Forest plots help only when read carefully
  11. Qualitative and quantitative interactions
  12. Meta-analysis adds another layer
  13. A subgroup average is not a personal forecast
  14. A practical credibility checklist
  15. References

A subgroup analysis asks whether an association or treatment effect differs across defined groups, such as age bands, disease severity, biomarker status, or baseline risk. In a randomized trial, the informative comparison is the treatment effect within each subgroup. What matters most is the difference between those effects.

One subgroup having p below 0.05 while another has p above 0.05 does not establish that treatment works in one and not the other. The proper statistical question is an interaction: is the difference between subgroup-specific effects larger than would be expected from sampling variation under a common-effect model?

Key takeaways#

Start with the question the trial was built to answer#

Most trials are designed and powered for an average treatment effect in the full analysis population. Randomization supports a fair comparison of assigned treatments across that population. A subgroup analysis asks a second question: does that average contrast vary according to a baseline characteristic?

That second question can matter. Treatment benefit may depend on a causal pathway, disease stage, baseline event risk, organ function, prior treatment, or ability to tolerate therapy. Absolute benefit can differ even when the relative effect is stable because higher-risk groups have more events to prevent. Safety can also vary across groups.

But the data available for the second question are thinner. Splitting 1,000 participants into two groups leaves roughly 500 per group, and an uneven or rare subgroup can be much smaller. The test must detect a difference between two noisy effect estimates, not merely whether one estimate differs from a null value, so an analysis can be scientifically important and still not be automatically reliable. That is what the protocol and the statistical analysis plan are for: they should state which effect modifiers matter, why, how they are defined, which effect scale will be used, and how multiplicity will be handled.[1][2]

The common error: comparing two p values#

Imagine you are handed an estimated risk ratio of 0.75 in participants under 65, with p equal to 0.03, and 0.82 in those 65 or older, with p equal to 0.12. It is incorrect to conclude that treatment benefits the younger group but not the older group. The point estimates are similar, and the second group may simply have fewer events or wider uncertainty.

“Significant” versus “not significant” is not itself evidence of a significant difference. A treatment-by-age interaction model directly compares the effects. Its confidence interval and p value address whether the difference between 0.75 and 0.82 is compatible with chance under the model.

The same logic applies when one estimate points toward benefit and the other is near the null. Direction alone is not enough if both intervals are wide. A direct interaction can remain highly uncertain.

The p-value explainer describes why threshold labels compress too much information. A forest plot should show subgroup estimates and intervals plus the interaction result, not invite you to count which horizontal lines cross the null.

What an interaction test does#

In a regression model, an interaction term combines assigned treatment with the subgroup variable. Its coefficient estimates how the treatment effect changes across subgroup levels on the model's scale.

Scale matters. A treatment can have a constant relative risk but different absolute risk reductions because baseline risks differ; that pattern is effect modification on the absolute scale without modification on the relative scale. Odds ratios, risk ratios, risk differences, mean differences, and hazard ratios do not answer identical questions.

The effect scale should match the clinical decision. Absolute effects are often useful for weighing benefit and harm. Relative effects can be more stable across baseline risk, and reporting both can reveal whether an apparent subgroup difference is mainly a consequence of risk rather than a changing biological response.

For more than two groups, an overall interaction test can ask whether any treatment-effect difference exists across levels. Pairwise claims made after an overall signal still need multiplicity and uncertainty addressed.

Interaction tests rely on model assumptions. A linear treatment-by-age term assumes a particular trend across age. Categories assume abrupt changes at their boundaries. Flexible models can reduce an unrealistic shape assumption but need more data and stronger safeguards against overfitting.

Prespecification makes the claim easier to trust#

A prespecified subgroup analysis is documented before treatment assignments are examined in relation to outcomes. The variable definition, cut point or functional form, endpoint, effect scale, model, direction, and priority should be recorded. “Age will be explored” is much weaker than a full analysis specification.

Prespecification reduces the opportunity to choose a favorable group, time point, model, or cutoff after seeing results. It does not make a finding true. A prespecified analysis can still be underpowered, biologically implausible, or contradicted by other evidence.

Prior rationale matters too. A subgroup hypothesized from pharmacology, pathophysiology, or consistent earlier studies is more credible than one selected only because its observed result is striking. The expected direction should be stated. An interaction opposite to the mechanism deserves more skepticism, even if its p value is small.[2][3]

Subgroup variables should usually be measured at baseline. Grouping by something that happens after randomization, such as adherence, response, rescue treatment, or treatment discontinuation, can break the randomized comparison. The post-randomization event may be caused by treatment and share causes with outcome. Specialized causal methods and explicit estimands are needed for those questions.

Multiple analyses generate patterns by chance#

A trial can divide results by age, sex, region, race or ethnicity, baseline severity, biomarker, organ function, prior therapy, comorbidity, and dozens of other factors. Each can be tested for several endpoints and time points. Even if the true treatment effect is constant, some interaction p values will look small in a large search.

This is the subgroup version of multiplicity. Reporting only the striking finding hides the number of opportunities that produced it. You need the planned subgroup list, the complete set examined, and the multiplicity strategy.

Formal adjustment can control a family-wise error rate or false discovery rate, but not every exploratory analysis needs to be converted into a confirmatory claim; another defensible approach is to label exploratory findings, report them completely, and require external confirmation.

Data-driven trees and machine-learning methods can search for heterogeneous treatment effects across many variables and combinations; they can discover useful hypotheses, but reuse of the same data for discovery and estimation exaggerates differences. Sample splitting, cross-validation, regularization, honest trees, and external validation can help. Clear reporting remains essential.

Power for the main effect is not power for interaction#

A study with 90 percent power for its primary average effect can have poor power for interaction. Each subgroup contains fewer participants and events, and the variance of a difference between effects includes uncertainty from both estimates.

An interaction that is clinically important can therefore have a wide interval and p above 0.05. Absence of a small p value is not proof that effects are equal. The interval around the interaction should be compared with the differences that would be large enough to change your decision.

Low power has another consequence: among the subgroup interactions that pass a selection threshold, estimated differences tend to be exaggerated. Selection favors results that were pushed away from the common effect by sampling variation. If treatment-effect heterogeneity is a primary objective, a trial can answer it properly by stratifying enrollment, oversampling important groups, prespecifying the interaction as a powered hypothesis, or pooling compatible trials, and all of that has to happen before data collection rather than after an overall result proves ambiguous.[2]

Randomization helps, but smaller groups can still be imbalanced#

Randomization is preserved when subgroups are defined by genuine baseline characteristics and treatment groups are compared within them. Yet chance imbalances become more likely as subgroup size decreases. Stratified randomization can improve balance for a limited set of important factors, but not for every exploratory category.

Adjusted interaction models can improve precision when baseline prognostic variables are prespecified and measured well. Adjustment should not be chosen only because it makes the interaction more favorable. Report adjusted and unadjusted reasoning clearly.

Missing outcomes can also differ by treatment and subgroup. If one subgroup has greater discontinuation or missingness, the apparent interaction may reflect data availability. Sensitivity analyses should address plausible missing-data mechanisms within the estimand of interest.

Measurement error in the subgroup variable usually weakens or distorts interaction. A biomarker near a cutoff can be misclassified. Assay differences across sites or delayed sample handling can create artificial regional patterns. The variable's measurement quality belongs in the credibility assessment.

Continuous variables should usually remain continuous#

Dividing age at 65, a laboratory value at its median, or risk score into “high” and “low” discards information. The chosen cutoff can create a sharp-looking contrast even when the true relationship changes gradually. Trying many cutoffs and reporting the best one further inflates false-positive risk.

Modeling a continuous interaction can show whether the effect changes smoothly across values. Restricted cubic splines or other flexible functions can capture nonlinearity when the dataset supports them. A graph with uncertainty bands is often more informative than several arbitrary boxes.

Clinically established thresholds can still be useful when they define treatment eligibility, disease categories, or an actionable decision. Their origin should be clear, and continuous sensitivity analysis can show whether the result depends on people just above or below the line.

Avoid assuming that a subgroup boundary identifies two biological types. A 64-year-old and a 65-year-old are unlikely to have categorically different responses solely because the label changed.

Forest plots help only when read carefully#

A subgroup forest plot displays effect estimates and confidence intervals across categories. A useful plot also shows you participant or event counts, subgroup definitions, the overall estimate, and interaction p values.

The plot can mislead you if you read it by checking whether each interval crosses the null. Wide intervals in small groups often create alternating “positive” and “negative” labels even when the estimates are compatible with one common effect. Axis truncation can exaggerate modest differences.

Repeated rows may not be statistically separate. Age categories partition the population, but rows for age, sex, severity, and region overlap. A single participant appears in one row of each subgroup family. Counting favorable rows as votes is invalid.

The overall effect should not be treated as another subgroup result. It is the combined estimate. Subgroup patterns need to explain why treatment effect varies around it, not compete with it for the smallest p value.

Qualitative and quantitative interactions#

A quantitative interaction means the treatment effect differs in magnitude but points in the same direction across groups. For example, absolute benefit may be larger in high-risk participants and smaller in low-risk participants.

A qualitative interaction means the treatment appears beneficial in one group and harmful in another. Such a pattern could change treatment choice and therefore has high clinical importance. It also deserves strong evidence because noise, multiplicity, or bias can produce apparent reversals.

Credibility increases when the interaction is large, prespecified, limited to a small number of hypotheses, supported by an interaction test, consistent across related outcomes, biologically plausible, aligned with the predicted direction, and replicated.[3]

Consistency should be assessed without requiring every study to cross a threshold. Estimates can support the same interaction while individual intervals remain wide. Conversely, one dramatic finding among otherwise similar studies may represent chance or a study-specific problem.

Meta-analysis adds another layer#

Subgroups in meta-analysis can use individual participant data or aggregate study-level data. Individual participant data allow a within-trial treatment-by-characteristic interaction to be estimated consistently across trials. Aggregate meta-regression can be vulnerable to ecological bias: a trial with an older average population does not reveal whether older participants within that trial had a different treatment effect.

Between-trial differences in dose, comparator, outcome, follow-up, and population can masquerade as participant-level effect modification. A meta-regression with few trials is especially unstable.

Combining interaction estimates from multiple trials can improve precision when definitions and analyses are compatible. Heterogeneity in the interaction itself should still be examined. A pooled subgroup result is not automatically transportable to a setting that differs from all contributing trials.

A subgroup average is not a personal forecast#

A subgroup estimate is an average comparison within a category. People in the group still differ in baseline risk, competing outcomes, preferences, adherence, comorbidity, and treatment susceptibility. Broad labels such as “older adults” can hide decades of variation.

Absolute benefit often depends on several prognostic factors at once. Risk-based analyses can be more informative than one-variable splits because they estimate how benefit changes across a multivariable baseline-risk distribution. Prediction models introduce their own calibration and transport questions, but they avoid pretending one characteristic determines all variation.

Subgroup evidence can identify a population for further study or support a differentiated decision when credible, and it should not turn a noisy average into certainty about the one person you are deciding for. The site's research approach preserves that boundary between a group-level estimate and an individual decision.

A practical credibility checklist#

Ask whether the subgroup variable was measured before randomization and defined in advance. Check whether the direction was predicted and supported by a clinical or biological rationale. Count how many subgroup variables, endpoints, cut points, time points, and models were searched before you were shown this one.

Then read the interaction estimate and confidence interval, not only the subgroup-specific p values. Check the effect scale and whether absolute and relative results tell different stories. Examine subgroup size, events, missing data, and balance.

Finally, look for consistency across related outcomes and other studies. Treat a large, replicated, prespecified interaction differently from one isolated post hoc split. If the result would reverse your practice, the evidential bar should reflect the consequences of being wrong.

References#

  1. Reporting of Subgroup Analyses in Clinical Trials
  2. Subgroup Analysis in Randomised Controlled Trials: Importance, Indications, and Interpretation
  3. Is a Subgroup Effect Believable? Criteria for Credibility

Questions and answers

Does significance in one subgroup but not another prove the effects differ?

No. Two separate threshold labels do not compare effects. A treatment-by-subgroup interaction test directly estimates whether the treatment contrast differs between groups and shows the uncertainty around that difference.

What does an interaction test assess?

It assesses whether the estimated treatment effect changes across subgroup levels on a chosen scale, such as a risk difference, risk ratio, or mean difference. The conclusion can depend on that scale.

Why are post hoc subgroup findings less credible?

Analysts can search many variables, cutoffs, outcomes, and models after seeing results. That creates many opportunities for chance patterns and selective reporting. Such findings are useful hypotheses but usually need confirmation.

Can a trial be powered for its overall effect but underpowered for subgroups?

Yes. Subgroups have fewer participants and events, and the interaction compares two uncertain treatment effects. A clinically meaningful difference can remain imprecise even in a trial with adequate overall power.

Can a subgroup result predict what will happen to one person?

Not directly. It estimates an average effect within a defined group. Individual baseline risk, comorbidity, treatment burden, preferences, and other characteristics can create substantial variation inside that category.