Most apparent subgroup differences should begin as hypotheses, not treatment rules. The correct statistical question is whether treatment effects differ between groups, tested through an interaction. ICEMAN then adds the design and context needed to decide whether that interaction is credible.
Key points#
- “Significant here, nonsignificant there” is not evidence that effects differ.
- An interaction test directly evaluates whether the treatment contrast changes across subgroup levels.
- Credibility rises when the subgroup variable was measured before treatment, the direction was predicted, few hypotheses were tested, and prior evidence supports the pattern.
- Small subgroups produce unstable estimates, so a dramatic point estimate can coexist with very weak evidence.
- Even a credible effect modifier may need confirmation and an absolute-benefit analysis before it changes decisions.
The comparison that headlines often miss#
Imagine a trial reporting a risk ratio of 0.70 in participants younger than 65, with p equals 0.03, and 0.85 in participants 65 or older, with p equals 0.20. It is tempting to say the treatment works only in younger people.
That conclusion compares each subgroup with a significance threshold, not with the other subgroup. The older group may simply have fewer events and a wider interval. A p value can cross 0.05 in one group and miss it in another even when the estimated effects are compatible.
The appropriate test includes a treatment-by-subgroup interaction term. Its null hypothesis is that the treatment effect is the same across subgroup levels on the chosen scale, and if the interaction is imprecise, the data do not establish a difference, regardless of the separate p values.
The scale matters. Effects may be constant on a relative scale but differ on an absolute scale because baseline risk differs: a constant 20 percent relative reduction prevents more events in a high-risk group than in a low-risk group. That is variation in absolute benefit, not necessarily biological interaction.
Why subgroup findings multiply so easily#
A trial may examine age, sex, disease severity, biomarker status, region, previous treatment, and several comorbidities. Each can be divided at multiple cut points and tested across multiple outcomes. If enough comparisons are run, some will look unusual by chance.
Selective visibility makes it worse. A report can highlight the striking subgroup while leaving the unsuccessful analyses in a supplement, or omitting them. What you see is one apparently focused hypothesis. What you do not see is the search that produced it.
Prespecification reduces this flexibility. It is strongest when the protocol or analysis plan names the subgroup variable, categories, outcome, statistical model, effect scale, and predicted direction before treatment assignments are analyzed; merely listing “subgroup analyses will be performed” is not a full specification.
What ICEMAN contributes#
ICEMAN stands for Instrument to assess the Credibility of Effect Modification Analyses. It was developed for claims made within randomized trials and across meta-analyses. The instrument does not convert a finding into a mechanical pass or fail. It organizes the considerations that should shape an overall credibility judgment.
Was a direction predicted?#
A rationale is stronger when it predicts which group should benefit more and explains why. Saying only that effects “may differ by age” can accommodate any result. Predicting reduced benefit with increasing disease duration because irreversible damage accumulates is more falsifiable, and the rationale should precede the data and ideally draw on clinical, pharmacologic, or prior empirical evidence. A story invented after the result can make almost any pattern sound plausible.
Was the variable measured at baseline?#
Baseline characteristics exist before randomization and cannot be caused by treatment. Post-randomization variables, such as adherence, biomarker response, or an adverse effect, can be consequences of assignment. Conditioning on them can break randomization and introduce selection bias. This does not make post-treatment questions unimportant. It means they require causal methods suited to intermediate variables rather than a routine subgroup comparison.
How many subgroup hypotheses were tested?#
One or two prespecified analyses create a different chance environment from dozens of exploratory comparisons. A credible report states the total set examined, not only the successful one, and multiplicity adjustment can be useful, but ICEMAN treats the number and planning of hypotheses as part of a broader judgment rather than relying on a single corrected p value.
Does the interaction support the claim?#
The interaction estimate and its interval tell you the magnitude and precision of the between-group difference; a very small interaction p value is more supportive than one barely below a threshold, but it cannot repair post hoc selection or an implausible rationale.
For continuous characteristics, preserving the continuous scale is often more informative than splitting at a convenient median; multiple cut points waste information and allow the apparent threshold to be chosen after inspection. Flexible modeling can show whether the treatment effect changes gradually or nonlinearly, with safeguards against overfitting.
Is the pattern supported elsewhere?#
Consistency across related outcomes, trials, and evidence sources can raise credibility when the comparisons are genuinely comparable, but repetition of the same analysis choices within one dataset is not independent confirmation. A subgroup effect that appears in one small trial but disappears in larger studies remains uncertain.
A worked credibility audit#
Suppose a cardiovascular trial claims greater benefit among people with a baseline inflammatory marker above a threshold.
Begin by locating the protocol. If the marker and cut point were specified with a predicted direction, that supports credibility. Next, confirm that the marker was measured before randomization and with the same assay across sites. Count the other biomarkers and cut points tested. Then inspect the treatment-by-marker interaction, not the within-group p values.
Ask whether the relationship persists when the marker is modeled continuously, whether it is driven by a few events, and whether absolute risks are reported. Check related outcomes for coherence without demanding identical significance. Finally, look for comparable evidence from another trial.
You may end up at high, moderate, low, or very low credibility. “Moderate” can be the honest conclusion when the hypothesis was planned and biologically sensible but the interaction remains imprecise.
Credibility is not the same as actionability#
A believable difference in relative effect does not automatically define who should receive treatment. Decisions also depend on baseline risk, absolute benefit, adverse effects, burden, alternatives, and whether the subgroup definition can be measured reliably in routine care.
Conversely, absence of proven interaction does not mean every person receives identical benefit. Trials are often underpowered for effect modification. The overall estimate may be the most stable guide while uncertainty about subgroup variation remains explicit. So the analysis earns its place when it narrows a credible question and points to the trial that would settle it. Its weakest use is turning a chance split into a confident personalized claim.
Sources and further reading
Questions and answers
If a treatment is significant in women but not men, does sex modify the effect?
Not by itself. The analysis must directly compare the treatment effects with an interaction test and consider prespecification, number of analyses, rationale, and precision.
Does a nonsignificant interaction prove equal effects?
No. It may reflect limited information. The interaction interval shows which differences remain plausible.
Is ICEMAN a statistical test?
No. It is a structured credibility assessment that combines statistical evidence with planning, measurement, prior rationale, and consistency.