“What is the treatment effect?” sounds like one question. In a clinical trial, it can hide several.
Should outcomes count after participants stop assigned treatment? What if they switch to the comparator, start rescue medicine, undergo surgery, or die before measurement? Is the goal the effect of assignment in practice, the effect if everyone could remain on treatment, or the effect among people who would tolerate either option?
Each question can be scientifically valid. They are not the same. The estimand framework pins the target effect down so that design, data collection, analysis, and interpretation all answer one coherent question, and so that you can tell which one you have been handed.
Estimand, estimator, and estimate#
An estimand is the precisely defined quantity of interest. For example: the difference between randomized groups in mean symptom change at 24 weeks, in all randomized adults, regardless of treatment discontinuation or rescue therapy.
An estimator is the rule or method applied to data. It might be a difference in sample means, a regression model, multiple imputation, inverse-probability weighting, mixed model, or survival model.
The estimate is the numeric result, such as a mean difference of minus 3.2 points with a 95% confidence interval.
Confuse the three and you get false debates. Two analysts can use different estimators for the same estimand. One estimator can also target different estimands under different assumptions. A precise method cannot repair an ambiguous target. The ICH E9(R1) addendum begins with the clinical question, then aligns the rest of the trial to it.
Why intention-to-treat was not enough#
The intention-to-treat principle generally analyzes participants according to randomized assignment. It preserves the comparison created by randomization and protects against selecting participants based on events after assignment.
Yet the phrase does not fully define how outcomes after intercurrent events are handled. If a participant stops medicine, should the trial continue collecting outcomes? If data are not collected, does the analysis impute a world in which treatment continued or estimate routine practice after stopping?
Both analyses can be labeled intention-to-treat in ordinary writing while targeting different effects. The estimand framework does not discard randomization or intention-to-treat. It asks authors to state the actual treatment strategy and the meaning of post-randomization events.
This matters because stakeholders want different information. A patient may ask what happens after choosing a treatment, including possible stopping. A regulator may also want the effect attributable to treatment under a particular adherence scenario. Both need explicit definitions.
The five attributes form one sentence#
The five attributes can be assembled into a sentence:
Among which population, what is the effect of which treatment condition compared with which alternative on which variable, when each intercurrent event is handled in a specified way, summarized by which population-level measure?
Every term constrains the others. A population defined by a post-treatment event may require a principal-stratum strategy. Death can make a symptom score undefined, changing the variable. A risk difference and hazard ratio summarize different aspects of the same event data. Put a table of estimands in the protocol and one vague objective can no longer support several incompatible analyses.
Attribute 1: treatment conditions#
Treatment conditions include the intervention and comparator being contrasted. They may be simple assignments, such as medicine A versus placebo, or strategies that allow dose adjustment, background care, adherence support, or rescue treatment.
The label “usual care” is insufficient if usual care varies substantially by site. The protocol should describe what care can occur and whether those variations are part of the condition.
Assignment to treatment is different from actually receiving every dose. A treatment-policy estimand may compare assignment strategies despite discontinuation. A hypothetical estimand may compare a world in which a specified event did not occur. Describe the treatment in enough detail that someone reading the result can interpret it and someone repeating the trial can reproduce it.
Attribute 2: population#
The population identifies the people targeted by the clinical question: it can be the full eligible trial population, a baseline-defined subgroup, or a principal stratum defined by potential post-randomization behavior.
Baseline subgroups such as disease stage, biomarker status, or prior treatment are observable before assignment. They should be scientifically motivated and adequately represented.
A principal stratum is more complex. It might include people who would survive to week 24 under either treatment or who would adhere under either assignment, and membership depends on potential outcomes under both conditions and is not fully observed. This creates strong identification assumptions. “Effect among people who tolerated treatment” based only on observed tolerators can break randomization because tolerability after assignment may be related to prognosis.
Attribute 3: variable or endpoint#
The variable is what would be measured for each participant to address the question: symptom change, event by a fixed time, time to hospitalization, biomarker value, quality of life, or a composite.
The definition includes timing, measurement method, adjudication, and how events are combined. “Cardiovascular outcome” is not enough. A composite of death, infarction, and hospitalization has different meaning from death alone.
Some intercurrent events can be incorporated into the variable, and if death is assigned the worst rank in a symptom endpoint, death becomes part of a composite variable strategy rather than missing data. Whatever is done, the variable has to keep meaning something to a patient: an endpoint that is mathematically convenient but cannot hold benefit and harm together coherently may be a poor choice.
Attribute 4: intercurrent events and their strategies#
An intercurrent event occurs after treatment initiation and affects either the interpretation or existence of the outcome measurement. It is not simply any post-baseline event.
Common examples include:
- Discontinuation of assigned treatment.
- Switching to another treatment.
- Use of rescue medicine.
- Dose change outside the assigned strategy.
- Surgery or another major procedure.
- Terminal event such as death before a planned measurement.
Missing data is not itself an intercurrent event. It is absence of a measurement. An intercurrent event may cause missingness, but the event should first be handled in the estimand and the missing value then addressed by an estimator, and different events in the same trial can use different strategies. Rescue medicine might use a treatment-policy strategy while discontinuation for toxicity uses a composite or hypothetical strategy.
Attribute 5: population-level summary#
The summary compares outcome distributions between treatment conditions. Examples include difference in means, risk difference, risk ratio, odds ratio, hazard ratio, restricted mean survival time difference, median difference, or odds of a favorable ordinal outcome.
The choice affects interpretation. A hazard ratio compares instantaneous event rates under proportional-hazards assumptions and is not a direct ratio of median survival, and a risk difference at 52 weeks provides an absolute effect tied to one time point.
For recurrent events, the summary might be a rate ratio or mean cumulative count. For clustered trials, it may target an individual-average or cluster-average effect. Endpoint and summary are chosen together, from the question and the expected data, and they are chosen before anyone has seen a result.
Treatment-policy strategy#
Under a treatment-policy strategy, the occurrence of the intercurrent event is considered part of what happens after assignment. Outcomes are used regardless of the event.
Suppose participants stop a medicine and receive other care. A treatment-policy effect can compare the outcome of assignment to the medicine strategy versus comparator in that trial context, including stopping and subsequent care.
This can approximate a practical “what happens if this treatment is chosen?” question. It requires collecting outcomes after discontinuation. If post-discontinuation data are mostly missing, the estimand may remain practical in words but depend heavily on extrapolation. The effect can be diluted by poor adherence, but “dilution” is not automatically a defect. It may be part of the real assignment strategy being evaluated.
Hypothetical strategy#
A hypothetical strategy asks what the outcome would have been in a world where a specified intercurrent event did not occur.
Examples include no treatment interruption caused by a drug shortage, no use of prohibited rescue medicine, or no discontinuation for a reason unrelated to treatment effect. The hypothetical world must be clinically meaningful and sufficiently defined.
Observed outcomes after the event may not answer that world. The estimator must model unobserved outcomes under assumptions about trajectories and reasons for the event.
The strategy can isolate a biological or attributable effect, but it can also create an unrealistic question: asking what would happen if nobody stopped a poorly tolerated medicine can ignore a central feature of the treatment.
Composite-variable strategy#
A composite strategy incorporates the intercurrent event into the endpoint. Treatment failure due to rescue medicine might be counted as nonresponse. Death can be included in a ranked composite with symptom outcomes.
This makes the event visible rather than imputing through it. The result depends on how components are combined. A responder endpoint can treat a small symptom miss and death as the same nonresponse, losing severity information.
Composite construction should reflect clinical priorities and avoid double counting, and look at the component results too, because the overall effect may be driven by the most frequent and least important one. The strategy earns its place when the intercurrent event is itself the failure of the treatment objective.
While-on-treatment strategy#
A while-on-treatment strategy focuses on outcomes before an intercurrent event, such as discontinuation. It answers what happens while participants remain on assigned treatment.
This can be useful for adverse events that occur during use or for an effect that is only meaningful while therapy is taken, but it does not answer the longer-term effect of choosing treatment when discontinuation is common.
Censoring at discontinuation can bias the estimate if the reason for stopping is related to the unseen outcome. Appropriate methods require assumptions about informative censoring. The time window also differs across people. A summary of “last value before stopping” can compare measurements taken at different disease stages and should not be mistaken for a common fixed-time endpoint.
Principal-stratum strategy#
A principal-stratum strategy targets the effect in people defined by how an intercurrent event would occur under both treatments; for example, the effect among people who would adhere under either assignment or survive under either assignment.
This can answer an important scientific question, especially when an outcome exists only for survivors. The difficulty is that each person is observed under only one treatment, so membership in the joint potential category is unknown.
Estimation requires assumptions, models, or auxiliary variables that cannot be fully verified. An observed subgroup such as actual survivors in each arm is not automatically the same principal stratum. The result can also be clinically selective. A favorable effect among people who would tolerate either treatment does not describe everyone deciding whether to start.
One fictional trial, five questions#
Imagine a 24-week randomized trial of a medicine for breathlessness. Some participants stop because of adverse effects, some start rescue therapy, and some die.
A treatment-policy estimand uses 24-week outcomes after stopping or rescue, comparing assignment strategies. A hypothetical estimand asks what symptoms would have been if rescue therapy had not been needed. A composite strategy counts rescue as treatment failure and incorporates death as the worst outcome.
A while-on-treatment estimand evaluates symptoms before discontinuation. A principal-stratum estimand estimates the effect among people who would survive to week 24 under either assignment.
All five can use the same randomized participants. Each can produce a different number, because each asks a different question. Reporting “the treatment effect” without saying which strategy was used hides that choice from you.
Intercurrent events should shape data collection#
If the primary estimand uses a treatment-policy strategy after discontinuation, the protocol must keep collecting outcomes. Stopping study medicine should not automatically mean leaving study follow-up.
Sites need training to distinguish withdrawal from treatment, withdrawal from particular procedures, and withdrawal from all follow-up. Consent materials should explain what data collection may continue.
Reasons and timing of discontinuation, rescue, switch, or dose modification should be recorded consistently. Otherwise, analyses cannot distinguish toxicity, lack of efficacy, logistics, and unrelated decisions. The estimand also affects visit schedules and endpoint ascertainment. A protocol designed only around on-treatment visits cannot later claim a treatment-policy effect.
Missing data comes after the estimand#
Once the outcome and event strategy are defined, missing data are the values that should exist for that estimand but were not observed.
For a treatment-policy estimand, a participant's week-24 outcome after stopping belongs in the target. If it is missing, the analysis needs an assumption about its distribution. For a hypothetical strategy, the target value may be counterfactual even when an observed post-rescue value exists.
Missing-at-random assumptions condition on observed information. Reference-based multiple imputation may assume that outcomes after discontinuation follow a comparator trajectory. Tipping-point analyses vary unseen outcomes until the conclusion changes. No method creates information without assumptions. What the estimand gives you is the ability to say which value is missing and why that matters.
Sensitivity and supplementary analyses#
A sensitivity analysis targets the same estimand as the main analysis but tests robustness to departures from assumptions, and if the main analysis assumes missing outcomes are predictable from observed data, a sensitivity analysis can impose worse unseen outcomes and find a tipping point.
A supplementary analysis provides additional insight, often by estimating another estimand or using a different perspective; a per-protocol effect beside an assignment-policy effect may be useful, but it does not test the same target.
Labeling matters. Presenting a different question as confirmation of the primary analysis can create false reassurance when numbers agree and apparent contradiction when they do not; prespecify the sensitivity analyses around the assumptions most capable of changing the interpretation.
Estimands and noninferiority trials#
Noninferiority trials are particularly sensitive to intercurrent events. Nonadherence and crossover can make groups look similar, potentially favoring a noninferiority conclusion under an assignment-policy effect.
A per-protocol or adherence-focused analysis can also be biased because treatment adherence is not randomized, and regulators often expect coherent evidence across estimands and analyses rather than a mechanical preference for one population.
The noninferiority margin belongs to the estimand. Historical evidence supporting the margin must refer to a comparable effect, population, endpoint, and treatment context. A trial should not choose whichever analysis happens to declare noninferiority. The question and margin come first.
Death requires special care#
Death can be an outcome, a competing event, an intercurrent event, or part of a composite, depending on the question. A quality-of-life score after death is undefined, not merely missing.
Assigning death the worst score creates a composite or ranked endpoint. A survivor-average analysis describes outcomes among observed survivors but can compare different survivor populations if treatment affects survival.
A principal-stratum effect among people who would survive under both treatments is scientifically coherent but hard to identify, and a while-alive estimand answers another question and may favor a treatment that increases survival with poor function. The best choice depends on patient priorities and the disease. Mortality and quality of life should often be shown separately even when combined.
Estimands in observational evidence#
Although ICH E9(R1) focuses on clinical trials, precise estimands also improve observational studies. A real-world evidence question should define population, treatment strategies, time zero, endpoint, later events, and summary.
FDA's non-interventional study guidance references estimands and asks investigators to address treatment changes, confounding, missing or misclassified data, and surveillance differences.
Without randomization, exchangeability and positivity join the assumptions; a beautifully stated estimand cannot remove unmeasured confounding, but it does put the causal target and the data gaps somewhere you can inspect them. Target-trial emulation and estimand specification are natural partners: the target trial states the protocol, and the estimand states the effect of interest within it.
How to read an estimand in a paper#
Find the objective, then go looking for a structured estimand table in the protocol or the statistical analysis plan. Name each intercurrent event and its strategy yourself rather than accepting a blanket “intention-to-treat” label.
Then ask whether the data collection supports the strategy. Were post-discontinuation outcomes collected? Did rescue use differ across arms? Was death treated coherently? Does the estimator target the stated summary?
Compare the primary result with sensitivity analyses of the same estimand, and treat any other estimand as a different perspective rather than a confirmation, and check that the abstract's wording still matches the primary target. Finally, turn the result into a sentence you could act on. “Assignment to treatment reduced symptoms regardless of later changes” means something different from “among people who could remain on treatment, symptoms would improve.”
The question-first conclusion#
The estimand framework disciplines the phrase “treatment effect.” It requires a trial to name who, which treatments, which outcome, what happens after important later events, and how the population result will be summarized.
That precision changes more than analysis. It determines what data sites collect, how discontinuation is handled, which assumptions are tested, and what the abstract can honestly claim.
The framework does not make difficult choices disappear. It makes them visible early enough for clinicians, patients, statisticians, and regulators to decide whether the chosen effect answers the question that matters.
References#
- International Council for Harmonisation. E9(R1) addendum on estimands and sensitivity analysis.
- FDA. E9(R1) final guidance for industry. 2021.
- International Council for Harmonisation. E9(R1) training material.
- Kahan BC, et al. The estimands framework: a primer. BMJ. 2024.
- FDA. Real-world evidence: considerations regarding non-interventional studies. 2024.
Questions and answers
Is an estimand just another word for endpoint?
No. The endpoint is one attribute. The estimand also defines treatments, population, handling of intercurrent events, and the population-level summary.
Is treatment policy the same as intention-to-treat?
They overlap in many trials, but intention-to-treat alone may not state how outcomes after discontinuation, rescue, switching, or death are handled. The estimand provides that precision.
Is missing data an intercurrent event?
No. An event such as discontinuation may lead to missing data. The estimand first defines the outcome that should count, then the estimator addresses whether it was observed.
Which intercurrent-event strategy is best?
There is no universal winner. The strategy should match a clinically meaningful question. Different events within one trial can require different strategies.
Can a trial have more than one estimand?
Yes. It can have primary and secondary estimands for different stakeholder questions. Their hierarchy and analyses should be prespecified and clearly labeled.