Evidence explainer

Evidence and research methods

The Counterfactual Idea in Causal Inference

A causal effect compares what happened with what would have happened. Only one of those is ever observed, so design and assumptions have to build the other one.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Potential outcomes give the idea notation
  2. Causal effects are contrasts, not properties
  3. From individual effects to population effects
  4. Why association can differ from causation
  5. Randomization builds a comparison in expectation
  6. Exchangeability in observational studies
  7. Positivity asks whether comparisons exist
  8. Consistency requires well-defined actions
  9. Interference challenges one-person-at-a-time thinking
  10. Causal diagrams make assumptions visible
  11. Time zero must align
  12. Selection and missingness create new counterfactual problems
  13. The estimand states the exact effect
  14. Target trials discipline observational work
  15. Negative controls and sensitivity analysis
  16. Transporting an effect to another population
  17. How to read a causal claim
  18. The central discipline
  19. References

Imagine two otherwise identical versions of you at the same moment. In one, you start treatment A. In the other, you do not. The causal effect for you is the difference between the two later outcomes.

Only one version can be observed. Once you start the treatment, the untreated outcome becomes counterfactual: what would have happened under the alternative. This missing comparison is the central problem of causal inference.

Data do not remove the problem by volume. A million treated records still show only treated outcomes for those people, and a study becomes causally informative when its design and assumptions justify using other observations to approximate the missing alternatives.

Potential outcomes give the idea notation#

Let Y(1) mean a person's outcome under action 1 and Y(0) the outcome under action 0. The individual causal effect is Y(1) - Y(0) for a numeric outcome.

If a person receives action 1, we observe Y(1) and not Y(0). If the person receives action 0, the reverse is true. This is sometimes called the fundamental problem of causal inference.

The notation does not solve the problem. What it does is force the question into view, and with it the ambiguity. What exactly is action 1? A prescription, actually taking a medicine, an invitation to a program, surgery by a particular method, or a policy applied at a hospital?

If the action is not well defined, its potential outcome is not well defined. “Exercise” could mean a walking offer, a supervised program, or achieving 150 minutes weekly. Each has a different comparator and different routes to an outcome.

Causal effects are contrasts, not properties#

A medicine does not have one context-free effect. Its effect is always relative to something: no treatment, placebo, usual care, delayed treatment, or another medicine. Change the comparator and the causal contrast changes.

The target population also matters. An average effect among people eligible for a trial may differ from the effect among everyone with a diagnosis, people who would choose treatment, or those treated in routine care.

Time matters too. A 30-day effect may differ from a five-year effect. Early harm and later benefit can coexist. Starting now versus never is different from starting now versus in six months.

The outcome must be precise. All-cause mortality, disease-specific mortality, a biomarker, symptoms, hospitalization, and quality of life are not interchangeable. A composite can change because of its most frequent component even when the outcome patients care about most does not change.

From individual effects to population effects#

Because individual counterfactuals are missing, researchers often target an average treatment effect, and the ATE compares the average outcome if everyone in the target population received action 1 with the average if everyone received action 0.

The average treatment effect among the treated, often called ATT, asks a different question: among people who actually received action 1, what was the average difference compared with what would have happened had those same people received action 0?

An effect may also be conditional on baseline characteristics. A conditional average treatment effect could differ by age, baseline risk, genotype, or disease severity. Reliable heterogeneity estimates require enough data, prespecified reasoning, and protection against multiple-testing noise.

An average can conceal variation. Some people may benefit, some may be harmed, and many may have little change; the data rarely identify each person's causal effect with certainty, even when the group average is estimated well.

Why association can differ from causation#

Suppose people who receive a new treatment have worse outcomes than those who do not. The treatment may be harmful. It may also be preferentially given to people who are sicker, a pattern called confounding by indication.

Now suppose treatment recipients do better. The treatment may work, or recipients may have better access, adherence, health literacy, baseline prognosis, or follow-up. A predictive association includes every pathway connecting treatment and outcome, not only the effect of treatment.

The causal question asks what the same target population would experience under alternative actions. This is why a highly predictive variable can be a poor intervention target. Gray hair predicts age-related risk, but dyeing hair does not reverse aging.

Prediction remains valuable. A model can triage risk without identifying a cause. Trouble begins when predictive importance is translated into “changing this feature will change the outcome” without a causal design.

Randomization builds a comparison in expectation#

In a randomized trial, assignment is determined by chance. Before treatment begins, the groups are exchangeable in expectation: the distribution of both measured and unmeasured baseline causes should be similar apart from random variation.

Randomization does not guarantee perfect balance in one finite trial. It does provide a known assignment mechanism that supports unbiased comparison under proper analysis, and allocation concealment prevents prediction of the next assignment, while blinding can reduce later differences in behavior, care, assessment, and reporting.

Adherence, crossover, loss to follow-up, missing outcomes, and competing events still complicate the effect. An intention-to-treat effect compares assignment strategies, not necessarily the biological effect of perfectly following treatment. The counterfactual framing is what makes that meaningful: you are comparing outcomes if everyone were assigned to one strategy versus the other, while allowing the post-assignment events defined by that strategy.

Exchangeability in observational studies#

Exchangeability means that, within the comparison being made, the treated and untreated groups could stand in for each other's missing counterfactual outcomes; in an observational study, treatment choice is not randomized, so exchangeability generally requires conditioning on sufficient pre-treatment causes of both action and outcome.

Investigators may use stratification, regression, matching, standardization, inverse-probability weighting, or other methods. These techniques adjust measured variables. None can guarantee removal of confounding by a variable that was not measured adequately.

The required assumption is often called no unmeasured confounding or conditional exchangeability. It should be defended with clinical knowledge, data provenance, causal diagrams, negative controls, validation, and sensitivity analysis. It should not be hidden behind a sophisticated estimator.

More covariates are not always better. Adjusting for a consequence of treatment can block part of the effect or create selection bias. Adjustment choices should follow the causal structure and time order.

Positivity asks whether comparisons exist#

Positivity means that each relevant type of person has a nonzero chance of receiving each action being compared, and if no one with a severe contraindication receives a medicine, the data cannot estimate the effect of giving that medicine to such people without unsupported extrapolation.

Structural violations arise when an option is impossible or prohibited. Practical violations occur when it is technically possible but exceedingly rare. Both can produce unstable weights, heavy model dependence, and extreme uncertainty.

Restricting the target population can restore a credible comparison. The resulting estimate then applies to a narrower group. This is often more honest than claiming a broad effect from a region where your data contain no overlap. Plots of propensity-score distributions, treatment frequencies within covariate groups, and weight diagnostics can reveal problems; a final average should not hide from you that much of it was inferred outside observed support.

Consistency requires well-defined actions#

Consistency links the observed outcome under the action actually received to the corresponding potential outcome. It sounds automatic, but it can fail when one label contains meaningfully different versions.

“Usual care” may vary by site. A surgery may differ by technique and operator. A policy may be adopted on paper but implemented with different staffing. A prescription does not specify dose changes or whether medicine was taken.

If versions have different effects, the intervention should be described with enough precision for the causal contrast. Sometimes the variation is intentionally part of a strategy, such as “initiate and manage according to a defined algorithm.” Then the algorithm, allowed adaptations, and comparator should be stated.

Measurement also matters. Misclassifying treatment can blur consistency and induce bias. A pharmacy fill, order, administration, and ingestion are different events.

Interference challenges one-person-at-a-time thinking#

Standard potential-outcome notation often assumes one person's outcome is unaffected by other people's treatment, and that assumption is called no interference and is part of the broader stable-unit treatment value framework.

It is implausible for vaccines, infectious disease, peer programs, hospital policies, environmental changes, and shared clinicians. Treating one person may protect another or consume a resource another needs.

Interference does not end causal inference. It changes the unit and estimand. Researchers may randomize clusters, define coverage levels, model networks, or estimate direct and spillover effects. Ignoring interference can misstate both benefit and harm. A vaccine's total population effect includes more than the direct protection among recipients.

Causal diagrams make assumptions visible#

A directed acyclic graph, or DAG, represents assumed causal relationships with arrows. It can help identify common causes that should be controlled and variables that should not.

A confounder causes both treatment and outcome. A mediator lies on the pathway from treatment to outcome. A collider is caused by two variables. Conditioning on a collider can create an association between its causes even when none existed.

For example, if both severe illness and treatment increase the chance of hospital admission, analyzing only admitted patients can make treatment and severity statistically related through selection, and adjustment cannot be decided from correlation alone.

A DAG is not proof that the arrows are correct. Its value is that it forces investigators to state assumptions you would otherwise never see. Competing plausible diagrams can guide sensitivity analyses and data collection.

Time zero must align#

Many observational biases arise because eligibility, treatment assignment, and follow-up do not begin at the same moment. A treated person may need to survive long enough to receive treatment, while an untreated person's follow-up begins earlier. The guaranteed survival time can create immortal-time bias.

Aligning time zero means identifying eligibility, assigning each person to a strategy based on information available then, and starting outcome follow-up together. A new-user design often helps by comparing people at treatment initiation rather than mixing prevalent users with nonusers.

Time-varying treatment and confounding require additional methods. A laboratory value can influence the next dose and also be influenced by prior dose. Standard adjustment for the latest value may remove part of the treatment effect. Marginal structural models and the parametric g-formula were developed for such structures under assumptions. Calendar time can also confound comparisons when treatments, testing, and background care change. Concurrent comparators are usually more credible than distant historical controls.

Selection and missingness create new counterfactual problems#

Loss to follow-up can make observed outcomes unrepresentative. If treatment and prognosis affect who remains measured, a complete-case comparison may be biased.

Censoring weights, multiple imputation, outcome models, and bounds can help under stated assumptions. No method recovers information without a model when missingness depends on unseen outcomes.

Survivor selection is especially important in studies of late outcomes. Conditioning on survival to a future landmark can select different mixtures of people across treatment groups, and competing events can prevent the outcome from occurring and require a precise estimand rather than automatic censoring. Data availability is itself a selection process: people who use one health system, own a wearable, answer a survey, or have complete laboratory data may differ from the target population.

The estimand states the exact effect#

ICH E9(R1) defines an estimand through five linked attributes: treatment conditions, target population, variable or endpoint, handling of intercurrent events, and population-level summary.

Intercurrent events occur after treatment starts and affect the existence or interpretation of the outcome. Examples include discontinuation, rescue medicine, treatment switching, death before measurement, or surgery after randomization.

Different strategies answer different questions. A treatment-policy strategy uses outcomes regardless of the event. A hypothetical strategy asks what would happen in a world where the event did not occur. A composite strategy incorporates the event into the outcome. A while-on-treatment strategy focuses on outcomes before the event. A principal-stratum strategy targets people defined by potential occurrence of the event.

These are not interchangeable missing-data techniques. They define different causal questions. The official ICH E9(R1) addendum aims to align objective, design, data collection, analysis, and interpretation.

Target trials discipline observational work#

Hernan and Robins recommend asking which randomized experiment an observational analysis is trying to emulate. The hypothetical target trial specifies eligibility, treatment strategies, assignment, start of follow-up, outcome, causal contrast, and analysis plan.

The real-world data then emulate those components as closely as possible. Assignment is not randomized, so adjustment is needed. Yet stating the target trial prevents several avoidable errors: vague strategies, misaligned time zero, selection based on future information, and mismatched follow-up.

Target-trial emulation is a design framework, not a magic label. It does not remove unmeasured confounding, measurement error, or poor overlap. A detailed emulation can still be wrong. Its strength is auditability. The open textbook *Causal Inference: What If* develops exchangeability, positivity, consistency, time-varying methods, and target trials with explicit examples.

Negative controls and sensitivity analysis#

A negative-control outcome is one that the treatment should not plausibly affect but that shares potential bias structures. An observed association can signal residual confounding or measurement problems. A negative-control action can serve a related role.

Sensitivity analysis asks how results change under plausible departures from assumptions, and it may quantify the strength of unmeasured confounding needed to explain an association, vary misclassification, test alternative lag periods, or compare analytic specifications.

Quantitative sensitivity analysis is more informative than saying “residual confounding is possible” at the end. Still, it only examines departures that were chosen. A robust estimate under one sensitivity model may remain vulnerable to another bias; triangulation compares evidence from designs with different likely biases, such as randomized trials, natural experiments, cohorts, and mechanistic studies. Agreement is stronger when the biases would not all push in the same direction.

Transporting an effect to another population#

Internal validity asks whether the study estimated its target effect without material bias. External validity asks whether that effect applies where you are.

Treatment effects may differ with baseline risk, disease severity, co-treatment, implementation, adherence, health-system capacity, or competing causes. A trial can be internally strong and not representative of every patient.

Transport methods reweight or model effect modifiers measured in both the study and target population. They require overlap and correct measurement. No statistical method transports an effect across a modifier that was never recorded or across an intervention that changes meaning.

Generalization should identify the destination: which people, setting, practice pattern, and calendar period. “Real world” is not one population.

How to read a causal claim#

Ask the study to complete this sentence: Among this population, what would the outcome have been over this period if everyone followed strategy A compared with if everyone followed strategy B?

Then examine:

  1. Were both strategies well defined and feasible?
  2. Did eligibility, assignment, and follow-up share a time zero?
  3. Could common causes of action and outcome be measured before action?
  4. Was there adequate overlap for both strategies?
  5. Were mediators or colliders adjusted inappropriately?
  6. How were adherence, switching, rescue treatment, death, and missing data handled?
  7. Did the estimator match the stated estimand?
  8. Were negative controls or sensitivity analyses informative?
  9. Does the target population match the people to whom the claim is applied?

A causal verb such as “reduced,” “prevented,” or “led to” owes you those answers. Statistical significance alone cannot supply them.

The central discipline#

The counterfactual idea turns “did treated people do better?” into a harder and more useful question: “what would these target people have experienced under each well-defined strategy?”

Because one alternative is always missing, credibility comes from design and assumptions. Randomization, exchangeability, positivity, consistency, time alignment, careful handling of later events, and transparent sensitivity analysis each protect a different part of the comparison.

No estimator abolishes the missing counterfactual. Good causal inference makes the substitute comparison clear enough for you to criticize, replicate, and improve.

References#

Questions and answers

Is a counterfactual something imaginary and therefore unscientific?

It is unobserved, not unconstrained. Study design, randomization, measured comparisons, and assumptions support inference about alternative outcomes. Making the missing comparison explicit improves scientific accountability.

Can a randomized trial estimate each person's causal effect?

Usually no. Randomization supports average effects across groups. Each participant still follows only one treatment path, so their individual alternative outcome remains unobserved.

Does adjusting for every available variable remove confounding?

No. Adjustment only addresses measured variables under a correct causal and statistical model. Controlling mediators, colliders, or poorly measured proxies can add bias.

Is target-trial emulation as strong as an actual randomized trial?

Not automatically. It improves design clarity, but observational assignment still requires assumptions about confounding and measurement that randomization can avoid.

Can prediction and causal inference use the same model?

They can use related methods, but their targets differ. Prediction estimates an outcome from available information. Causal inference estimates a contrast under alternative actions and requires causal assumptions.