A propensity score is the estimated probability that a person receives the treatment or condition being compared, given a set of measured baseline characteristics. Researchers use that probability to construct groups that are more comparable before estimating an outcome difference.
The method can be valuable when randomization is impossible, but its promise is often overstated. A propensity score does not turn an observational study into a randomized trial. It balances only variables that were measured and modeled adequately, only among people with enough overlap to support a comparison, and only for the target population implied by the chosen method.
Define the causal question before fitting a model#
Consider a study comparing outcomes in people who received treatment A with outcomes in people who did not. Treatment choice may depend on age, disease severity, kidney function, prior events, frailty, clinician preference, and access to care. Those same factors may affect the outcome. A crude outcome difference therefore combines the treatment effect with baseline differences.
Before choosing a statistical method, the study needs a clear target trial: who would have been eligible, what treatment strategies are compared, when follow-up begins, which outcome is measured, and over what period. It also needs an estimand, the exact effect the analysis seeks.
Common targets include:
- the average treatment effect, or ATE, for the full eligible population;
- the average treatment effect among the treated, or ATT;
- the average treatment effect among people whose treatment could plausibly vary, sometimes approached through overlap weighting.
These are not interchangeable. A method can be technically correct and still answer a different population question from the one you assumed when you read the abstract.
What the score represents#
For each person, the propensity score can be written as:
P(treatment = 1 | measured baseline covariates)
It compresses a vector of baseline variables into one estimated probability. Two people with similar scores had a similar modeled chance of receiving treatment, even if their individual characteristics are not identical. Within groups of people with the same true propensity score, the measured baseline covariates should be distributed similarly between treatment groups.
The word “true” matters. Researchers only have an estimated score from a chosen model. Omitted nonlinear relationships, missing interactions, measurement error, and poor data can leave imbalance. The credibility of the analysis therefore rests on what happens after the score is used, not on the elegance of the equation used to produce it.
Four common ways to use propensity scores#
Matching#
Each treated person is paired with one or more untreated people who have a similar score. Calipers limit how far apart matches may be. Matching can make the comparison intuitive and remove people with no comparable counterpart.
Its target often resembles the effect among treated people who could be matched. A paper will show how many people were excluded and who was left, and if many high-risk treated patients have no match, a clean-looking matched cohort may answer a narrower question than the original study.
Inverse probability weighting#
Weighting creates a pseudo-population by assigning more influence to people whose observed treatment was less probable, and for an ATE analysis, treated people commonly receive weights related to 1/score, while untreated people receive weights related to 1/(1-score).
Scores near zero or one generate very large weights. A few observations can then dominate the estimate. Stabilization, trimming, or truncation may improve precision, but each choice changes the effective population and should be reported with sensitivity analyses.
Stratification#
The sample is divided into score groups, often fifths, and treatment effects are estimated within groups before being combined. Stratification is simple but can leave residual imbalance within broad score bands, particularly in the tails.
Covariate adjustment#
The propensity score is included as a variable in an outcome regression. This is compact but depends on the form of the outcome model and may be less transparent than a well-diagnosed design stage. It should not be treated as proof that confounding disappeared.
Variable selection should follow causal logic#
A useful propensity model includes baseline causes of the outcome and variables related to treatment choice that help achieve exchangeable groups, and selection based only on statistical significance or automated prediction can omit clinically important confounders.
Variables measured after treatment starts are usually inappropriate. A post-treatment biomarker, adherence measure, complication, or dose change may lie on the pathway from treatment to outcome. Adjusting for it can remove part of the effect of interest or open a biased path through a collider.
An instrument-like variable that strongly predicts treatment but is unrelated to the outcome except through treatment can worsen precision and amplify hidden bias in some settings, and a variable affected by both treatment and an unmeasured cause of the outcome can also create bias. A causal diagram and subject knowledge are often more useful than a maximal list of every available field.
Missing data require their own plan. Treating “missing” as a harmless category can distort relationships. Complete-case analysis may select a different population. Multiple imputation should respect treatment, outcome, timing, and analysis structure, and uncertainty from imputation should be carried into estimates.
Balance is the core diagnostic#
The propensity model does not need to classify treatment perfectly. In fact, perfect separation can signal that there is no valid comparison for parts of the population. The central question is whether measured baseline covariates are balanced after matching or weighting.
Standardized mean differences compare covariate distributions without being driven by sample size in the way a P value is. Values close to zero are desirable. A threshold such as 0.1 is often used as a practical flag, not a guarantee. Continuous variables should be checked beyond their means because similar means can hide different spread or tails. Categorical levels, nonlinear terms, and important interactions may also need review.
A Love plot can show standardized differences before and after adjustment across many covariates. Good reporting also includes score distributions by treatment group, the number matched or discarded, effective sample size under weighting, weight distributions, and any trimming rule, and it evaluates balance in the final analytic sample using the final weights. If the report offers only a high c-statistic for the score model, it has not shown balance.
Overlap defines where a comparison is possible#
Positivity means every type of eligible person has some chance of receiving each treatment strategy, and if all people with a certain severity level receive treatment and none remain untreated, the data cannot show what would have happened to an untreated counterpart in that region.
Lack of overlap appears as separated score distributions, failed matches, or extreme weights. The proper response may be to narrow the target population, change the estimand, or acknowledge that the question is not supported. Statistical smoothing cannot create counterfactual information where no comparison exists.
Trimming people at score extremes can produce a more stable estimate, but the result then applies to those retained. Authors should describe that population in clinical terms, not only as a numeric score interval.
What propensity scores cannot fix#
Unmeasured confounding#
If frailty, disease severity, socioeconomic constraint, clinician judgment, or another common cause is absent or poorly measured, the score cannot balance it. A table showing excellent measured balance tells you nothing direct about the variables nobody recorded.
Confounding by indication#
Sicker patients often receive more intensive treatment. Detailed baseline measures may reduce this bias, but a coarse diagnosis code rarely captures the reasoning that drove treatment selection.
Time-related bias#
If treatment status is defined using information accrued after follow-up begins, immortal time or selection bias can arise. Aligning eligibility, treatment assignment, and time zero is a design requirement, not a feature the score repairs.
Measurement and classification error#
Incorrect treatment dates, incomplete outcomes, or noisy confounders can bias an otherwise polished analysis. Propensity machinery does not improve source data.
Model fishing#
Trying many matching calipers, trimming thresholds, or covariate sets and reporting only the favorable result introduces analytical selection. Prespecification and a structured sensitivity set make the result more credible.
Stronger analyses make residual uncertainty visible#
Useful sensitivity work includes alternative estimands, matching and weighting specifications, different reasonable trimming rules, and outcome models that adjust again for remaining covariate imbalance. Doubly robust estimators combine a treatment model and an outcome model, but the name does not mean immunity to hidden confounding; their protection depends on at least one model being correctly specified under the other causal assumptions.
Negative-control outcomes or treatments can reveal residual bias when a relationship appears where none should exist. Quantitative bias analysis can show how strong an unmeasured confounder would need to be to explain the estimate. Instrumental-variable or target-trial methods may address different biases, but they bring separate assumptions that need evaluation.
A reader's audit checklist#
- Is the target population, treatment strategy, time zero, outcome, and estimand explicit?
- Were clinically important pre-treatment confounders measured with adequate detail?
- Were post-treatment variables excluded from the propensity model?
- Is overlap shown, and are exclusions or extreme weights reported?
- Is balance demonstrated after adjustment for every important covariate?
- Does the analysis account for matching or weighting when estimating standard errors?
- Are effect estimates given with confidence intervals and absolute as well as relative scales when useful?
- Do sensitivity analyses address model choices and unmeasured confounding?
- Does the conclusion stay observational, or does it claim randomization-level certainty?
Sources and further reading
- Austin, An Introduction to Propensity Score Methods, Multivariate Behavioral Research (2011)
- Austin and Stuart, Best Practice for Inverse Probability of Treatment Weighting, Statistics in Medicine (2015)
- Elze and colleagues, Comparison of Propensity Score Methods, Journal of the American College of Cardiology (2017)
- BMJ, Propensity Score Analysis in Observational Studies (2019)
- Greifer and Stuart, Choosing the Estimand When Matching or Weighting, Observational Studies (2023)
Questions and answers
Is a high propensity-model c-statistic desirable?
Not necessarily. Predictive separation can indicate poor overlap. The goal is covariate balance in a population where both treatment strategies are plausible.
Does matching remove all confounding?
No. It can reduce imbalance in measured baseline variables. Hidden or poorly measured causes can remain.
Is a propensity score always better than conventional regression?
No. Performance depends on the question, data, overlap, models, and diagnostics. Its practical advantage is often the separation of a transparent design stage from outcome estimation.
Why can two valid propensity analyses give different answers?
They may target different populations, retain different participants, use different assumptions, or respond differently to limited overlap. The estimand and analytic population must be compared before treating the estimates as contradictory.