Observational treatment studies begin in a world where clinicians and patients make choices. A medicine is prescribed because symptoms worsened, a laboratory value crossed a threshold, another treatment failed, a contraindication narrowed the options, or a patient preferred one tradeoff; those reasons are often related to the outcome the study later measures.
That is the setting for confounding by indication. A simple comparison can mistakenly attribute the treated group's starting prognosis to the treatment itself. The distortion can run in either direction. A treatment reserved for the sickest patients may look harmful even when it helps. Preventive therapy given to health-conscious patients with better access to care may look unusually protective even if much of the observed advantage comes from those background differences.
The treatment decision carries prognostic information#
Suppose a database study asks whether an inhaled medicine increases hospitalization among people with lung disease, and clinicians may prescribe that medicine when symptoms, prior attacks, oxygen measures, or rescue-medication use signal greater severity. Those same features raise the chance of hospitalization. If the analysis compares recipients with everyone who did not receive the medicine, the treated group may be on a worse trajectory before the first dose.
The indication does not have to be a formal diagnosis. It can include duration of disease, response to previous therapy, perceived frailty, pregnancy plans, kidney function, adherence history, insurance coverage, or a clinician's judgment that never enters a structured field. Treatment contraindications create a related problem: people who cannot receive a drug may differ systematically from those who can.
Confounding is defined in relation to a specific causal question. A variable can be a confounder for one treatment comparison and not another. Disease severity measured before treatment may confound the relation between treatment and outcome, while a biomarker changed by treatment is instead a mediator and usually should not be handled as if it were a baseline confounder.
Why direction cannot be guessed from the label#
The phrase sometimes suggests that treated people are always sicker, but clinical selection is more varied. Confounding by severity often biases toward apparent harm. Confounding by contraindication can make a treatment appear safer because people at highest risk were steered away from it; healthy-user bias can favor preventive medicines when recipients also exercise, attend screening, and follow other recommendations. Channeling bias occurs when a newer drug is directed toward patients with particular histories.
The net direction depends on which forces dominate and how the outcome is defined. Even careful subject-matter intuition may not predict the magnitude. A study should therefore identify the treatment-assignment process rather than assert that residual confounding must have favored a preferred conclusion.
Negative control outcomes or exposures can help reveal some implausible patterns. If a treatment appears to prevent an outcome that it cannot reasonably affect, shared health behavior or care access may be responsible. A negative control is a diagnostic, not a universal correction; it requires its own assumptions.
Design the observational study as a trial question#
Target-trial thinking asks you to write down the randomized trial you would have run if it were feasible: eligibility criteria, treatment strategies, assignment, time zero, follow-up, outcome, causal contrast, and analysis plan. The observational analysis then tries to align each element.
This exercise prevents several errors that masquerade as confounding. If eligibility is assessed after treatment starts, the groups may be selected using future information, and if follow-up begins before a treated person actually receives treatment, the period in which that person must survive can create immortal-time bias. If treatment status is defined using future persistence, early outcomes can be excluded selectively.
The estimand matters too. An intention-to-treat-like effect of starting treatment differs from a per-protocol effect of continuing it. The latter needs methods for time-varying treatment and confounders affected by earlier therapy. Merely excluding people who stop a medicine can introduce selection bias.
New users make baseline interpretable#
A prevalent-user design includes people already taking a treatment. Those users have survived, tolerated, and perhaps benefited from earlier use. Early adverse events and early discontinuations are missing. Their measured health at study entry may already have been changed by therapy.
A new-user design identifies treatment initiation and measures covariates before that moment. It gives the comparison a visible time zero and captures early events. A washout period can help distinguish new use from ongoing use, though the period must be suited to the medicine and available data.
New-user design does not itself make treatment random. The choice to start one therapy rather than another remains confounded. It does make the assignment process easier to study and reduces biases created by conditioning on successful past use.
Active comparators create a fairer clinical fork#
Comparing a treatment starter with a person receiving no treatment often contrasts different disease stages, participation in care, and prescribing thresholds; an active comparator is an alternative intervention used for the same clinical purpose at the same point in care.
For example, comparing initiators of two second-line medicines can be more credible than comparing initiators of one medicine with every person who has the diagnosis, because both groups have reached a decision to intensify therapy. The design can align disease activity, prior treatment failure, healthcare contact, and willingness to take medicine.
Comparator choice requires clinical knowledge. Two drugs may share an indication on paper but be used in different ages, kidney-function ranges, or stages of illness. Restriction to patients eligible for either option can improve exchangeability, but it may narrow the population to which results apply.
Measure confounders before the decision#
Good covariate measurement reconstructs what informed treatment selection. Relevant domains can include disease severity, trajectory, prior outcomes, comorbidities, previous medicines, laboratory results, frailty, healthcare use, clinician or site, socioeconomic constraints, and calendar time.
Timing is critical. A laboratory value obtained after treatment starts may reflect treatment response. A diagnosis recorded during follow-up may have been discovered because one group received closer monitoring. Covariates should be anchored to a prespecified baseline window, with attention to whether missingness itself signals care patterns.
Proxy variables can help when severity is not recorded directly. Recent emergency visits, escalating doses, specialist encounters, and diagnostic testing may capture parts of clinical concern. Proxies can also create noise or open biased pathways. Their inclusion should follow a causal model and domain reasoning rather than an automatic search for variables associated with the outcome.
What regression and propensity scores actually do#
Outcome regression estimates treatment association conditional on included covariates. Propensity scores summarize the estimated probability of receiving treatment given measured baseline variables. Researchers can match, stratify, weight, or adjust using the score.
Neither method discovers unrecorded clinical judgment. Both rely on no important unmeasured confounding, correct time ordering, adequate model specification, and a positivity condition: patients with a given covariate pattern need a meaningful chance of receiving either treatment.
Balance after weighting or matching should be shown with standardized differences and distributions, not only a model's c-statistic. A propensity model can predict treatment well yet leave important imbalance. Conversely, extreme prediction can reveal poor overlap. Large weights from near-certain treatment choices can make an estimate unstable and target a thinly represented hypothetical population.
Covariate selection should prioritize causes of the outcome and treatment assignment measured before initiation. Selecting only variables with a statistically significant treatment association can omit important prognostic factors. Adjusting for an instrumental variable that strongly predicts treatment but has no relation to outcome except through treatment can sometimes increase bias from unmeasured confounding and reduce precision.
High-dimensional methods still need a causal design#
Large claims and electronic-record datasets permit automated identification of diagnostic, procedure, and prescription codes. High-dimensional propensity-score approaches can select proxies for unmeasured clinical characteristics. Machine-learning models can represent nonlinearities and interactions.
These tools improve modeling capacity, not identification by decree. They cannot recover a severity measure that leaves no trace in the data. They may learn healthcare processes that differ across sites or time. If treatment and outcome definitions leak future information, a flexible model can encode the error more efficiently.
Separate design from outcome inspection where possible, document your variable construction, evaluate overlap and balance, and test transportability. Reproducible code lists and data provenance matter because an opaque phenotype can hide which parts of the clinical decision were captured.
Sensitivity analyses should challenge the conclusion#
A useful sensitivity analysis changes an assumption that could plausibly explain the result. You might use a narrower active comparator, vary the new-user washout, restrict to people with complete baseline testing, add disease-trajectory variables, apply negative controls, or quantify how strong an unmeasured confounder would need to be.
Lagging the treatment definition can reduce reverse causation for some questions, but arbitrary lags can exclude genuine early effects. Quantitative bias analysis requires transparent ranges for confounder prevalence and associations. An E-value summarizes one dimension of unmeasured confounding but does not address selection, measurement error, positivity, or time-related bias.
Convergence across different designs is more persuasive than repeated versions of one model. Randomized evidence, natural experiments, self-controlled designs, instrumental-variable analyses, and active-comparator cohorts each have different weaknesses. Agreement can strengthen a conclusion when the biases are genuinely distinct.
Reporting that lets readers audit the comparison#
RECORD-PE extends reporting guidance for studies using routinely collected health data. A clear report names the data source, coding algorithms, treatment episodes, washout, eligibility, follow-up, censoring, confounder definitions, missing-data handling, and analysis population.
A baseline table should show clinically meaningful severity and healthcare-use measures before and after weighting or matching. Authors should report exclusions, overlap, weight distributions, attrition, event counts, absolute risks, and uncertainty. A relative estimate without event rates can conceal whether a large ratio concerns a rare outcome.
The discussion should identify likely residual confounders and their possible directions. Phrases such as “associated with” are usually appropriate unless the study's assumptions and design support a causal contrast, so watch which verb you are being handed. No adjustment method deserves a causal verb solely because it is sophisticated.
A reader's worked audit#
Begin with the clinical fork: why would one patient receive treatment A and another receive B or nothing? List those reasons before turning to the regression table. Then ask whether the database measures them before initiation.
Check that eligibility, treatment assignment, and follow-up share one time zero. Look for new users and a comparator that represents the same decision stage. Inspect whether groups overlap and whether balance is reported for the variables that matter clinically, not only billing-code counts.
Finally, compare the main result with absolute outcomes and sensitivity analyses. If a small modeling change reverses the estimate, the conclusion is fragile and you should say so. If several well-motivated designs converge, the evidence is more reassuring, but uncertainty about unmeasured clinical judgment remains.
Confounding by indication is not a reason to discard all real-world evidence. It is a reason to build the comparison around how care is actually chosen. The best observational studies do not pretend that treatment assignment was random. They make the treatment decision visible, align the groups as closely as the data allow, and state what uncertainty survives.
References#
- RECORD-PE reporting guideline for pharmacoepidemiology
- ENCePP methodological guide for pharmacoepidemiology
- Using big data to emulate a target trial
- The active-comparator new-user study design
- Variable selection for propensity-score models
- FDA framework for a real-world evidence program
It does not provide medical advice or determine whether a treatment is appropriate for any person.*
Questions and answers
What is confounding by indication?
It occurs when the clinical reason for prescribing a treatment also predicts the outcome. The treated and comparison groups then differ in prognosis before treatment, so an unadjusted association mixes that difference with any treatment effect.
Can statistical adjustment remove the problem completely?
Not necessarily. Regression, matching, and weighting can address measured, correctly timed variables when models and overlap are adequate, but they cannot guarantee control of unrecorded severity, clinician judgment, measurement error, or an incorrectly designed time zero.
Why is an active comparator often better than no treatment?
People starting alternative treatments for the same condition are more likely to share an indication, disease stage, care access, and willingness to use therapy. The comparator still needs clinical scrutiny because prescribing channels and contraindications can differ.
Does a large database protect against confounding by indication?
No. A large sample can produce a very narrow confidence interval around a biased estimate. Design quality, covariate measurement, overlap, and sensitivity analyses determine whether greater precision is informative.
What should a reader look for first?
Check the clinical reason for treatment, eligibility, new-user status, active comparator, common time zero, baseline severity, post-adjustment balance, overlap, absolute risks, and analyses that probe residual confounding.