Observational comparisons struggle when clinicians choose treatment based on prognosis that the dataset does not fully capture. Instrumental-variable analysis seeks a different source of treatment variation: a variable that changes the probability of treatment but is otherwise disconnected from the outcome. Examples include random assignment with nonadherence, clinician prescribing tendency, distance to a treatment center, policy rules, and genetic variants.
The appeal is substantial. Under valid assumptions, an instrument can identify a causal effect even with unmeasured treatment-outcome confounding. The danger is equally substantial. A plausible-sounding natural experiment can swap one unmeasured-confounding problem for assumptions you will find even harder to defend.
The instrument separates assignment from treatment received#
In a randomized trial, assignment is generated without regard to baseline prognosis. If everyone follows assignment, the intention-to-treat comparison estimates the effect of being assigned. When some participants do not adhere, random assignment can act as an instrument for treatment received.
Assignment predicts treatment, but treatment is not perfectly determined. If assignment affects the outcome only through treatment and other assumptions hold, the assignment-outcome effect can be scaled by the assignment-treatment effect to estimate an effect among participants whose receipt follows assignment.
Observational IV studies look for naturally occurring variation with similar structure. A clinician who tends to prescribe drug A may make a patient more likely to receive A, and the analysis tries to use the tendency rather than the patient's actual choice as the quasi-random source of variation.
Relevance is the first, testable assumption#
The instrument must be associated with treatment. This is the first-stage relationship. A preference measure that changes prescribing by one percentage point contains far less identifying information than one that changes it by 30 points.
A report will show treatment rates across instrument values, the first-stage coefficient and confidence interval, and an appropriate strength statistic; the familiar F-statistic threshold of 10 is a rough heuristic from particular linear settings, not a universal certificate. Multiple instruments and clustered data need suitable diagnostics.
Very strong prediction does not establish validity. A hospital rule can dictate treatment while also changing staffing, monitoring, or access to follow-up. It passes relevance but may violate exclusion.
Independence means no uncontrolled common cause#
The instrument should be as-if randomly distributed with respect to causes of the outcome. In a diagram, no open backdoor path should connect instrument and outcome. Random assignment supports this by design; natural instruments require an argument.
Distance to a specialist center may predict treatment, but people living nearby can differ in income, urbanicity, environmental factors, and healthcare access, and a calendar policy change may coincide with other changes. Physician preference may be associated with specialty, hospital resources, and patient mix.
Baseline covariate balance can detect some problems. If age, disease severity, or socioeconomic measures differ across instrument groups, independence is less plausible. Balance on measured variables does not prove balance on unmeasured ones. The report should explain why remaining pathways are unlikely, not simply announce that P values exceed 0.05.
Exclusion restricts every path to run through treatment#
The instrument must affect the outcome only by changing the intervention being studied. This is often the hardest assumption. A clinician who prefers drug A may also order more monitoring, prescribe related therapies, or manage risk factors differently, and distance to a high-volume surgical center can change postoperative care as well as which surgery is received.
The exclusion restriction applies to the precisely defined treatment. If the intervention is “receive drug A,” but the instrument also changes dose, adherence support, or timing, those may be part of a broader treatment package rather than violations. The estimand must match the pathway.
Empirical data can falsify some exclusion stories. Showing that the instrument changes a co-intervention linked to outcome is evidence against validity. Failure to find such a path does not prove none exists.
The simple Wald ratio shows the method's logic#
For a binary instrument and treatment with a continuous outcome, a simple IV estimator divides the instrument-outcome difference by the instrument-treatment difference: if assignment increases treatment by 20 percentage points and improves an outcome by two units, the ratio estimates ten units per full treatment change under the assumptions.
The denominator reveals why weak instruments are dangerous. A small first-stage difference makes the ratio unstable. A tiny instrument-outcome imbalance from bias or chance becomes magnified.
Two-stage least squares generalizes this logic for linear models. The first stage predicts treatment from the instrument and covariates. The second relates the predicted treatment component to outcome. Using predicted treatment in an ordinary second-stage logistic or Cox model does not automatically produce a valid nonlinear IV estimator; methods and estimands differ by outcome scale.
Monotonicity defines a local population#
With binary assignment and treatment, participants can be described by how treatment would respond to instrument level. “Compliers” take treatment when encouraged and not when unencouraged. Always-takers and never-takers do not change. Defiers would move in the opposite direction.
Under relevance, independence, exclusion, and monotonicity, the usual estimate is a local average treatment effect among compliers. Monotonicity assumes the instrument does not cause anyone to move opposite the intended direction. For randomized encouragement this may be plausible; for physician preference, one doctor's preference can affect different patients in complex ways.
The complier group is latent. Researchers cannot identify each person as a complier; they can sometimes describe its expected characteristics, but the effect should not be generalized automatically to people whose treatment is unaffected by the instrument.
Effect homogeneity offers a different identifying route#
Some IV methods assume the treatment effect is constant across people or that effect modification has a specified structure, which can identify an average effect beyond compliers, but it is often biologically demanding. Treatment effects commonly vary with disease severity, age, contraindications, and adherence.
Newer frameworks use assumptions such as no simultaneous heterogeneity, which still require careful explanation. A methods label does not eliminate the need to state the fourth point-identifying assumption.
The BMJ 2024 guide emphasizes that the three core assumptions alone often identify broad bounds rather than one precise point. Applied papers has to say which additional assumption narrows those bounds, and what population the result represents.
Physician preference is convenient but fragile#
A common instrument uses the physician's previous prescription or proportion of recent prescriptions for a drug, and it can reduce confounding by individual patient indication because preference influences which option is chosen.
Patients are not randomly assigned to physicians. Referral, insurance, geography, and disease complexity affect clinician selection. Preference may correlate with specialty, experience, follow-up intensity, or concurrent treatment. It can also change over time after new evidence, safety alerts, or formulary shifts. The analysis should demonstrate temporal construction, first-stage strength, patient covariate balance, clinician clustering, and robustness to alternative preference definitions. It should test whether preference predicts other care processes linked to outcome.
Distance and facility instruments package more than travel#
Relative distance to a facility offering a procedure can influence receipt. It may be compelling when geography was established before illness and nearby facilities otherwise provide similar care. In practice, distance captures transport, rurality, income, referral networks, and facility quality.
An instrument comparing differential distance to two facility types can reduce some residential confounding, but exclusion remains: hospital choice can change the entire care team. Restricting to patients clinically eligible for both options can improve comparability while narrowing generalizability. Geospatial instruments should report residential data timing, mobility, emergency transport, border effects, and whether patients bypass the nearest facility. Straight-line distance may misrepresent travel time and access.
Policy and calendar instruments risk simultaneous change#
Coverage thresholds, formulary rules, age cutoffs, and regional policies can create discontinuities in treatment. Regression-discontinuity designs near an arbitrary threshold can offer strong quasi-random variation if people cannot manipulate assignment and other determinants change smoothly.
A before-after policy indicator is weaker when infection waves, secular trends, staffing, coding, or concurrent programs change. Calendar time affects many outcomes directly. The policy date can satisfy relevance while violating independence and exclusion. Graphical checks around thresholds, bandwidth sensitivity, manipulation tests, and falsification outcomes help. Interpretation applies near the cutoff and under the local policy context.
Mendelian randomization is an IV design, not genetic randomization magic#
Mendelian randomization uses genetic variants associated with a modifiable trait as instruments. Alleles are assigned at conception, reducing some reverse causation. The core assumptions remain relevance, independence, and no pathway to the outcome except through the trait.
Horizontal pleiotropy violates exclusion when a variant affects outcome through another biological pathway. Population stratification, assortative mating, dynastic effects, linkage disequilibrium, and selection into the genetic sample can undermine independence. Weak variant-trait associations create weak-instrument bias. STROBE-MR recommends transparent reporting of variant selection, sample overlap, harmonization, instrument strength, pleiotropy analyses, population structure, and the causal estimand. Sensitivity methods have different assumptions and should not be treated as a vote count.
Binary and time-to-event outcomes require care#
Two-stage least squares has a straightforward additive interpretation for continuous outcomes under linear assumptions. With binary outcomes, an additive risk difference may be estimated in some settings, but predicted values can fall outside zero to one. Two-stage residual inclusion, structural mean models, and other estimators target different effects.
For survival outcomes, censoring, competing risks, treatment timing, and noncollapsible hazard ratios complicate IV analysis. A causal hazard ratio is not obtained merely by placing first-stage predicted treatment into a Cox model. Look for an explicit justification of the estimator, the model, and the censoring assumptions. Absolute-risk and survival-curve estimates can be more decision relevant, but their IV identification may require stronger modeling. Statistical consultation is not optional for a novel instrument-outcome combination.
Weak instruments distort more than precision#
Weak instruments produce wide confidence intervals, but finite-sample estimates can also drift toward the confounded conventional association. Selecting instruments because they happened to be strong in the same data can create winner's-curse bias.
With many weak instruments, overfitting the first stage can worsen bias. Limited-information maximum likelihood and weak-instrument-robust confidence procedures can help under particular models. None repairs a violation of independence or exclusion. A very large IV point estimate compared with conventional adjustment may reflect a local effect, weak-instrument amplification, or invalidity. The difference deserves explanation, not automatic preference for the IV result.
Standard errors must match the assignment structure#
If the instrument varies by clinician, hospital, region, or policy period, outcomes within clusters are correlated. Standard errors and effective sample size must reflect the level at which quasi-random variation occurs. Thousands of patients treated by ten clinicians do not provide thousands of independent instrument assignments.
Two-stage estimation also propagates first-stage uncertainty. Bootstrapping, robust sandwich estimators, or analytic methods should match the design. Confidence intervals that ignore either clustering or first-stage uncertainty can look narrower to you than they are.
Missing treatment, instrument, or outcome data create selection. Complete-case restriction is not automatically compatible with IV independence. The missingness process needs its own assumptions and sensitivity analysis.
Overidentification tests do not certify all instruments#
When several instruments are available, statistical tests can assess whether their estimates are mutually consistent under model assumptions. Passing an overidentification test does not prove that every instrument is valid; all can share the same violation. Failing can signal heterogeneity or misspecification rather than identify which instrument is wrong.
Instruments should be selected from a causal rationale before outcome analysis. Presenting only the combination that yields the desired result turns a design argument into model search.
Negative control outcomes and instrument associations with pre-treatment variables can uncover pathways. Quantitative sensitivity analysis can show how small violations affect the estimate, especially when the instrument is weak.
A credible report lets readers audit assumptions#
Draw a directed acyclic graph connecting instrument, treatment, outcome, measured confounders, and suspected alternative paths. Define timing: when the instrument is set, when treatment occurs, and when follow-up begins. Show first-stage distributions, strength, overlap, and treatment rates.
Report baseline variables by instrument, but emphasize effect sizes and clinical relevance rather than significance tests. Examine co-interventions and outcomes the treatment cannot cause. Name the fourth identifying assumption, target population, effect scale, and estimator.
Compare IV results with conventional adjustment and other designs without declaring agreement as proof. Triangulation is strongest when methods have different plausible biases. A valid instrument can be highly informative; a weakly defended one can make confounding harder for you to see.
References#
- BMJ guide and checklist for instrumental-variable studies
- Instrumental-variable methods for causal inference
- Physician prescribing preference as an instrument
- How to report instrumental-variable analyses
- STROBE-MR reporting guideline
- Understanding IV assumptions in observational studies
It does not provide medical advice or establish that any proposed instrument is valid.*
Questions and answers
What is an instrumental variable?
It is a variable associated with treatment, with no uncontrolled common cause shared with the outcome, and no route to the outcome except through treatment; those conditions are assumptions about a specific causal question.
Can data prove that an instrument is valid?
No. First-stage relevance is measurable and some violations can be detected through imbalance or alternative pathways. Independence and exclusion cannot be proven from observed data alone and need subject-matter justification.
What is a weak instrument?
It changes treatment probability only slightly. Dividing by that small first-stage effect produces imprecision, finite-sample bias, and sensitivity to even small violations of independence or exclusion.
What does local average treatment effect mean?
Under monotonicity and the core assumptions, it is the average effect among people whose treatment choice is changed by the instrument, called compliers, and it need not apply to always-treated, never-treated, or the full population.
Is Mendelian randomization automatically unbiased?
No. Pleiotropy, population structure, selection, assortative mating, linkage, sample overlap, and weak instruments can violate assumptions. Genetic assignment does not make every variant a valid instrument.