Evidence explainer

Diabetes and metabolic health

What Mendelian Randomization Can and Cannot Show

Mendelian randomization asks whether genetic differences in a modifiable factor track an outcome in a pattern consistent with cause. What it can support depends on whether its assumptions are credible.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The natural-experiment intuition
  2. Instruments connect genes, factors, and outcomes
  3. Assumption one: relevance
  4. Assumption two: independence
  5. Assumption three: exclusion restriction
  6. MR-Egger and other sensitivity methods
  7. Colocalization asks whether the signal is shared
  8. Reverse direction and bidirectional analysis
  9. Binary factors and liability
  10. Lifelong genetic effects and treatment effects differ
  11. Nonlinearity and effect modification
  12. Multivariable Mendelian randomization
  13. Selection and two-sample compatibility
  14. Multiple testing and automated scans
  15. Reading an MR paper step by step
  16. Triangulation gives the result meaning
  17. The disciplined conclusion
  18. Sources

Mendelian randomization uses inherited genetic variants as instrumental variables to investigate causal questions. If certain variants reliably influence a factor such as LDL cholesterol, and those variants also influence coronary disease in proportion to their effect on LDL, the pattern can support a causal role for lifelong LDL differences.

The method addresses some limitations of ordinary observational research. Genotype is fixed before many later behaviors and illnesses, so reverse causation is less likely, and many social and clinical confounders are also less strongly related to genotype than to a measured behavior.

Those advantages are conditional. Genetic variants can affect several pathways, ancestry structure can create associations, and the genetic proxy may not match a treatment. Mendelian randomization is best treated as one form of causal evidence, not a genetic trial or a machine that turns association into proof.

The natural-experiment intuition#

At conception, a child receives one allele from each parent at many genetic loci, and under idealized conditions, this segregation creates groups that differ slightly in a biological factor but are otherwise comparable with respect to later confounders.

The intuition resembles randomization, but the analogy has limits. Mate choice is not random across all traits. Ancestry and geography shape allele frequencies and outcomes. Parents' genes can influence the family environment. Survival and study participation can select genotypes and risks.

Genetic variants are also not assigned by an investigator to answer one clinical question. Their biological effects may be broad, small, lifelong, and developmentally compensated. What the design buys you is that the assumptions are explicit and, where possible, testable.

Instruments connect genes, factors, and outcomes#

An instrumental variable is associated with the factor of interest, sometimes called the risk factor or trait; researchers estimate how strongly each variant changes the trait and how strongly it relates to the outcome.

With one valid instrument, the ratio of variant-outcome association to variant-trait association estimates the outcome change per unit difference in the trait, and with many variants, methods combine ratio estimates, often through inverse-variance weighting.

Modern two-sample Mendelian randomization commonly obtains trait associations from one genome-wide association study and outcome associations from another. This increases scale and permits many questions. But it moves the weight onto harmonization, ancestry, and allele alignment. It moves weight onto sample overlap and compatible populations. Whatever number comes out, it estimates the variation those instruments represent, under the method's assumptions, and nothing wider.

Assumption one: relevance#

The variants must predict the trait. Weak associations produce unstable estimates and can magnify bias. Instrument strength is often summarized with F statistics or variance explained.

Genome-wide significant variants reduce false-positive selection but do not guarantee adequate strength. Winner's curse can inflate the discovery association. A polygenic score can improve strength while adding more variants with uncertain biology and greater pleiotropy risk.

In one-sample analyses, weak-instrument bias often moves estimates toward the confounded observational association. In nonoverlapping two-sample analyses, it often moves toward the null under common conditions. Sample overlap can change the direction. A paper should therefore report how the instruments were selected, how strong they are, which sample they were discovered in, and whether the same participants contributed to both associations.

Assumption two: independence#

The instrument should not share a cause with the outcome. Because genotype precedes later disease, many ordinary confounders are less likely to cause genotype. Independence can still fail.

Population stratification occurs when ancestry-related genetic differences align with outcome differences caused by environment or social structure. Principal components and ancestry-restricted analyses can reduce but not eliminate it. Results in one ancestry may not transfer to another.

Assortative mating, in which partners correlate on traits, can create genetic-environmental patterns. Dynastic effects occur when parental genotype shapes a child's environment. Geographic structure and participation selection can also connect variants with outcomes outside the intended pathway, and a within-family design can reduce several of these biases, though it takes large samples and often gives lower precision.

Assumption three: exclusion restriction#

The variant must influence the outcome only through the factor being studied. If a variant changes the outcome through another biological route, horizontal pleiotropy violates the assumption.

Vertical pleiotropy is different. If the variant changes factor A, which changes mediator B, which changes the outcome, B lies on the causal pathway and does not automatically invalidate the instrument for A.

Horizontal pleiotropy can be balanced, with positive and negative effects averaging near zero, or directional. Many-variant methods often rely on additional assumptions about its distribution. No statistical test can prove that every instrument meets exclusion. Biological annotation, phenome-wide scans, and robust estimators can increase or decrease credibility. So can negative controls and consistency across instrument sets.

MR-Egger and other sensitivity methods#

Inverse-variance weighting is efficient when all variants are valid or pleiotropic effects satisfy favorable conditions, and weighted median estimation can remain consistent when at least half the weight comes from valid instruments under its assumptions. Mode-based methods seek the largest cluster of similar estimates.

MR-Egger regression allows a nonzero intercept that can indicate average directional pleiotropy and estimates a slope under the InSIDE assumption, meaning instrument strength is independent of direct effects; it often has low precision and is vulnerable to weak instruments.

MR-PRESSO and outlier-removal methods can identify unusual variants, but an outlier can represent genuine biology rather than error. Deleting variants until results agree is not valid sensitivity analysis; read all of these as diagnostics run under different assumptions, not as a vote in which the preferred estimate wins.

Colocalization asks whether the signal is shared#

Suppose variants in one genomic region associate with a protein and a disease. The associations may be driven by the same causal variant, supporting a shared mechanism. They may instead come from two variants in linkage disequilibrium that happen to travel together.

Colocalization methods estimate support for shared or distinct causal signals. Fine-mapping and credible sets help when several variants could explain a region. Results depend on priors, variant coverage, ancestry, and the possibility of multiple causal variants.

For drug-target Mendelian randomization using variants near a gene, colocalization is especially important. A strong ratio estimate with no shared causal signal can be misleading. Colocalization firms up what is happening at that locus. On its own it does not establish the rest of the causal chain.

Reverse direction and bidirectional analysis#

Reverse causation is reduced because genotype is fixed before most outcomes, but the causal direction between two traits can still be unclear when variants influence both or when phenotype associations are complex.

Bidirectional Mendelian randomization uses instruments for trait A to test outcome B and instruments for B to test A. Steiger-type approaches compare variance explained to ask which phenotype appears closer to the genetic signal.

Measurement error can distort direction tests. A well-measured downstream marker may appear to explain more genetic variance than a noisy upstream trait, and disease liability can also differ from measured disease status. So before you accept a direction, look for biology, temporality, and instrument specificity. Look for agreement with other designs, rather than one diagnostic statistic.

Binary factors and liability#

When the factor is binary, such as having a disease diagnosis, genetic variants usually influence an underlying liability rather than switching the disease on or off, and the Mendelian randomization estimate may represent a change in outcome per unit of genetic liability, not the effect of treating diagnosed disease.

Scaling an estimate per doubling of odds can improve reporting but remains abstract. It should not be interpreted as the effect of preventing one clinical case unless strong additional assumptions hold.

Similarly, genetic proxies for behaviors such as smoking or alcohol use may reflect initiation or quantity. They may reflect metabolism or dependence. The instrument's phenotype defines the question. So a liability estimate should not be read back out as a claim about treating diagnosed disease.

Lifelong genetic effects and treatment effects differ#

Genetic differences often act from conception, while clinical interventions begin later and last for a shorter period, and a small lifelong difference can accumulate a large effect that does not predict the effect of a short trial with the same unit change.

A variant may alter a target's expression in specific tissues or developmental stages. A drug may inhibit the protein systemically, incompletely, or through a different mechanism. Compensation during development can blunt or redirect genetic effects.

Mendelian randomization can prioritize a target and suggest direction. It generally cannot establish formulation, dose, or treatment duration. It cannot establish reversibility, adverse effects, or benefit-harm balance. Randomized clinical trials remain necessary before recommending a therapy.

Nonlinearity and effect modification#

Standard Mendelian randomization often estimates an average linear effect. The true relationship may have thresholds, plateaus, or different effects by baseline level.

Stratifying participants by the measured trait can induce collider bias because genotype contributes to that trait. Specialized residual-stratification and doubly ranked methods attempt nonlinear estimation under additional assumptions.

Subgroup analyses by age, sex, disease, or medication can also be biased by selection and low power, and a difference between significant and nonsignificant subgroup estimates is not evidence of interaction. A nonlinear claim needs prespecified methods, adequate samples, graphical uncertainty, and replication before you take it seriously, and even then it is rarely strong enough by itself to define a clinical threshold.

Multivariable Mendelian randomization#

Traits such as LDL cholesterol, triglycerides, and apolipoprotein B share genetic architecture. Multivariable Mendelian randomization includes genetic associations with several factors to estimate their direct effects conditional on one another.

This can help distinguish correlated pathways, but it needs instruments strong enough jointly and assumptions about every included factor. Measurement scale, collinearity, weak conditional strength, and omitted pathways can destabilize results.

The direct effect may answer a narrower question than readers expect: changing one factor while holding the others fixed, if such an intervention is biologically possible. So look for the conditional strength, and for a stated reason each variable belongs in the model. Adding factors mechanically can exchange one bias for another.

Selection and two-sample compatibility#

Biobanks are not random samples. Participation can depend on health, education, geography, survival, and access. Conditioning on participation can create associations between genotype and outcome.

Case-control studies can enrich disease efficiently but need compatible definitions and sampling. Survival bias matters for late-onset outcomes when a genotype affects earlier mortality.

Two-sample analysis assumes the genetic associations apply to compatible underlying populations. Differences in ancestry, age, sex, phenotype definition, measurement, covariate adjustment, or calendar period can violate that assumption, so a paper should say the cohort composition, the overlap, the harmonization, and the selection mechanisms. A large sample size does not correct systematic incompatibility.

Multiple testing and automated scans#

Platforms can run thousands of trait-outcome pairs quickly. This enables discovery and increases false-positive risk. Correlated traits and overlapping datasets make simple correction challenging.

Hypotheses, primary analyses, and sensitivity plans should be registered when feasible. Discovery findings need replication with nonoverlapping data and biologically justified instruments. Results should not be selected only because they pass a threshold.

Automated pipelines can also harmonize the wrong allele, include palindromic variants with ambiguous frequency, or mix genome builds, and quality control and code review remain essential, because an attractive causal diagram does not verify the data feeding it.

Reading an MR paper step by step#

Start with the causal question and the instrument biology. Then check the trait and outcome definitions, the ancestry, the samples, the overlap, and the strength, and confirm that the alleles were harmonized and that the estimates use compatible scales.

Review the core estimate with its interval. Then examine heterogeneity, pleiotropy diagnostics, and robust estimators. Examine leave-one-out results, colocalization, direction tests, and within-family or replication evidence where available.

Look for negative controls and comparison with observational and trial evidence. Ask whether the genetic proxy matches the proposed intervention in timing, tissue, and mechanism.

Finally, read the conclusion for calibration. “Supports a causal role” can be justified; “proves this drug will work” usually cannot.

Triangulation gives the result meaning#

Mendelian randomization is strongest when its biases differ from other evidence. Prospective cohorts may have confounding, trials may be short, natural experiments may have policy-specific effects, and genetic analyses may have pleiotropy. Agreement across them is harder to explain by one shared flaw.

Disagreement is informative. It can reveal timing, nonlinearity, intervention differences, selection, or an invalid instrument. Investigators should investigate the mismatch rather than rank methods by preference.

If you are the patient or the clinician, the final action should still rest on clinical trials, guidelines, and your own situation; genetic causal evidence can explain why a target was pursued without becoming personal medical advice.

The disciplined conclusion#

Mendelian randomization can move a question beyond ordinary correlation by using inherited variants and explicit instrumental-variable logic. It can identify likely causal targets, challenge implausible claims, and guide trials.

Its value depends on transparent assumptions. When relevance, independence, exclusion, population compatibility, and biological correspondence are uncertain, the conclusion should be equally uncertain. The method is a powerful part of causal inference precisely when it is not mistaken for randomization.

Sources#

The metadata sources include STROBE-MR, current methods reviews, practical appraisal guidance, MR-Base, and the foundational MR-Egger sensitivity method.

Sources and further reading

  1. STROBE-MR Reporting Guideline
  2. Sanderson and colleagues, Mendelian Randomization
  3. Davies, Holmes, and Davey Smith, Reading Mendelian Randomization Studies
  4. Burgess and colleagues, Guidelines for Mendelian Randomization Investigations
  5. Hemani and colleagues, MR-Base Platform
  6. Bowden and colleagues, MR-Egger Regression for Detection of Pleiotropy

Questions and answers

Is Mendelian randomization the same as a randomized clinical trial?

No. Genetic variants are assigned at conception, but they are not a randomized treatment, and the method depends on untestable assumptions, population structure, biology, and data quality.

What are the three core instrumental-variable assumptions?

The variants must predict the factor of interest, share no cause with the outcome, and affect the outcome only through that factor rather than through another pathway.

What is horizontal pleiotropy?

It occurs when a genetic variant affects the outcome through a pathway other than the factor being studied, violating the exclusion assumption and potentially biasing the causal estimate.

Can Mendelian randomization predict the effect of a specific drug dose?

Usually not directly. Genetic differences often act from early life and may influence a biological target differently from a medicine started later, so trials are needed for dose, benefit, and harm.

What makes a Mendelian randomization result more credible?

Strong biologically justified instruments, preregistered analysis, ancestry-aware quality control, minimal sample overlap, concordant robust methods, colocalization where relevant, negative controls, replication, and triangulation with other designs.