Evidence explainer

Diabetes and metabolic health

Linkage and Association in Genetics: Different Evidence, Different Scale

Linkage asks whether a region and a trait co-segregate through a pedigree. Association asks whether alleles and traits occur together more often than expected in a sampled population.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Begin with the unit of comparison
  2. Recombination supplies the map
  3. What a linkage statistic is testing
  4. Association uses statistical co-occurrence
  5. Linkage is not linkage disequilibrium
  6. A compact comparison
  7. Why complex traits changed the balance of methods
  8. An association can fail without bad statistics
  9. Moving from a locus to a mechanism
  10. How to read a genetics headline

Begin with the unit of comparison#

Two study results can both point to a chromosome and still rest on different evidence.

In linkage analysis, the comparison happens across inheritance events inside families. Investigators ask whether a marker or chromosomal region is transmitted with a trait through a pedigree more often than expected if they were unlinked.

In genetic association analysis, the comparison concerns allele and trait frequencies in sampled people. Investigators ask whether a variant is statistically related to a disease or measured trait after accounting for the study design and relevant sources of bias.

That distinction is more dependable than the shortcut you will usually hear, that linkage is “for rare disease” and association is “for common disease.” Those patterns often describe historical strengths, but modern methods can use association in families, sequence rare variants, or combine designs. What a study can detect comes down to the sampling, the inheritance model, and the variant frequency. It comes down to the effect size and the analysis.

This article explains research methods. A linkage peak or association result is not, by itself, a clinical genetic test result or an estimate of one person's risk.

Recombination supplies the map#

During formation of egg and sperm cells, paired chromosomes can exchange corresponding DNA segments; these recombination events separate loci that are farther apart more often than loci that are close together.

NHGRI defines linkage as the tendency of nearby sequences on the same chromosome to be inherited together, and in a pedigree, researchers can observe which marker alleles travel with a phenotype and use recombination across relatives to infer a likely region. The result is a location estimate based on inheritance, not necessarily an identified gene or causal change.

The number and structure of informative families matter. A large, well-characterized pedigree with many affected relatives may carry considerable information for a strongly penetrant locus. Small pedigrees, uncertain diagnoses, or genetic heterogeneity can weaken the signal. So can reduced penetrance, phenocopies, or many small genetic effects.

What a linkage statistic is testing#

A linkage analysis compares how well the observed family data fit different recombination fractions or inheritance models. A logarithm of the odds, or LOD, score is one familiar summary. Its interpretation depends on the model, marker information, pedigree structure, and whether researchers specified the analysis before seeing the data.

A linkage peak generally covers a region rather than a single base. The region may contain many genes and regulatory elements. Fine mapping, sequencing, and segregation checks may be needed. So may functional work and evidence from other families. They help you move from “this interval travels with the trait” to a defensible causal explanation.

Association uses statistical co-occurrence#

An association study can compare people with and without a disease, analyze a quantitative trait, or use other designs, and a genome-wide association study tests variants across the genome without limiting the search to one candidate gene. NHGRI describes GWAS as scans of genomic markers in many individuals to identify variants statistically associated with disease or traits.

The core estimate asks whether genotype predicts the outcome under a specified model, often while adjusting for age, sex, and ancestry-related structure. The adjustment also covers study site or other prespecified variables. Large samples are often necessary because many common-variant effect sizes are modest and genome-wide testing imposes a stringent threshold to control false positives.

Replication in an independent sample helps you tell a durable signal from chance or study-specific bias. Meta-analysis can increase power, but only if phenotype definitions, samples, and analysis choices are sufficiently compatible and duplicated participants are handled correctly.

Linkage is not linkage disequilibrium#

The similar names cause persistent confusion.

Linkage concerns co-transmission of loci within families because they are physically near one another on a chromosome.

Linkage disequilibrium, commonly shortened to LD, is a population-level correlation between alleles at different sites. It reflects recombination history along with mutation, drift, selection, admixture, and population structure. LD patterns vary across populations.

GWAS commonly detects a tag variant correlated through LD with one or more nearby variants. So the variant you see named in a report may be causal, may be one of several plausible causal variants, or may simply mark a haplotype containing the relevant change. The NHGRI GWAS fact sheet explicitly cautions that associated variants may be traveling with causal variants rather than causing disease themselves.

A compact comparison#

Article data table
QuestionLinkage analysisAssociation analysis
Primary evidenceCo-segregation through pedigreesStatistical relationship between genotype and trait in a sample
Historical strengthLoci with substantial effects in informative familiesVariants detectable through population or family association designs
Mapping scaleOften a broader chromosomal intervalOften a smaller LD-defined locus, though not necessarily one variant
Major vulnerabilitiesPedigree errors, phenotype misclassification, model assumptions, locus heterogeneityPopulation structure, batch effects, phenotype variation, multiple testing, selection bias
Typical next workSequence and prioritize the interval, test segregation, replicate in other familiesFine-map, replicate, study diverse populations, connect variants to genes and function

The table summarizes tendencies rather than fixed rules.

Why complex traits changed the balance of methods#

For a highly penetrant Mendelian condition in a sufficiently informative family, one locus can create a recognizable co-segregation pattern. Many common diseases and quantitative traits have a different architecture: numerous variants, small effects, environmental contributions, gene interactions, diagnostic heterogeneity, and more than one biological route to a similar phenotype.

Under that architecture, a family may carry too little information for conventional linkage to detect any one small effect. Large association studies aggregate evidence across many people instead. This shift contributed to the prominence of GWAS for traits such as type 2 diabetes.

The shift did not make linkage obsolete. Pedigree analysis remains important for disease-gene discovery, variant interpretation, and families with suggestive inheritance. Nor did GWAS make causality automatic. Each method provides a different layer of evidence.

An association can fail without bad statistics#

Several problems can produce or obscure an association:

Quality control, principal components, and mixed models address parts of these problems, not every possible bias. So do replication and sensitivity analyses.

Moving from a locus to a mechanism#

A credible path may include statistical fine-mapping, sequencing, expression or protein data, chromatin and regulatory annotation, colocalization, perturbation experiments, relevant cellular or animal models, and evidence that the proposed mechanism fits human biology.

The nearest gene is not automatically the effector gene. A regulatory variant can act over distance or in a particular cell state, and the GWAS Catalog also distinguishes author-reported and mapped genes, a practical reminder that database mapping is annotation rather than proof.

Converging evidence is more persuasive than a single dramatic experiment. A functional effect in an artificial system may be real without explaining the human association; a strong human association may be reproducible before its molecular pathway is known.

How to read a genetics headline#

When a genetics result reaches the news, these questions tell you what you are actually looking at:

  1. Was the study linkage, association, sequencing, or a combination?
  2. What family or population was sampled, and how was the phenotype defined?
  3. Is the result a broad interval, a lead variant, a credible set, or an experimentally supported causal variant?
  4. Was it replicated?
  5. What is the effect size and uncertainty, not just the p-value?
  6. Does the evidence connect the locus to a gene and mechanism?
  7. Is there any validated clinical use, or is the result still research?

Sources and further reading

  1. NHGRI Linkage definition (accessed 2026-07-15)
  2. NHGRI Genome-Wide Association Studies Fact Sheet (accessed 2026-07-15)
  3. NHGRI Genome-Wide Association Studies glossary entry (accessed 2026-07-15)
  4. NHGRI-EBI GWAS Catalog documentation (accessed 2026-07-15)
  5. NCBI Bookshelf Genetics and Health, linkage and association overview (accessed 2026-07-15)

Questions and answers

Does the strongest GWAS variant cause the trait?

Not necessarily. It may be correlated through LD with the relevant variant. Fine-mapping and functional evidence are needed to distinguish possibilities.

Can linkage identify one exact mutation?

Usually linkage identifies a region. Sequencing and segregation analysis can then evaluate candidate variants within that region.

Does a small p-value mean a large genetic effect?

No. Statistical significance depends on effect size, sample size, variability, and the testing framework. Very large studies can detect small effects.

Can a research association predict my individual risk?

Not by itself. Clinical validity and utility require separate evidence, and risk estimates may depend on ancestry, phenotype, environment, family history, and the model used. Personal interpretation belongs with qualified genetics and clinical professionals.