A genome-wide association study, or GWAS, scans a large set of genomic variants and tests whether any are statistically associated with a trait, and the trait might be a disease diagnosis, laboratory value, height, drug response, imaging feature, or another measured phenotype.
The result is a map of associations in a studied sample. An associated marker can point toward biology, support target research, or contribute to prediction, but it is not automatically the causal DNA change, the responsible gene, or an explanation of one person's outcome. Those conclusions require further evidence.
The basic design#
Researchers define a phenotype and assemble participants with genomic data. In a case-control study, allele frequencies are compared between people with and without a disease. For a quantitative trait, such as fasting glucose, regression tests whether genotype dosage is associated with higher or lower values.
Modern arrays directly genotype selected single-nucleotide variants. Imputation uses reference panels and local correlation patterns to estimate additional untyped variants. Sequencing-based studies can assess rarer variation more directly, although quality control and statistical power remain challenging.
Each variant is tested while accounting for prespecified covariates. For a binary outcome, a result may be an odds ratio per copy of an allele; for a continuous trait, it may be a mean difference in standardized units, and the effect applies under the study model and phenotype definition.
GWAS is often called hypothesis-free because it does not start with one candidate gene. It is not assumption-free. Researchers still choose the phenotype, inclusion criteria, and genotyping platform. They choose the imputation panel, covariates, genetic model, and significance rules. Those choices shape what you can find.
Why the significance bar is high#
Testing millions of variants creates many opportunities for chance associations. A conventional p-value of 0.05 would generate an enormous number of false-positive signals. GWAS commonly uses a genome-wide threshold around 5 x 10^-8, derived as a practical correction for the number of effectively separate common-variant tests in many datasets.
A p-value estimates compatibility of the data with a statistical null under the model. It does not tell you the probability that a finding is true, the size of the effect, or its clinical importance. Very large samples can make tiny effects highly significant.
Confidence intervals and effect estimates belong beside p-values. A variant with an odds ratio of 1.05 can be biologically informative at population scale while contributing little to individual prediction, while a rare variant with a larger estimate may have wide uncertainty if few carriers are observed.
Replication in new participants helps distinguish robust signals from chance, technical artifact, or study-specific bias; a replication should match the allele, phenotype, direction, and analysis closely enough to test the same claim. Combining discovery and replication without transparent separation can overstate evidence.
Quality control happens before the headline#
Sample checks identify duplicates, unexpected relatedness, and sex-chromosome discrepancies. They identify contamination, low call rates, and ancestry outliers under the analysis plan. Variant checks examine call quality, missingness, and allele frequency. They examine imputation quality and departures that may signal error.
Batch effects can occur when cases and controls are genotyped on different platforms, plates, laboratories, or dates. If technical conditions track the phenotype, a machine artifact can look genetic. Balanced processing and covariate adjustment reduce but do not eliminate this risk.
Phenotype quality is equally important. A billing code may capture a different disease spectrum from adjudicated clinical criteria, combining subtypes can dilute or mix mechanisms, and medication use can alter laboratory traits. Misclassification usually reduces power but can also introduce bias when it differs systematically.
Related participants violate simple independence assumptions. Mixed models or kinship adjustment can use family structure while controlling false positives. Study reports should show these choices, exclusions, and sensitivity analyses rather than treating quality control as an invisible preprocessing step.
Population structure can create false association#
Allele frequencies vary across populations because of demographic history. Disease prevalence and measured traits can also vary because of environment, healthcare, geography, and social conditions. If those structures align, a variant can appear associated without causing the trait.
Principal components, mixed models, stratified analyses, and careful sampling help control population structure. Genomic inflation statistics can reveal broad residual signal, though polygenic traits and huge samples complicate interpretation.
Genetic ancestry describes patterns of DNA similarity to reference data. It is not interchangeable with race, ethnicity, culture, or lived environment. These variables can each carry different scientific and social information. Collapsing them can produce poor adjustment and misleading biological claims. Within-family analyses can reduce some population and family-level confounding, though they may have less power and answer a somewhat different question. Agreement across population and family designs can strengthen inference; disagreement is informative rather than a nuisance to hide.
A lead variant usually identifies a region#
Nearby variants are often correlated through linkage disequilibrium, abbreviated LD. A genotyped marker can be associated because it travels with another variant that affects biology. The strongest p-value in a region is therefore not necessarily the causal variant.
LD patterns differ across populations. This can make the lead marker change even when the underlying biology is shared. Multi-ancestry analysis can improve mapping by using different correlation structures, provided models handle heterogeneity and representation carefully.
Fine-mapping estimates a credible set of variants that could explain a signal. Functional annotation asks whether candidates alter protein sequence, gene regulation, or chromatin. It asks whether they alter splicing or expression in relevant tissues. Expression quantitative trait loci and colocalization can test whether trait and gene-expression signals plausibly share a cause.
The nearest gene is a convenient label, not a conclusion. Regulatory variants can act over long distances, affect multiple genes, or operate only in a particular cell state. Experimental perturbation can support mechanism, but laboratory systems may not reproduce the human disease context fully.
Association is not causation#
Genotype is fixed at conception, which reduces some forms of reverse causation, but a GWAS association can still be noncausal. LD, population structure, and selection can contribute. So can participation bias, phenotype error, and assortative mating. So can indirect parental effects and technical artifacts.
Mendelian randomization uses genetic variants as instruments to study whether a modifiable factor may causally affect an outcome. It adds a design, not certainty. Instruments must be associated with the factor, not share causes with the outcome, and affect the outcome only through the proposed pathway. Pleiotropy and weak instruments can violate these assumptions.
Triangulation is stronger than relying on one method. Consistent evidence from GWAS, fine-mapping, and colocalization can build a causal case. So can molecular assays, animal or cellular models, and pharmacology. So can natural experiments and clinical trials. Each source has different biases.
Even a causal molecular target does not guarantee a safe, effective therapy. Direction, timing, and tissue matter. So do developmental effects, dose-response, and off-target consequences. Lifelong small genetic differences are not identical to starting a potent drug in adulthood.
Common diseases are usually polygenic#
Many common traits reflect thousands of variants with small effects plus environment, development, behavior, and chance. GWAS repeatedly shows that genetic contribution is distributed rather than confined to one “disease gene.”
Heritability estimates the proportion of trait variation associated with genetic differences in a particular population and environment. It does not tell you what fraction of one person's trait is genetic, whether a trait can be changed, or whether group differences are genetic.
GWAS heritability may be lower than estimates from family designs because arrays capture variants imperfectly, rare and structural variants are missed, interactions are difficult to estimate, and family estimates can include shared environment or other assumptions. “Missing heritability” is a gap between models, not evidence for a single missing gene.
Environmental change can alter disease prevalence even when genomes are stable. Genetic effects can also depend on diet, medicines, age, sex, or other contexts. Most GWAS are not powered to map all interactions reliably.
From GWAS to polygenic scores#
A polygenic score sums trait-associated alleles after weighting them by effect estimates from GWAS or related models. It can stratify relative genetic liability within a population. It is not a separate measurement of destiny.
Development choices include which GWAS to use, LD adjustment, and variant selection. They include ancestry composition, phenotype, and tuning data. Validation must be kept separate. A score can rank people reasonably while its absolute probabilities are miscalibrated.
Incremental value matters. A diabetes score should be compared with age, family history, and body measures. It should be compared with laboratory data and existing risk models, not with no information. Decision-curve or impact analysis may be needed to show whether reclassification changes beneficial action.
Clinical utility goes beyond prediction. A score needs a defined use, effective intervention, and acceptable harms. It needs understandable communication and evidence that acting on it improves outcomes. Reclassification can also lead to anxiety, unnecessary testing, false reassurance, or inequitable access.
Why ancestry representation matters#
GWAS discovery datasets have historically overrepresented people with European genetic ancestry. This limits discovery and can make polygenic scores less accurate in other populations. Different allele frequencies and LD patterns mean that weights learned in one dataset do not transfer uniformly.
A 2023 Nature study showed polygenic-score accuracy varying continuously across genetically inferred ancestry rather than along a few neat categories, and that finding argues against treating broad labels as uniform calibration groups.
Increasing diversity improves science, but sample size alone is not enough. Community partnership, consent, and governance matter. So do benefit sharing, phenotype quality, and suitable reference panels. So do local researchers and careful interpretation. Historical misuse of genetics makes trust and accountability essential. Multi-ancestry methods can share information across populations and improve power, but they still require transparent evaluation in the population you intend to use them in, including people with mixed ancestry and groups too small for stable averages.
Reading a Manhattan plot#
A Manhattan plot places genomic position along the horizontal axis and -log10(p) on the vertical axis. Each point is a tested variant. Tall clusters are regions with strong statistical evidence. The alternating chromosome colors help you keep track of where you are.
One tower can contain many correlated points rather than many separate causal variants. Conditional analysis asks whether more than one association remains after accounting for the lead signal. Locus plots add LD, genes, and annotations.
A quantile-quantile plot compares observed with expected p-values. Early broad deviation can indicate population structure or technical inflation; upward deviation only in the tail is more compatible with true polygenic signals, though interpretation depends on sample size and architecture. Plots are summaries. Credibility still depends on phenotype, sample flow, and quality control. It depends on ancestry methods, effect estimates, replication, and availability of summary statistics.
Reporting and reproducibility#
STREGA extends observational reporting guidance for genetic association studies. Useful reports specify participant selection, phenotype, and genotyping and quality control. They specify population structure, relatedness, and the multiple-testing approach. They specify replication and limitations.
The NHGRI-EBI GWAS Catalog curates eligible published associations and, where available, summary statistics. A catalog entry records what a study reported; it does not certify causality or clinical utility. You still have to inspect the original study and its metadata.
Open summary statistics enable replication, meta-analysis, fine-mapping, and score development. But privacy and consent constraints may limit sharing. Harmonization must preserve allele orientation, genome build, ancestry, sample overlap, and phenotype definitions.
Prepublication analysis plans can distinguish confirmatory tests from exploration. Code, variant filters, and complete results make analytic flexibility visible. A celebrated lead signal is more informative when the full search space and null findings are available.
What a careful claim sounds like#
“Variant X is genome-wide significantly associated with trait Y in this meta-analysis, with replication in a prespecified dataset” is a statistical claim. “The locus implicates pathway Z” is a biological inference that should cite fine-mapping or functional work. “Targeting Z will prevent disease” is a therapeutic hypothesis requiring much more evidence.
For prediction, state population, score version, and comparator. State discrimination, calibration, and decision effect. For mechanisms, state the causal chain and which links remain uncertain. For group comparisons, separate genetic ancestry, social conditions, and sampling.
GWAS has transformed understanding of complex-trait architecture and uncovered unexpected pathways. Its power comes from scale and systematic scanning. Its discipline comes from refusing to turn association into a gene story, individual diagnosis, or causal claim before the necessary evidence exists.
Sources#
- NHGRI fact sheet on genome-wide association studies
- NHGRI-EBI GWAS Catalog overview
- EMBL-EBI introduction to GWAS and linkage disequilibrium
- STREGA reporting statement for genetic association studies
- Study of polygenic-score accuracy across the ancestry continuum
- NHGRI fact sheet on participation in genomic research
Questions and answers
Does a genome-wide significant variant cause the trait?
Not necessarily. The marker may tag a correlated causal variant or reflect residual bias. Fine-mapping, replication, functional evidence, and other causal methods are needed before assigning mechanism.
Why do GWAS use such a strict significance threshold?
Millions of tests would otherwise generate many chance findings. A threshold near 5 x 10^-8 is commonly used for common-variant scans, with quality control, replication, and effect estimates still required.
Is a GWAS result the same as a polygenic risk score?
No. GWAS produces association estimates. A polygenic score combines many estimates into a predictive model and needs separate tuning, validation, calibration, comparison, and clinical-utility evidence.
Why can genetic prediction differ across ancestry groups?
Discovery representation, allele frequencies, LD patterns, phenotype definitions, environmental context, and model assumptions differ. Accuracy can vary continuously, so every intended population needs direct evaluation.
Can GWAS results guide an individual's treatment?
Usually not alone. Clinical use requires a specific validated test and action pathway that adds value beyond existing information, performs fairly, and improves meaningful outcomes with acceptable harms.