Evidence explainer

Evidence and research methods

Understanding Selection Bias in Health Research

Selection bias is not about whether a sample resembles everyone. It is about whether who entered, who stayed, and what you conditioned on changed the comparison.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Selection changes the comparison
  2. Representativeness is a related but separate question
  3. Recruitment can build in distortion
  4. Collider bias in plain language
  5. Loss to follow-up and missing outcomes
  6. Conditioning on survival or diagnosis
  7. Administrative and electronic health-record data
  8. Trial eligibility and postrandomization selection
  9. Screening and diagnostic studies
  10. Clues in a research report
  11. Design and analysis responses
  12. A disciplined interpretation
  13. References

Selection bias is a distortion created by who enters an analysis, who remains under observation, or which subset is conditioned upon, and it can make two groups differ for reasons generated by the study process rather than by the treatment or factor under study.

A sample does not have to resemble an entire country to yield a valid comparison. A specialized clinic can support an internally valid study of its patients; the problem begins when selection changes the relation the study is trying to estimate, or when a result from one selected group is transported beyond what the data support.

Selection changes the comparison#

Imagine a study that compares a treatment with no treatment. People enter a specialty clinic partly because their illness is severe and partly because they have access to referral. If treatment and comparison groups come through different referral routes, severity, access, and clinician concern may differ before treatment begins.

Statistical adjustment can help when those determinants are measured accurately and modeled appropriately. It cannot guarantee repair when important determinants are absent, measured after selection, or represented by weak proxies.

Selection also occurs after enrollment. Participants may stop responding because they improve, worsen, move, lose insurance, dislike the assigned care, or experience an adverse event. Restricting analysis to those with complete data can change group composition.

The result is bias when the selected data no longer preserve the comparison needed for the target estimate. Direction is not automatic. Selection can exaggerate, attenuate, or even reverse an association.

Internal validity asks whether the comparison is credible for the studied population. External validity asks whether it transports to another population, setting, or period. A sample can do well on one and poorly on the other.

A randomized trial may enroll a narrow population and still estimate the randomized contrast without selection bias among participants, assuming conduct and follow-up remain sound. Applying that result to older patients with more comorbidity is a transport question.

Conversely, a database may include nearly every resident of a region and still produce a biased treatment comparison if treatment initiation, eligibility, and follow-up are misaligned. Population coverage does not replace causal design. So ask two separate questions: did selection distort the comparison inside the data, and does the estimate that came out of it apply to the people whose decision you are trying to inform?

Recruitment can build in distortion#

Volunteer studies can overrepresent people with time, health interest, digital access, or trust in research, and this becomes a selection-bias problem when willingness to participate depends on variables related to both the factor and outcome being compared.

Hospital control groups illustrate another risk. Patients admitted for a condition can differ from the source population that produced cases. If the reason for admission is related to the risk factor under study, the control group's frequency of that factor may be distorted.

Referral filters can create similar patterns. People reaching a tertiary center have passed through symptoms, clinician decisions, geography, insurance, and prior test results. Associations inside that center can differ from those in primary care because the selection pathway has conditioned the sample.

Clear source-population definitions help. Cases and controls should arise from a population in which any member who developed the outcome could have been identified as a case. Comparison groups in cohort studies should enter through compatible eligibility and observation processes.

Collider bias in plain language#

A collider is a variable caused by two other variables in a causal structure. Conditioning on it can make its causes appear associated.

Suppose both severe disease and strong social support increase the chance of entering an intensive rehabilitation program. Among program participants, people with less severe disease may appear more likely to have weaker support, because either severity or support helped them enter. Restricting the analysis to the program has connected two causes that need not have been associated in the source population.

Conditioning includes more than adding a variable to a regression. Restricting a sample, matching, stratifying, selecting complete cases, or choosing who appears in a dataset can all condition on a variable.

Not every selected variable is a collider, and not every collider creates large practical bias. The structure depends on the question and causal relations. A causal diagram can make assumptions visible, but drawing arrows does not validate them. Clinical and operational knowledge are still needed.

Loss to follow-up and missing outcomes#

Attrition is often summarized as a percentage. That percentage is incomplete. Five percent missing can be serious if it consists mainly of poor outcomes in one group. Twenty percent can sometimes have less effect if missingness is well characterized and plausibly unrelated to outcome after measured information.

In a randomized trial, excluding participants after randomization can erode the baseline comparability randomization created. Reasons for missing outcomes should be reported by group, so you can see who is gone and why, and the analysis should follow the estimand. Intention-to-treat preserves assigned groups but still needs a principled approach to missing outcomes.

Complete-case analysis assumes conditions that may not hold. Multiple imputation models missing data from observed information and propagates some uncertainty, but it cannot verify that unobserved outcomes behave as assumed. Inverse-probability weighting models the chance of remaining observed. Both approaches depend on model specification, positivity, and measured predictors of missingness. Sensitivity analyses can then ask how conclusions change under departures from missing-at-random assumptions, and a range of plausible outcomes for the missing participants tells you more than a claim that an imputation method solved the attrition.

Conditioning on survival or diagnosis#

Some studies include only people who survived long enough to be measured. If the studied factor affects survival and other causes of survival affect the outcome, restricting to survivors can distort associations.

Studies limited to people with a diagnosed disease can also create bias when both a factor and another cause influence diagnosis. Testing intensity is part of this process. People tested more often have more opportunities to receive the label, so an analysis of diagnosed cases can reflect detection pathways as well as disease.

Prevalent-user designs include people already taking a treatment. Those who had early intolerance, rapid failure, or an event may have stopped and disappeared before cohort entry. New-user designs often improve alignment by beginning follow-up at treatment initiation, although they do not remove all confounding or selection. That is why time zero is a design element: eligibility, assignment, and follow-up should begin coherently, rather than at convenient but different points for each group.

Administrative and electronic health-record data#

Large routine datasets are selected by care itself. A laboratory value exists because someone ordered the test. Follow-up is visible when a person returns to the same system. Medication records can miss cash purchases, samples, outside prescriptions, or actual use.

Requiring a recorded measurement can select patients based on clinician suspicion, illness, access, and prior results. Restricting to people with complete covariates may create a different population and association.

Linkage can help recover outside outcomes, but failed linkage may not be random. Death registries, claims, pharmacy data, and patient-reported information cover different processes and periods. Data provenance should say who can appear, when the records start, what causes capture, and what happens when care leaves the network.

Machine-learning scale does not remove this problem. A model can predict accurately within the selected system while failing for people missing from it. The external-validity guide separates performance in the captured population from transport to another setting.

Trial eligibility and postrandomization selection#

Before randomization, eligibility criteria shape the trial population. That may be scientifically appropriate. The transport claim should match the enrolled population.

After randomization, analyses restricted by adherence, treatment received, rescue therapy, or a postbaseline biomarker can break comparability. People who adhere may differ in prognosis: a per-protocol effect can be a legitimate estimand, but estimating it requires methods that handle time-varying confounding and selection rather than simply comparing adherers.

Composite outcomes and competing events add another layer. Restricting analysis to those free of a competing event can select a subgroup influenced by treatment and prognosis, so the question should name how competing events are handled, rather than have them deleted silently.

Cochrane's Risk of Bias 2 tool considers missing outcome data and deviations from intended interventions among its domains. ROBINS-I includes bias in participant selection and several other domains for nonrandomized intervention studies. These tools structure judgment; they do not convert incomplete reporting into low risk.

Screening and diagnostic studies#

Verification bias occurs when the reference standard is more likely to be performed after a positive index test or in sicker patients. Cases with negative tests may never receive definitive verification. Sensitivity and specificity can then be distorted.

Spectrum effects arise when a test is evaluated in clear cases and healthy controls rather than the mixed, ambiguous population where it will be used. This is often framed as applicability, but selection can also bias accuracy estimates when inclusion depends jointly on test results and disease.

Screening programs attract participants who may differ in baseline health behavior, access, and risk. Comparing screened volunteers with nonparticipants can therefore mix the effect of screening with selection. Randomized invitation designs or carefully designed quasi-experiments can provide stronger evidence.

Clues in a research report#

Start with the flow diagram. Count the people assessed, eligible, enrolled, assigned, followed, and analyzed, and at each transition ask why participants were lost and whether the reasons differ across groups.

Look for exclusions made after outcomes were known, complete-case restrictions, different data-availability requirements across groups, and follow-up that starts at different points. A table comparing retained and lost participants can help, but similarity on measured variables does not establish similarity on unmeasured ones.

Check whether eligibility depends on a variable influenced by treatment or prognosis. Ask where the dataset came from and which people could never appear in it. Then read the protocol or registration to see whether the analytic restrictions were planned.

STROBE improves transparent reporting of observational studies, but reporting compliance is not proof of unbiased design. The article on verification bias covers one diagnostic pattern, and confounding by indication addresses a distinct reason treated and untreated patients differ.

Design and analysis responses#

The strongest response is to prevent unnecessary selection. Recruit comparison groups from the same source process, align time zero, minimize burden, maintain contact, use multiple outcome sources, and collect reasons for nonparticipation and loss.

Prespecify the target population and estimand. Measure determinants of selection when possible. Sampling weights can reconnect a selected sample to a target population if selection probabilities are estimated adequately. Inverse-probability-of-censoring weights can address measured predictors of loss. Standardization and transport methods can help with external validity.

Every method needs assumptions. Extreme weights can produce unstable estimates, positivity fails when some types of people have essentially no chance of being observed, and missing determinants leave residual bias behind, so diagnostics, uncertainty, and sensitivity analyses should travel with the main estimate. Quantitative bias analysis can then show how strong an unmeasured selection process would have to be to change a conclusion, and several plausible scenarios are more credible than one corrected number.

A disciplined interpretation#

Do not ask only whether selection bias is possible. It almost always is. Ask which selection event could connect the factor and the outcome you care about, whether the direction is predictable, and how large the effect might reasonably be.

Then separate the claims. A study can estimate a valid association in one selected population without supporting a causal claim or broad transport, and it can also have excellent population coverage but a poorly aligned comparison.

What you are reading for is a traceable path from the source population to the analysis set. Each filter should come with a reason, a count, and an assessment of how it could alter the target estimate. The site's research overview places this audit beside measurement, confounding, and reporting checks.

References#

  1. Hernan, Hernandez-Diaz, and Robins: A structural approach to selection bias
  2. ROBINS-I tool for nonrandomized intervention studies
  3. Cochrane Risk of Bias 2 tool
  4. STROBE reporting guidance
  5. Using big data to emulate a target trial
  6. Sackett: Bias in analytic research

Questions and answers

Is selection bias the same as having a nonrepresentative sample?

No. Poor representativeness mainly limits transport to another population. Selection bias occurs when selection distorts the comparison or association being estimated within the analysis.

Can a large dataset remove selection bias?

No. More records reduce random error but do not repair a selection mechanism that creates a systematically distorted comparison.

What is collider bias?

Collider bias can arise when analysis conditions on a common consequence of two variables, creating an association between them even when none existed before conditioning.

Does loss to follow-up always cause selection bias?

No. Bias depends on why outcomes are missing and how that missingness relates to the compared groups and outcomes. The amount missing alone does not determine the direction or size of bias.

How can researchers reduce selection bias?

They can define eligibility and time zero clearly, recruit comparison groups through the same pathway, retain participants, measure reasons for missingness, avoid harmful conditioning, and use prespecified sensitivity analyses.