An association can be completely real, measured with care and reproduced many times, and still point to the wrong explanation. The most common reason is confounding: a third factor that is linked to the suspected cause and also drives the outcome on its own, so the two look connected when neither is moving the other. Getting into the habit of asking "what else differs between these groups" is one of the most useful skills for reading health research, because it separates a finding worth acting on from a coincidence dressed as a cause.
The definition worth memorizing#
A confounder is a factor that is associated with the suspected cause, affects the outcome on its own, and is not merely a step on the causal path between the two. Ignore it, and it can manufacture an association from nothing, hide one that is genuinely there, or make a true effect look larger or smaller than it is. It is the mechanism by which correlation gets to pose as cause.
Notice the three parts, because a factor has to satisfy all of them. It must connect to the thing you suspect, it must have its own line to the outcome, and it must not sit on the causal chain (if it does, it is a mediator, and adjusting for it would be a mistake rather than a fix).
A clean example: lighters and lung cancer#
Suppose you collected data on thousands of adults and found that people who carry a cigarette lighter develop lung cancer far more often than people who do not. The association is real and it will replicate. It is also useless as a cause, because a lighter has never harmed a lung. What links the two is smoking. Smokers are the ones carrying lighters, and smoking is what raises lung cancer risk. Smoking is associated with the suspected cause (the lighter), it affects the outcome on its own (cancer), and it sits off to the side rather than on any chain from flint to tumor. Split your data so you compare smokers only against other smokers, and the lighter signal disappears.
A second version is even harder to see because the confounder is age. Older people are more likely to have gray hair, and older people are more likely to die in any given year, so gray hair correlates with mortality. Nobody blames the hair. Real confounders in health research are seldom this obvious, and they braid themselves into nearly every comparison that matters.
Why careful research trips on this too#
The dangerous confounders in clinical work live inside the reasons people ended up in one group rather than another. Imagine an observational study reporting that patients taking a particular medicine do better than patients not taking it. The people who are offered a treatment, accept it, and keep filling the prescription tend to differ from those who do not: they may have more contact with the health system, more money for co-pays, or milder illness to begin with. Each of those traits improves outcomes on its own, so the raw comparison cannot tell you whether the medicine helped, did nothing, or caused mild harm. The groups were already unequal before the first dose. This is not a story about sloppiness; the question is genuinely hard, because you can never observe the same patient in the parallel world where they skipped the drug.
A sharper form, called confounding by indication, is where the very reason a treatment was chosen also predicts the outcome. The sickest patients get the most aggressive drugs, so in raw data those drugs can look harmful, not because they hurt anyone but because they were given to people already in trouble. The apparent effect can flip direction entirely. This is a central problem in pharmacoepidemiology, the study of how medicines behave once they leave the trial and enter real populations, and it is one reason a drug that looked dangerous in a database can look neutral or protective once the sicker patients are accounted for. A common repair is the active-comparator design, which compares one treatment against another drug used for the same condition rather than against no treatment at all, so the two groups start out more alike.
The defenses, from strongest to most fragile#
Randomization#
The most powerful protection is to decide who receives the intervention by chance. In a randomized controlled trial, something equivalent to a coin flip assigns each participant to a group. Because assignment is random, the groups end up balanced on average across every confounder, including the ones nobody thought to measure and the ones nobody has even named. That last point is what makes randomization special: adjustment can only handle confounders you listed, while randomization neutralizes the unknown ones too, by construction. It is also why researchers randomize instead of comparing, say, clinics that volunteered for a program with clinics that did not, since volunteer clinics differ in funding, culture, and baseline performance in ways no one can fully list.
Statistical adjustment#
When randomizing would be impossible or unethical, researchers measure the confounders they can anticipate and account for them with methods like regression, stratification, or matching. Done well, this removes a lot of bias. The limit is worth stating bluntly: adjustment only neutralizes confounders you thought of and measured accurately. One you missed, or one you measured crudely, passes straight through and leaves residual bias behind, sometimes dressed in the language of rigor. So when a study rests entirely on adjustment, the useful question is which confounders it could not measure, and which direction those would push the result.
Design choices made before any data arrive#
Some of the best protection is built into the blueprint. Restricting a study to one narrow group removes a confounder by holding it fixed. Matching each person who has the risk factor to a similar person who does not builds balance in from the start. A more inventive approach, Mendelian randomization, uses a genetic variant fixed at conception as a stand-in for a risk factor, on the logic that the genes you are born with cannot be shaped by the habits you later adopt, so they are shielded from most confounding.
What to do as a reader#
You do not need to run any statistics to read defensively. When you meet a claim that A affects B, first ask whether the comparison came from a randomized trial or from observation, because that single fact changes how much trust the association has earned. If it was observational, ask what the authors adjusted for and what they admit they could not. A study that names its likely confounders and discusses the ones it could not rule out is being honest with you. One that reports a strong association and stops there has handed you a correlation and invited you to supply the causation yourself.
The goal is not to distrust every association, because plenty of them are causal. A real, repeatable link is where an investigation starts, not where it ends, and the whole craft is in asking what else was in the room.
Sources and further reading
Questions and answers
Is a confounder the same as bias?
Confounding is one type of bias, but not the only one. Bias is any systematic error that pushes a result away from the truth. Selection bias (who got into the study) and measurement bias (how accurately things were recorded) are separate problems that adjustment for confounders will not fix.
If a study adjusted for many variables, is the result trustworthy?
Not automatically. A long list of adjusted variables shows effort, not completeness. What matters is whether the important confounders were among them and measured well, and whether any major one was left out. A short, well-chosen list can beat a long, careless one.
Can you ever prove causation without a randomized trial?
Sometimes, when the evidence converges: a consistent effect across different populations and methods, a dose-response pattern, a plausible mechanism, and results from designs like Mendelian randomization that are hard to confound. No single observational study proves cause on its own, but a body of them can build a strong case.