Evidence explainer

Evidence and research methods

The Table 2 Fallacy: One Model, Many Wrong Questions

Regression software gives every predictor a coefficient. That symmetry is mathematical, not causal: a model built to answer one question can mislead you on every other row.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Begin with the estimand, not the software output
  2. A simple example: exercise, blood pressure, and stroke
  3. Confounders belong to relationships, not variables
  4. Colliders can manufacture an association
  5. Mediators change the question
  6. Noncollapsibility can change coefficients without confounding
  7. Prediction has a different contract
  8. Statistical significance cannot assign meaning
  9. Interactions require defined scales
  10. Better reporting separates jobs
  11. A reader's audit of any coefficient table
  12. References

A multivariable regression table has a seductive visual order. Each row names a variable, and each row receives a coefficient, confidence interval, and p value. Nothing stops you from reading the rows as parallel answers: the adjusted effect of treatment, the adjusted effect of age, the adjusted effect of smoking, and so on.

The Table 2 fallacy is the mistake of giving those rows the same causal meaning. A model is usually designed around one focal question. The other variables were included to control confounding, improve precision, encode design, or support prediction. Their coefficients may be adjusted for the wrong variables, conditioned on their consequences, or stripped of part of the effect.

Begin with the estimand, not the software output#

An estimand is the quantity the study seeks to learn. Examples include the average effect of assigning one treatment rather than another over one year, the total effect of smoking on stroke risk, or the direct effect of an intervention not operating through blood pressure.

Those are different questions. The total effect of a variable should not be adjusted for mediators on its causal pathway, and a controlled direct effect requires a defined intervention on the mediator and stronger assumptions. A prediction model may not seek any causal effect at all.

Regression software does not know which quantity is intended. It estimates parameters under a mathematical specification. The analyst supplies causal meaning through design, timing, assumptions, and variable selection.

Westreich and Greenland named the fallacy in 2013 to show why estimates for covariates in a model cannot be interpreted automatically as adjusted causal effects, and the problem is not the physical table number. It is the transfer of one model's adjustment logic to every row.

A simple example: exercise, blood pressure, and stroke#

Suppose the focal question is the effect of an exercise program on stroke. Age, baseline disease, and smoking may be included to address confounding. Blood pressure measured after the program begins may lie on the pathway from exercise to stroke.

If the model adjusts for follow-up blood pressure, the exercise coefficient no longer represents the total effect operating through all pathways. It is conditional on a postintervention variable and may also acquire bias if unmeasured causes affect both blood pressure and stroke.

Now look at the blood-pressure row. Its coefficient is not automatically the causal effect of lowering blood pressure. Exercise can cause blood-pressure change and may affect stroke through other routes. Conditioning on exercise and the other covariates creates a different quantity. Medication, diet, disease severity, and measurement timing may confound the blood-pressure question in ways the exercise model did not address. So the two rows sitting one above the other, in the same typeface, under the same column headings, are answers to questions that needed different models. The formatting hides that.

Confounders belong to relationships, not variables#

A variable is not inherently “a confounder.” It is a confounder for a particular causal contrast in a defined setting. Age may confound one association, mediate another through aging-related processes, or modify an effect.

Directed acyclic graphs, or DAGs, make the proposed structure visible. Nodes represent variables, arrows represent assumed causal directions, and the graph helps identify paths that should be blocked or left open. The useful work occurs before the diagram is drawn: define time order, distinguish common causes from consequences, and state which pathways the estimand includes.

A minimally sufficient adjustment set blocks noncausal backdoor paths without conditioning on unnecessary descendants or colliders. Adding every measured predictor is not a safer default. More adjustment can increase rather than reduce bias.

Colliders can manufacture an association#

A collider is caused by two other variables. Conditioning on it can create an association between its causes even if they were otherwise unrelated.

Imagine studying the relation between disease severity and treatment response only among people admitted to a specialty center. Admission may be caused by severe disease and by strong referral access. Restricting to admitted patients can associate severity with access. If access also affects follow-up or outcome, a distorted relation appears.

In regression, conditioning includes stratification, restriction, matching, and placing the collider in the model. The coefficient table will not tell you that a noncausal path was opened. A small standard error may make the biased result look especially convincing. The article on confounding and causation develops related problems with additional examples.

Mediators change the question#

A mediator carries part of an effect. If treatment lowers blood pressure and lower blood pressure reduces stroke, blood pressure mediates some treatment benefit. Adjusting for it removes or distorts that pathway depending on the design and assumptions.

You will often see the remaining treatment coefficient described as “the effect independent of blood pressure.” That phrase is vague. Is it a controlled direct effect under a hypothetical intervention fixing blood pressure? A natural direct effect comparing mediator values that would have occurred under another treatment? A conditional association among people with the same measured value? These quantities differ. Mediator-outcome confounding, measurement error, interactions, and treatment-induced confounders complicate direct-effect analysis. A standard regression with the mediator added is rarely sufficient to justify a broad mechanistic claim.

Noncollapsibility can change coefficients without confounding#

Odds ratios and hazard ratios can differ between adjusted and unadjusted models even when the added covariate is not a confounder. This property, called noncollapsibility, means a coefficient change is not direct evidence that confounding was removed.

For logistic regression, the conditional odds ratio among people with the same covariates can differ from the marginal population odds ratio, and both may be mathematically valid and answer different questions. Comparing coefficients across models as though they differ only because one is “better controlled” can mislead. Risk differences and risk ratios may be easier to map to population decisions, though they still require sound design. Standardizing predicted outcomes to the target population can produce interpretable marginal contrasts.

Prediction has a different contract#

In a prediction model, the goal may be accurate estimation of future risk rather than causal explanation, and a variable can improve prediction even if it is a consequence of disease or a proxy for another process. Removing it because it is not causal may lower predictive performance.

Predictor coefficients still should not be read casually. Correlated predictors can redistribute weight, penalization changes estimates, nonlinearities matter, and transport to another population can fail. Variable importance is model- and data-dependent.

The error occurs when a model optimized for prediction is presented as evidence that changing a predictor will change the outcome, and a risk marker can identify who is at risk without being a useful intervention target.

Statistical significance cannot assign meaning#

A p value tests compatibility between data and a statistical model under a null hypothesis; it does not test whether adjustment was causally appropriate, whether the model measured the intended construct, or whether bias is small.

Selecting “independent predictors” because p is below 0.05 turns a continuous and design-dependent measure into a causal label. Confidence intervals describe sampling uncertainty under assumptions, not uncertainty from unmeasured confounding, selection, misclassification, or selective analysis.

Greenland and colleagues recommend reporting estimates, interval bounds, and the assumptions needed for interpretation instead of sorting results into significant and nonsignificant; a wide interval can contain important benefit and harm. A narrow interval around a biased estimate remains biased.

Interactions require defined scales#

An interaction means that the effect of one variable differs across levels of another on a chosen scale. Additive interaction concerns differences in absolute risk; multiplicative interaction concerns ratios. A product term in logistic regression tests one particular scale and functional form.

The coefficient for a main term in a model with an interaction is conditional on the other interacting variable's reference value, often zero. It is not the overall effect. Centering and clinically meaningful contrasts improve interpretation.

Effect modification is a property of the causal effect in a target population, while statistical interaction also depends on modeling choices. Tables should report stratum-specific or standardized estimates rather than asking you to reconstruct them from several coefficients.

Better reporting separates jobs#

First, state the focal question and estimand. Explain why each adjustment variable was included, ideally with a causal diagram or design rationale. Report the focal estimate in an interpretable measure with absolute risks where possible.

Second, label other rows honestly. If they are adjustment coefficients with no causal interpretation, omit them from the main table or place them in a supplement, and if they support prediction, describe them as model parameters and report validation metrics.

Third, when a secondary variable has substantive causal importance, analyze it as a new question. Redefine baseline, time zero, intervention, comparator, confounders, and outcome. The new question may require a different sample or design.

Lederer and colleagues recommend explicit reporting of confounder control and caution against selecting adjustment variables solely by statistical criteria. STROBE reporting can improve transparency, but a complete report cannot rescue a conceptually wrong model.

A reader's audit of any coefficient table#

Circle the one row the study was designed to estimate. Write the causal question in words. For each other row, ask why the variable entered the model and whether it occurs before or after the intervention you would have to imagine for that row. Look for mediators, colliders, proxies, and interactions. Determine whether the measure is conditional or standardized.

Then read the methods for selection, missing data, functional form, and sensitivity analysis. The related guide on confounding and causation provides a broader framework.

The Table 2 fallacy is avoided by refusing to let visual symmetry replace causal reasoning. One regression can contain many numbers. It does not automatically contain many valid causal answers.

References#

  1. The Table 2 fallacy
  2. Causal Inference: What If
  3. Statistical tests, p values, confidence intervals, and power
  4. Control of confounding and reporting in causal studies
  5. Scoping review of the Table 2 fallacy
  6. Directed acyclic graphs and adjustment sets

Questions and answers

Why is it called the Table 2 fallacy?

Many articles place adjusted coefficients in a second table and interpret every row causally, even though the model and covariates were chosen for only one focal question.

Are coefficients for adjustment variables always useless?

No. They may be useful descriptively or for prediction, and sometimes a covariate has a valid causal interpretation. That interpretation needs its own estimand and adjustment reasoning.

Does statistical significance tell whether a row is confounded?

No. A small p value does not repair a wrong adjustment set, and a large p value does not prove that no meaningful association exists.

Can one regression model answer several causal questions?

Sometimes, but only if the causal structure makes the required adjustment set valid for each question. That condition should be argued rather than assumed.

How should authors report secondary predictors?

Label the focal estimand, separate descriptive or predictive coefficients, and fit question-specific causal models when secondary effects are genuinely of interest.