Evidence explainer

Evidence and research methods

How to Interpret an E-Value Without Overclaiming

An E-value asks how strong a hidden confounder would have to be to explain away a result. It is a stress test, not evidence that none exists.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The question in plain language
  2. Why the confidence-limit E-value matters
  3. An E-value has no universal grading scale
  4. What the calculation assumes
  5. A hidden confounder is not just any predictor
  6. Common misreadings
  7. How to audit an E-value report
  8. A worked interpretation

Observational studies compare groups that were not assigned at random. Even after adjustment, an omitted cause may differ between the groups and also affect the outcome. The E-value does not remove that concern. It puts one boundary around it, asking how strong those two relationships would have to be to explain away an estimated association under the method's assumptions.

The question in plain language#

Suppose an observational analysis reports that a treatment is associated with a risk ratio of 2.0 for an outcome after measured covariates are adjusted. The E-value asks: what minimum relationship would a hidden confounder need with treatment and with outcome, beyond the measured covariates, to reduce that observed result to no association?

For a risk ratio of 2.0, the point-estimate E-value is about 3.4. Under the bounding framework, an unmeasured factor would need risk-ratio relationships of at least 3.4 with both treatment and outcome, or a combination at least that strong, to fully explain the estimate. The calculation does not assert that the relationships are equal in reality. It summarizes a worst-case boundary in one number.

For risk ratios above 1, the common formula is the observed risk ratio plus the square root of that ratio multiplied by one less than the ratio. Protective estimates below 1 are first placed on the reciprocal scale. Software can handle other effect measures through conversions, but those conversions add assumptions you should be shown.

Why the confidence-limit E-value matters#

The point-estimate E-value asks how much confounding would move the center of the estimate all the way to no effect. A study's uncertainty may already place one confidence limit close to no effect. In that case, much weaker confounding could make the result statistically compatible with the null.

Reporting an E-value for the confidence-limit nearest the null shows how much hidden confounding would be needed to move the interval across that boundary. If the estimate's E-value is 3.4 but the confidence-limit E-value is 1.3, describing the finding as robust would omit an important part of the picture.

The threshold need not always be exactly no effect. A decision may hinge on whether an association remains above a clinically important risk ratio. Sensitivity analysis should align with the actual claim, not mechanically use one target.

An E-value has no universal grading scale#

A value of 2 may be difficult to explain away in one setting and ordinary in another. Consider an omitted variable such as disease severity. In a treatment study, severity may be strongly related to treatment choice and outcome. In another setting, available measurements may already capture severity well, making a remaining association of that magnitude less plausible.

Benchmarking helps. Compare the E-value with relationships observed for measured covariates and ask whether any realistic omitted factor could be stronger than the ones you can see. This comparison is not definitive because measured covariates may be imperfect proxies, and several modest confounders can act together. Still, it is more informative than calling an E-value "large" without context.

Prevalence matters to plausibility even though the headline E-value does not report it. A very rare hidden factor cannot generally explain a common outcome pattern in the same way as a prevalent factor with similar relative associations. More detailed bias analyses can specify prevalence, direction, and multiple confounders rather than compressing them into one bound.

What the calculation assumes#

The E-value is most naturally interpreted when the study estimates a causal risk-ratio contrast under a clearly defined intervention and covariate set. Several assumptions sit behind that interpretation:

A high value cannot rescue a study that fails these basics. If outcome coding differs between treatment groups, the relevant bias is measurement, not hidden confounding. If early symptoms influence treatment choice, reverse causation may dominate. If conditioning on follow-up creates selection bias, the E-value answers the wrong problem.

A hidden confounder is not just any predictor#

To confound the association, an omitted variable must relate to both the study factor and the outcome after the measured covariates are considered. A strong predictor of the outcome that is evenly distributed between treatment groups does not create confounding. A factor strongly related to treatment but unrelated to outcome does not either.

This two-sided requirement is the E-value's central insight. It pushes you to name a candidate variable and ask about both links. "Unmeasured lifestyle" is too vague. A useful argument would identify a specific behavior, explain why it differs between groups, quantify its association with the outcome, and assess whether existing covariates already capture part of it.

Direction also matters. Some unmeasured factors could strengthen rather than weaken the observed association. The standard E-value focuses on the amount of confounding needed to attenuate a result toward a threshold, not every possible bias direction.

Common misreadings#

"The E-value proves causation." It does not. It quantifies sensitivity to one class of unmeasured confounding under assumptions.

"No confounder this strong exists." The calculation cannot establish that. Subject knowledge and empirical benchmarks must support the claim.

"A value above 2 is robust." There is no universal cutoff. Context determines whether 2 is demanding or plausible.

"The measured covariates are weaker, so hidden confounding is impossible." An omitted construct may be measured more strongly than its recorded proxies, and several factors may combine.

"The E-value fixes a weak association." Because the E-value grows mechanically with the effect estimate, it does not supply new data. It helps interpret what confounding would be required.

How to audit an E-value report#

First, read the study design and directed causal assumptions. Decide which omitted causes are realistic before you look at the number. Then check:

  1. Is the original effect measure compatible with the method, and is any conversion explained?
  2. Are values reported for both the point estimate and the confidence limit nearest no effect?
  3. Is the target threshold clinically meaningful?
  4. Are measured covariates used as empirical benchmarks?
  5. Do authors discuss prevalence, measurement quality, and combinations of omitted factors?
  6. Are selection, misclassification, reverse causation, and missingness assessed separately?
  7. Does the conclusion remain cautious about causality?

The most persuasive use is embedded in a broader bias analysis. Negative-control outcomes, alternative definitions, quantitative scenarios, and analyses addressing missing data may each probe threats the E-value cannot.

A worked interpretation#

Suppose an adjusted estimate is 1.60, with a 95 percent confidence interval of 1.10 to 2.33. The E-value for the point estimate is about 2.58, while the value for the lower confidence limit is about 1.43. A balanced interpretation would say that moving the point estimate fully to no association requires stronger joint confounding than moving the interval to include no association. A modest omitted factor could be enough for the latter.

Next ask what variables could create those relationships. If a measured severity score changes both treatment selection and outcome by risk ratios near 2, a residual or poorly measured severity construct is not far-fetched. If all plausible covariates have much weaker relationships and are measured well, the result is less sensitive. The evidence comes from that comparison, not from the number alone.

Sources and further reading

  1. VanderWeele and Ding, introducing the E-value, Annals of Internal Medicine (2017)
  2. Haneuse, VanderWeele, and Arterburn, using the E-value in observational studies, JAMA (2019)
  3. Greenland, limitations and misinterpretations of E-values, Annals of Internal Medicine (2019)
  4. Ding and VanderWeele, sensitivity analysis without assumptions, Epidemiology (2016)

Questions and answers

Is a higher E-value always better?

It indicates that stronger unmeasured confounding would be needed for the chosen threshold. It does not evaluate other biases or study quality, so "better" is too broad.

Why report two E-values?

The point-estimate value addresses complete attenuation of the central estimate. The confidence-limit value shows how much confounding would move the statistical interval to the threshold and is often smaller.

Can several weak confounders matter together?

Yes. The E-value bound can represent an unmeasured construct that summarizes combined confounding, but it does not tell readers which factors, prevalences, or pathways create it. Scenario-based analyses provide that detail.