Evidence explainer

Evidence and research methods

What a Confidence Interval Is Not

A confidence interval shows how precisely a study estimated its effect, conditional on its data, model and assumptions. That is less than most readings assume.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key takeaways
  2. What a confidence interval does say
  3. Why 95 percent does not mean a 95 percent probability of truth
  4. A confidence interval is not a flat probability strip
  5. The null value depends on the effect scale
  6. Crossing the null does not prove no effect
  7. Excluding the null does not prove importance
  8. A narrow interval is not a quality certificate
  9. Replication is not expected to land inside the first interval
  10. Confidence levels do not fix analysis multiplicity
  11. A five-question reading method
  12. What to report alongside the interval
  13. References

A confidence interval is an estimate plus an account of sampling uncertainty under stated assumptions. It is not the probability that the scientific claim is true, not a promise that a replication will land inside the same limits, and not proof that values outside the interval are impossible. A 95 percent confidence procedure is designed to cover the target parameter in about 95 percent of repeated samples when the procedure's assumptions hold.

Once a particular interval has been calculated, the usual frequentist framework treats the parameter as fixed and the interval as the result that happened to be observed, and the interval either covers that parameter or it does not. The familiar “95 percent” describes the method's long-run performance, not a personal confidence score attached to this one range.

Key takeaways#

What a confidence interval does say#

Begin with the target. A study may estimate a mean difference, risk difference, risk ratio, odds ratio, hazard ratio, diagnostic sensitivity, prevalence, correlation, or model coefficient. The point estimate is the single value produced by the sample. A confidence interval places limits around it using a statistical procedure chosen for that target and design.

The interval is often described as a range of values reasonably compatible with the data under the model.[1] “Compatible” is useful because it avoids pretending that the limits are a probability distribution. It also keeps the assumptions visible. Compatibility can change when the outcome model, variance estimate, clustering, missing-data method, or adjustment set changes.

Width is a practical signal of precision. Larger samples, lower variability, more events, and efficient designs often produce narrower intervals. Small samples, rare events, heterogeneous measurements, and cluster dependence often produce wider ones, though the relationship is not simply “more participants equals precision.” In a survival analysis, the number and timing of events may matter more than the total enrolled. In a cluster trial, 2,000 participants in ten clinics do not provide the same information as 2,000 unrelated observations.

Why 95 percent does not mean a 95 percent probability of truth#

Imagine you repeated the entire sampling process many times from the same population, calculating an interval by the same recipe each time: some intervals would sit higher, some lower, some narrower, and some wider. Under the assumed model, about 95 percent of those intervals would contain the fixed target parameter, and about 5 percent would miss it.

The observed interval is one member of that sequence. Frequentist probability describes the random procedure before the sample is observed. After calculation, the limits are fixed numbers. Saying “there is a 95 percent chance that the true value is between 4 and 10” adds a probability statement that the standard interval did not calculate.

Everyday language makes this difficult because “confidence” ordinarily describes belief. Statistical terminology uses it as the name of a coverage procedure. Short teaching phrases often blur the distinction to make it memorable, but the blur can matter in decisions. You can read 95 percent as strong confirmation of an effect you already favor, even when the study design is biased or the model assumptions are doubtful.

A Bayesian credible interval answers a different question. Given a specified prior distribution, likelihood, data, and model, it can assign posterior probability to a parameter range.[2] That does not make it automatically preferable or objective. The prior and model need scrutiny. The important point is that confidence and credible intervals are not interchangeable labels for the same calculation.

A confidence interval is not a flat probability strip#

Another common picture treats all values inside the limits as equally likely and every value outside as impossible. A conventional confidence interval supplies neither conclusion. It reports a set defined by the procedure, not a uniform probability distribution over the parameter.

For many familiar symmetric models, values close to the point estimate fit the observed data better than values near the boundaries. Values just outside a 95 percent interval are not transformed into impossibilities. If the confidence level changed from 95 to 90 percent, the interval would narrow; at 99 percent it would widen. Nature did not change. The analyst selected a different long-run coverage target.

This is why quoting only the most favorable endpoint misleads you. Suppose a treatment's estimated risk difference is 4 percentage points with a 95 percent interval from 1 to 7 points. “The treatment may help by as much as 7 points” names the optimistic boundary but omits the estimate and the less favorable boundary; the complete statement preserves all three numbers and the population, outcome, and time horizon.

The null value depends on the effect scale#

For a difference, the null value is usually zero: a mean difference of zero, or a risk difference of zero, represents no difference between groups on that scale. For a ratio, such as a risk ratio, odds ratio, or hazard ratio, the null is one, because a ratio of one means equal modeled rates or odds, so the scale decides where the line sits.

An interval can therefore change visual meaning when the scale changes. A risk ratio interval from 0.70 to 0.95 excludes one and favors the numerator group. A risk difference interval from minus 8 to minus 1 percentage points also excludes zero, but it communicates absolute effect. Both may be correct descriptions of the same data. The absolute scale often makes the likely practical impact easier for you to judge.

Logarithmic scales are commonly used for ratios because ratio uncertainty is multiplicative and values cannot fall below zero. A symmetric interval on the log scale becomes asymmetric after conversion back to the ratio scale. That asymmetry is not an error. It reflects the geometry of the measure.

Crossing the null does not prove no effect#

Suppose an estimated risk reduction is 10 percentage points with a 95 percent confidence interval from a 2-point increase to a 22-point reduction. The interval includes zero, so a two-sided test at the corresponding conventional level would not reject the null. But the data remain compatible with clinically meaningful benefit, a small harm, and values between them. “No effect” is one allowed value, not the only conclusion.

This pattern often indicates limited precision. A small study can fail to exclude the null because it did not collect enough information, even when the point estimate is important. Calling such a result “negative” erases the width and endpoints. A genuinely informative null result requires an interval narrow enough to rule out the effects that would have mattered to you.

The same logic supports equivalence and noninferiority designs. Those studies define margins before the results are known. The relevant question is whether the confidence interval stays within the prespecified acceptable region, not simply whether it crosses zero or one. An ordinary superiority study that fails to reject a null cannot be relabeled as proof of equivalence.

The related article on statistical significance versus clinical importance explains why the threshold and the practical decision are separate.

Excluding the null does not prove importance#

With enough information, an interval can be narrow enough to exclude the null around a tiny effect. A risk reduction of 0.2 percentage points with an interval from 0.1 to 0.3 may be estimated precisely and still be too small to change a decision after burden, cost, adverse effects, or alternatives are considered.

Statistical precision answers “how tightly did this analysis estimate its target?” Clinical or policy importance asks “would values in this range alter what you should do?” The latter requires a meaningful-effect threshold, absolute risk, outcome severity, time horizon, patient priorities, and consequences of error.

Nor does null exclusion establish causation. A precisely estimated observational association can reflect confounding. A randomized trial can be biased by missing outcomes, unblinding, measurement problems, or selective analysis. The interval quantifies uncertainty within the selected statistical calculation. It does not automatically incorporate every source of bias.

A narrow interval is not a quality certificate#

Precision and validity are different. A bathroom scale that is consistently five pounds high may give nearly identical readings each morning. Its repeatability is high; its accuracy is poor. Statistical estimates can behave similarly.

Consider several routes to a narrow but untrustworthy interval:

Confidence intervals are conditional. They usually represent sampling variability if the design, data, model, and analysis behave as assumed. They do not contain a universal allowance for mistakes, bias, model search, or selective publication.

Replication is not expected to land inside the first interval#

A 95 percent confidence interval concerns a population parameter under a procedure. It is not a prediction interval for the estimate from the next study. A replication has its own sample, standard error, implementation, population, and random variation. Its point estimate can fall outside the first interval even when both studies are compatible with the same underlying effect.

Prediction intervals address future observations or study effects under a specified model. In a meta-analysis, a confidence interval around the pooled mean estimates uncertainty in the average effect, while a prediction interval tries to describe where the underlying effect in a new setting might fall. Heterogeneity can make the prediction interval much wider than the pooled confidence interval.

For visual context, how to read a forest plot explains how study intervals, pooled estimates, and weights appear together. The key is to identify which quantity each line describes.

Confidence levels do not fix analysis multiplicity#

A single 95 percent interval has its advertised coverage under its assumptions. If researchers calculate twenty unrelated 95 percent intervals, the chance that every one covers its target is less than 95 percent. Looking across many outcomes, time points, subgroups, and models raises the chance that at least one result appears unusually favorable by chance.

Simultaneous confidence procedures, multiplicity adjustments, hierarchical models, and prespecified analysis plans can address parts of this problem. Simply printing “95% CI” beside each selected result does not. Ask how many comparisons were possible, which were planned, and whether the interval was chosen after somebody looked at the data. The American Statistical Association has emphasized that scientific conclusions should not rest on a threshold alone and that transparent reporting, study design, context, and numerical summaries all matter.[4] Confidence intervals support that richer reading only when their full context is reported.

A five-question reading method#

When a paper hands you an estimate and an interval, ask:

  1. What is the estimand? Name the population, treatment or characteristic, outcome, time point, comparison, and handling of events after baseline.
  2. What scale is used? Difference, ratio, odds, hazard, correlation, or another measure changes the null and interpretation.
  3. What do both endpoints mean in practice? Translate the least and most favorable compatible values into absolute consequences when possible.
  4. What produced the width? Consider events, sample size, variability, clustering, missing data, and model choice.
  5. What uncertainty is missing? Bias, multiplicity, model selection, measurement error, transport to another population, and future-study variation may not be included.

Read the estimate before asking whether the interval crosses a threshold. Then compare the entire interval with a clinically meaningful region. This avoids reducing continuous uncertainty to a binary label.

Related guides cover reading antidepressant effect sizes, forest plots, and clinical versus statistical importance. The site's research overview places interval interpretation within a broader evidence-appraisal framework.

What to report alongside the interval#

A useful report gives the point estimate, confidence level, lower and upper limits, units, effect scale, analysis population, and method. It also provides counts or baseline risk when an absolute interpretation matters. Predefined meaningful-effect thresholds, sensitivity analyses, and missing-data information help the interval answer a decision rather than merely decorate a result.

The honest conclusion can then be specific: the analysis estimates a modest benefit but remains compatible with little benefit; it rules out the prespecified harmful range; or it is too imprecise to distinguish important benefit from harm. Those statements preserve uncertainty without pretending that nothing can be learned.

References#

  1. Statistical Tests, P Values, Confidence Intervals, and Power: A Guide to Misinterpretations
  2. Confidence and Credible Intervals Around Effect Estimates
  3. The Fallacy of Placing Confidence in Confidence Intervals
  4. American Statistical Association Statement on Statistical Significance and P Values

Questions and answers

Is there a 95 percent probability that the true value is inside a 95 percent confidence interval?

Not under the standard frequentist definition. The 95 percent describes how often the interval procedure would cover the fixed parameter across repeated samples under its assumptions. A Bayesian credible interval can support a posterior probability statement, but it comes from a different framework.

Does an interval that crosses the null prove no effect?

No. It means the analysis did not exclude the null at the corresponding confidence level. The interval may also include meaningful benefit or harm. Evidence for absence requires enough precision to rule out effects that would matter.

Is every value inside a confidence interval equally supported?

No. A confidence interval is not a probability distribution. In many common models, values nearer the point estimate fit the observed data better than values near the limits, while values just outside are not impossible.

Does a narrow confidence interval guarantee a valid result?

No. Narrow limits indicate high model-based precision. Bias from design, measurement, missing data, confounding, selective reporting, or incorrect dependence assumptions can remain and may not be represented in the interval.

How is a credible interval different?

A credible interval comes from Bayesian analysis. Conditional on the prior, likelihood, data, and model, it can state a posterior probability that the parameter lies in a range. Its interpretation depends on those inputs, which should be reported and evaluated.