A nonsignificant result is usually an unresolved result, not a demonstration that nothing happened. To know what a study actually excludes, read the estimated effect and its confidence interval, decide what size of effect would matter to you, and only then look at the p value.
Key points#
- Statistical significance asks whether the data are difficult to reconcile with a specified null model. It does not decide whether an effect exists.
- A confidence interval shows the range of effect sizes reasonably compatible with the data under the model and assumptions used.
- A wide interval that spans meaningful benefit and harm is inconclusive, even if its p value is above 0.05.
- Evidence that an important effect is absent requires a narrow interval relative to a prespecified margin, often through an equivalence design.
- Study quality, missing data, outcome choice, and adherence can add uncertainty that a confidence interval alone does not capture.
Begin with the clinical question, not the threshold#
Imagine a trial comparing two ways to prevent hospital readmission. The estimated risk difference is 2 percentage points in favor of the new program, with a 95 percent confidence interval from 5 points worse to 9 points better, and the corresponding test is not statistically significant.
It would be inaccurate to conclude that the programs work equally well. The data remain compatible with modest harm, no difference, and a worthwhile benefit. The study did not settle the question because its estimate is too imprecise.
Now imagine the same estimate with an interval from 0.5 points worse to 4.5 points better. That interval is much tighter. Whether it rules out an important difference depends on what patients, clinicians, and health systems considered meaningful before seeing the results. If a 3-point improvement would justify the program's burden and cost, the second result still leaves an important benefit possible. If only a 10-point improvement would matter, it rules that larger benefit out. So clinical importance has to be defined on the outcome scale before the interval can be read at all, because a significance threshold cannot supply that judgment for you.
What a p value actually contributes#
A conventional p value is calculated under an assumption such as no average treatment difference; it asks how unusual the observed data, or data more extreme, would be if that assumption and the statistical model were correct. A p value above 0.05 says the observations were not unusual enough to cross the chosen threshold.
It does not say there is a 95 percent chance that the null hypothesis is true. It does not measure the probability that the treatment is ineffective. It does not distinguish a precisely estimated trivial effect from an imprecisely estimated important one.
Two studies can therefore produce the same nonsignificant p value while carrying very different information: a large, well-run trial may tightly exclude the effects that matter, while a small trial may allow almost every clinically plausible outcome. Labeling both simply “negative” erases the distinction. One word, two very different studies.
Precision is the missing dimension#
The confidence interval puts the estimate and its uncertainty on the same scale as the outcome. For a risk difference, that may be percentage points. For a blood-pressure outcome, it may be millimeters of mercury. For a time-to-event outcome, it may be a hazard ratio or a difference in restricted mean survival time.
A useful interval reading has three steps:
- Locate the point estimate. This is the single effect most consistent with the fitted model, not a guarantee.
- Inspect both interval limits. Ask whether they include meaningful benefit, meaningful harm, or both.
- Compare the limits with a prespecified minimum important difference. The question is not merely whether zero lies inside the interval.
The interval is still conditional on the study's design and model, and it reflects sampling uncertainty but does not automatically account for selective reporting, outcome misclassification, loss to follow-up, protocol deviations, or a biased comparator. A narrow interval around a biased estimate can be precisely wrong.
Why small studies often cannot answer “no difference”#
Power is the long-run probability that a planned test will detect an effect of a specified size when that effect is present, assuming the design and analysis behave as planned. If a study has 80 percent power for its target effect, it still has a 20 percent chance of missing that effect under those assumptions. Smaller true effects are even harder to detect.
This makes nonsignificance unsurprising in a small study. Adding participants usually narrows the confidence interval because it reduces random uncertainty, and that may reveal a real effect, or it may establish that any remaining effect is too small to matter. Either outcome is more informative than the original wide interval.
Post hoc calculations of “observed power” based on the study's own effect estimate usually add little. They are largely a transformation of the p value and can encourage circular reasoning. The interval already displays the relevant precision. The more productive questions you can put to the paper are whether the planned sample size was reached, what event count was assumed, how much information was lost, and which effect sizes remain compatible with the result.
How to test for practical sameness#
A conventional superiority trial is designed to find evidence of a difference. Failing that test does not reverse the burden of proof and establish sameness. If the real question is whether two options are similar enough for practical purposes, the design should say so before data collection.
An equivalence trial defines lower and upper margins that represent the largest differences considered unimportant, and evidence for equivalence requires the confidence interval to fall entirely inside those margins. The margins need clinical justification, not convenience chosen to make the result pass.
A noninferiority trial asks a one-sided question: is the new option no worse than the comparator by more than an acceptable margin? Such trials require particular care because weak adherence, crossover, and an ineffective comparator can make treatments appear similar. Both intention-to-treat and per-protocol analyses are often informative, and agreement between them is more reassuring than either alone.
Bayesian methods and Bayes factors can also quantify evidence favoring a region of negligible effect, provided the prior assumptions and decision region are transparent. The common principle is that evidence for absence must be designed and analyzed as its own question.
A disciplined way to read a null result#
When an abstract tells you “there was no difference,” replace that sentence with a short audit:
- What outcome and effect measure were used?
- What was the estimate, not just the p value?
- Which clinically important effects remain inside the interval?
- Was the sample size based on a credible event rate and target difference?
- Did the study reach its planned enrollment and event count?
- Were missing data, crossover, or measurement error likely to dilute a real effect?
- Was this a superiority, equivalence, or noninferiority question?
- Are the conclusion and title more definite than the interval permits?
The 1995 warning by Altman and Bland remains useful because it is a rule of logic, not a statistical fashion. Not detecting a signal may reflect a small effect, a noisy measure, too little information, or a genuinely negligible effect. The interval and the design are what let you separate those possibilities.
Sources and further reading
Questions and answers
Does p greater than 0.05 mean the treatments are equal?
No. It means the data did not cross the study's chosen significance threshold under the specified test. Equality requires a separate argument based on precision and a clinically justified margin.
Can a very large study find a statistically significant but trivial effect?
Yes. With enough information, a tiny difference can yield a small p value. That is why effect size and clinical importance must be read alongside statistical evidence.
When is a nonsignificant result genuinely informative?
It is informative when the interval is narrow enough to exclude effects that would change decisions, the outcome and analysis are credible, and important biases are unlikely. In that setting, the study can support evidence of no meaningful effect, even though it cannot prove an effect is exactly zero.