A trial reports that treatment lowered an outcome by 0.8 points with a p-value of 0.03. The result meets a conventional statistical threshold. Whether 0.8 points is noticeable, durable, worth the burden, or applicable to a particular person remains unanswered.
Another trial estimates a 15 percent relative reduction with a p-value of 0.08. Calling it "negative" also leaves out essential information. The confidence interval may include both a meaningful benefit and no important benefit, making the result inconclusive rather than proof of no effect.
Statistical significance and clinical importance answer different questions. Good interpretation keeps them separate, then connects effect size, uncertainty, outcome meaning, harms, and context.
What a p-value actually says#
The American Statistical Association defines a p-value in relation to a specified statistical model. If the null hypothesis and model assumptions were true, it measures how unusual the observed data, or more extreme data, would be for the chosen test.
For a simple comparison, the null might state that two population means are equal, and a p-value of 0.03 says that data at least this incompatible with that null would occur about 3 percent of the time under the model. It does not say there is a 3 percent probability that the null is true. That reverse probability would require a different framework and additional assumptions.
The p-value also does not say that chance was the only possible explanation. Bias, confounding, measurement error, missing data, protocol deviations, and model misspecification can produce small p-values. Randomization strengthens causal inference when preserved, but statistical significance does not diagnose whether randomization, blinding, follow-up, or outcome measurement was sound.
Why 0.05 is not a natural boundary#
The common threshold of 0.05 is a convention, not a biological law. Results with p-values of 0.049 and 0.051 generally provide very similar statistical information. Labeling the first a success and the second a failure creates an artificial cliff.
Prespecified thresholds can support error control in confirmatory research, especially when regulators or decision-makers need defined rules. The error rate belongs to the testing procedure across repeated studies under its assumptions, not to a claim that every result below the threshold is true. A threshold should therefore be reported as part of the design, while interpretation uses the estimate, precision, prior evidence, and consequences of error. Moving the threshold after results are seen undermines its error-control role.
Effect size is the starting point for importance#
An effect estimate describes the size and direction of a difference under the analysis. It may be a mean difference, risk difference, risk ratio, odds ratio, hazard ratio, or another measure. The choice affects interpretation.
For a symptom scale, a mean difference of 0.8 points has no inherent meaning without the scale range, direction, variability, baseline score, timing, and evidence about what change patients notice. A statistically precise laboratory change can remain clinically unimportant if it does not alter how people feel, function, or survive.
For events, relative and absolute measures answer complementary questions. Suppose risk falls from 2 percent to 1 percent over five years. The relative risk is 0.50, a 50 percent relative reduction. The absolute risk difference is 1 percentage point, or 10 fewer events per 1,000 people over five years, and if baseline risk were 20 percent and the same relative effect applied, the absolute reduction would be 10 percentage points. The relative effect can look stable while absolute benefit moves with baseline risk. You usually need both.
Confidence intervals show precision, with limits#
A confidence interval gives a range generated by a statistical procedure. Under repeated sampling and correct assumptions, a stated proportion of such intervals would contain the true parameter, and the usual 95 percent interval should not be read as a 95 percent probability that this realized interval contains the truth under a frequentist interpretation.
Intervals are useful because they show which effect sizes are compatible with the data and model, and a narrow interval around a trivial difference suggests that an important effect may be unlikely. A wide interval crossing no difference may include meaningful benefit and harm, signaling uncertainty.
Confidence intervals do not include every uncertainty. They typically reflect sampling variation under the model, not bias from missing outcomes, misclassification, nonadherence, selective reporting, or poor external validity. A precise biased estimate remains biased. The relation between a two-sided 95 percent interval and a p-value threshold near 0.05 can make the interval look like another binary test, when the better use is to hold the full range against effects you would call meaningful.
Non-significant is not the same as no effect#
When an interval includes the null, the data may be compatible with no difference. They may also be compatible with benefit or harm. The width and location determine what has been learned.
A small trial can fail to reach significance because few events occurred, not because groups are equivalent. Writing "there was no effect" overstates such a result. More accurate language reports the estimate and interval, then says whether important effects remain possible.
Equivalence and noninferiority studies ask different questions and require prespecified margins and suitable analyses. Failure to show superiority is not proof of equivalence. Likewise, a noninferiority conclusion depends on whether the margin is clinically acceptable, trial conduct preserved assay sensitivity, and analyses support the claim. Absence of evidence and evidence of absence converge only when the data are precise enough to exclude the differences that matter.
Large and small samples#
As sample size grows, standard errors often shrink. A very large study can produce a small p-value for a difference too small to matter. This is not a defect in the calculation. It reflects that the test addresses departure from the null, not importance.
A small study can estimate a large effect but remain imprecise. Early trials are especially vulnerable to unstable estimates, selective publication, and exaggerated apparent effects. A dramatic point estimate with a wide interval should not be treated as a promise.
Sample-size planning should be connected to the primary outcome and a target difference. ICH E9 and trial-reporting guidance emphasize prespecification of objectives, outcomes, and analysis. Power is the probability of meeting the statistical criterion under a specified alternative and assumptions. It is not the probability that a completed nonsignificant study missed a true effect. Event-driven studies add another layer: information depends heavily on the number of events, not only the number enrolled. Long follow-up and complete ascertainment can matter as much as recruitment.
What counts as a clinically important difference#
A minimum important difference is the smallest change considered meaningful to patients or decision-makers in a particular context; it can be estimated using patient ratings, external anchors, distributional methods, prior trials, and expert or stakeholder judgment. No method makes it universal.
An individual-level change threshold and a between-group mean difference are not interchangeable. If 60 percent of one group and 45 percent of another reach a meaningful personal improvement, the response-rate difference may be informative even when the mean shift looks modest. Conversely, dichotomizing a continuous outcome can discard information.
The FDA's patient-focused guidance on clinical outcome assessments emphasizes whether a measure is fit for its intended purpose and context of use; the outcome concept, population, measurement properties, and interpretation of change all matter. A questionnaire can be validated for one use but not another.
Importance also depends on duration. A short improvement may matter for an acute condition and be inadequate for a chronic one, and a benefit that fades after treatment ends has a different meaning from a sustained effect.
Harms, burden, cost, and patient values#
Clinical importance is a benefit-harm judgment, not a benefit number alone. A small benefit may be worthwhile when an intervention is safe, simple, and inexpensive. The same benefit may be unattractive when treatment is invasive, burdensome, costly, or associated with serious harm.
Harms often have fewer events than efficacy outcomes, so estimates are less precise. A nonsignificant harm comparison does not prove safety. Trials may be too short or selective to detect rare or delayed events. Safety interpretation can require postmarket and observational evidence.
Patient values can vary. One person may prioritize avoiding hospitalization; another may prioritize cognitive function, fertility, fatigue, or daily treatment burden. Shared decisions use evidence to define reasonable choices, then incorporate those priorities.
Cost-effectiveness and affordability are separate analyses. A clinically important effect can be poor economic value at a certain price, and a cost-effective option can still strain a budget. Statistical significance settles neither.
Multiplicity changes the false-positive risk#
If a study tests many outcomes, time points, subgroups, models, and definitions, some p-values may fall below 0.05 by chance under null effects. Prespecification and multiplicity adjustments help control defined error rates.
The distinction between confirmatory and exploratory analysis matters. A prespecified primary outcome carries different weight from a favorable subgroup found after examining many possibilities. Exploratory findings can generate hypotheses, but they need that label and further testing.
Outcome switching can occur when a registered primary outcome is omitted, changed, or displaced by a more favorable result. CONSORT 2025 asks trial reports to identify prespecified outcomes and any changes with reasons. Compare the protocol, the registration, the statistical plan, and the publication yourself. Multiplicity is not cured by reporting only the successful test: selective reporting hides the number of opportunities for a chance finding and hands you a deceptively simple paper.
Model choices and estimands#
Every estimate answers a defined question. An intention-to-treat estimand can compare assignment strategies regardless of later adherence. A per-protocol estimand asks about following the protocol and requires assumptions and methods for deviations. Missing-data strategies can target different hypothetical outcomes.
Hazard ratios can change over time and do not directly state an absolute risk difference. Odds ratios can look farther from one than risk ratios when outcomes are common. Adjusted and unadjusted estimates answer related but not always identical questions. A p-value detached from the estimand is difficult to interpret, so ask: effect of what strategy, compared with what, in whom, over what time, for which outcome, under which handling of intercurrent events?
A worked interpretation#
Imagine a randomized trial in which an intervention reduces a five-year event risk from 12 percent to 10 percent. The risk ratio is 0.83, and the absolute difference is 2 percentage points, or 20 fewer events per 1,000 participants. The 95 percent interval for the absolute difference runs from 0.2 to 3.8 percentage points, and the p-value is 0.03.
The result is statistically significant under the prespecified analysis. The estimated benefit is 20 fewer events per 1,000 over five years, with compatible values from 2 to 38 fewer. Whether that is important depends on outcome severity, treatment harms, burden, cost, adherence, and applicability.
Now imagine the same point estimate with an interval from 1 percentage point more events to 5 percentage points fewer and a p-value of 0.18, and the study has not established benefit at the chosen threshold. It also has not ruled out a meaningful benefit or small harm. The correct conclusion is uncertainty, not equivalence. In both cases the numbers, rather than the adjective "positive" or "negative," carry the information you need.
A reading sequence#
Begin with the question and prespecified primary outcome. Confirm the population, comparator, time horizon, and estimand. Read the effect size on an interpretable scale, then its confidence interval. Compare that range with a justified important-difference threshold.
Next examine design validity, missing data, multiplicity, adherence, and outcome switching. Review harms and burden with the same seriousness as benefit. Translate relative effects into absolute effects using an appropriate baseline risk.
Finally, place the result beside prior evidence and patient priorities. Statistical significance is one feature of the analysis. It is not the final clinical judgment.
References#
- American Statistical Association statement on p-values
- CONSORT 2025 statement
- CONSORT 2025 explanation and elaboration
- CONSORT guidance on primary and secondary outcomes
- ICH E9 Statistical Principles for Clinical Trials
- FDA patient-focused guidance on fit-for-purpose clinical outcome assessments, 2025
Questions and answers
Does p < 0.05 mean the result is probably true?
No. It describes data compatibility with a null model under assumptions. It does not give the probability that the hypothesis is true or account automatically for bias, selective reporting, or model error.
Can a tiny effect be statistically significant?
Yes. A large sample can estimate a very small difference precisely enough to cross a statistical threshold. Importance depends on the effect's size and consequences, not the p-value alone.
Does a nonsignificant result prove there is no effect?
No. The confidence interval may include important benefit or harm. A precise interval excluding meaningful differences supports absence more strongly than a wide inconclusive interval.
Is a minimum important difference fixed for a scale?
Usually not. It depends on the population, condition, outcome, time, method, and decision context. Individual change and between-group differences also require separate interpretation.
Why report both absolute and relative effects?
Relative effects show proportional change, while absolute effects show expected event differences at a stated baseline risk and time. The same relative effect can yield very different absolute benefit across risk groups.