Evidence explainer

Evidence and research methods

What a P Value Really Means

A p value is calculated by assuming the null hypothesis. It describes how well the data fit that assumption, not whether a claim is true.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key takeaways
  2. Start with the conditional probability
  3. What “at least as extreme” means
  4. Why p below 0.05 is not proof
  5. Why a large p value does not prove no effect
  6. Effect size and uncertainty answer different questions
  7. Sample size can dominate the result
  8. Multiple testing changes the landscape
  9. Model assumptions and data quality remain upstream
  10. Prior evidence affects how surprising a result should be
  11. A better reporting pattern
  12. References

A p value answers a conditional question: if a specified null hypothesis and all other assumptions in the statistical model were true, how probable would the observed test result, or a result at least as incompatible with that hypothesis, be? It does not tell the probability that the null hypothesis is true. It does not tell the probability that the result happened by chance, and it does not measure the size or importance of an effect.

The definition has several moving parts. The null hypothesis must be stated. The test statistic and one-sided or two-sided alternative must be chosen. The model includes assumptions about sampling, measurement, distributions, adjustment variables, missing data, and analysis decisions. A p value is therefore not a property of the data alone.

Key takeaways#

Start with the conditional probability#

Suppose a randomized trial compares mean blood pressure under two assigned treatments. The null hypothesis states that the relevant mean difference is zero. A statistical test summarizes the observed difference relative to its uncertainty. If its two-sided p value is 0.03, the careful reading is: under the zero-difference hypothesis and the test's other assumptions, a result at least this far from zero in either direction would occur with probability 0.03.

The sentence does not run backward. It is incorrect to say there is a 3 percent probability that the null hypothesis is true, and the calculation assumed the null so that it could describe the distribution of possible test results. To assign a probability to a hypothesis you need another framework, such as a Bayesian model with stated prior assumptions.

It is also misleading to say there is a 3 percent probability that the finding is “due to chance.” Chance is already represented through the probability model, but bias, confounding, measurement error, selective reporting, model misspecification, and protocol deviations are separate possibilities. The p value does not allocate probabilities among those explanations.

What “at least as extreme” means#

The test statistic defines how disagreement with the null is measured. For a two-sided test, results in both directions count. For a one-sided test, only the prespecified direction ordinarily counts. A one-sided question can be appropriate in narrow settings, but choosing it after seeing the direction of the data changes the error properties.

Different valid tests applied to the same broad research question can yield different p values because they use different statistics or assumptions. A rank-based test, a t test, a regression model, and a permutation test do not define extremeness identically. Adjustment for prognostic variables can also change precision.

Exact and asymptotic methods differ too. Some tests calculate a probability from the exact finite-sample distribution under the null. Others rely on large-sample approximations, and small samples, sparse data, clustering, repeated observations, or violated distributional assumptions can make an approximation unreliable. So “what test produced this number?” is part of interpreting it, not a technical afterthought.

Why p below 0.05 is not proof#

The familiar 0.05 threshold is a convention, not a law of nature, and if an analysis plan defines an alpha level of 0.05, p below that level can trigger a decision to reject the null within the specified testing procedure. That decision rule controls a long-run error rate under its assumptions. It does not transform the result into certainty.[1][2]

Values on opposite sides of the line can carry almost the same information. P equals 0.049 and p equals 0.051 are not substantively different simply because one receives the label “statistically significant.” Reporting them as a discovery and a failure can exaggerate random variation.

The threshold also says nothing about practical importance. A very large study may estimate a tiny average difference with great precision and produce a small p value, and the effect can still be too small to matter clinically, operationally, or personally. Conversely, a smaller study can estimate an important difference but yield p above 0.05 because the interval remains wide.

The American Statistical Association emphasized that scientific conclusions and decisions should not be based only on whether a p value passes a particular threshold. It also warned that p values do not measure effect size or importance and cannot supply a complete measure of evidence by themselves.[1][2]

Why a large p value does not prove no effect#

A large p value means the observed statistic is not unusually incompatible with the null under the selected model. It does not demonstrate that the null is correct. Many nonzero effects can also be compatible with the data when precision is limited.

Consider an estimated risk difference of 8 percentage points with a 95 percent confidence interval from -5 to 21. A two-sided test of zero may have p above 0.05. The data are compatible with no difference, modest harm, and potentially meaningful benefit. Calling the result “no effect” hides that uncertainty.

To support a claim of similarity, researchers need an equivalence or noninferiority design with a justified margin, suitable analysis, and adequate precision. Failing to reject a zero-effect null is not the same as showing that effects of concern are absent.

This distinction matters in systematic reviews. You will see one study report a statistically significant estimate while another does not, yet the two estimates may be statistically compatible with each other. The correct question is whether their effects differ, assessed through an interaction or heterogeneity analysis, not whether their individual labels differ.

Effect size and uncertainty answer different questions#

The effect estimate describes the observed magnitude and direction on a chosen scale: a mean difference, risk ratio, odds ratio, hazard ratio, correlation, or another measure. A confidence interval describes the range of parameter values that remain reasonably compatible with the data and model under its repeated-sampling construction.

For a conventional two-sided test and matching 95 percent confidence interval, the null value is excluded exactly when p is below 0.05, subject to compatible methods; this relationship does not make the interval a disguised hypothesis probability. A 95 percent interval is not a 95 percent probability that this one fixed interval contains the true value. It comes from a procedure that would cover the target parameter 95 percent of the time over hypothetical repetitions under its assumptions.

Intervals are usually more informative than a significance label because they display magnitude and precision. The guide to what a confidence interval is not examines those limits. Ask, too, whether the interval includes effects that would change your decision, not only whether it includes the null.

Sample size can dominate the result#

A test statistic often divides an estimate by its standard error. Increasing the effective sample size usually reduces that standard error. If a small nonzero difference persists, its standardized distance from the null grows and the p value can shrink.

This is why a small p value does not imply a large effect. It can reflect a modest estimate measured with high precision. You also cannot compare p values across studies as if the smaller number always identified the larger or more credible effect. Studies can differ in size, variance, outcomes, models, and data quality.

Small studies create a different problem. Their estimates tend to be noisy and their confidence intervals wide, and among the small studies that cross a publication threshold, effect estimates can be exaggerated because unusually large observed effects are more likely to pass the line. Selection for “positive” results can therefore distort the literature you get to see.

Power is related but not interchangeable with evidential strength. Before a study, power is the probability of rejecting the null under a specified alternative and design. After data collection, an observed p value and interval are more useful than “observed power,” which largely repackages the same result.

Multiple testing changes the landscape#

If twenty unrelated true null hypotheses are each tested at alpha 0.05, the expected number of false rejections is one, and the probability of at least one false rejection is about 64 percent if the tests are statistically independent. Real analyses may have correlated outcomes, but the principle remains: more opportunities to test create more opportunities for small p values.

Multiplicity arises from many outcomes, time points, subgroups, models, transformations, stopping rules, and repeated looks at accumulating data. A prespecified primary analysis and an appropriate multiplicity strategy help preserve interpretable error rates. Exploratory analyses can be valuable, but they should be labeled and replicated rather than presented as if they were the sole planned test.

Selective reporting is especially damaging because you may see one small p value without the many analyses that produced larger ones, and a number that looks compelling in isolation can be routine within the full search process. Protocols, registrations, statistical analysis plans, and complete results make that search space more visible.

The article on what a subgroup analysis shows extends this reasoning to treatment-effect variation.

Model assumptions and data quality remain upstream#

Every p value inherits the quality of the design and measurements. Randomization can support causal interpretation when assignment, follow-up, and analysis are sound. In observational research, adjustment depends on measured variables and model choices. A small p value cannot remove residual confounding or repair an inappropriate comparator.

Misclassification can bias an effect toward or away from the null; missing data can change the target population the estimate represents; clustered or repeated observations require methods that reflect dependence; nonlinear relationships can be hidden by an oversimplified model; and outliers can dominate some statistics. Any one of them can make a numerically precise p value answer the wrong question.

Data errors are another upstream risk. A transposed coding label, duplicated records, unit mismatch, or incorrect denominator can produce a calculation that is mathematically flawless but scientifically false. Reproducible code, data checks, sensitivity analyses, and transparent exclusions are therefore part of statistical interpretation.

Prior evidence affects how surprising a result should be#

The same p value can have different evidential implications in different settings. A well-powered confirmatory trial testing a plausible, prespecified effect is not equivalent to one small exploratory analysis selected from thousands of unlikely hypotheses. The p value alone does not encode biological plausibility, previous studies, measurement reliability, or the size of the search space.

This does not mean a popular hypothesis should receive a pass or a novel one should be dismissed. It means evidence accumulates across design quality, replication, coherence, and prior knowledge. Extraordinary claims usually need stronger and more reproducible support than one threshold crossing.

A Bayesian analysis makes prior assumptions explicit and updates them with a likelihood; that can yield a posterior probability for a parameter or hypothesis, but the answer depends on the prior and model. A Bayes factor is also not a p value. Each tool asks a different question.

A better reporting pattern#

Begin with the research question, target population, comparison, outcome, and time frame. Report how the data were generated and which analysis was prespecified. Then present the effect estimate with its confidence interval and exact p value, where useful, rather than only “significant” or “not significant.”

Also report the number of participants and events, missing data, protocol deviations, multiplicity approach, adjusted variables, sensitivity analyses, and material limitations. Distinguish confirmatory from exploratory findings. If a decision threshold was set in advance, explain why it was appropriate.

Interpret magnitude before the label. Ask whether the interval includes benefits or harms large enough to change what you would do. Compare the estimate with prior evidence and consider bias or model failures that the p value cannot detect. The site's research approach treats the number as one part of a traceable evidence argument.

The most accurate summary is often calibrated: the data are more or less compatible with specified values under a model, while design and external evidence determine what that compatibility can support.

References#

  1. American Statistical Association Statement on Statistical Significance and P-Values
  2. The ASA's Statement on P-Values: Context, Process, and Purpose
  3. Sifting the Evidence: What Is a P-Value?

Questions and answers

Is a p value the probability that the null hypothesis is true?

No. The p value is calculated under the assumption that the null hypothesis and statistical model are true. It describes the probability of the observed test statistic or one at least as incompatible with the null, not the probability of the hypothesis.

Does p below 0.05 prove an important effect?

No. It crosses a conventional threshold within a testing procedure. It does not measure effect size, clinical relevance, causal validity, freedom from bias, or replication probability. Read the estimate, interval, design, and analysis plan.

Does a large p value prove there is no effect?

No. A large p value can arise when the null is compatible with the data, but also when the study is imprecise, variable, small, or insensitive. Equivalence requires a justified margin and an analysis designed to test it.

Why can a tiny effect have a small p value?

With a large effective sample and low standard error, a small estimate can lie many standard errors from the null. The p value can then be small even when the effect would not change a practical or clinical decision.

What should be reported with a p value?

Report the effect estimate, confidence interval, sample and event counts, outcome definition, prespecified analysis, missing data, multiplicity, sensitivity analyses, and study limitations. Exact p values can supplement that context but should not replace it.