Give an honest dataset to two careful analysts and they can reach opposite conclusions, both defensible, because a single study contains dozens of small forks in the road. P-hacking is what happens when those forks are chosen, consciously or not, to nudge a p-value below the .05 line. In a 2011 paper in Psychological Science, Joseph Simmons, Leif Nelson, and Uri Simonsohn put a number on the danger: exploiting just four of these ordinary choices can raise the false-positive rate from the advertised 5% to roughly 61%. The problem is rarely fraud, and the fix is rarely better intentions. It is transparency about the choices, reinforced by preregistration.
Key points#
- A p-value summarizes one analysis, but most studies could have been analyzed many defensible ways.
- These forks (which outcome, when to stop collecting data, which cases to exclude, which covariate to add) are called researcher degrees of freedom.
- Combining four common forks in a single study pushed a simulated false-positive rate to about 61%.
- The main remedy is disclosure of every choice; preregistration goes further by fixing the plan before the data arrive.
- When reading a paper, ask whether the analysis was planned in advance and whether alternative versions are shown.
Why one p-value tells you so little#
Think of a study as a branching path rather than a straight line. At each junction the analyst faces a reasonable decision with no single correct answer. Collect twenty participants, or keep going until the pattern firms up? Treat an unusually fast response as valid data or as noise to discard, and if discard, at what cutoff? Report the reaction-time measure or the accuracy measure? Adjust for baseline age, or sex, or nothing? Simmons and colleagues named this latitude researcher degrees of freedom, and their unsettling point was that almost any choice can be defended after the results are in view. When the person at the junction is hoping for significance, a well-documented tendency toward motivated reasoning means the defensible choices tend to line up on the significant side, with the analyst sincerely believing the evidence led the way.
The label p-hacking became popular only later. The lasting contribution of the 2011 paper was to show, in hard numbers, how much a well-meaning analyst can distort a result without ever touching a fabricated data point.
The Beatles study, rebuilt from its choices#
To make the risk impossible to wave away, the authors ran a real experiment engineered to "prove" something that cannot be true. Twenty undergraduates listened either to the Beatles track "When I'm Sixty-Four" or to a control song. Each then reported a birth date, and the analysis adjusted for the father's age. On paper the result looked clean: listeners were nearly a year and a half younger after the Beatles song than after the control, F(1, 17) = 4.92, p = .040. A pop song had, in the statistics, altered people's actual age.
Read literally, every clause there is an honest analysis honestly reported. What the write-up left out were the roads not shown. The team had gathered several other variables and presented only father's age as the covariate. It ran the model with that covariate and never displayed the version without it. And it watched the p-value climb as participants trickled in, stopping the moment the number dipped under .05, with no target sample size set at the start. Rebuild the same study with everything on the table and the effect evaporates: drop the covariate and significance disappears. The finding was an artifact of undisclosed flexibility, nothing more.
Turning a hunch into a number#
The prank study is memorable; the simulations are the evidence. Simmons and colleagues generated thousands of datasets containing no real effect whatsoever, then let a hypothetical analyst reach for the usual tools. Testing two related outcomes instead of committing to one nearly doubled the false-positive rate. Peeking at the data and continuing to add subjects until significance appeared, a practice called optional stopping, produced a false hit about 22% of the time on its own. Slipping in a covariate, or dropping one of three conditions, each added still more. Stack all four moves into one study and the false-positive rate climbed to roughly 61%. Put plainly, an analyst using these everyday techniques on pure noise was more likely to announce a discovery than to correctly report that nothing was there.
That is why a lone significant p-value, presented without its history, carries so little weight. A few years earlier, John Ioannidis had warned that results grow less trustworthy wherever there is "greater flexibility in designs, definitions, outcomes, and analytical modes" (PLoS Medicine, 2005). The False-Positive Psychology paper converted that warning into a measured probability.
From disclosure to preregistration#
The remedy the authors proposed is modest by design. They set out six requirements for authors: decide the data-collection stopping rule before starting and report it; collect at least twenty observations per cell or justify a smaller number; list every variable measured; report every condition, including the ones that failed; if any observations are excluded, show the result with them retained as well; and if a covariate is used, also show the result without it. Four companion guidelines ask reviewers to enforce that openness and to stop rewarding results that look too tidy to be real. None of this bans exploration. It only insists that exploratory work be labeled as exploratory rather than repackaged as a prediction confirmed.
Disclosure has one gap the authors named frankly: it cannot surface the studies that were run and then set aside. Closing that gap is the job of preregistration. As Brian Nosek and colleagues explain (PNAS, 2018), recording the hypothesis and the analysis plan before the outcomes are known draws a clean line between genuine prediction and after-the-fact narrative, so a confirmatory claim can be distinguished from an exploratory one. Registered reports and public preregistration have since spread through psychology, medicine, and clinical trials for that reason.
A short checklist for reading a study#
If you are weighing a headline result, a handful of questions carry most of the load. Was the analysis registered in advance, and does the paper point to that plan? Does the reported outcome match the one the study was built to measure, or a stand-in that surfaced along the way? Are the sample size, the exclusions, and the covariates stated openly, with the alternative versions on view? A paper that answers these questions in plain sight deserves more confidence than a neater one that stays silent. The takeaway from the Beatles experiment is not that researchers are dishonest. It is that a p-value means little until you know how many paths were open on the way to it.
Sources and further reading
Questions and answers
Is p-hacking the same as fraud?
Usually not. Fraud involves inventing or altering data. P-hacking is choosing among genuinely defensible analysis options in ways that favor a significant result, often without the analyst recognizing the bias. The data are real; the selective reporting is the problem.
Does preregistration prevent all of this?
It removes a large share of the flexibility by fixing the hypothesis and analysis plan before the results are seen, and it exposes studies that would otherwise vanish unpublished. It does not replace careful design, honest reporting, or replication, and analysts can still deviate from a plan, which is why the deviations themselves should be disclosed.
What should a reader do with a single significant study?
Treat it as a lead rather than a verdict. Check whether the analysis was planned in advance, whether the outcome was the pre-specified one, and whether the finding has held up in independent replication before giving it much weight.