Evidence explainer

Evidence and research methods

What Large Replication Projects Actually Found

Large replication projects did not prove that whole fields are false. They showed that repeated effects come out smaller, and that a great many methods sections cannot be followed.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Reproducibility and replicability
  2. What counts as successful replication
  3. The Reproducibility Project: Psychology
  4. Why the famous 36 percent needs context
  5. High-profile social science experiments
  6. Many Labs and variation across sites
  7. The cancer-biology project met a different obstacle
  8. A replication can disagree for several reasons
  9. Effect sizes are more informative than votes
  10. Selection and the winner's curse
  11. Better reporting reduces practical failure
  12. Replication is not the only credibility check
  13. How to read a replication headline
  14. What the projects changed
  15. Sources

Large replication projects were designed to replace anecdotes about irreproducible science with systematic evidence. Teams selected groups of published experiments, wrote protocols, contacted original authors, repeated procedures with new samples, and compared results using prespecified criteria.

The projects found genuine reasons for concern. Replication effects were often smaller, fewer repeated studies crossed a conventional significance threshold, and the methods that mattered were sometimes impossible to reconstruct. They also showed why “the replication rate” is not a property of an entire discipline. The number changes with the sample, outcome, rule, power, and meaning assigned to consistency.

A careful reading treats a replication as another study within a cumulative process. It neither excuses a weak original nor declares a research field bankrupt after one mismatch.

Reproducibility and replicability#

These words are used inconsistently across fields. The U.S. National Academies adopted a helpful distinction. Reproducibility means obtaining consistent computational results using the same data, code, methods, and conditions; replicability means obtaining consistent results in a new study that addresses the same scientific question and collects new data.

Under this vocabulary, rerunning an analysis script tests reproducibility. Recruiting new participants and repeating an experiment tests replicability. Both can fail for different reasons.

Computational failure may reveal missing code, undocumented preprocessing, software differences, or errors. Empirical mismatch may reflect sampling, measurement, procedure, context, analysis, or an incorrect scientific claim. A paper can be computationally reproducible yet empirically fragile. It can also contain a durable finding even when the archived code is incomplete. Check which meaning a project is using before you compare its percentage with anyone else's.

What counts as successful replication#

Suppose an original study reports a positive effect with p below 0.05. A replication could be judged by whether its p-value is also below 0.05 in the same direction. That binary rule is familiar but unstable near the threshold.

Another rule asks whether the replication effect points in the same direction. A third examines whether the original estimate falls within the replication confidence interval. A fourth combines original and replication data in a meta-analysis. Researchers can also compare whether the estimates differ more than expected by sampling error.

These questions are related but not interchangeable. A well-powered replication can estimate a smaller positive effect precisely without crossing a chosen success rule, and a small replication can point in the same direction while remaining too imprecise to distinguish a meaningful effect from zero. A pooled analysis can be positive even when the second study alone is not. A responsible report shows several of these metrics with their uncertainty rather than picking the most dramatic one.

The Reproducibility Project: Psychology#

The Open Science Collaboration coordinated replications of 100 experimental effects from papers published in 2008 in three psychology journals. Teams used structured protocols, sought original materials, registered plans, and generally used samples with high planned power to detect the original effect size.

Ninety-seven of the original effects had been statistically significant. Thirty-six replications were statistically significant in the same direction. Forty-seven percent of original effect sizes fell within the replication confidence interval, and 39 percent of effects received a subjective rating of replication success; replication effect sizes were, on average, roughly half the original magnitude.

These results demonstrate substantial attenuation under the project's criteria. They do not mean that 36 percent of psychology findings are true. The sample covered specified journals, a publication year, experiments that could be attempted, and selected effects within papers. The project itself presented multiple indicators because no single rate captured every aspect.

Why the famous 36 percent needs context#

A p-value is affected by effect size, sample size, variability, and analysis. Requiring both studies to cross 0.05 creates paradoxes, and an original estimate just above the line and a replication just below it can be nearly identical, while two significant estimates can differ materially.

Selection matters too. Published literatures contain more statistically significant findings than all studies conducted. Conditional on passing a threshold, an estimate is likely to overstate its underlying magnitude, especially in a small sample. This is often called the winner's curse.

If a replication is powered for the inflated original estimate, it may still be underpowered for the smaller underlying effect. That does not erase the discrepancy; it changes what can be inferred from a binary outcome. What you should take from the project is the combination: fewer significant results and smaller estimates, read against its sampling limits.

High-profile social science experiments#

Camerer and colleagues repeated 21 experimental social-science studies published in Nature and Science from 2010 through 2015; the replication teams used larger samples, on average, and worked with original authors on protocols.

Thirteen of 21 replications produced a significant effect in the same direction, and replication effect sizes averaged about half the original size. Prediction markets and surveys of researchers also forecast which studies would replicate better than chance, suggesting that experts could detect some credibility signals.

The selection was systematic within the stated journals and period, not representative of all social science. High-profile publication can intensify selection because novel and surprising effects attract attention. The study showed that prestige is not a validity test, and an earlier project in experimental economics repeated 18 laboratory studies and found a higher proportion meeting its replication criteria, again with attenuation in effect size. Differences across projects caution against one universal rate.

Many Labs and variation across sites#

Many Labs projects coordinated the same protocols across many laboratories and countries. This design separates average replication from variation by setting more directly than a single-site repeat can.

Many Labs 2 tested 28 classic and contemporary psychology findings in samples from dozens of countries, and about half showed a significant effect in the same direction overall under the project's criteria, while effect sizes and heterogeneity varied. Some effects were robust across settings; others were weak or absent.

Context did not provide a blanket explanation for nonreplication. Measured cultural and setting variables explained limited variation for many effects. At the same time, a standardized mass replication can differ from the original context in subtle ways. The project illustrates a constructive question: not only “is the effect real?” but “how large is it, how much does it vary, and under which conditions?”

The cancer-biology project met a different obstacle#

The Reproducibility Project: Cancer Biology planned to repeat selected experiments from 53 high-impact papers. It used registered reports, in which protocols were peer-reviewed before results were known.

The effort could not complete its original scope. Missing details, unavailable reagents, unclear protocols, biological-material differences, cost, and time made many experiments infeasible or altered plans. Experiments associated with a smaller subset of papers were completed and synthesized.

Across repeated effects, replication estimates were generally weaker than original estimates, and categorical conclusions were mixed. Yet the inability to execute a faithful protocol was itself a major result, and if skilled teams working with substantial resources cannot determine exactly how an experiment was done, the published record is not sufficient for cumulative science. Preclinical biology adds challenges such as cell-line identity, reagent lots, animal strains, microbiome, laboratory conditions, and complex multistep procedures. Transparency is necessary but does not make living systems identical.

A replication can disagree for several reasons#

Ordinary sampling variation ensures that estimates differ even when studies are perfect. Low power widens that variation. Measurement reliability, participant recruitment, exclusions, treatment fidelity, and analytic choices add more.

The original result may be biased by selective reporting, undisclosed flexibility, missing data, or publication selection. The replication can also contain error or depart from the construct it intended to repeat. Context may modify an effect, although context should be specified and tested rather than invoked after any unwanted result.

Scientific claims range from narrow to broad. “This exact procedure changed this measure in this population at this time” is easier to repeat faithfully than a general theory. A direct replication tests whether conditions believed sufficient produce a similar result. A conceptual replication tests the broader theory with different operations. Failure in each has a different meaning. Fraud, meanwhile, cannot be inferred from disagreement alone. It requires evidence about conduct, not just statistics.

Effect sizes are more informative than votes#

Counting significant and nonsignificant studies treats evidence as ballots. A better synthesis compares estimates and uncertainty. How large was the original effect? How large was the replication? Are intervals wide? Are differences consistent with sampling? What does a pooled estimate show, and is pooling justified?

An exact replication estimate near zero with a narrow interval can strongly challenge a large original effect. A small positive estimate with a wide interval may remain inconclusive, and an effect that consistently shrinks from 0.8 to 0.2 can still exist while having very different theoretical and practical importance.

Prediction intervals can show how much future results might vary. Heterogeneity models can estimate between-study variation, but they need enough studies and credible measurement. A random-effects model cannot explain why estimates differ. The question to end on is what claim the evidence still supports once the new results are in.

Selection and the winner's curse#

Journals, researchers, and media tend to favor novel positive results. If many teams study weak or null effects, the few estimates that pass a threshold are more likely to be published. The literature you can see then exaggerates prevalence and size.

Within a study, trying several outcomes, exclusions, transformations, covariates, or stopping points can create the same selection. The final paper you read may present one coherent path even though many paths were available.

Preregistration can record hypotheses, primary outcomes, exclusions, and analyses before results are known. It does not guarantee quality and does not forbid exploratory work. It separates confirmation from discovery. Registered reports go further by making publication decisions before outcomes are available, based on the importance of the question and quality of the protocol, which reduces outcome-dependent publication while retaining peer review.

Better reporting reduces practical failure#

Methods need enough detail to implement, not merely understand. Materials, questionnaires, software versions, code, stimuli, reagent identifiers, protocols, and data dictionaries can be shared when ethics, consent, privacy, and law permit.

FAIR data principles aim for findable, accessible, interoperable, and reusable research objects. Access may be controlled rather than public for sensitive data, but conditions and governance should be explicit.

Analysis containers and archived environments can preserve computational dependencies. Versioned protocols record deviations. Reporting guidelines improve completeness for particular designs. None substitutes for careful study design. Original authors can help identify ambiguities, but replication teams must retain analytic autonomy and report disagreements. Collaboration should improve fidelity without turning the process into approval by the original result's proponents.

Replication is not the only credibility check#

Triangulation asks whether different designs with different biases converge. Randomized experiments, longitudinal cohorts, natural experiments, mechanistic studies, and qualitative work can contribute distinct information.

Robustness analysis tests reasonable alternative specifications. Negative controls can reveal residual bias. Multiverse analysis shows how conclusions vary across defensible analytic choices. Prospective multisite studies test generalization.

Systematic reviews evaluate the full evidence base, including unpublished or registered studies where possible. Meta-analysis can increase precision but inherits the quality and selection of its inputs. For clinical decisions, replication is one part of certainty alongside risk of bias, consistency, directness, precision, harms, feasibility, and patient values. A replicable laboratory effect may still have no meaningful health benefit.

How to read a replication headline#

Identify the target. Was the project repeating one effect, a paper, a method, or a theory? Check the sampling frame and how many planned studies were completed. Note whether protocols were registered and whether original authors reviewed fidelity.

Then inspect the success rule. Did the headline quote significance, direction, confidence-interval overlap, subjective judgment, or meta-analysis? Look for effect-size attenuation and uncertainty. Ask whether the replication was powered for a realistic effect rather than only the original estimate.

Finally, separate inability to complete from a completed null result. Both reveal problems, but they are not the same evidence about the claim. A credible report makes that distinction visible to you instead of converting one project into a verdict on all science.

What the projects changed#

Large replication projects supplied evidence that publication systems can overstate effects and that methods are often too incomplete for efficient repetition. They also demonstrated collaborative solutions: shared protocols, large samples, multisite coordination, registered reports, open materials, and forecast evaluation.

Their deeper message is not that replication produces a final stamp. It is that scientific confidence should be earned cumulatively. Initial studies generate estimates and hypotheses. Replications update them. Differences trigger investigation. Synthesis states what remains credible and uncertain.

That process can reduce confidence in a famous result while increasing trust in the system's ability to correct itself.

Sources#

The metadata sources link the major psychology, social-science, economics, multisite, and cancer-biology projects and the National Academies terminology report. Percentages should be read from those methods, not detached from them.

Sources and further reading

  1. Open Science Collaboration, Estimating the Reproducibility of Psychological Science
  2. Camerer and colleagues, Replicability of Social Science Experiments
  3. Camerer and colleagues, Replicability of Laboratory Experiments in Economics
  4. Klein and colleagues, Many Labs 2
  5. Errington and colleagues, Replicability of Preclinical Cancer Biology
  6. National Academies, Reproducibility and Replicability in Science

Questions and answers

Did the psychology replication project find that only 36 percent of psychology is true?

No. Thirty-six percent referred to one significance-based result among 100 selected effects under that project's rules. Other metrics differed, and the sample did not represent every study or claim in psychology.

Does a failed replication prove the original result was fraudulent?

No. Fraud is one possible but uncommon explanation. Sampling variation, low power, bias, analytic choices, procedural differences, context, measurement, and an incorrect original claim can also contribute.

Is reproducing an analysis the same as replicating a finding?

Not under the National Academies distinction. Reproducibility obtains consistent computational results from the same data and code, while replicability obtains consistent results across new studies addressing the same question.

Why were replication effect sizes often smaller?

Selective publication, winner's curse, small original samples, flexible analyses, measurement error, context differences, and ordinary sampling variation can all make original estimates larger than later estimates.

What practices make evidence easier to trust?

Clear protocols, adequate power, validated measures, preregistration when appropriate, complete reporting, shared materials and code, registered reports, direct replication, and synthesis across studies.