Evidence explainer

Physician-scientist and medical humanities

Why Replication Studies Are Hard

Copy an original study too closely and you may preserve its bias. Change too much and you have tested a different question. The hard part is deciding which features have to stay fixed.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The words are used inconsistently
  2. The claim must be specified before it can be repeated
  3. Methods sections omit tacit knowledge
  4. Reagents and materials are not stable labels
  5. Context may be part of the mechanism
  6. Exact copying can preserve the original flaw
  7. Too much change creates an escape route
  8. Original effect sizes are often too large
  9. Power is not a ritual percentage
  10. Significance is not a replication verdict
  11. Measurement reliability constrains replicability
  12. Researcher skill can interact with treatment
  13. Preregistration protects the replication question
  14. Data and code sharing reveal a different class of problem
  15. Multisite replication separates signal from local luck
  16. Failed replication is not evidence of misconduct
  17. Incentives make replication harder to sustain
  18. How to interpret a replication pair

"Repeat the experiment" sounds like a simple instruction. A published study rarely contains every decision, material property, and environmental condition needed to recreate it. It rarely contains every software dependency and piece of tacit knowledge. Even if it did, a second study would face a deeper choice: how closely should it copy the first?

A close replication asks whether a result returns under methods intended to match the original. A conceptual replication changes methods while testing the same underlying claim, and the close version can reveal fragility, but may also reproduce the same hidden bias. The conceptual version can show broader robustness, but a different result may reflect a changed mechanism rather than a failed claim.

Replication is difficult because research findings are conditional. They depend on design, population, and measurement. They depend on implementation, analysis, context, and chance. Good replication makes those conditions visible and narrows uncertainty. It is not a ceremonial attempt to declare an earlier paper right or wrong.

The words are used inconsistently#

Different fields swap the meanings of reproducibility and replicability. The National Academies chose a clear convention: reproducibility is obtaining consistent computational results with the same data, code, methods, and analysis conditions; replicability is obtaining consistent results in a new study that answers the same scientific question with new data.

The distinction prevents a common overclaim. Successfully rerunning code shows that the reported numbers can be generated from the available files, but it does not show that new participants, specimens, laboratories, or observations will produce a similar effect. So check which of the two a report means before its verdict. Otherwise two teams can argue about a "replication failure" while describing different tasks.

The claim must be specified before it can be repeated#

A paper can contain a broad theory, several hypotheses, and one primary analysis. It can contain many secondary analyses and an attention-grabbing conclusion. Which claim is the replication testing?

The target should name the population, intervention or predictor, and comparator. It should name the outcome, timeframe, direction, and magnitude that matter. If the original paper called one analysis exploratory, promoting it later to the central claim changes the test, and replicators should document why the target matters and how it connects to the original evidence. Consultation with original authors can clarify ambiguity, but those authors should not have veto power over a fair test.

Methods sections omit tacit knowledge#

Laboratory protocols may not state how quickly a sample is moved, what a healthy cell culture looks like, how firmly a device is positioned, or when an operator discards a questionable run. Behavioral studies may omit the room setup, recruiting script, platform filters, or exact timing.

Experienced staff learn these details through demonstration and judgment. The omission may be innocent, yet it makes transfer difficult. Video protocols, checklists, shared code, annotated materials, and pilot verification can reduce the gap. Too much informal help from the original team creates another concern: a result that only insiders can produce is less general than the publication implied.

Reagents and materials are not stable labels#

A named antibody can differ by lot. Cell lines drift, become contaminated, or are misidentified. Animal microbiomes vary by facility. Software libraries change defaults. Survey platforms alter participant pools, and clinical standards evolve.

Authentication, lot records, and version pinning are therefore part of the scientific method. So are reference materials and quality-control thresholds. "Same reagent" on an invoice does not guarantee the same functional material. NIH rigor guidance emphasizes authentication of key biological and chemical resources because an unrecognized material difference can make both the original and replication uninterpretable.

Context may be part of the mechanism#

A social intervention can depend on language, norms, or economic conditions. It can depend on institutional trust or what participants already know. A clinical comparison depends on background care and available rescue treatment. An ecological study can depend on season and prior disturbance.

A different result across contexts can weaken a universal claim while supporting a conditional one; the right conclusion may be "the effect occurs under these conditions" rather than "the replication failed."

It is better to measure the moderators you think plausible than to invent them once you have seen the results; multisite studies can vary context systematically and estimate heterogeneity instead of treating one laboratory as the universe.

Exact copying can preserve the original flaw#

Suppose the first study used a biased assay, an invalid outcome, or a confounded comparator. A perfect copy may reproduce the numerical result without validating the scientific interpretation.

Close replication is useful for checking whether the original phenomenon recurs; robustness tests then alter questionable assumptions, measures, exclusions, or models, and conceptual replication uses a different operationalization of the same theory. Confidence grows when several methods with different biases converge; agreement among copies that share one hidden error is weaker.

Too much change creates an escape route#

When a replication differs, defenders can point to population, setting, or materials. They can point to training, dose, or analysis. Sometimes the difference truly matters. Sometimes it becomes an unfalsifiable list of reasons that protects a claim from every result.

A fair protocol distinguishes essential from incidental features before collecting data. It can include a high-fidelity arm plus planned variants. The original authors may review the protocol, while the replication team retains control and records disagreements. Naming those features in advance turns a vague argument about who is right into a moderation question somebody can test.

Original effect sizes are often too large#

Small studies produce noisy estimates. When journals and researchers preferentially select results that cross a significance threshold, published estimates are enriched for upward fluctuations. This is sometimes called the winner's curse.

Flexible stopping, outcome selection, and subgroup analysis add opportunities for a favorable result. So do covariate choices and model variations. Even without misconduct, the published effect can overstate the underlying effect.

So if you power a replication to the exact original estimate, you can end up with a second underpowered study. Plan instead around the smallest scientifically important effect, shrinkage, meta-analytic evidence adjusted for bias where that is defensible, or a conservative range.

Power is not a ritual percentage#

Power depends on effect size, sample size, and variability. It depends on design, attrition, and clustering. It also depends on measurement reliability and analysis. Declaring 80 percent power based on an optimistic input does not make a study informative.

Replications often need much larger samples than originals because they seek precise confirmation or contradiction. Multisite designs must account for site and cluster effects. Rare events may require a different design or pooled program. Precision can be more useful than a binary target: a confidence interval that excludes effects large enough to support the original claim can be informative even if it includes a small nonzero effect.

Significance is not a replication verdict#

The original study can have p below 0.05 and the replication p above 0.05 even when their effect estimates are statistically compatible. "Significant" and "not significant" are not necessarily significantly different.

So when you set the two studies side by side, compare the effect sizes and confidence intervals, test the difference or the interaction where that is appropriate, and check whether the replication estimate lies within a prespecified prediction interval. Direction, magnitude, and uncertainty all matter. Bayesian analyses can quantify support for specified effect ranges, but they depend on model and prior choices. No method rescues an imprecise or biased study by relabeling it.

Measurement reliability constrains replicability#

A noisy measure attenuates effects and increases uncertainty. If reliability differs across studies, effect estimates can differ even when the underlying relation is stable.

Construct validity matters more deeply. Two scales with the same label may measure different aspects of anxiety, trust, pain, or function, and a biomarker assay can be precise while measuring a process only loosely related to the theory. Replications should report instrument versions, scoring, and language adaptation. They should report calibration, reliability, and evidence that the measure functions comparably in the new population.

Researcher skill can interact with treatment#

Some procedures require surgery, behavioral coding, psychotherapy, microscopy, or complex preprocessing. Training and competence influence fidelity. If the replication team performs the method poorly, a negative result may not test the treatment.

On the other hand, a method that works only with its inventors may not be ready for broad claims, and competence should be demonstrated through objective criteria, blinded quality review where feasible, and a transparent learning phase. Reporting failed runs and protocol deviations helps distinguish treatment failure from implementation failure without discarding inconvenient data after the fact.

Preregistration protects the replication question#

A preregistration records hypotheses, outcomes, and exclusions before results are known. It records sample size, stopping, and analysis. It does not guarantee good design or honest conduct. It makes changes visible.

Registered Reports go further. A journal reviews the question and method before data collection and offers in-principle acceptance based on rigor rather than outcome. This reduces publication bias against null or contradictory replications. Exploratory analyses remain valuable when labeled and separated from the confirmatory test. Discovery and confirmation are different jobs.

Data and code sharing reveal a different class of problem#

Computational reproduction can uncover missing files, undocumented cleaning, or software drift. It can uncover coding errors or discrepancies between a protocol and final analysis. Containerized environments, versioned code, data dictionaries, and executable workflows improve durability.

Privacy, consent, proprietary rights, and security can limit open data. Alternatives include controlled access, synthetic examples, audited analysis, and complete metadata. "Data unavailable" should not be the end of planning when verification is central. A successful computational reproduction does not validate data collection. A failed one should be diagnosed before launching an expensive new-data replication.

Multisite replication separates signal from local luck#

One original site and one replication site leave site differences confounded with study status, and coordinated studies across several laboratories can use common protocols, shared analysis, and local measurements to estimate both an average effect and heterogeneity.

Sites should not be treated as interchangeable copies. Random-effects or hierarchical models can represent variation, while prespecified site-level moderators explore why effects differ. Central coordination reduces some variation but can introduce shared errors. Local autonomy increases realism but requires stronger documentation and quality assurance.

Failed replication is not evidence of misconduct#

Fraud, fabrication, and falsification can make findings irreproducible, but most discrepancies do not establish misconduct. Chance, weak design, and hidden conditions are common alternatives. So are measurement error, analysis flexibility, and honest mistakes.

Accusations require separate evidence and due process. A replication paper should critique claims and methods without speculating about motives; the reverse reading is just as easy to get wrong: a successful replication does not certify every part of the original conduct either. It supports a finding under the conditions that were tested.

Incentives make replication harder to sustain#

A novel positive finding can be easier to publish, easier to fund, and easier to put on a promotion case than a careful verification. A replication team can rebuild someone else's method from an incomplete methods section and still be told the work is unoriginal.

NIH launched a broader replication and reproducibility initiative in 2026 that explicitly addresses incentives, infrastructure, and coordinated research; that policy recognition matters because individual rigor cannot solve a system that rewards only novelty. Funders and institutions can support replication grants, data stewardship, and methods staff. They can support negative-result publication and credit for reusable materials and validation work.

How to interpret a replication pair#

Read both protocols, not just the two p-values. Start where the two studies were most likely to diverge: who was studied, how the method was actually carried out, and what counted as the outcome. Then check power, exclusions, and attrition. Check the analysis, the context, and every deviation from plan. Look at the absolute estimates and the uncertainty around them rather than the labels attached to them.

Ask whether the replication was close enough to test repeatability and different enough to add knowledge, and work out whether a moderator was predicted in advance or invented after somebody saw the result.

The judgment you reach may be strong confirmation, partial confirmation, informative inconsistency, or unresolved uncertainty. Science benefits when those four categories replace victory language.

Sources and further reading

  1. National Academies report Reproducibility and Replicability in Science, 2019
  2. NIH resource on replication and reproducibility, reviewed June 2026
  3. NIH guidance on rigor and transparency
  4. Registered Reports framework from the Center for Open Science
  5. Multilaboratory replication of social-science experiments
  6. Statistical guidance on interpreting p-values

Questions and answers

Is reproducibility the same as replication?

Under the National Academies convention, no. Reproducibility reruns analysis with the same data and code; replication collects new data for the same question.

Does a p-value above 0.05 mean the replication failed?

No. Interpret effect magnitude, confidence interval, power, direction, and the direct comparison with the original estimate.

Should replicators contact the original authors?

Often yes, to clarify methods and obtain materials. The exchange should be documented, and original authors should not control publication of the outcome.

Is a conceptual replication weaker than an exact copy?

It answers a different question. A close copy tests repeatability under matched methods; a conceptual replication tests whether the claim generalizes across methods.

Can one replication settle a scientific claim?

Rarely. Confidence usually comes from multiple studies, methods, populations, and a synthesis that accounts for quality and heterogeneity.