Scientific news often arrives as a single result: a food is associated with longer life, a test detects disease earlier, a drug changes a biomarker, or a social program improves an outcome. The result may be careful and important. It may also be a false positive, an overestimate, a finding limited to the study population, or a true effect described with more confidence than the data support.
One study is not “just one study” in the sense that it should be ignored. A decisive trial, natural experiment, or observation can overturn an old belief. The limitation is logical: one estimate cannot reveal all the ways its own design, data, analysis, and setting shaped the answer.
A useful appraisal asks you two questions at once:
- What does this study contribute?
- How much should the conclusion change when this result is placed beside everything else you already know?
Every result is conditional#
A study result belongs to a question defined by population, intervention or factor, and comparator. The question is also defined by outcome, time horizon, and method. Change any of these and the answer may change.
A blood-pressure trial in selected adults followed for twelve weeks may provide a precise estimate of short-term blood-pressure change. It does not necessarily establish long-term cardiovascular benefit, safety in pregnancy, effects in frail older adults, adherence in routine care, or comparative value against every alternative.
An observational association between a behavior and an outcome may hold up under every adjustment the measured data allow. It can still differ across age, geography, and culture. It can differ across baseline risk, coexisting illness, measurement method, and calendar time. Confounding can survive adjustment, because unmeasured variables, imperfect measurements, and incorrect functional forms are all possible, and none of them will announce itself in the table you are reading.
The appropriate conclusion mirrors the design. “In this trial, assignment to treatment reduced the measured outcome at six months” is stronger than “the treatment works for everyone,” but it is also more useful because its boundaries are visible.
Chance is not eliminated by significance#
Sampling creates uncertainty. If many equivalent samples were drawn, their estimates would vary. Confidence intervals show a range of values compatible with the data and model assumptions; they do not guarantee that the true value lies inside a particular interval.
A p value summarizes how incompatible the observed data, or more extreme data, are with a specified statistical model that includes a null hypothesis. It is not the probability that the hypothesis is true, the probability the result will replicate, or the size and importance of an effect.
Conventional thresholds create an artificial border. A p value of 0.049 and one of 0.051 are nearly the same evidence, yet binary language can turn the first into a discovery and the second into no effect. Effect estimates and intervals give you more to work with.
Small studies usually have wider uncertainty and more unstable estimates. When only striking small-study results become visible, published effects can be exaggerated. A large sample reduces random error, but it does not repair systematic bias. A very large biased study can estimate the wrong quantity with great precision.
Multiplicity creates more chances to be surprised#
A dataset may contain many outcomes, time points, and subgroups. It may contain many model specifications, definitions, and transformations. It may contain many exclusions and stopping rules. If analysts try enough combinations, some will cross a threshold by chance.
Multiplicity is not inherently misconduct. Exploratory analysis is essential for discovery. The problem arises when exploratory choices are reported as if they were one prespecified confirmatory test.
Registration, protocols, and statistical-analysis plans create a record of what was planned before results were known. They help readers identify changes, additions, and omissions. CONSORT 2025 strengthens randomized-trial reporting around registration, protocol access, and data sharing. It strengthens reporting of intervention description, effect estimates, harms, and the flow of participants.
Prespecification is not a ritual that makes an analysis correct. A planned method can still be poorly chosen. Nor should a sensible correction be prohibited when an assumption fails. The key is transparency: distinguish planned from post hoc analysis, explain deviations, report all important outcomes, and describe the uncertainty created by flexibility.
Bias can move the estimate#
Bias is a systematic departure from the target answer. It can enter before recruitment, during measurement, through missing data, in analysis, or during publication.
Selection bias occurs when inclusion, allocation, retention, or analysis creates groups that differ in outcome-relevant ways. Measurement bias occurs when outcomes or predictors are observed differently between groups. Confounding mixes the effect of a measured factor with other causes. Deviations from assigned intervention can change the meaning of a trial estimate. Missing outcomes can bias results when missingness is related to prognosis or treatment response.
Blinding can reduce some biases, but it is not always possible and it is not a universal quality label. Allocation concealment protects randomization before assignment; participant or assessor masking addresses what happens afterward. Objective outcomes can still be affected by missingness, timing, processing, or analytic choice. What a risk-of-bias assessment does is ask how each of those mechanisms applies to the specific result, which is why a checklist score that sums unrelated items can hide the one flaw that matters.
The chosen outcome may not be the outcome that matters#
Studies often use intermediate or surrogate endpoints because they can be measured sooner and with fewer participants. Blood pressure, tumor response, viral load, laboratory concentrations, or imaging findings may be informative. A change in a surrogate does not always translate into better survival, function, symptoms, or quality of life.
Surrogate validity depends on context. It is not enough for the marker to correlate with outcome risk, and the treatment's effect on the marker must reliably predict its effect on the clinical outcome across relevant interventions and settings. Off-target harms or alternate biological pathways can break that link.
Composite outcomes can increase event counts, but the components may differ in severity, frequency, and susceptibility to judgment; a favorable composite driven mainly by a frequent, less important component should not be described as if every component improved equally.
Outcome timing matters too. An early benefit can fade, a delayed harm can emerge, and treatment discontinuation can alter later effects. One follow-up window cannot show the entire trajectory.
Subgroup findings need restraint#
Researchers often ask whether effects differ by age, sex, or disease severity. They ask whether effects differ by biomarker status, geography, or prior treatment. These questions can identify real heterogeneity, but subgroup analyses multiply opportunities for chance findings.
The relevant statistical question is not whether treatment was significant in one subgroup and nonsignificant in another. It is whether the effect estimates differ from each other beyond expected sampling variation, usually assessed with an interaction test.
Credibility rises when a subgroup hypothesis was prespecified, limited in number, supported by a biological or contextual rationale, measured at baseline, demonstrated by a credible interaction, consistent across related outcomes, and replicated elsewhere. A dramatic forest-plot pattern alone is weak evidence.
Subgroups can also reduce applicability. A trial may exclude people with common comorbidities, polypharmacy, limited mobility, language barriers, or reduced access. Strong internal validity within the enrolled sample does not guarantee the same absolute benefit, feasibility, or harm balance in broader practice.
Rare harms and long latency are easy to miss#
A trial sized for a common efficacy endpoint can be too small or too short to identify a rare adverse event. Eligibility criteria and close monitoring may further reduce the harms seen during the study.
Post-authorization surveillance, registries, and claims data can add information. So can electronic health records, pharmacovigilance reports, and long-term extensions. Each has limitations: spontaneous reports lack a clear denominator, administrative codes may misclassify events, and observational comparisons face confounding.
The answer is not to choose one “best” source for every safety question. It is to combine designs while respecting what each can estimate; a signal may begin with case reports, become measurable in a database, and receive a more causal test through a randomized comparison or natural experiment.
Publication creates a selected record#
The public literature is not a random sample of completed analyses. Positive, surprising, and favorable results are often easier to publish, promote, and cite. Entire studies can remain unavailable, while outcomes within a published study may be omitted or reframed.
Trial registries, regulatory submissions, and protocols help you see the missing record. So do analysis plans, preprints, data repositories, and requests for unpublished results. Funnel plots and statistical tests can sometimes suggest small-study effects, but they do not diagnose publication bias on their own.
Funding and author interests do not automatically invalidate a study. They create reasons to examine design, comparator choice, and outcome hierarchy. They create reasons to examine analysis, data access, writing, and publication rights. Transparency makes scrutiny possible; it does not substitute for scrutiny.
What Ioannidis's 2005 argument did and did not show#
The essay “Why Most Published Research Findings Are False” used a mathematical framework to show how credibility depends on prior plausibility, statistical power, the number of relationships tested, bias, and competition among teams. In fields with low prior probability, small studies, flexible analysis, and selective publication, a nominally positive result can have a high chance of being false.
The title is often repeated as if one empirical audit proved that most papers in every discipline are wrong. That is not what the paper established. Its value lies in the structure of the argument: a p value cannot be interpreted apart from the research environment that generated it. The practical response is better design, adequate sample size, and registered protocols. It is complete reporting, data and code access when appropriate, replication, and evidence synthesis. Cynicism is not a method.
Reproducibility and replication answer different questions#
The National Academies distinguishes reproducibility from replicability in a useful way.
Computational reproducibility means obtaining consistent results using the same data, code, methods, and analysis conditions. It detects coding errors, undocumented data changes, software dependencies, and ambiguities in the reported method.
Replicability means obtaining consistent results in a new study that collects new data to address the same scientific question. It tests whether the finding survives new sampling and implementation.
A result can be reproducible but not replicate. The original code can perfectly regenerate an estimate that arose from chance, bias, or a setting-specific effect. A result can also appear not to replicate because the new study targeted a different population, dose, outcome, or estimand. Replication needs conceptual alignment, not merely a similar title.
Exact duplication is not always the strongest test. A direct replication probes the original conditions. A conceptual replication changes methods while testing the same underlying proposition. Convergence across both types is more informative than repetition of one pipeline.
A body of evidence is not a vote count#
Counting positive and negative studies ignores their size, precision, design, bias, and target question. Five small weak studies do not automatically outweigh one large rigorous study, and a pooled estimate does not erase the defects of its inputs.
Systematic review begins with a defined question and an attempt to find all eligible evidence. Reviewers assess study limitations, extract comparable results, and decide whether pooling is sensible. Heterogeneity is not merely a statistic to remove; it can reveal real differences in population, intervention, outcome, follow-up, or method.
GRADE evaluates certainty across a body of evidence using domains that include risk of bias, inconsistency, indirectness, imprecision, and publication bias. Other factors can raise or lower certainty depending on design and context. The final rating expresses confidence in an effect estimate for a particular outcome and question, not a permanent grade for a topic.
A meta-analysis can improve precision when studies estimate sufficiently related quantities. It can also produce a polished average of incomparable or biased estimates. Protocol, search strategy, and inclusion rules matter as much as the diamond at the bottom of a forest plot. So do outcome choice, model, and sensitivity analyses.
Triangulation strengthens causal reasoning#
Different designs fail in different ways. Randomized trials reduce confounding by indication but may be short and selective. Cohort studies can represent broader practice but face confounding. Case-control studies efficiently examine rare outcomes but depend on selection and measurement. Mechanistic research supports plausibility but may not predict clinical effects. Natural experiments can exploit external variation but require strong assumptions.
Triangulation asks whether evidence produced through different bias structures points toward a coherent conclusion. If a randomized trial, a well-controlled observational study, a dose-response pattern, and a plausible mechanism agree, no single shared error may explain all of them. If they conflict, the pattern can reveal effect modification, measurement problems, adherence differences, or hidden bias.
Triangulation is not permission to assemble only supportive evidence. The value comes from specifying how each design could be wrong and asking whether those errors would plausibly generate the observed pattern.
A practical reading sequence#
Begin with the research question and prespecified primary outcome. Identify the comparison, time horizon, and analysis population.
Read the effect estimate and the confidence interval before the abstract's adjectives. Compare absolute and relative effects. A large relative change can represent a small absolute difference at low baseline risk.
Inspect participant flow, exclusions, and missing data. Inspect protocol deviations, adherence, and harms. Check the registration and the protocol for outcome switching, or for analysis changes nobody explained.
Ask whether the comparator represents current care and whether the setting resembles the decision you are facing. Separate evidence of efficacy under study conditions from feasibility and effectiveness in routine use.
Then locate related studies, systematic reviews, guidelines with transparent methods, and later updates. Note whether the new study confirms, narrows, contradicts, or merely adds another uncertain estimate. Then end with a calibrated conclusion: what seems established, what remains uncertain, and what evidence would change your judgment.
The cumulative-evidence conclusion#
A single study is a structured observation, not a final verdict. Its value comes from the clarity of its question, the protection of its design, the completeness of its data, and the honesty of its analysis.
Science becomes dependable through correction and accumulation. Reanalysis checks the computation, replication tests new data, longer follow-up finds the delayed outcomes, broader populations test transportability, systematic review places every eligible result in context, and triangulation asks whether methods that fail differently converge anyway.
The goal is neither instant belief nor reflexive dismissal. It is calibrated updating: let a strong study change the conclusion, but require the conclusion to survive the limits that no single study can examine by itself.
References#
- National Academies of Sciences, Engineering, and Medicine. Reproducibility and Replicability in Science. 2019.
- Higgins JPT, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Current version.
- GRADE Working Group. GRADE Handbook.
- Hopewell S, et al. CONSORT 2025 statement. BMJ. 2025.
- Ioannidis JPA. Why most published research findings are false. PLoS Medicine. 2005.
- EQUATOR Network. Reporting guidelines for health research.
- Sterne JAC, et al. RoB 2: a revised tool for assessing risk of bias in randomized trials. Cochrane.
Questions and answers
Can one large randomized trial ever be enough?
It can provide high-certainty evidence for a well-defined question, especially with a large effect, strong design, complete reporting, and supportive context. Confirmation, longer follow-up, safety evidence, and applicability may still be needed.
Does a statistically significant result mean it will replicate?
No. Replication also depends on effect size, precision, bias, prior plausibility, analytic flexibility, publication selection, and whether the new study targets the same question.
Is a meta-analysis always better than one study?
No. A systematic synthesis can be stronger, but pooling biased, selectively reported, or clinically incompatible studies can give a precise yet misleading average.
What is the difference between reproducibility and replication?
Reproducibility usually means regenerating a result with the same data and analysis. Replication means testing the scientific question with newly collected data.
Should an early surprising study be ignored?
No. Treat it as evidence that can update a belief, with the size of the update matched to design quality, precision, prior evidence, multiplicity, and the need for confirmation.