Evidence explainer

Evidence and research methods

The File Drawer Problem

Your literature search can be flawless and still return a distorted sample of the research that was actually done. The distortion happened long before you started searching.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What exactly goes into the drawer?
  2. Why the bias changes the answer
  3. A documented example from antidepressant trials
  4. Publication bias is not the same as ordinary missing data
  5. What a funnel plot can and cannot tell you
  6. Look for records that exist before results
  7. How a careful reviewer searches for missing evidence
  8. How to read a meta-analysis when the drawer may be open
  9. Why null and uncertain results have value
  10. The evidence-completeness conclusion
  11. References

Suppose ten teams test the same treatment. Chance, sampling variation, and genuine differences among studies produce a spread of results. Two studies look strikingly favorable, five look uncertain, and three lean against benefit. If only the striking studies reach journals, the published record no longer describes the research program. It describes the selected winners.

That is the file drawer problem. Robert Rosenthal used the phrase in 1979 to capture a simple threat to inference: studies with results judged uninteresting or unfavorable may remain in researchers' files while more exciting findings become visible. Modern evidence science uses broader terms such as publication bias, non-reporting bias, and bias due to missing evidence.

The metaphor remains useful, but the real system is more complicated than a literal drawer. A study may appear years late, only as a conference abstract, in a registry without a journal article, or in a paper that omits an unfavorable outcome. The missing unit can be a whole study, a time point, an analysis, a subgroup, or a harm.

What exactly goes into the drawer?#

Several different processes can remove evidence from view.

Whole-study nonpublication#

A completed study never receives a full public report. Investigators may decide the result is not worth writing, a sponsor may not prioritize submission, or journals may decline it; the motives differ, but the inferential problem is the same when the decision correlates with the findings.

Delayed publication#

An unfavorable study may eventually appear, but years after favorable studies. During that interval, reviews, guidelines, and decisions rely on an incomplete record. Time-lag bias can therefore matter even when the final publication rate looks acceptable.

Selective outcome reporting#

A trial measures pain, function, quality of life, adverse events, and several laboratory outcomes, but the article emphasizes only results that crossed a statistical threshold. The study itself is visible; the prespecified outcome set is not.

Selective analysis reporting#

Researchers can define populations, covariate adjustments, time points, or statistical models in multiple defensible ways. Choosing which analysis to present after seeing the results makes the published estimate a selected result rather than a neutral one.

Location and language bias#

Results may be more likely to reach prominent journals or English-language publications when they look favorable, and a search you restrict to the major databases then samples a filtered evidence stream.

These mechanisms can coexist. A trial can be registered, publish one article, omit its primary outcome, and delay adverse-event results. Calling the study “published” would miss most of the problem.

Why the bias changes the answer#

A meta-analysis estimates an average from the studies it can include. Its mathematics cannot tell you whether the input set is a result-dependent sample of all the studies performed, and if favorable estimates have a better chance of entering the data set, the pooled effect can move away from the truth.

Imagine a treatment whose true average effect is zero. Small studies will still scatter around zero. By chance, some will have P values below 0.05, and if those papers are visible while null-looking studies disappear, a review can find an apparently consistent benefit even though selection created the pattern.

The distortion is not limited to whether an effect exists. It can affect:

The direction is not always toward benefit. In fields where surprising harms attract attention, unusual adverse associations may be overrepresented. The defining feature is result-dependent visibility, not a fixed direction.

A documented example from antidepressant trials#

One unusually informative analysis compared journal articles with regulatory records. Erick Turner and colleagues examined 74 studies of 12 antidepressants involving 12,564 participants that had been reviewed by the US Food and Drug Administration.

The FDA judged 38 studies positive; 37 of them were published as positive. Among 36 studies judged negative or questionable, only three were published in a way that agreed with that assessment, and the rest were unpublished or were written in a way the investigators considered positive despite the FDA judgment.

From journal publications alone, 94% of the visible trials appeared positive. In the FDA record, 51% were positive, and the mean weighted effect size was 0.37 in published studies and 0.15 in unpublished studies, and the published literature inflated the overall effect-size estimate by 32%.

This example does not imply that every field has the same bias or magnitude. It shows what becomes measurable when an external record identifies studies regardless of their results. Without that denominator, the unseen research is much harder to count.

Publication bias is not the same as ordinary missing data#

Missing participant outcomes within a study can bias an effect estimate, especially when loss to follow-up depends on prognosis or treatment response. Publication bias operates at a different level: an entire study result may be absent from the evidence synthesis.

The remedies differ. Multiple imputation may address some missing participant data under explicit assumptions. It cannot identify an unknown trial that was never registered or disclosed.

It is also useful to separate non-reporting from poor study quality. A perfectly randomized, well-conducted trial can contribute to bias if its result stays hidden; conversely, a published trial may have high risk of bias even if every outcome is reported. Visibility and validity are separate dimensions.

What a funnel plot can and cannot tell you#

A funnel plot places each study's effect estimate against a measure of its precision or size, and in an idealized setting, smaller studies scatter more widely around the pooled effect, forming a roughly symmetric inverted funnel. If small unfavorable studies are missing, one side may look sparse.

Asymmetry is a clue, not a diagnosis. Other causes include:

Cochrane therefore warns against treating funnel-plot asymmetry as proof of nonpublication, and statistical tests for asymmetry also have limited power when few studies exist and can detect small-study effects that have other explanations.

“Trim and fill” and selection models explore how conclusions might change under assumptions about missing studies. They do not reconstruct known facts. Their output should be labeled as sensitivity analysis, not corrected truth.

Rosenthal's fail-safe number asked how many null studies would be required to overturn statistical significance. The intuition is accessible, but the method assumes a particular distribution of unseen results and keeps the focus on a threshold. A large fail-safe number does not establish that the effect estimate is unbiased or clinically important.

Look for records that exist before results#

The strongest defense creates a public trace before investigators know the answer.

Prospective trial registration records the research question, treatment groups, primary outcomes, and planned time points. A registry lets reviewers search for a study even if no journal article appears. A dated protocol and statistical analysis plan make later changes visible. Registry summary results can disclose findings without waiting for journal publication.

Registration alone is not enough. A vague outcome such as “symptoms” leaves room to choose among many scales and time points. Retrospective registration occurs after researchers have seen some data. Records may not be updated, and results may remain overdue. The useful question is not merely “Was there a registration number?” but “Was the study registered prospectively, specifically, and completely?”

WHO's 2025 registry-results guidance identifies a minimum set of summary items intended to make trial findings interpretable, and in the United States, applicable trials have results-reporting duties under FDAAA 801 and its final rule. NIH policy reaches NIH-funded clinical trials whether or not they fall under the statutory category. These systems create discoverable records, but compliance and enforcement still matter.

How a careful reviewer searches for missing evidence#

A journal database is a starting point, not the whole search.

Reviewers can search ClinicalTrials.gov, WHO's International Clinical Trials Registry Platform, regional registries, regulatory reviews, health-technology assessment reports, conference proceedings, dissertations, preprints, and sponsor result portals. They can compare publication outcomes with the registry and protocol. They can contact investigators for unavailable results and map multiple articles back to one underlying study.

The study-flow diagram should account for registered and otherwise identified studies that lack usable results; a risk-of-bias assessment should distinguish missing whole studies from selective non-reporting of an outcome in a known study. Review authors should explain why evidence may be missing and test how plausible assumptions affect conclusions. Even so, search breadth cannot guarantee completeness, because a never-registered study may leave no searchable trace at all. That is why prospective transparency is more powerful than retrospective detective work.

How to read a meta-analysis when the drawer may be open#

Ask seven questions:

  1. Was the review protocol registered before the analysis?
  2. Did the search include trial registries and regulatory sources?
  3. Were unpublished data eligible, or were studies excluded for lacking a journal article?
  4. Do registry outcomes match the published outcomes and time points?
  5. Is there qualitative evidence that results are missing because of their direction?
  6. Are funnel plots used only when enough studies make them interpretable?
  7. Do sensitivity analyses change the estimate or the decision?

A symmetrical funnel plot does not close the case. If all studies are large, if reporting bias affects large and small studies similarly, or if the number of studies is low, asymmetry may not appear. Direct comparisons between protocols, registrations, regulatory files, and publications will usually tell you more.

Why null and uncertain results have value#

A well-designed study that does not find a clear difference can narrow the range of plausible effects. It can rule out a large benefit, document harms, show that an intervention fails in a particular setting, or reveal that an endpoint is too noisy. Its information depends on precision and design, not on whether P is below 0.05.

Public reporting also respects participants. People accept research burdens to contribute to generalizable knowledge. Withholding interpretable findings wastes that contribution and may cause future participants to repeat an unproductive question.

The cultural remedy is to reward rigorous questions and complete reporting rather than dramatic answers; registered reports, in which journals review methods before results exist, separate publication decisions from outcome direction. Data-sharing policies, results deadlines, and funder monitoring add structural support.

The evidence-completeness conclusion#

The file drawer problem is a sampling problem hidden inside the research record. You see a set of studies and may assume it represents everything the investigators learned. That assumption is unsafe when visibility depends on results.

No single graph resolves the problem. The best response combines prospective registration, specific protocols, complete and timely results, broad searching, comparisons across records, and transparent sensitivity analyses.

The central reading habit is simple: before you ask what the published studies say, ask which studies and outcomes are missing, and why.

References#

Questions and answers

Is the file drawer problem the same as publication bias?

The file drawer problem is a vivid form of publication bias in which studies remain unpublished because their results seem uninteresting or unfavorable, and publication bias is broader and includes delayed, lower-visibility, or selectively framed publication.

Can a meta-analysis fix missing studies statistically?

No method can know the contents of evidence that was never disclosed. Models can test assumptions about missing studies, but registries, protocols, regulatory records, and direct result reporting provide stronger evidence.

Does a symmetrical funnel plot rule out publication bias?

No. Funnel plots have low information with few studies and can miss bias that does not depend on study size. Symmetry is reassuring only within its limited assumptions.

Are nonsignificant studies necessarily null studies?

No. A nonsignificant result may be imprecise and compatible with meaningful benefit or harm. The effect estimate and confidence interval are more informative than the label “negative.”

What is the most practical safeguard for a new trial?

Register it prospectively with specific outcomes and time points, make the protocol and analysis plan public, and report summary results on time regardless of the findings.