Evidence explainer

Evidence and research methods

Single-Blind, Double-Blind, and Open Peer Review: What Trials Show

Masking author identities or disclosing reviewer names barely moved rated review quality. Those trials tested formats, not whether peer review works.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. First define the intervention
  2. What “review quality” meant in the trials
  3. Double-blind review: the 1998 masking trial
  4. Open identity: the 1999 randomized trial
  5. Open reports: the 2010 experiment
  6. What systematic reviews add
  7. Format trials are not a verdict on peer review
  8. How to read a journal's policy
  9. A more useful research agenda
  10. References

Peer review labels describe who can see whose identity and, in some systems, who can read the reports. Single-blind review usually means reviewers know the authors but authors do not know the reviewers. Double-blind review usually attempts to conceal both sides. Open review can mean signed reports, public reports, author identities known to reviewers, or several of these features together.

Randomized trials at biomedical journals tested some of these design choices. Masking author identity did not measurably improve average review-quality scores in a major 1998 trial, and masking often failed. Asking reviewers to sign their reports did not materially change rated quality in a 1999 trial, but more invitees declined, and a 2010 trial found that possible publication of signed reports did not meaningfully change quality either. These findings are informative and narrow. They do not prove that peer review itself is effective, ineffective, unbiased, or interchangeable across disciplines.

First define the intervention#

The familiar labels can mislead. “Single blind” may conceal the reviewer from the author while leaving the author's identity visible. “Double blind” may remove names and affiliations from the manuscript, yet reviewers can infer identity from citations, topic, data set, writing style, or a preprint. “Open peer review” may reveal names only to the author, publish the reports, publish author replies, reveal the editor, or allow public commenting after publication.

Two journals can therefore claim the same model while running different processes, so a policy worth trusting tells you:

Without that detail you are comparing medicines by the color of the box rather than by the active ingredient.

What “review quality” meant in the trials#

Trials need measurable outcomes. Several journal experiments used versions of a Review Quality Instrument. Editors or authors rated whether a report addressed the research question, originality, methods, presentation, interpretation, constructiveness, and support for comments. Scores were averaged.

That outcome captures useful features, but it is not truth. A polite, detailed report can miss a fatal analytic error. A short report can identify the decisive flaw. Authors may rate a recommendation for rejection differently from editors: the 1999 BMJ trial found that author ratings were more favorable for reviews recommending publication, illustrating how outcome assessment can depend on a participant's stake in the decision. A quality score is also only one outcome among many, alongside time spent, turnaround, reviewer refusal, recommendation, number and type of errors detected, changes made to the manuscript, editorial decision, civility, diversity, conflicts, retractions, and the reliability of published claims. No single early trial measured all of them.

Double-blind review: the 1998 masking trial#

The 1998 JAMA trial coordinated work at several journals. Manuscripts were sent to paired reviewers under masked or unmasked conditions, and review quality was scored. Concealing author identity did not produce a meaningful average improvement in quality.

The trial also examined whether masking worked. Reviewers were asked to guess authorship, and success varied. Familiarity with the author reduced masking success. A specialized field, self-citation, recognizable methods, institutional data, conference presentations, and circulating preprints can all reveal provenance despite removed names.

The correct conclusion is not that double-blind review has no value. The experiment found no average quality gain under its implementation, journals, manuscripts, reviewers, and measurement tool. Double-blind review may still influence perceived fairness, willingness to submit, prestige bias, geographic bias, gender bias, or early-career participation. Those outcomes need their own credible studies.

It is also possible for masking to carry costs. Removing clues about a research group can make it harder to recognize duplicate publication, undisclosed reuse of a cohort, or a conflict of interest. Editors can retain author information while masking reviewers, but the process must be designed deliberately.

Open identity: the 1999 randomized trial#

The BMJ randomized reviewers to remain anonymous or to be asked to reveal their identity to the authors. Review quality scores were nearly the same: the anonymous group averaged 3.06 and the identified group 3.09 on a 1-to-5 scale. Recommendations and time taken also did not differ significantly.

Reviewer recruitment did change. Thirty-five percent of reviewers asked to be identified declined, compared with 23% in the anonymous group, a 12 percentage-point difference. The estimate had uncertainty, but it identified a practical tradeoff. A format can leave the reports of participating reviewers unchanged while altering who agrees to participate.

Selection after invitation matters. The signed-review group consists only of people willing to review under that condition. More junior scholars, researchers in small fields, people reviewing powerful authors, or people raising sensitive concerns may make different choices, and a refusal-rate effect can change workload and representation even if average report scores among completed reviews look similar.

Open identity can support accountability and allow credit. It may also discourage frank criticism or create fear of retaliation. The trial did not establish that either mechanism dominates in every field. It showed feasibility at one large medical journal and no important mean quality loss among returned reviews.

Open reports: the 2010 experiment#

A later BMJ trial tested another feature. Reviewers had already agreed to sign their reports. Eligible manuscripts were randomized so reviewers were told that the signed report might be posted online with the published article or would be shared only with the author.

Editors and authors found no meaningful difference in review quality, though the possibility of public posting was associated with a high refusal rate earlier in recruitment and more time spent on reviews. Because all arms already involved signed review, this was not a clean comparison of fully anonymous review with fully public review. It isolated the additional prospect of public reports.

That distinction is central. “Open” is a bundle, so when a journal tells you it runs open review, ask which part it means. A trial of signed identity does not test public reports, and a trial of public reports among willing signers does not test mandatory openness for everyone. Conclusions should match the randomized contrast.

What systematic reviews add#

Systematic reviews of interventions to improve biomedical peer review have pooled small, heterogeneous experiments. One 2016 review found that open review had, at most, a small favorable association with quality scores while increasing reviewer time and reducing rejection recommendations. Training and checklists produced mixed results. The authors emphasized the limited evidence base.

The Cochrane review of editorial peer review reached an even broader caution: direct experimental evidence that editorial peer review improves the quality of biomedical reports is sparse. This does not mean peer review contributes nothing. It means the counterfactual is difficult to study. Journals rarely randomize publishable manuscripts to no external review, then follow errors, corrections, and downstream use.

The evidence also ages. Preprints, data repositories, code sharing, registered reports, online discussion, and AI-assisted screening change what identities can be concealed and what review is expected to do; a trial conducted before widespread preprinting cannot answer whether present-day masking succeeds.

Format trials are not a verdict on peer review#

Several causal questions are often collapsed:

  1. Does external review improve a manuscript compared with no external review?
  2. Does concealing author identity change reviewer behavior?
  3. Does revealing reviewer identity change report content or recruitment?
  4. Does publishing reports help readers evaluate the editorial process?
  5. Does a review system reduce social bias or redistribute it?
  6. Does the system detect invalid work before or after publication?

The classic trials mainly address questions 2 and 3, with one study addressing part of question 4. They do not directly answer question 1, which is the one you are relying on whenever you treat “peer reviewed” as a mark of soundness. They also study average effects, which can conceal important subgroup differences and rare harms.

A review model is a governance choice as well as an intervention. A journal may value transparency even if a quality score is unchanged. Another may prioritize safe candid criticism in a small or hierarchical field. Empirical evidence should inform those values, not pretend to replace them.

How to read a journal's policy#

Look past the label when you are sizing up a journal. Does it publish review histories? Are reviewer names optional or mandatory? Can authors appeal? Are statistical, methodological, patient, or data reviews available when relevant? Are conflicts declared? Does the journal describe how editors select reviewers and manage competing interests?

Then separate transparency from validity. Public reports let you see some of the process, but a visible process can still miss errors. Anonymous review can be rigorous, and signed review can be superficial. Publication in a peer-reviewed journal is a reason for you to inspect the methods, not a substitute for inspecting them.

Related articles explain who qualifies as an author under ICMJE and CRediT, how reporting guidelines such as CONSORT work, and corrections, retractions, and expressions of concern. The site's research overview connects publishing methods to broader evidence appraisal.

A more useful research agenda#

Future studies should prespecify which open or masked features are changing and measure more than a composite quality score. Outcomes could include specific error detection, inappropriate citation requests, civility, reviewer diversity, author revision, decision reliability, time, refusal, appeals, and later corrections. Trials should report invitations as well as completed reports so selection is visible.

Cluster randomization may be needed when editors or reviewers handle multiple manuscripts. Qualitative work can explain why a format changes behavior. Audits of bias require carefully designed simulated manuscripts or matched real submissions, because author identity is entangled with topic, institution, methods, and resources.

The strongest conclusion from the existing trials is modest. Review format changes are testable. The classic experiments did not find large average quality improvements from masking authors or naming reviewers, and they identified feasibility and recruitment tradeoffs. That is better than relying on intuition, but it is not the last word, and it is not enough for you to read a journal's label as a verdict on the paper under it.

References#

  1. Masking author identity and review quality, randomized trial
  2. Open reviewer identity and review quality, randomized trial
  3. Possible public posting of signed reports, randomized trial
  4. Systematic review of interventions to improve biomedical peer review
  5. Cochrane review of editorial peer review

Questions and answers

Is double-blind peer review proven to be better?

No. A major randomized biomedical-journal trial found no average improvement in rated review quality, and author masking was often incomplete. Other outcomes, such as fairness or bias, require separate evidence.

Did open peer review make reviewers less critical?

The 1999 randomized trial found no significant difference in publication recommendations or mean quality scores. It was not large enough to exclude every change in candor or every subgroup effect.

Why did more reviewers decline signed review?

Possible reasons include workload, discomfort with disclosure, hierarchy, conflict risk, or principled opposition. The trial measured the refusal difference but did not fully establish each person's reason.

Does a public review report prove an article is reliable?

No. It improves process visibility, but readers still need to appraise the study design, data, analysis, reporting, and the issues that reviewers may have missed.

Have trials proved that peer review works better than no peer review?

Not convincingly. Most randomized studies compare versions of peer review among manuscripts already selected for review. Direct trials of the entire editorial system against no external review are rare and difficult.