Evidence explainer

Evidence and research methods

What Blinded Outcome Review Adds to an Open-Label Trial

A PROBE trial randomizes openly and then judges outcomes without knowing who got what. That protects the classification, not everything that happened before it.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Why a partial blind exists
  3. The journey from event to endpoint
  4. Where adjudication adds the most value
  5. What empirical studies of assessor blinding show
  6. Adjudication can help without changing the primary result
  7. How to appraise a PROBE trial
  8. Match the claim to the protected outcome

A blinded endpoint committee can make an open-label randomized trial more reliable, but only for the part of the evidence it actually controls. It can prevent knowledge of treatment from coloring whether a suspected event meets a predefined endpoint. It cannot change what participants reported, what clinicians ordered, or which events were sent to the committee.

Key points#

Why a partial blind exists#

Full blinding is not always feasible. A procedure leaves an incision, a home blood-pressure strategy changes what equipment a participant uses, and a complex service redesign looks nothing like routine care. Even medication trials can have dosing schedules or recognizable effects that reveal assignment.

Researchers then face a choice. They can conduct an entirely open trial and accept that outcome judgment may be influenced, or they can separate treatment delivery from endpoint classification. PROBE is one version of the second approach; participants are randomized prospectively, the intervention is delivered openly, and a committee that does not know the assignment judges whether endpoint criteria were met.

The design is not a lesser form of randomization. Allocation can still be centrally concealed and the analysis can follow a prespecified plan. The compromise concerns knowledge after assignment.

The journey from event to endpoint#

An outcome does not appear in a dataset fully formed. Consider a possible stroke in a cardiovascular trial. Several steps occur:

  1. A participant notices symptoms and decides whether to report them.
  2. A clinician decides which examination, imaging, or consultation to obtain.
  3. Trial staff identify the episode as a potential endpoint.
  4. Records are assembled, redacted, and sent to adjudicators.
  5. The committee applies the protocol definition.
  6. The classified event enters the analysis.

Blinded adjudication directly protects step 5. It may also improve consistency in the evidence packet and definitions. It does not automatically protect steps 1 through 4.

If clinicians know a participant received the intervention, they may investigate ambiguous symptoms more or less intensively. If participants expect benefit, they may report subjective symptoms differently. If trial staff apply different thresholds for forwarding events, the adjudication committee sees a selected set; a perfectly blinded final decision cannot classify an event it never receives, and nothing in the report will show the ones that never arrived.

Where adjudication adds the most value#

The method is well suited to endpoints with an objective core and a judgment boundary. Examples include classifying stroke from symptoms and imaging, confirming myocardial infarction from clinical records and biomarkers, judging whether a hospitalization meets a disease-specific definition, or determining cause of death from a standardized dossier.

The committee should work from prespecified definitions and comparable records. Treatment names, device images, distinctive laboratory patterns, and narrative clues that reveal allocation should be removed when feasible. More than one reviewer, conflict resolution, documented queries, and an audit trail make the process reproducible.

All-cause mortality usually requires little adjudicative judgment, although completeness of follow-up remains essential. At the other end of the spectrum, pain, fatigue, mood, and quality of life originate with the participant. A blinded committee cannot recreate an uninfluenced symptom report after an unblinded participant has supplied it. For those outcomes, lack of participant blinding remains a central limitation.

What empirical studies of assessor blinding show#

Meta-epidemiologic studies have compared blinded and nonblinded assessors judging the same outcomes in the same randomized trials, and a 2012 BMJ review found that nonblinded assessment tended to exaggerate intervention effects for subjective binary outcomes. An updated 2025 analysis similarly reported an average exaggeration of about 29 percent when judgment was involved.

That percentage is not a correction to apply mechanically to every open trial. The included trials, endpoints, and assessment processes varied. The finding instead establishes that observer knowledge can move outcome classification enough to alter an effect estimate. The practical implication is to blind assessors whenever you can, and to scrutinize judgment-dependent outcomes whenever nobody did.

Adjudication can help without changing the primary result#

Endpoint committees are sometimes justified as a way to “improve accuracy,” but the effect on a treatment comparison depends on the errors they correct. If local investigators misclassify outcomes similarly in both arms, adjudication may change event counts while leaving the relative effect nearly unchanged, and if misclassification differs by treatment because local assessors know allocation, blinded review may materially change the comparison.

Committees can also create new problems. Definitions may be so restrictive that clinically relevant events are excluded. Evidence packets may differ across sites. Missing documents can force an “unable to determine” category. Repeated requests can delay database lock. Central consistency is useful only when the rules, source material, and missingness are transparent.

How to appraise a PROBE trial#

Start before outcome assessment:

Then inspect event capture:

Finally, inspect adjudication:

A report that merely says “events were blindly adjudicated” leaves you too much to guess. CONSORT 2025 asks trial reports to specify who was blinded and how. For a PROBE design, the route by which events were captured is just as important to you.

Match the claim to the protected outcome#

Suppose an open trial finds fewer adjudicated strokes but no difference in self-reported function. The blinded committee makes the stroke classification more credible. It does not confer blinding on the functional outcome. Conversely, if an open lifestyle trial improves a laboratory value measured by an automated central laboratory, the objectivity of the measurement reduces one concern, while differences in medication adjustment or follow-up may remain.

Each endpoint carries its own bias pathway. “The trial was PROBE” is therefore a starting description, not a global quality score.

Sources and further reading

  1. Blood Pressure, original description of the PROBE design
  2. BMJ, observer bias in randomized trials with binary outcomes
  3. Journal of Clinical Epidemiology, updated empirical analysis of blinded and nonblinded assessors
  4. CONSORT 2025 explanation and elaboration on blinding and outcome assessment

Questions and answers

Is PROBE the same as a double-blind trial?

No. In a PROBE trial, participants and treating clinicians generally know the assigned intervention. Only endpoint assessment is designed to be blind.

Does blinded adjudication make subjective outcomes objective?

Not when the outcome originates in an unblinded participant's report. A committee may apply scoring rules consistently, but it cannot remove expectation effects already present in the source information.

Are hard outcomes immune to open-label bias?

Not entirely. Death is difficult to misclassify, but cause of death can require judgment. Stroke or hospitalization can depend on whether symptoms were investigated. Completeness and equal event detection still matter.