Evidence explainer

Evidence and research methods

How to Read a RoB 2 Risk-of-Bias Assessment

RoB 2 does not award a quality score to a trial. It asks whether bias could materially affect one specified result, across five domains tied to the effect of interest.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. A result, not a paper, is the unit of assessment
  2. Choose the effect of interest
  3. Domain 1: bias arising from the randomization process
  4. Domain 2: bias due to deviations from intended interventions
  5. Domain 3: bias due to missing outcome data
  6. Domain 4: bias in measurement of the outcome
  7. Domain 5: bias in selection of the reported result
  8. How signalling questions and algorithms work
  9. The overall judgment is not an average
  10. Risk of bias is not every aspect of evidence quality
  11. A practical reading sequence

Randomization can create comparable groups, but a randomized label does not guarantee an unbiased estimate. Allocation may be predictable, participants may deviate from assigned treatment, outcomes may be missing or measured differently, and investigators may select one favorable result from several analyses. Cochrane's RoB 2 tool organizes appraisal around those pathways for a specific result.

A result, not a paper, is the unit of assessment#

A trial can have low risk of bias for objectively measured mortality at 30 days and high risk for a selectively reported subjective symptom score at 12 months. Missingness, assessor knowledge, analysis choices, and deviations differ by outcome and time.

Before the colored traffic lights, locate the exact result assessed. State the intervention and comparator, outcome definition, time point, effect measure, analysis population, and whether the target is the effect of assignment or the effect of adhering to intervention. An assessment that says only “Study X: low risk” is not RoB 2 as intended. It hides the result-specific reasoning you need.

Choose the effect of interest#

For many pragmatic treatment questions, reviewers want the effect of assignment to intervention, analogous to an intention-to-treat estimand; this includes consequences of ordinary deviations after assignment and asks what happens under a policy of assigning each option.

Other questions concern the effect of adhering to intervention as specified. That requires addressing deviations and nonadherence differently and often needs causal methods beyond a per-protocol comparison. Simply excluding participants who did not comply can destroy randomization because reasons for adherence may predict outcome. RoB 2's second domain has pathways tailored to the chosen effect. Reviewers should not switch targets after seeing which analysis looks cleaner.

Domain 1: bias arising from the randomization process#

This domain asks whether the allocation sequence was random, whether it was concealed until enrollment and assignment, and whether baseline differences suggest a randomization problem.

Sequence generation and concealment are distinct. A computer-generated sequence can still be subverted if recruiters can see upcoming assignments. Central allocation or an appropriately controlled randomization service can protect concealment. Alternation, birth date, record number, or an open list is predictable and not truly randomized.

Baseline p-values should not be used to test whether randomization “worked.” Chance creates some imbalances. Instead, look for patterns unlikely under chance, important prognostic imbalance combined with weak concealment, or exclusions that occurred after allocation; a single imbalance does not prove bias, and perfectly balanced tables do not prove concealment.

Domain 2: bias due to deviations from intended interventions#

For the effect of assignment, the central concerns are deviations caused by trial context and whether they were unbalanced and likely to affect outcome, and participants learning their allocation may seek extra care, clinicians may co-intervene, or staff may apply different thresholds. Blinding can reduce these deviations but is not always feasible.

The analysis should preserve participants in randomized groups. Excluding people after assignment, analyzing “as treated,” or using an inappropriate per-protocol set can introduce bias. A “modified intention-to-treat” label should make you ask two questions: what exactly was excluded, and why?

For the effect of adhering, reviewers consider deviations from the regimen and whether appropriate analysis adjusted for them. Time-varying adherence, treatment switching, and prognostic reasons for stopping make naive comparisons misleading. Blinding alone does not determine the rating. An open-label trial of an objective outcome can remain credible if co-interventions and behavior were balanced; a nominally blinded trial can fail if side effects reveal allocation.

Domain 3: bias due to missing outcome data#

Count how many randomized participants are missing the assessed outcome in each group and why. The proportion alone is not enough. Even 5% missing can seriously bias a small effect if missingness is strongly related to unobserved outcomes; a larger proportion may matter less under convincing evidence that missingness is unrelated.

Complete follow-up or a valid analysis that accounts for missingness can support low risk, and reasons such as adverse effects, worsening disease, lack of efficacy, or inability to attend can depend on outcome. Similar dropout percentages across groups do not guarantee similar missing outcomes.

Last observation carried forward and simple mean imputation rarely solve the problem and can understate uncertainty. Multiple imputation also rests on assumptions about what predicts missing values. Look for sensitivity analyses under plausible departures from missing-at-random assumptions.

Do not confuse missing outcome data with selective nonreporting of an entire measured outcome. The former belongs mainly here; the latter belongs in domain 5.

Domain 4: bias in measurement of the outcome#

Ask whether the measurement method was inappropriate, differed between groups, or could have been influenced by knowledge of intervention. The answer depends on the outcome.

All-cause mortality is difficult to manipulate but can still be missing or misclassified in some systems. Imaging interpretation, clinical event adjudication, clinician-rated scales, self-reported symptoms, and treatment decisions involve more judgment. If participants know assignment, their reporting may change; if assessors know assignment, ambiguous events may be classified differently.

Blinded independent adjudication can help when criteria and source data are adequate. It does not repair an invalid outcome definition or differential surveillance. If one group attends more visits, more events may be detected even with blinded adjudicators.

Patient-reported outcomes are not inherently biased. They are the correct way to measure symptoms and lived function, and risk arises when intervention knowledge plausibly changes reporting beyond the construct of interest, or the instrument and administration are unsuitable.

Domain 5: bias in selection of the reported result#

This domain compares the reported result with a prespecified plan finalized before unblinded data were available. Investigators may have multiple eligible scales, time points, subscales, thresholds, covariate adjustments, missing-data methods, and analysis populations. Choosing among them after seeing results can bias the estimate.

Trial registration helps but often lacks analysis detail. The protocol and the dated statistical analysis plan are what you want. Check amendments and timing. A plan posted after data lock cannot prove prespecification.

Outcome switching is one form, but selective analysis is subtler. The trial may report the promised outcome while choosing the most favorable model. A result can be completely reported and still be selected from many candidates.

Absence of a public plan does not prove selection occurred, but it may justify some concerns when multiple analyses were possible and authors cannot clarify; reviewers should distinguish lack of information from evidence of high risk.

How signalling questions and algorithms work#

Each domain contains factual or inferential questions with response options such as yes, probably yes, probably no, no, and no information. These feed an algorithm that proposes low risk, some concerns, or high risk.

“Probably” should reflect a reasoned judgment based on available evidence, not discomfort. Reviewers may override the algorithm, but should document why. Free-text support is the audit trail; colored symbols without quotations, page references, registry checks, and logic are not reproducible.

At least two trained reviewers commonly make separate assessments and resolve differences. Calibration on sample trials helps align interpretation. RoB 2 is detailed because casual judgment is unreliable.

The overall judgment is not an average#

Overall low risk generally requires low risk in every domain, and some concerns in at least one domain usually leads to overall some concerns, unless multiple concerns collectively justify high risk. High risk in one domain generally produces high overall risk.

Do not assign numbers such as low=1, some concerns=2, high=3 and calculate a mean. Domains are not interchangeable and the categories are not equally spaced. A severe selective-reporting problem cannot be canceled by excellent randomization.

The direction of bias is sometimes predictable, but not always. Reviewers can note whether bias likely favors an intervention, moves toward no effect, exaggerates magnitude, or is unpredictable. This helps interpret rather than merely label evidence.

Risk of bias is not every aspect of evidence quality#

RoB 2 addresses systematic error in a randomized-trial result. It does not grade imprecision, indirectness, inconsistency across trials, publication bias across the evidence base, relevance, feasibility, or ethics. Those belong elsewhere in evidence assessment.

A low-risk result can still be too imprecise to guide practice. A high-risk result can occasionally point in the correct direction, but the design does not let you rely on it confidently. Risk of bias is not a claim of research misconduct.

Reporting quality also differs from bias. A well-conducted trial may be rated some concerns because essential details are unavailable. That is appropriate uncertainty for you, not proof the concealed method was poor.

A practical reading sequence#

First, verify that the assessment names a specific result and RoB 2 version or variant. Second, confirm the effect of interest. Third, read supporting judgments for each signalling question rather than the summary color.

Check source documents: article, supplement, registry, protocol, analysis plan, regulatory review, and author clarification. Examine whether judgments follow the correct design variant and whether reviewers handled missing information consistently.

Finally, see how risk-of-bias judgments affected synthesis: sensitivity analyses, subgrouping, certainty ratings, or interpretation. If every trial receives low risk despite missing plans, large attrition, and unblinded subjective outcomes, the tool may have been applied mechanically.

Sources and further reading

  1. Cochrane Handbook, Chapter 8, Assessing Risk of Bias in a Randomized Trial
  2. Cochrane Methods, Risk of Bias 2 Tool and Full Guidance
  3. Sterne and colleagues, RoB 2, A Revised Tool for Assessing Risk of Bias in Randomised Trials, BMJ (2019)
  4. Cochrane, About Risk of Bias 2
  5. Minozzi and colleagues, Empirical Evidence on Blinding and Randomized-Trial Effect Estimates, Cochrane Database of Systematic Reviews

Questions and answers

Is an unblinded trial automatically high risk?

No. The impact depends on deviations, co-interventions, outcome measurement, and the effect of interest. Some objective outcomes can remain at low risk in open-label trials.

Does intention-to-treat analysis guarantee low risk in domain 2?

No. It supports the effect-of-assignment analysis, but trial-context deviations and incorrect handling of post-randomization events may still matter.

How much missing data is acceptable?

There is no universal percentage. The likely outcomes among missing participants, reasons, group differences, effect size, and sensitivity analyses determine risk.

Can a trial have different RoB 2 ratings for two outcomes?

Yes. That is expected when missingness, measurement subjectivity, time point, or selective-reporting opportunities differ.

Is “some concerns” the same as medium quality?

No. It is a domain-based bias judgment for a result, not a generic quality tier or numerical midpoint.