A nonrandomized study can estimate an intervention effect without random assignment, but its numerical precision does not guarantee a causal answer. Treatment choices, eligibility, timing, missingness, measurement, and selective reporting can create a convincing association that differs from the effect the study set out to estimate.
ROBINS-I provides a structured way to judge those threats, and its defining move is to compare the study with a hypothetical pragmatic randomized trial addressing the same question, called the target trial. Reviewers then assess risk of bias for a particular result, not the prestige of the journal or the study as a whole.
A version choice is now essential. The original 2016 ROBINS-I has seven bias domains. A draft ROBINS-I V2 was released on November 20, 2025 with six domains, revised signaling questions, algorithms, and separate variants for intention-to-treat and per-protocol effects. V2 remains a draft subject to change as of July 15, 2026. A protocol should name and preserve the chosen version.
Write the target trial first#
The target trial should specify:
- eligibility criteria;
- intervention strategies and comparator;
- assignment procedures in the ideal randomized design;
- time zero, when eligibility, assignment, and follow-up begin;
- outcome and follow-up period;
- causal effect of interest, such as assignment to intervention or adherence to intervention;
- analysis that would estimate that effect.
This exercise turns a vague review question into a testable one. “Did the drug improve survival?” is incomplete, and until you say which comparison you mean the answer can differ between initiating treatment now versus not initiating, current use versus nonuse, continuous adherence versus discontinuation, or one dose strategy versus another.
Time zero deserves special attention. Eligibility, treatment classification, and follow-up should be aligned. If treatment is defined using events after follow-up begins, patients must survive and remain event-free long enough to be classified as treated. That guaranteed period is immortal time.
For example, a study may call anyone who fills a prescription within 90 days after discharge “treated” while counting deaths from discharge. A person who dies on day 10 cannot enter the treated group. The design has assigned early deaths preferentially to the comparator, even if the medicine has no effect.
Decide which effect is being estimated#
Randomized trials often distinguish an effect of assignment from an effect under adherence. Nonrandomized studies need the same clarity.
An intention-to-treat-type effect compares initiation or assignment strategies regardless of later deviation; a per-protocol-type effect compares sustained strategies and therefore requires appropriate handling of time-varying adherence, switching, and factors that affect both adherence and outcome.
ROBINS-I V2 provides separate confounding variants for these estimands. This is not cosmetic. Adjusting only baseline variables may be reasonable for an initiation effect but inadequate for an adherence effect when health changes over time influence both subsequent treatment and prognosis.
Domain 1: bias due to confounding#
Confounding occurs when causes of intervention choice also affect the outcome. In routine care, sicker patients may preferentially receive a treatment, creating confounding by indication. The reverse can happen when clinicians reserve an intervention for healthier patients likely to tolerate it.
Predefine the important confounding domains from subject knowledge, causal diagrams, prior evidence, and the target trial, before you have seen any results. Age, disease severity, prior treatment, comorbidity, socioeconomic access, calendar time, and healthcare-seeking may matter, but the required set is question-specific.
Then ask whether the study measured those domains accurately, at the right time, and analyzed them appropriately. “Adjusted analysis” is not enough. A database may contain a diagnosis code without capturing severity. A laboratory value measured after treatment starts may be affected by the treatment. Broad propensity-score overlap cannot repair an unmeasured cause.
Adjustment can also cause harm. Conditioning on a mediator blocks part of the effect. Conditioning on a collider can open a noncausal path. Automated selection based only on statistical significance is a weak substitute for causal reasoning. Residual confounding should be judged in direction and magnitude where possible. Negative-control outcomes, active comparators, quantitative bias analysis, sensitivity analysis, and triangulation can strengthen interpretation, but none mechanically proves that confounding is absent.
Domains 2 and 3: classification and selection#
Bias in classification of interventions concerns whether the study correctly assigns the strategies being compared. Prescription orders do not prove dispensing, dispensing does not prove ingestion, and procedure codes can be incomplete. Differential misclassification is especially concerning if knowledge of prognosis or outcome affects classification.
Timing can turn classification into immortal-time bias. A time-fixed “ever treated” variable is often inappropriate when treatment begins during follow-up. New-user designs, aligned time zero, time-varying treatment definitions, cloning-censoring-weighting approaches, or sequential trial emulation may be needed, depending on the question.
Selection bias concerns who enters the study or analysis and whether inclusion depends on intervention and outcome-related factors. Requiring survival until a later landmark, including only people with post-baseline measurements, or conditioning on continued membership in a health system can distort comparison. The draft V2 explicitly places immortal-time questions in both classification and selection because the mechanism can arise in either location; what a reviewer has to identify is the actual pathway, rather than attach the label and move on.
Domain 4: bias due to missing data#
Missingness can affect intervention status, confounders, outcomes, or analytical inclusion. The percentage missing is not enough to tell you whether there is bias. The central question is whether missingness depends on the true value or prognosis in a way that differs across intervention groups.
Complete-case analysis is valid only under assumptions that are often implausible. Multiple imputation can improve efficiency and reduce bias when its model includes relevant predictors and assumptions are defensible, though it does not recover information that was never recorded or make data missing not at random become harmless. Examine the reasons for the missingness and its timing, the differences between groups, the outcome ascertainment after discontinuation, the variables in the imputation model, and the sensitivity analyses for departures from the assumptions.
Domain 5: bias in measurement of outcomes#
Outcome measurement can be biased when assessors know intervention status, definitions differ between groups, surveillance intensity differs, or instruments have unequal validity, and an intervention that increases clinical visits can make asymptomatic events more likely to be detected without changing their true incidence.
Objective is not synonymous with unbiased. Death records can be incomplete; laboratory thresholds can change; administrative algorithms can have different positive predictive values across settings. Blinded adjudication helps for judgment-dependent outcomes, but only if ascertainment and referral to adjudication are also comparable. Specify whether errors are likely related to intervention and whether they could move the estimate toward or away from the null, remembering that nondifferential misclassification does not always bias toward the null, especially with multiple treatment categories, imperfect specificity, or adjusted models.
Domain 6: bias in selection of the reported result#
A dataset can support many outcomes, definitions, time points, models, subgroup restrictions, lag periods, and missing-data methods. Reporting the most favorable estimate after seeing the results introduces selection bias even when every individual analysis, taken on its own, is technically defensible. That is the pattern you are looking for.
Compare publications with protocols, registrations, statistical-analysis plans, regulatory reports, and methods described before outcome analysis. Look for unexplained switching, multiple eligible estimates, selective omission, or a model chosen because it crossed a significance threshold.
The V2 draft includes an analysis-plan signaling question and revised algorithms. Absence of a public prespecified plan raises concern, but a dated internal plan or transparent presentation of all analyses can still be informative. Registration alone does not guarantee fidelity.
What changed from the original tool#
The original ROBINS-I evaluated seven domains: confounding; selection of participants; classification of interventions; deviations from intended interventions; missing data; measurement of outcomes; and selection of the reported result.
The draft V2 has six domains and no separate deviations-from-intended-intervention domain. It reorganizes intervention-effect questions through distinct confounding variants, revises signaling responses to strong and weak forms of yes and no, provides algorithms that suggest domain judgments, and adds triage routes for studies clearly at critical risk. Ratings produced with the two versions are therefore not directly interchangeable, and a review should not begin with the original tool, switch domain labels midway, and report the result as “ROBINS-I” without qualification.
Make and report the judgment#
Domain judgments in ROBINS-I range from low through moderate, serious, and critical risk, with no information available when evidence is insufficient. Overall judgment follows the most concerning domain rather than averaging domains. A critical problem cannot be cancelled by five reassuring domains.
Low risk is demanding because the comparison is a well-conducted randomized trial addressing the same question. “Moderate” does not mean mediocre study quality; it indicates some bias relative to that benchmark but no serious problem identified.
Record supporting quotes or data locations, rationale, likely direction of bias, and unresolved information. Duplicate independent assessment and consensus can reduce idiosyncratic judgment. Pilot the tool on a small sample so that you and your co-reviewer agree on the target trial and the confounders before you work through the full set.
Do not convert labels into points and sum them. A score assumes that unlike biases are commensurable and that small problems can offset a fatal one. Also keep ROBINS-I separate from GRADE or another certainty framework: risk of bias is one contributor to certainty, alongside inconsistency, indirectness, imprecision, publication bias, and other considerations.
References#
- ROBINS-I V2 draft resources and version status
- Original 2016 ROBINS-I article
- Cochrane Handbook Chapter 25
Questions and answers
Is ROBINS-I a checklist for all observational studies?
No. It was designed for nonrandomized studies of intervention effects. Other questions, such as prevalence, diagnosis, prognosis, or the effects of environmental factors, may require different tools.
Should a review use ROBINS-I V2 now?
V2 is publicly available but remains a draft as of July 15, 2026. Review teams should follow their protocol and funder or handbook requirements, name the exact version, train reviewers, and avoid changing versions after seeing results.
Can propensity-score matching produce a low-risk confounding judgment?
Not by itself. The judgment depends on whether important confounders were identified, measured well, and controlled appropriately, with adequate overlap and no damaging adjustment.
Is a study at serious risk of bias useless?
Not necessarily. Its estimate may still contribute with appropriate caveats, sensitivity analysis, or triangulation. Critical-risk results are generally too compromised to provide a useful effect estimate for the target question.
Does ROBINS-I determine the certainty of a systematic-review conclusion?
No. It assesses risk of bias in specific study results. A separate evidence-synthesis framework evaluates the broader body of evidence.