Evidence explainer

Evidence and research methods

QUADAS-2 and QUADAS-3: How to Judge Bias in Diagnostic Accuracy Studies

QUADAS-2 organizes diagnostic-study bias around selection, the index test, the reference standard, and flow and timing. QUADAS-3 replaced it in February 2026.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Start with a precise review question
  2. Domain 1: patient selection
  3. Domain 2: the index test
  4. Domain 3: the reference standard
  5. Domain 4: flow and timing
  6. Turn signaling questions into reasoned judgments
  7. Do not add the domains into a score
  8. What QUADAS-3 changed
  9. QUADAS and STARD do different jobs
  10. A reader's rapid diagnostic-study audit

QUADAS-2 became the standard framework for judging risk of bias and applicability in diagnostic accuracy studies, and it asks whether patient selection, conduct of the index test, the reference standard, and participant flow could make sensitivity or specificity misleading. It does not produce a quality score, certify a test, or repair incomplete reporting.

As of February 2026, QUADAS-3 is the current version recommended by the tool's developers; the four-domain logic of QUADAS-2 remains useful for reading older systematic reviews, but new reviews should use the current tool. QUADAS-3 shifts judgments from a whole-study label to individual accuracy estimates, defines the synthesis question more explicitly, and compares studies with an ideal test accuracy trial.

Start with a precise review question#

Diagnostic accuracy is conditional. A test can perform differently in symptomatic and screened populations, primary and specialist care, early and advanced disease, or with different thresholds, so the review should define participants, index test, target condition, reference standard, and intended role in the pathway.

The intended role can be replacement, triage, or add-on. A triage test may prioritize sensitivity to avoid missed disease, while a confirmatory test may prioritize specificity to prevent false positives, and pooling studies across roles can create an average that fits no decision.

Thresholds also matter. Sensitivity and specificity move together as a threshold changes, and a study evaluating a prespecified cutoff is different from one that selects the cutoff giving the most favorable result in the same sample. QUADAS judgments make sense only against this defined question. “Applicable” does not mean broadly realistic. It means sufficiently aligned with the population, test conduct, and target condition that the estimate can inform the planned synthesis.

Domain 1: patient selection#

The first domain asks whether the enrolled sample could distort accuracy. Consecutive or randomly selected eligible patients usually resemble a clinical stream better than a convenience sample. Inappropriate exclusions can remove difficult-to-diagnose cases and make the test look cleaner.

A classic case-control accuracy design enrolls obvious disease cases and healthy controls. The groups differ in more than disease status: they differ in severity, symptoms, comorbidity, and setting, so the test ends up distinguishing extremes rather than the borderline patients who create uncertainty in practice. This spectrum effect often inflates both sensitivity and specificity.

Useful questions include:

Applicability asks whether the enrolled population matches the review. A low-bias study in a tertiary referral clinic may still have high applicability concern for community screening.

Domain 2: the index test#

The index test is the test under evaluation. Its result should be interpreted without knowledge of the reference-standard result. If readers know the final diagnosis, ambiguous index findings may be unconsciously pushed toward agreement.

Thresholds should be specified before results are analyzed. Choosing a cutoff that maximizes sensitivity and specificity in the same dataset creates optimism. A test with a continuous result can support several prespecified thresholds, but each estimate needs clear labeling.

The report should describe specimen handling, device version, reader training, number of readers, adjudication, image quality rules, and treatment of indeterminate results. Excluding uninterpretable tests can overstate real-world performance if failures are clinically consequential. Applicability can be compromised when a laboratory assay, scanner, reader procedure, or threshold differs materially from the version intended for use, because a brand name alone may conceal software and hardware changes.

Domain 3: the reference standard#

The reference standard is the method used to decide whether the target condition is present. It should classify the condition accurately enough that disagreements are not simply reference errors.

Some conditions have a strong pathological standard. Others rely on expert adjudication, follow-up, imaging, or a composite. Composite standards can be appropriate, but their components and decision rules need to be prespecified. A reference that includes the index test creates incorporation bias because the test helps define its own correctness.

Reference interpretation should be masked to the index test where interpretation is subjective; if the reference result is objective and automated, lack of masking may matter less, but the reason should be stated rather than assumed.

Applicability asks whether the reference defines the same target condition as the review. A study of any cancer is not automatically applicable to a review of clinically significant cancer. Different disease definitions can change apparent false positives and false negatives.

Domain 4: flow and timing#

Every participant should ideally receive both the index test and the same appropriate reference standard within an interval short enough that disease status is unlikely to change, and departures create several biases.

Partial verification occurs when only some participants receive the reference, often those with a positive index test. False negatives then remain undiscovered and sensitivity looks better. Differential verification occurs when positive tests receive an invasive standard while negative tests receive follow-up or a weaker standard: accuracy may change because the standard changes with the index result.

Long delays allow disease progression, treatment, or recovery between tests. Excluding participants with missing, indeterminate, or discordant results can also produce a favorable complete-case sample. Reconstruct the flow diagram yourself: how many were eligible, tested, verified, excluded, and analyzed at each threshold. Percentages without denominators make that impossible.

Turn signaling questions into reasoned judgments#

QUADAS-2 signaling questions are usually answered yes, no, or unclear, with “yes” framed to suggest low concern. They lead to domain judgments of low, high, or unclear risk of bias. The final rating requires judgment; it is not determined by counting answers.

Review teams should tailor signaling guidance before assessing studies. For example, the acceptable test-to-reference interval depends on the target condition. Hours may matter in acute infection, while months may be reasonable for a stable genetic condition. Tailoring after seeing study results risks inconsistent standards.

Two reviewers commonly assess each study and resolve differences through discussion or a third reviewer. Agreement statistics can be reported, but high agreement does not guarantee correct judgments. The rationale for each rating tells you more than a colored icon alone. An unclear rating should mean that information needed for judgment is missing, not that the study is probably acceptable, and contacting authors, checking protocols, and comparing related publications can resolve some of that uncertainty.

Do not add the domains into a score#

A total such as “three out of four domains passed” gives false arithmetic. Biases differ in direction and importance. Partial verification can severely inflate sensitivity, while a minor timing concern may have little effect. Two studies with the same total can have very different credibility.

Systematic reviews should display domain-specific judgments and examine whether pooled estimates change when high-risk studies are excluded or modeled separately, and sensitivity analysis is useful, but removing studies after seeing favorable results can itself introduce selection. The approach should be prespecified.

What QUADAS-3 changed#

The University of Bristol released QUADAS-3 as the current recommended tool in 2026. Several changes address recurring problems with QUADAS-2:

Estimate-level assessment matters because one study can report several thresholds, populations, readers, or reference standards. One estimate may be low risk while another from the same study is not. A single whole-study label loses that structure.

The ideal-test concept asks reviewers to specify how a strong study would answer the synthesis question, then judge departures; this makes hidden assumptions visible and helps separate a study's reporting detail from its actual fit. Older QUADAS-2 assessments do not become worthless. They should be interpreted as products of the earlier tool, and updates should explain whether reassessment with QUADAS-3 could change conclusions.

QUADAS and STARD do different jobs#

STARD is a reporting guideline for diagnostic accuracy studies. It helps authors report participant flow, test methods, thresholds, cross-tables, and uncertainty. QUADAS is a risk-of-bias and applicability tool used mainly in evidence synthesis.

A thoroughly reported study can show that its design is biased. A poorly reported study can be methodologically sound but impossible for you to judge. STARD adherence should not be converted into a low-risk QUADAS rating, and a QUADAS judgment should not substitute for a reporting assessment.

A reader's rapid diagnostic-study audit#

  1. Does the population match the intended testing point?
  2. Were participants enrolled consecutively or randomly without inappropriate exclusions?
  3. Was the index threshold prespecified and interpreted without reference knowledge?
  4. Can the reference standard classify the exact target condition?
  5. Was reference interpretation protected from index-test knowledge?
  6. Did all participants receive the same reference within a suitable interval?
  7. Are indeterminate, missing, and excluded results shown?
  8. Are sensitivity and specificity reported as paired estimates with confidence intervals?
  9. Does the review use the current QUADAS-3 tool for new assessments?
  10. Are applicability and risk of bias kept separate?

Sources and further reading

  1. Whiting and colleagues, QUADAS-2, Annals of Internal Medicine (2011)
  2. University of Bristol, QUADAS-3 Current Tool and Resources
  3. Whiting and colleagues, QUADAS-3, Annals of Internal Medicine (2026)
  4. Cochrane, Introducing QUADAS-3 (2026)
  5. EQUATOR Network, STARD 2015 Reporting Guideline

Questions and answers

Is QUADAS-2 still the recommended tool?

No. As of 2026, the developers recommend QUADAS-3 for new systematic reviews. QUADAS-2 remains important for understanding older reviews and its four-domain framework.

Can QUADAS determine whether a test is clinically useful?

No. It evaluates bias and applicability of accuracy estimates. Clinical utility also depends on prevalence, consequences, downstream management, harms, cost, and patient outcomes.

Is “unclear” between low and high risk?

No. It means reporting is insufficient for a justified judgment. It should not be averaged as a middle score.

Does perfect reporting mean low risk of bias?

No. Complete reporting can make a serious design problem easy to identify. Transparency and validity are related but distinct.