Sensitivity and specificity sound as if they are properties that a diagnostic test carries everywhere. In a study, they are calculated by comparing an index test with a reference standard that classifies who truly has the target condition.
The calculation assumes the reference classification is correct. Often it is not.
Culture can miss infection after antibiotics. Histopathology depends on sampling and interpretation. Imaging can miss small lesions. Clinical diagnosis can change with follow-up. No single interview perfectly classifies a mental-health condition. If the benchmark makes errors, the new test's apparent errors are a mixture of index-test mistakes and reference mistakes.
That is the imperfect reference standard problem.
The two-test table hides an assumption#
A classic diagnostic accuracy study gives participants both:
- the index test being evaluated;
- the reference standard used to define the target condition.
The cross-tabulation has four cells:
- true positive: both say disease is present;
- false negative: the reference says disease, the index says no disease;
- false positive: the reference says no disease, the index says disease;
- true negative: both say no disease.
Sensitivity is true positives divided by everyone classified as diseased by the reference. Specificity is true negatives divided by everyone classified as not diseased by the reference. The words “true” and “false” belong to the reference standard's classification. They are not direct observations of an inaccessible biological truth.
A gold standard and a reference standard are not synonyms#
STARD defines a clinical reference standard as the best available method for establishing the target condition. A gold standard would classify it without error.
The distinction matters because the phrase “gold standard” can stop you thinking: a biopsy, a culture, an expert diagnosis, or an older assay is easy to read as error-free merely because it is established.
A reference can fail through:
- sampling error, such as a biopsy missing the lesion;
- biological timing, such as pathogen burden below detection early in illness;
- analytical error or an unsuitable cutoff;
- reader variability;
- an incomplete disease definition;
- treatment between the index test and reference procedure;
- inability to observe the target condition directly.
The error may differ across disease severity, patient subgroups, sites, specimen types, or operators.
How reference error changes apparent accuracy#
Consider an index molecular test compared with culture for an infection. Culture is highly specific but misses some true infections. The molecular test detects one of those missed cases.
Against biological truth, the molecular result is a true positive. Against culture, it is recorded as a false positive. The new test's estimated specificity falls even though the disagreement may reflect greater sensitivity.
Now consider a reference standard that sometimes labels non-disease as disease. If the index test correctly returns negative in one of those people, the two-by-two table calls it a false negative. Estimated sensitivity can fall.
The exact bias is not fixed. It depends on:
- the reference standard's sensitivity and specificity;
- disease prevalence and spectrum;
- the index test's accuracy;
- whether the tests make errors in the same people;
- how indeterminate results are handled;
- which participants receive verification.
Agreement with an imperfect reference is not the same as accuracy against the target condition.
Conditional dependence is the hidden complication#
Many corrections imagine that index and reference errors are unrelated once true disease status is held constant. That assumption often fails.
Two nucleic-acid tests may miss the same low-load specimens. Two readers may use the same imaging feature. A composite standard may include a test that measures the same biomarker as the index test. Shared biology and shared interpretation create conditional dependence.
If both tests make the same errors, their agreement can exaggerate apparent accuracy. If they make complementary errors, agreement can look worse than either test's true performance. So a model that assumes the two are conditionally independent should say why that holds here, and show what changes if it does not.
Partial verification bias#
Sometimes only participants with a positive or suspicious index test receive the definitive reference procedure, because a biopsy may be too invasive for everyone, or clinicians may judge it unnecessary after a negative screen.
This creates partial verification, also called workup bias. The chance of learning the reference result depends on the index test. Negative participants who actually have disease can disappear from the accuracy table, commonly inflating sensitivity and altering specificity.
The solution is not simply to exclude unverified participants. Exclusion preserves the selected sample. Better options include a design in which all participants receive the reference, random verification of a subset with appropriate weighting, or prespecified statistical correction with defensible missingness assumptions. Whichever route a study takes, you can find who went unverified, why, and what their index results and characteristics were.
Differential verification bias#
In differential verification, every participant may receive some reference, but not the same one. Positive index tests might receive biopsy while negative tests receive clinical follow-up.
This can be reasonable when one reference is invasive, though it can also bias accuracy if the standards differ in their ability to detect disease, and the verification pathway itself may depend on the index result and clinical risk. A team that has thought about it reports results separately by index-reference combination, explains the assignment rule, and shows how plausible differences in reference accuracy move the estimates.
Incorporation bias#
Incorporation bias occurs when the index test forms part of the reference definition.
Suppose an expert panel diagnoses a syndrome using symptoms, imaging, laboratory data, and the index biomarker. Agreement is built into the outcome. Sensitivity and specificity can be inflated because the test helps define the truth against which it is judged.
The cleanest prevention is to exclude the index result from the reference classification. If clinical reality requires the result, investigators can create a research adjudication without it, report both definitions, or study patient outcomes rather than claim circular accuracy.
Review bias and access to clinical information#
Reference assessors who know the index result may reinterpret borderline evidence to agree. Index-test readers who know the reference result can do the reverse. STARD therefore asks whether each reader had access to the other's result and to clinical information.
Complete blinding is not always clinically appropriate. A radiologist may normally need history. The study should prespecify which information reflects intended use and which information would create circularity. Reproducibility also matters. If the reference depends on expert interpretation, report training, criteria, number of readers, disagreement resolution, and inter-reader agreement.
Timing can create disagreement without either test being wrong#
Disease can emerge, resolve, or change between tests. Treatment may lower pathogen load. A lesion may grow. An immune response may appear later.
If the index test and reference are separated by too much time, the disagreement you see can reflect a real change in target status, and the acceptable interval depends on the disease and tests. STARD asks for the time interval and any interventions between them. “Same patient” does not guarantee “same disease state.”
Use the best available reference carefully#
When a credible reference exists but is imperfect, a strong design can still reduce avoidable error:
- define the target condition and intended use precisely;
- use the same reference in all participants where feasible;
- apply it without knowledge of the index result;
- train readers and measure agreement;
- collect specimens at clinically appropriate times;
- resolve indeterminate reference results by a prespecified rule;
- report reference limitations and expected direction of bias;
- perform sensitivity analyses using plausible reference accuracy.
The reference standard should match the claim. Histology may establish tissue abnormality but not whether using a test improves health. Clinical follow-up may establish outcome but be too late for an early diagnostic-use claim.
Composite reference standards#
A composite combines several sources by a fixed rule. Disease might be present if culture or PCR is positive, or if two of three clinical criteria are met.
Composites can increase sensitivity when no single test captures every case. They also create new problems:
- an “OR” rule may accumulate false positives;
- an “AND” rule may miss cases;
- components may be dependent;
- missing components require another rule;
- different studies use different composites;
- the index test can enter the composite and create incorporation bias.
Adding more imperfect tests does not necessarily approach truth. A 2018 BMJ analysis showed that composite performance can worsen and vary with prevalence; the study should justify each component and rule, estimate component performance where possible, and compare alternative definitions.
Expert panels#
An adjudication panel can combine history, examination, imaging, laboratory results, and follow-up. This is useful when diagnosis is inherently clinical.
Panels are not error-free. They may disagree, use implicit criteria, or be influenced by index results. Stronger panels use:
- a written target-condition definition;
- standardized case packets;
- index-result masking where possible;
- multiple readers;
- an explicit disagreement process;
- reported agreement and indeterminate rates;
- sensitivity analysis across panel definitions.
Consensus can reduce random disagreement while preserving shared systematic error.
Clinical follow-up and deferred verification#
Follow-up can reveal whether disease declares itself, resolves, or remains absent. It is common when immediate definitive testing is invasive or unavailable.
Follow-up must be sufficiently long and complete. Treatment started because of the index result can prevent the outcome used to verify disease, creating a paradox. Loss to follow-up may depend on symptoms and risk, and outcome assessors can be influenced by the original result. None of that is resolvable after the fact, so a prespecified algorithm has to say in advance what events count, how treatment is handled, and how uncertain cases are classified.
Latent class models#
Latent class analysis treats true disease as an unobserved class and uses patterns across multiple imperfect tests to estimate prevalence and test accuracy.
This can be valuable when no test deserves error-free status. It does not discover truth without assumptions. Models may require:
- a defined number of latent classes;
- conditional independence or an explicit dependence structure;
- enough tests, populations, or prior information for identification;
- sensible prior distributions in Bayesian analyses;
- stable performance across modeled groups.
A systematic review of 64 latent-class diagnostic studies found wide variation in model specification and incomplete reporting of fit and assumptions. Model fit, identifiability, priors, alternative dependence structures, and sensitivity analyses should all be shown. Latent classes may also reflect statistical patterns that do not map neatly to a clinical disease definition.
Bounds and sensitivity analyses#
When you do not know how accurate the reference is, one corrected estimate can create false precision. A paper can instead give a plausible range for the reference's sensitivity and specificity and show the range of index accuracy that follows from it.
Partial-identification methods report bounds compatible with stated assumptions. These intervals may be wide, but width is honest when the data cannot isolate the truth, and the analysis that reports them should separate what the observed data taught it from what its assumptions and outside evidence supplied.
Sometimes accuracy is the wrong target#
If no credible reference classification exists, asking whether the new test improves decisions or outcomes may be more informative.
A randomized test-impact study can assign diagnostic strategies and measure time to correct treatment, unnecessary procedures, complications, symptoms, or health outcomes, and such a study answers whether using the test helps, not its pure sensitivity and specificity.
Test-impact trials have their own challenges. Clinicians may not follow results, downstream treatment may vary, and large samples may be needed. They can nevertheless bypass an unresolvable truth label when the actual decision is whether to adopt a testing pathway.
How to read an accuracy study#
Ask:
- What exactly is the target condition?
- Why was this reference standard chosen?
- What errors can the reference make, and in which patients?
- Did every participant receive the same reference?
- Were index and reference readers masked appropriately?
- Was the index test part of the reference definition?
- What was the interval, and did treatment occur between tests?
- How were indeterminate and missing results handled?
- Were alternative reference assumptions tested?
- Does the study population match intended use?
QUADAS-2 organizes related concerns into patient selection, index test, reference standard, and flow and timing. STARD makes visible the reporting you need in order to answer them.
The honest-reference conclusion#
Diagnostic accuracy is always accuracy relative to a target definition and a method for observing it. When the reference makes errors, a two-by-two table can mislabel correct index results and reward shared mistakes.
The solution is not to pretend the benchmark is gold. Define its limitations, protect its interpretation, verify participants consistently, disclose timing and missingness, and analyze plausible alternative truths. When truth cannot be classified credibly, the most useful question may change from “Does the new test agree?” to “Does using this testing strategy improve decisions and outcomes?”
References#
- Deeks JJ, Bossuyt PM, Leeflang MM, Takwoingi Y, editors. Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy. Version 2.0, 2023.
- Cohen JF, et al. STARD 2015 explanation and elaboration. BMJ Open. 2016.
- Whiting PF, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Annals of Internal Medicine. 2011.
- van Smeden M, et al. Latent class models in diagnostic studies without a reference standard. American Journal of Epidemiology. 2014.
- Collins J, Huynh M. Estimation of diagnostic test accuracy without full verification. 2014.
- Naaktgeboren CA, et al. Concerns about composite reference standards. BMJ. 2018.
Questions and answers
What is an imperfect gold standard?
It is better called an imperfect reference standard: the best available method for classifying the target condition, but one that still produces false positives or false negatives.
Can a new test be more accurate than the reference standard?
Yes. A new test may detect true cases the reference misses. Conventional agreement analysis can then count those results as false positives unless another design or model addresses reference error.
Does combining several tests create a perfect reference?
No. An explicit composite may improve classification, but its rule can accumulate component errors and dependence. It needs its own justification and sensitivity analysis.
What is verification bias?
It occurs when receipt or type of reference verification depends on the index result or related clinical factors. The verified subset then differs systematically from everyone tested.
Can latent class analysis solve the problem?
It can estimate accuracy without declaring one test perfect, but results depend on class structure, dependence, identifiability, priors, and model fit. It is an assumption-based tool, not an automatic cure.