A biomarker is a defined characteristic measured as an indicator of normal biology, disease, or a response to an intervention. Blood pressure, troponin, and viral RNA can all serve as biomarkers in particular contexts. So can a gene variant, tumor imaging, and a wearable-derived signal.
The broad definition explains why usefulness cannot be inferred from the category. A measurement can be analytically precise but clinically irrelevant. It can predict an outcome yet add nothing to information already available. It can separate groups statistically while producing no safe action. It can respond to a drug without mediating benefit.
The right question is not “Is this a good biomarker?” It is “Is this measurement fit for this exact use, in this population, at this time, with this decision and consequence?”
Begin with the context of use#
Context of use is the complete statement of what a biomarker will do. It names the disease or biological state, population, and specimen. It names timing, measurement method, role, and decision. It may also specify a threshold and consequence.
“Use protein X in cancer” is not a context. “Use plasma protein X measured before treatment to enrich a phase 2 trial for adults with metastatic tumor Y who are more likely to progress within six months” is closer.
The same marker can be useful for one role and harmful for another, and a marker that helps select participants at higher event risk might shorten a trial, yet be too inaccurate for screening healthy people. A marker that monitors treatment may not diagnose disease. A prognostic marker does not automatically predict treatment response. Evidence travels only as far as the context supports. Changing specimen, assay, threshold, population, or decision can require new validation.
Analytical validity: can the measurement be trusted?#
Analytical validity concerns the test process. Precision asks whether repeated measurements agree. Accuracy or trueness asks whether results match an accepted reference. Other features include limit of detection, linearity, and specificity. They include interference, stability, carryover, and reportable range.
Preanalytical conditions can dominate. Time from collection to processing, tube type, and temperature may change a result. So may fasting, posture, exercise, and freeze-thaw cycles. Imaging depends on acquisition and reconstruction. Wearables depend on placement, skin contact, movement, and algorithms.
A test can perform well in the developer's laboratory and poorly across routine sites. Multisite reproducibility, lot variation, instrument calibration, quality control, and operator training matter. Analytical validity is necessary. A beautifully calibrated assay is still useless if it measures something unrelated to the intended question.
Biological and clinical validity#
Biological validity asks whether the marker reflects the process claimed. Clinical validity asks how the result relates to a clinical state or outcome in the target population.
Association is a starting point. Researchers should examine temporality, dose-response patterns, and mechanisms. They should examine confounding, reverse causation, and consistency. For diagnosis, sensitivity and specificity depend on threshold and reference standard. For prognosis, risk estimates need calibration across time.
Validation should occur in data not used to discover the marker or choose its cutoff. Reusing the development sample creates optimism. Cross-validation helps during development but does not replace external validation in different sites, times, and patient groups. Spectrum matters. A marker that separates advanced disease from healthy volunteers may fail in the real clinical population where early disease competes with similar conditions.
Clinical utility is the action test#
Clinical utility asks whether using the biomarker improves decisions and outcomes relative to current practice. This is harder than showing prediction.
A marker can be accurate yet nonactionable. If all patients should receive the same care, additional classification adds cost without changing management, and if the action triggered by a positive result is ineffective or harmful, better prediction can worsen outcomes.
Utility includes benefits, false-positive procedures, and missed disease. It includes anxiety, treatment toxicity, and delay. It includes cost, feasibility, and equity. It also includes whether clinicians and patients can understand and use the result, and randomized biomarker-strategy trials can compare care guided by the marker with care guided by the existing pathway. Other designs may be appropriate, but they must preserve the link between information, action, and outcome.
Biomarker categories describe roles#
The FDA-NIH BEST resource distinguishes several categories. Diagnostic biomarkers detect or confirm a disease or identify a subtype. Monitoring biomarkers are measured over time. Pharmacodynamic or response biomarkers change in response to an intervention. Predictive biomarkers identify people more likely to experience a favorable or unfavorable treatment effect.
Prognostic biomarkers estimate the likelihood of an event regardless of treatment. Susceptibility or risk biomarkers indicate potential to develop a condition. Safety biomarkers signal toxicity. Each role carries a different design and comparison.
The prognostic-predictive distinction is often blurred. A marker can identify poor prognosis in every treatment group without indicating who benefits more from one therapy, and prediction of treatment effect requires an interaction, or other credible evidence that relative benefit differs by marker status. Naming the category tells you what is being claimed. It does not tell you whether the claim has been validated.
Discrimination is not enough#
For a binary outcome, discrimination measures how well a model or marker ranks people who experience the event above those who do not. Area under the receiver operating characteristic curve, or AUC, is common.
A high AUC can coexist with poor calibration, meaning predicted probabilities do not match observed frequencies. It can also produce unacceptable false positives at the threshold needed for practice, and a modest AUC improvement may matter if it changes a critical decision, while a statistically significant improvement may be practically irrelevant.
Calibration should be assessed overall and across risk ranges, sites, times, and relevant groups. Sensitivity, specificity, predictive values, and likelihood ratios should be reported at prespecified thresholds. Predictive values depend on prevalence. A test can perform well in a referral center and produce many false positives in low-prevalence screening.
Incremental value over existing information#
A new biomarker should be compared with the actual clinical baseline, not with no information. Age, symptoms, and examination may already capture much of its signal. So may routine laboratories, imaging, and validated risk scores.
Adding a marker can improve fit without changing decisions. Investigators can examine changes in calibration, discrimination, and classification. Net reclassification statistics need careful interpretation because results depend on risk categories and can look favorable without clear clinical benefit.
Decision-curve analysis weighs true and false positives across threshold probabilities, translating performance into net benefit under stated assumptions. It is useful, not definitive. Thresholds must correspond to real actions and plausible values. The simplest question remains powerful: how many people change categories, what happens to them, and is that change beneficial?
Thresholds convert numbers into consequences#
Most biomarkers are continuous. A threshold divides a smooth distribution into labels such as positive and negative; the choice trades sensitivity against specificity and embeds the relative cost of missed disease and false alarms.
One threshold may be suitable for ruling out a dangerous condition, where sensitivity is prioritized. Another may be needed to rule in disease before an invasive procedure. An intermediate zone may appropriately trigger a second test rather than a binary verdict.
Data-driven cutoff selection in a small sample overfits. Thresholds should be prespecified, validated, and tied to a care pathway, and values near the line carry uncertainty and biological variation. Showing the continuous result, its reference context, and that uncertainty can be more honest than presenting the cutoff as a natural boundary.
Surrogate endpoints face a higher bar#
A surrogate endpoint substitutes for a clinical outcome such as survival, symptoms, or function in a trial. Surrogates can shorten studies and reduce sample requirements. They can also mislead.
A marker can be prognostic without capturing the causal pathway through which treatment affects the outcome. A drug may improve the marker but cause offsetting harm elsewhere. Conversely, a treatment may help through a pathway the marker does not measure.
Validation therefore requires treatment-level evidence across trials and, often, intervention classes, and the central question is whether the effect of treatment on the surrogate reliably predicts its effect on the clinical outcome in the proposed context. FDA tables distinguish validated surrogate endpoints, reasonably likely surrogates used in accelerated approval, and candidate endpoints. The labels carry different certainty and postapproval obligations.
Qualification is specific, not universal#
FDA's Biomarker Qualification Program can accept a biomarker for a specified context of use in drug development, and once qualified, it can be used under that context across CDER development programs without reconfirming its utility for every individual product.
Qualification is not approval of a clinical test. FDA explicitly notes that qualification of a biomarker does not imply that the measurement device is cleared or approved for patient care. Conversely, authorization of a test device does not automatically qualify the biomarker for a drug-development use.
Qualification packages address the need, context, and benefits and risks. They address analytical performance, evidence, and remaining gaps. The evidence bar rises with the consequence of error. Ask “qualified for what?” before taking qualification as a general badge.
Study design should mimic intended use#
Diagnostic studies should enroll consecutive or appropriately sampled people in whom the test would actually be considered; the reference standard should be valid and applied without knowledge of the index result where possible. Verification bias occurs when only selected participants receive the reference.
Prognostic studies need a defined starting point, representative cohort, adequate follow-up, blinded outcome assessment, and proper handling of censoring and competing risks; predictive biomarkers need randomized treatment comparisons or strong causal designs.
Specimens collected for one purpose may not reflect routine processing. Retrospective analyses of trial biobanks can be rigorous when protocols, missingness, and assay plans are prespecified, but convenience samples invite selection. Whatever the source, the validation data have to look like the places and people the marker will meet in use, including the uncommon conditions that matter most when they appear.
Bias and equity can enter at every stage#
Who gets tested, whose specimen is adequate, how the reference diagnosis is made, and who receives follow-up can all create bias. A marker may appear less accurate in a group because disease definition, care access, or specimen handling differs.
Biological differences are possible, but broad demographic categories are poor substitutes for measured causal factors. Claims should avoid turning group averages into fixed individual biology.
Utility also depends on access. A marker that directs patients to a specialist who is not available, or a therapy they cannot afford, may widen disparities, and false positives can burden groups already facing diagnostic delay or mistrust. So look for performance estimated in each relevant subgroup, with enough data to say how uncertain it is, and then for an investigation of why it differs rather than a parity claim.
Monitoring after adoption#
Performance can drift when disease prevalence, instruments, laboratory lots, referral patterns, or treatments change. A threshold validated before a new therapy may no longer have the same predictive value afterward.
Quality systems should track invalid samples, turnaround, and calibration. They should track distribution shifts, false results, and downstream actions. They should track adverse events and outcomes. Algorithms that combine biomarkers need version control and change governance.
Postmarket monitoring is especially important when implementation differs from trials. Clinicians may order a test in populations never studied or repeat it at unsupported intervals. Stopping rules are part of responsible adoption. A marker should be revised or withdrawn when evidence shows poor performance, no utility, or net harm.
A practical appraisal sequence#
Write the context of use in one sentence. Name the decision you make today and what the biomarker would change about it. Check the assay, the specimen, the units, the reproducibility, and the preanalytical requirements.
Then find independent validation in the population you care about. Look past AUC to calibration, thresholds, predictive values, missing data, and subgroup uncertainty. Compare it with the information you already have, and quantify how many people change category.
Map every possible result to an action. Estimate the benefits, harms, costs, burden, and equity of each. Where the stakes justify it, look for direct evidence that the biomarker-guided strategy improved outcomes. Last, get the regulatory status right and plan the monitoring. A promising association can be worth researching long before it is ready for routine use.
The shortest definition of useful#
A useful biomarker measures reliably, means what it claims in a defined context, and improves a real decision. Remove any link and the chain fails.
This standard is demanding, because biomarkers can move people toward treatment, reassurance, or invasive testing. They can move people toward trial eligibility or a drug approval. Precision in the laboratory deserves equal precision in the claim.
Sources#
The metadata sources include the FDA-NIH BEST definitions, FDA and ICH qualification frameworks, National Academies evaluation, and foundational surrogate-endpoint methods.
Sources and further reading
Questions and answers
Is a biomarker the same as a clinical outcome?
No. A biomarker is a measured characteristic of biology or response. A clinical outcome describes how a person feels, functions, or survives, although a validated surrogate may substitute in a defined trial context.
Does a statistically significant association make a biomarker useful?
No. Association can support biological validity, but usefulness also requires reliable measurement, adequate discrimination and calibration, incremental information, an actionable pathway, and evidence that using it helps.
What is a biomarker's context of use?
It is the precise purpose and setting for interpretation, including population, disease state, role, decision, specimen, timing, and conditions. Evidence for one context does not automatically transfer to another.
Is an FDA-qualified biomarker approved for routine patient testing?
Not automatically. Qualification supports a defined use in drug development, while authorization of a test device and adoption in patient care are separate questions with separate evidence.
How should a new biomarker claim be evaluated?
Define the intended decision, inspect assay performance, validate in representative independent data, compare with current information, assess thresholds and harms, and test whether biomarker-guided action improves meaningful outcomes.