Clinical trials often want an answer years before enough strokes, fractures, disability, or deaths occur to measure direct benefit. A surrogate endpoint offers an earlier or easier signal, such as a laboratory value, imaging measure, or physical sign, and the shortcut is scientifically valuable only if changing that signal predicts what treatment will do to the outcome you actually care about.
A biomarker is not automatically a surrogate#
A biomarker is a measurable characteristic of a biological process or response. It may help diagnose disease, estimate prognosis, select treatment, monitor response, or detect toxicity. Those uses do not by themselves make it a valid substitute for a clinical outcome.
To serve as a surrogate, the endpoint must stand in for how a person feels, functions, or survives in a particular treatment evaluation. The key inference runs from "the intervention changed the marker" to "the intervention will change the clinical outcome." That inference requires evidence beyond the marker's association with prognosis.
High blood pressure can identify people at greater stroke risk, but the surrogate question is more specific: across relevant interventions, does treatment-induced change in blood pressure predict treatment-induced change in stroke? FDA describes blood-pressure reduction as a validated surrogate in defined contexts because extensive trial evidence supports that link. The same logic cannot be transferred automatically to every marker.
Why correlation can fail#
A marker and an outcome may move together because both reflect underlying disease severity. Changing the marker need not change the disease. The marker could also lie on only one of several causal pathways.
The classic warning involved ventricular arrhythmias after myocardial infarction. Certain drugs suppressed an abnormal rhythm associated with sudden death but increased mortality. The marker moved in the desired direction while off-target or parallel effects harmed patients. The trial showed why biological plausibility and prognostic correlation cannot substitute for treatment-effect evidence.
Three failure patterns recur:
- The surrogate is correlated with outcome but is not on the causal pathway.
- The treatment changes the surrogate but reaches the clinical outcome through additional beneficial or harmful pathways.
- The surrogate captures one component of disease while treatment changes another component that matters more to patients.
Safety outcomes create a further limit. Even a valid efficacy surrogate cannot capture an unforeseen serious harm. Benefit-risk assessment still requires direct safety evidence.
The Prentice framework#
Prentice formalized a demanding definition: a treatment comparison using the surrogate should provide a valid test of the treatment effect on the true endpoint. Operational criteria include an effect of treatment on the surrogate, an effect on the true outcome, association between surrogate and outcome, and full capture of the treatment's outcome effect by the surrogate.
Full capture means that once the surrogate is known, treatment assignment adds no information about the clinical outcome. This is difficult to establish. A nonsignificant residual treatment effect does not prove it is zero, especially in a study too small for the clinical outcome, though the framework remains conceptually important because it centers the right question. Its strict statistical implementation in one trial is rarely enough, which led validation toward evidence across trials.
Individual-level and trial-level association#
Meta-analytic validation separates two levels.
Individual-level association asks whether participants with more favorable surrogate values tend to have better clinical outcomes within trials. This is prognostic information. It can be strong even when a treatment that changes the surrogate does not improve the outcome.
Trial-level association asks whether trials producing larger treatment effects on the surrogate also produce larger treatment effects on the clinical outcome. This directly addresses prediction of treatment benefit. Researchers estimate how well the surrogate effect forecasts the clinical effect across multiple randomized comparisons and quantify prediction uncertainty.
Both levels are informative, but trial-level evidence is the harder and more relevant test for replacing a clinical endpoint, and it requires several trials with variation in treatments and effect sizes. A near-perfect fit from only a few similar trials may be unstable, and it may not transport to the new mechanism.
Validation needs a context of use#
Calling a biomarker "validated" without qualification is too broad. The relationship may depend on:
- disease stage and patient population;
- treatment class or mechanism;
- clinical outcome and time horizon;
- direction and magnitude of surrogate change;
- measurement method and schedule;
- background therapy and care setting.
A surrogate validated for one drug class may fail for another if the newer treatment has additional pathways. A short-term marker may predict a near-term event but not long-term function. A threshold useful in adults may not transfer to children. Regulatory qualification and an endpoint's use in a specific application also differ. FDA's public table shows endpoints that supported approvals and specifies disease, population, mechanism, and whether the use supported traditional or accelerated approval, but presence in the table is not permission to use the marker for every product.
Three regulatory categories#
FDA resources describe three broad levels of support.
A candidate surrogate endpoint is still being evaluated for its ability to predict clinical benefit.
A reasonably likely surrogate endpoint has strong mechanistic or epidemiologic support but insufficient clinical data for full validation. In serious conditions with unmet need, such an endpoint may support accelerated approval when the statutory and evidentiary requirements are met.
A validated surrogate endpoint has a clear mechanistic rationale and clinical data providing strong evidence that treatment effects on it predict a specific clinical benefit in the stated context; it may support traditional approval in that context. The three labels reflect different amounts of uncertainty: "reasonably likely" should not be rewritten as "proven," and "candidate" should not be presented as an established replacement for a patient-relevant outcome.
Accelerated approval borrows against later evidence#
FDA's accelerated approval pathway can allow earlier access for drugs addressing serious conditions and unmet need based on a surrogate or intermediate clinical endpoint reasonably likely to predict benefit, and the evidence of effect still must come from adequate and well-controlled studies.
Clinical benefit then has to be verified and described through confirmatory studies. Current law and FDA policy provide authorities concerning the timing and completion of these trials. If benefit is verified, the indication may convert to traditional approval. If trials fail to verify sufficient benefit or are not completed with due diligence, FDA can pursue changes or withdrawal through applicable procedures. The structure makes the uncertainty explicit: accelerated approval is a regulatory decision that earlier access is justified despite residual uncertainty, not a declaration that the surrogate is fully validated.
How to read a surrogate-based trial#
Start with the endpoint's status in the exact context you are reading about. Ask:
- Is it a candidate, reasonably likely, or validated surrogate?
- Which patient-relevant outcome is it meant to predict?
- Does evidence include treatment-effect associations across several randomized trials?
- Were treatments in validation studies similar to the new intervention's mechanism?
- How wide is the prediction interval for clinical benefit?
- Could off-target effects bypass the surrogate?
- What magnitude and duration of change were observed?
- Is direct evidence on symptoms, function, survival, and harms available?
- If approval is accelerated, what confirmatory study is required and what is its status?
A statistically persuasive change in an unvalidated marker remains evidence that the marker moved. It does not automatically become evidence that patients benefited.
A current example of context-specific qualification#
In December 2025, FDA announced qualification of percentage change in total hip bone mineral density at 24 months as a validated surrogate endpoint for phase 3 trials of investigational therapies in postmenopausal women with osteoporosis at risk for fracture. The statement is deliberately specific about measure, timing, population, and use.
That example shows what validation looks like in practice. It does not say every bone-density measure at every time point predicts fractures for every population and mechanism. The context is part of the endpoint.
Sources and further reading
- U.S. Food and Drug Administration, surrogate endpoint resources for drug and biologic development
- U.S. Food and Drug Administration, table of surrogate endpoints used for approval or licensure
- Prentice, definition and operational criteria for surrogate endpoints, Statistics in Medicine (1989)
- Fleming and DeMets, why surrogate endpoints can mislead, Annals of Internal Medicine (1996)
- Buyse and colleagues, meta-analytic validation of surrogate endpoints, Biostatistics (2000)
- U.S. Food and Drug Administration, accelerated approval program
Questions and answers
If a biomarker predicts prognosis, is it a valid surrogate?
Not necessarily. Prognostic association concerns differences among people. Surrogacy requires evidence that treatment-induced changes in the marker predict treatment-induced changes in the clinical outcome.
Does FDA's surrogate table mean every listed endpoint is fully validated?
No. The table includes context and the type of approval supported. Some endpoints are reasonably likely surrogates used for accelerated approval, while others have support for traditional approval.
Why are confirmatory trials needed after accelerated approval?
The initial decision may rely on an endpoint reasonably likely, but not yet proven, to predict clinical benefit. Confirmatory studies test whether the anticipated patient benefit actually occurs.