Evidence explainer

Evidence and research methods

Why Overall Survival Is the Gold Standard in Cancer Trials

Living longer is the outcome that needs the least translation in an oncology trial. It is also slow to measure, and crossover, later treatment, and missing follow-up can distort it.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What overall survival actually measures
  2. Why the endpoint is so persuasive
  3. Why overall survival can take so long
  4. Crossover changes the question
  5. Subsequent treatment can dilute a difference
  6. Progression-free survival answers an earlier question
  7. Response rate is narrower still
  8. Symptoms and function belong in the endpoint picture
  9. The hazard ratio is not a duration
  10. A nonsignificant OS result is not always reassurance
  11. How to read an oncology survival claim

An oncology trial can measure a tumor on a scan within months. It may take years to learn whether the same treatment helps people live longer. That timing difference creates a persistent tension: the most direct endpoint often arrives last.

Overall survival, usually abbreviated OS, is the time from randomization until death from any cause. It needs no translation. It is often called the gold standard because death is unambiguous, survival matters directly to patients, and the outcome captures both benefit and fatal harm. No statistical surrogate has to translate an OS benefit into clinical meaning.

The phrase can still be misunderstood. FDA does not require OS to be the primary endpoint in every cancer trial, and in some diseases, waiting for mature survival results would be impractical or could delay access to a treatment with a large effect. Progression-free survival, durable tumor response, symptoms, function, and quality of life can all be important. The sound question is not whether one endpoint always wins. It is whether the chosen endpoint answers the claim and whether the total evidence shows a favorable balance of benefit and risk.

What overall survival actually measures#

For a randomized trial, the clock usually starts at random assignment, not at diagnosis, first dose, or documented response. The event is death from any cause. Participants who are still alive at the analysis cutoff are censored at the last date they are known to be alive.

All-cause mortality avoids the uncertain task of assigning one cause to every death. A person with advanced cancer may die from tumor progression, infection, a cardiovascular event, treatment toxicity, or several interacting causes. Classifying one as the definitive cause can invite disagreement and bias. Counting every death gives a reproducible outcome.

Randomization is essential to interpretation. A longer median survival in a treated group is persuasive only if the comparison group began with a similar prognosis and was followed under a comparable protocol. A single-arm survival curve gives you no counterfactual at all. Historical controls may differ in diagnostic methods, supportive care, disease classification, and access to later treatment.

Why the endpoint is so persuasive#

OS is patient-centered without relying on an assumption about what a laboratory value or image means, and a confirmed survival improvement shows that, within the trial's conditions, the treatment strategy delayed death on average relative to the comparator.

The endpoint is also resistant to some common measurement biases. A scan can be read differently across observers. Tumor assessments occur at scheduled intervals, and a missed visit can move a recorded progression date. Knowledge of treatment assignment may influence when a clinician orders an unscheduled scan. Death dates are generally less subjective, provided vital status is collected completely.

OS integrates competing effects. A treatment might control cancer better but cause fatal toxicity. Another may have modest antitumor activity yet improve supportive care enough to extend life. The survival endpoint records the net result. That is one reason FDA's August 2025 draft focuses on OS as an important safety endpoint even when another measure is primary.

Why overall survival can take so long#

A survival analysis needs enough deaths, not merely enough enrolled participants. If a cancer has a favorable prognosis, the necessary events may take years, and treatments received after progression can extend life in both groups and reduce the apparent difference caused by the randomized treatment.

Large benefits can sometimes be detected with fewer events, while smaller differences need more. Event timing, expected dropout, allocation ratio, interim analyses, and the desired precision all affect sample size. A trial with immature OS may show a hazard ratio whose confidence interval includes meaningful benefit and meaningful harm. Calling it neutral at that point is stronger than the data allow.

Long follow-up creates operational demands. Sites close, participants move, and consent for later contact may change; a trial needs durable methods for confirming survival status, and you should check that it did not follow one group more carefully than the other.

Crossover changes the question#

In many cancer trials, participants assigned to control may receive the investigational treatment after progression. Crossover can be ethically attractive when early evidence is strong and other options are limited. It also makes the groups more similar after progression.

An intention-to-treat analysis still compares assignment to the original strategies. If many control participants later receive the study drug, the estimate may understate the biological effect of receiving it, and statistical adjustment methods rely on assumptions that may not hold, especially when switching depends on prognosis.

This is not a flaw that can be erased by choosing a favorable model. The protocol should define crossover, follow participants afterward, and present the main randomized estimate alongside justified sensitivity analyses. Ask whether the claim concerns the initial treatment strategy, the treatment received, or access to a whole sequence of therapies.

Subsequent treatment can dilute a difference#

Cancer care often changes during a multiyear trial. New therapies become available, and their use may differ by country, insurance system, site, or study group. If effective later treatment is used more often after one assigned therapy, OS reflects the whole pathway rather than the first drug alone.

That feature is clinically realistic but complicates attribution. Progression-free survival may isolate the period before later treatment more clearly. OS may better represent the consequence of beginning with one strategy in the care system that existed during the trial. Both can be informative if the estimand is explicit. Reports should set out the post-progression treatments by randomized group, not merely state that later therapy was allowed, because large imbalances deserve explanation and can decide whether the results transfer to a setting with different access.

Progression-free survival answers an earlier question#

Progression-free survival, or PFS, measures time until radiographic or clinical progression or death, according to prespecified rules; it often produces more events earlier than OS and is less affected by therapies given after progression.

PFS has its own vulnerabilities. Scan schedules create interval censoring because progression happens between observations. Missing assessments, informative dropout, inconsistent imaging, and unblinded review can shift the recorded date. Central review may reduce reader variation but does not repair a biased schedule or missing scans.

A PFS benefit can matter by delaying disease worsening, treatment change, or symptoms. It does not always predict longer life. Surrogacy is empirical and context-specific. Evidence from one cancer, disease stage, or treatment class cannot automatically validate PFS for another.

Response rate is narrower still#

Objective response rate is the proportion of participants whose tumors shrink by a prespecified amount. Duration of response asks how long that shrinkage lasts. These measures can show strong antitumor activity quickly, including in a single-arm study when spontaneous major responses are rare.

Response does not count stable disease as success, and tumor shrinkage may not improve symptoms or survival; a high response rate with brief duration can be less valuable than a smaller but durable benefit. The assessment criteria, confirmation rules, evaluator masking, and denominator all matter.

FDA's accelerated approval pathway can use an endpoint reasonably likely to predict clinical benefit for serious conditions with unmet need. Approval is contingent on required studies that verify benefit. Project Confirm makes the status and outcomes of oncology accelerated approvals easier to inspect.

Symptoms and function belong in the endpoint picture#

Living longer is not the only outcome that matters to the person being treated. Pain, breathlessness, fatigue, physical function, ability to work, treatment burden, cognition, and time outside a hospital can distinguish two strategies even when OS is similar.

Patient-reported outcomes need validated instruments, prespecified timing, adequate completion, and methods for missing data. Attrition is especially concerning when the sickest participants stop completing questionnaires. A favorable score among remaining respondents may not describe everyone randomized. Read the quality-of-life data alongside the survival data rather than as decoration: a therapy that extends life with severe persistent toxicity presents a different tradeoff from one that extends both life and time without symptoms.

The hazard ratio is not a duration#

Cancer trial reports often summarize OS with a hazard ratio. It compares instantaneous event rates over follow-up under model assumptions. A hazard ratio of 0.75 does not mean every participant lives 25 percent longer, and it does not directly state the absolute number of months added.

Kaplan-Meier curves show how survival probabilities separate over time. Medians describe when each curve crosses 50 percent but ignore the rest of the distribution. Landmark survival at clinically chosen times and restricted mean survival time can add interpretable absolute information.

Nonproportional hazards are common when curves separate late, cross, or show a long-term surviving subgroup. In that setting, one hazard ratio can obscure changing effects. The analysis plan should anticipate plausible patterns and present complementary measures.

A nonsignificant OS result is not always reassurance#

If a trial was powered for PFS, its OS analysis may be descriptive or underpowered. Wide confidence intervals can remain compatible with clinically important mortality harm. Multiplicity rules may also limit formal claims after testing several endpoints.

The FDA's August 2025 oncology document was draft guidance at the research date, and it recommends planning OS collection, follow-up, analyses, and safety monitoring when OS is not the primary efficacy endpoint. It also discusses adequate baseline risk characterization and continued follow-up after treatment discontinuation. The principle is simple: an earlier efficacy signal should not end observation of a later vital outcome, and safety review needs the whole randomized population with balanced ascertainment.

How to read an oncology survival claim#

Start with the population and the comparator. Disease stage, biomarker status, prior treatments, performance status, and available care determine baseline prognosis, so confirm which endpoint was primary and whether the testing hierarchy protected you against a false-positive claim.

Check the number of deaths, median follow-up, confidence interval, and survival curves. Look for crossover, missing vital status, imbalanced withdrawal, and post-progression treatment. Ask whether the result is mature enough for the conclusion.

Finally, combine survival with serious adverse events, patient-reported outcomes, and treatment burden before you settle on a view. OS is powerful because it is direct. Good oncology interpretation remains multidimensional because survival alone cannot describe how that time was lived.

Sources and further reading

  1. FDA final guidance on clinical trial endpoints for approval of cancer drugs and biologics, December 2018
  2. FDA draft guidance on assessment of overall survival in oncology trials, August 2025
  3. FDA Oncology Center of Excellence Project Endpoint
  4. FDA Accelerated Approval Program
  5. FDA Oncology Center of Excellence Project Confirm
  6. ICH E9 R1 guideline on estimands and sensitivity analysis

Questions and answers

Does every cancer approval require an overall survival benefit?

No. Depending on the setting, FDA may consider progression-free survival, response and response duration, symptoms, or a surrogate under accelerated approval. The endpoint must support the proposed claim, and the total benefit-risk evidence still matters.

What is the difference between overall survival and cancer-specific survival?

Overall survival counts death from any cause. Cancer-specific survival requires assigning a cause of death, which can be uncertain and open to classification bias.

Can a drug improve progression-free survival without improving overall survival?

Yes. Later therapies, crossover, limited follow-up, treatment toxicity, or weak surrogacy can produce that pattern. It does not make PFS meaningless, but it limits what can be concluded about longevity.

Why continue survival follow-up after a participant stops treatment?

Stopping treatment can relate to toxicity or worsening disease. Losing those participants can bias the comparison and conceal delayed benefit or harm.

Is median survival the amount of time a person should expect to live?

No. It is the time when the estimated survival curve reaches 50 percent in a study group. Individual prognosis varies, and the median does not describe the entire curve.