Evidence explainer

Evidence and research methods

Reading an Objective Response Rate in a Cancer Study

Objective response rate is the share of analyzed participants whose tumors shrank enough to count. It is an early signal of activity, and on its own it says nothing about survival.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Define the numerator and denominator
  2. RECIST 1.1 is common, not universal
  3. Best overall response depends on time
  4. Confirmation protects against noise
  5. Independent central review can reduce some bias
  6. Confidence intervals reveal information size
  7. Historical controls are fragile
  8. Response does not equal clinical benefit automatically
  9. Accelerated approval comes with unfinished evidence
  10. Immunotherapy needs special care
  11. A practical appraisal sequence

Objective response rate, commonly abbreviated ORR, is a tumor-based endpoint used in many early oncology trials. Under a stated response system, it is the percentage of participants whose best response is complete response or partial response, and a high rate can be compelling when spontaneous shrinkage is rare, but the number is interpretable only once you have its denominator, measurement rules, confirmation, duration, review process, and disease context.

Define the numerator and denominator#

The numerator usually contains participants with a best overall response of complete response or partial response. Stable disease is not part of ORR. “Disease control rate” often adds stable disease, but its meaning depends heavily on the minimum duration and natural history, and you should not substitute it for ORR.

The denominator may be all treated participants, all randomized participants, everyone meeting measurable-disease criteria, or an “evaluable” subset, and excluding people who stopped early, deteriorated, died, missed imaging, or had an uninterpretable scan can inflate the rate you are quoted. A rigorous report defines nonevaluable participants and states how they were counted.

Suppose 20 of 80 treated participants respond. ORR is 25%. If investigators exclude 10 people without follow-up imaging and report 20 of 70, it becomes 28.6%. The clinical data did not improve; the denominator changed.

RECIST 1.1 is common, not universal#

For many solid tumors, RECIST 1.1 identifies measurable target lesions at baseline, with no more than five total and two per organ: most target lesions use longest diameter; pathologic lymph nodes use short-axis measurement. Their measurements form a baseline sum.

A partial response generally requires at least a 30% decrease in the target-lesion sum relative to baseline. Progressive disease generally requires at least a 20% increase relative to the smallest sum recorded, plus an absolute increase of at least 5 mm, or a new lesion or unequivocal progression of non-target disease. Complete response requires disappearance of target lesions, with pathologic nodes reduced below the specified short-axis threshold, and resolution of non-target disease under the criteria.

Stable disease falls between response and progression. It can be meaningful for a rapidly progressing cancer, but it is sensitive to scan schedule and minimum duration. Brain tumors, lymphomas, prostate cancer, multiple myeloma, leukemia, mesothelioma, ovarian cancer, and immunotherapy studies may use disease-specific or modified criteria. Do not assume that “partial response” means RECIST in every paper you read.

Best overall response depends on time#

Best overall response is the best category recorded from treatment start until progression or another protocol-defined cutoff, while applying confirmation rules, and more frequent imaging gives more opportunities to detect a transient response and can identify progression sooner. Comparing ORR across trials with different schedules can therefore mislead.

Time to response indicates how quickly shrinkage first appears. Duration of response runs from first qualifying response until progression or death under the specified definition; median duration should include a confidence interval and a description of censoring and follow-up.

An ORR of 60% with a median response of two months can be less valuable than 35% with multi-year durable responses, depending on cancer and available therapy. If many responders remain in response, a median may be immature or not reached; report follow-up and landmark proportions rather than treating “not reached” as infinite.

Confirmation protects against noise#

Measurement variability, inflammation, necrosis, and scan timing can create apparent shrinkage. In single-arm trials where ORR is the primary endpoint, RECIST 1.1 generally requires confirmation of complete or partial response at a later assessment, commonly at least four weeks later under protocol rules.

Randomized trials may not require confirmation because the control group helps interpret measurement fluctuations, but protocols can still require it. Reports should distinguish confirmed from unconfirmed ORR.

Confirmation can introduce survivor bias if only participants who remain well enough for another scan qualify. That is appropriate to a durable-response definition but should be transparent. Early deaths and progression belong in the denominator.

Independent central review can reduce some bias#

Open-label investigators know treatment and clinical course. Independent blinded central review can apply imaging criteria consistently and is particularly important for a single-arm trial supporting a regulatory decision.

Central review does not make imaging objective in an absolute sense. Readers still choose target lesions, judge new lesions, handle missing scans, and distinguish progression from artifact. Discordance between investigators and central reviewers should be shown and investigated. Informative censoring can affect blinded review if scans stop after clinical deterioration. Central reviewers cannot assess images that were never obtained.

Confidence intervals reveal information size#

ORR is a binomial proportion and should have a confidence interval. A rate of 50% in 10 people is compatible with a wide range; 50% in 500 is much more precise. The method used for small samples matters.

Single-arm phase 2 designs often test whether a rate exceeds an uninteresting historical benchmark, and the benchmark, type I error, power, interim stopping rule, and planned sample size should be prespecified. Choosing a weaker benchmark after results are known makes a drug look more active. Multiplicity arises when investigators test several doses, biomarker groups, response definitions, or time points, so a striking subgroup rate based on a few people is a hypothesis unless the design controlled the claim.

Historical controls are fragile#

In a single-arm trial, every participant receives the investigational treatment; a response is less confounded than a survival comparison when untreated tumors virtually never shrink, but its magnitude still depends on who enrolled and how measurement occurred.

Modern molecular selection can identify a more responsive population. Better imaging, central review, supportive care, prior therapy, disease burden, and eligibility can differ from the historical study. If the older control used investigator assessment and the new study uses intensive central imaging, rates may not be comparable. Randomization remains the strongest way to distinguish treatment effect from prognosis and measurement. FDA guidance encourages randomized approaches where feasible, including in accelerated-development programs.

Response does not equal clinical benefit automatically#

Tumor shrinkage may relieve pain, obstruction, bleeding, or neurologic compromise, making response directly meaningful in some contexts. In others, shrinkage is a surrogate expected to predict later benefit.

ORR does not measure participants whose disease remains stable for a long time, delayed immunotherapy effects, toxicity, symptom burden, function, or survival. A treatment can improve ORR without improving overall survival if responses are brief, resistant clones emerge, later therapy differs, or toxicity offsets benefit. Conversely, an effective therapy may improve survival with little conventional shrinkage. Endocrine therapy, cytostatic agents, and immunotherapy can challenge simple size-based endpoints.

Accelerated approval comes with unfinished evidence#

FDA's accelerated approval pathway can rely on a surrogate or intermediate clinical endpoint reasonably likely to predict clinical benefit in a serious condition. Oncology approvals have often used ORR and duration of response, especially for large effects in biomarker-defined populations with unmet need.

Accelerated approval is approval under specific legal criteria, not a declaration that survival benefit is proven. Required confirmatory trials are intended to verify benefit. FDA's Project Confirm tracks oncology accelerated approvals, conversions, ongoing requirements, and withdrawals. Read whether the indication is accelerated or traditional, whether confirmatory trials were underway, and whether later results verified benefit. Evidence status can change after the original headline.

Immunotherapy needs special care#

Immune therapies can occasionally produce apparent progression from inflammation or delayed response. iRECIST introduced unconfirmed progressive disease followed by confirmation for clinical trials. It standardizes data collection rather than directing individual patient care.

Pseudoprogression is not common enough to assume every growing lesion is benign. Continuing therapy after apparent progression requires protocol criteria and clinical judgment. ORRs using RECIST 1.1 and iRECIST should be labeled separately.

A practical appraisal sequence#

Identify the disease, the line of therapy, the biomarker, the intervention, and the response criteria, then confirm the measurable disease and the analysis population. Recalculate the numerator and the denominator yourself, with the nonevaluable participants put back in. Record complete and partial responses separately, and inspect the confidence interval.

Check scan schedule, confirmation, independent review, missing imaging, duration of response, follow-up, and competing events. For a single-arm trial, evaluate the historical benchmark and natural response rate. Then integrate progression-free and overall survival, symptoms, quality of life, and adverse effects.

Finally, check regulatory status and confirmatory evidence. ORR is a valuable signal, not a complete benefit-risk assessment.

Sources and further reading

  1. U.S. Food and Drug Administration, Clinical Trial Endpoints for the Approval of Cancer Drugs and Biologics
  2. Eisenhauer and colleagues, Revised RECIST Guideline Version 1.1, European Journal of Cancer (2009)
  3. U.S. National Cancer Institute, RECIST 1.1 Guidelines
  4. U.S. Food and Drug Administration, Project Confirm
  5. Seymour and colleagues, iRECIST Guidelines for Immunotherapeutics Trials, Lancet Oncology (2017)

Questions and answers

Is stable disease counted in ORR?

No. ORR generally includes complete and partial responses. Stable disease may be included in a separate disease-control rate with a defined duration.

Does a 30% decrease mean the tumor is 30% smaller by volume?

No. RECIST uses selected linear dimensions. A 30% decrease in their sum is not the same as a 30% reduction in volume or total body tumor burden.

Why does response need confirmation?

Repeat imaging helps distinguish sustained shrinkage from measurement variation or a transient change, particularly when a single-arm trial relies on ORR.

Is a high ORR enough for full approval?

Sometimes response can support approval depending on context, magnitude, durability, disease, and benefit-risk. Accelerated approval based on a surrogate generally requires confirmatory evidence of clinical benefit.

Can two trials' ORRs be compared directly?

Only cautiously when population, line of therapy, criteria, imaging schedule, review method, denominator, and follow-up are comparable. A randomized head-to-head estimate is stronger.