Evidence explainer

Evidence and research methods

How to Read a Responder Analysis

A responder analysis can make a trial result easier to picture, but one threshold erases differences within both groups. Read it beside the full outcome distribution.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What a responder percentage answers
  2. Dichotomizing throws away distance
  3. The threshold needs a clinical argument
  4. Prespecification protects the inference
  5. Missing outcomes can change the answer
  6. Baseline change scores need care
  7. Relative measures can make effects look larger
  8. When responder analysis adds real value
  9. A practical appraisal sequence
  10. References

A responder analysis takes a measured outcome, often a symptom, function, laboratory, or weight-change scale, and classifies each participant as a responder or nonresponder. The rule might be “at least a 30% reduction in pain” or “an improvement of 10 points or more.” The result is a proportion you can picture.

The price of that simplicity is lost information. Someone who improves by 9.9 points and someone who worsens by 20 points receive the same nonresponder label under a 10-point rule. Someone who improves by 10.1 points and someone who improves by 40 points are both responders. The threshold preserves one distinction and discards all the others.

What a responder percentage answers#

Suppose a randomized trial measures pain on a 0-to-100 scale and defines response as at least a 20-point improvement at week 12. If 54% of the treatment group and 39% of the comparison group respond, the absolute difference is 15 percentage points, and the corresponding number needed to treat is about seven, if the endpoint, follow-up, and causal assumptions are sound.

That statement answers a threshold question: how did assignment affect the probability of crossing 20 points? It does not say that the average treatment effect was 20 points, that every responder benefited because of treatment, or that nonresponders received no benefit. Some people would have crossed the threshold under either assignment, and some classified as nonresponders still improved, though the binary result can be valuable when the threshold represents a real decision or a well-supported patient-important state. Remission, freedom from rescue treatment, or achievement of a functional goal may be intrinsically categorical. The concern is greatest when an essentially continuous outcome is split only to make the result sound intuitive.

Dichotomizing throws away distance#

Continuous analysis uses how far each observation lies from the others. Binary analysis retains only which side of one line it falls on; the discarded distance contains statistical information, so the binary comparison generally has a wider confidence interval or needs more participants to attain similar power.

The cost is not merely a technical loss. A cutoff can conceal very different patterns. Two treatments may produce the same responder proportion even if one shifts the entire distribution modestly and the other helps a smaller group dramatically. Conversely, a small shift in many participants can create a large difference in responder rates when the cutoff sits near the dense middle of the distributions.

Snapinn and Jiang showed that power depends strongly on the chosen threshold and on the shape and variability of the outcome. There is no fixed conversion between a mean difference and a responder-rate difference. The entire distribution matters.

The threshold needs a clinical argument#

A credible responder definition names the amount of change, direction, time window, measurement instrument, baseline requirements, and handling of intercurrent events, and it should be justified in the target population using patient input and suitable anchor-based evidence where possible.

An external anchor asks whether score change tracks a separate, interpretable judgment or event, such as a clearly understood global assessment of change. The anchor itself must be reliable and sufficiently related to the construct. Distribution-based quantities such as a standard deviation fraction or standard error of measurement can describe scale and noise, but they do not tell you what people consider important.

Different thresholds answer different questions. A minimal detectable change asks whether change exceeds measurement error, a meaningful within-person change asks whether one person's change is important, and a between-group mean difference asks about the average causal effect. Treating these as interchangeable creates a false sense of precision. Thresholds may also vary by baseline severity, direction of change, duration, and population, so a 10-point improvement may represent a different experience for someone starting at 25 than for someone starting at 90.

Prespecification protects the inference#

Trying several cutoffs and reporting the most favorable one is a form of outcome selection. Even if every cutoff sounds plausible after the fact, the resulting confidence interval and p-value no longer have their usual interpretation unless multiplicity is addressed.

Look for the threshold in the protocol, statistical analysis plan, or trial registration. Check whether the time point, analysis population, covariate adjustment, and missing-data rule were also prespecified. If several responder endpoints were tested, ask how they fit the endpoint hierarchy.

A useful sensitivity analysis shows responder differences across a range of credible cutoffs. An empirical cumulative distribution function displays the proportion of participants at every observed level. It shows whether one chosen threshold represents a broad separation or an isolated crossing.

Missing outcomes can change the answer#

A participant who discontinues treatment, starts rescue therapy, dies, or misses the visit does not arrive with an automatic responder label. Classifying every missing value as nonresponse is conservative in some situations and biased in others. Last observation carried forward can also create implausible outcomes.

ICH E9(R1) asks trialists to define the estimand, the precise treatment-effect question, which includes the population, outcome variable, summary measure, and strategy for events such as treatment discontinuation or rescue medication. A responder endpoint should therefore state whether it addresses outcomes regardless of stopping treatment, outcomes while treatment is used, a composite that counts an event as failure, or another clearly defined question.

Missing-data assumptions should align with that question and be challenged in sensitivity analyses. Compare withdrawal and follow-up across groups. A clean responder percentage can conceal substantial unobserved data.

Baseline change scores need care#

Many responder rules use change from baseline. Measurement error at baseline can create regression to the mean: participants selected during an unusually bad measurement may improve on repeat testing without an intervention effect. Randomization balances that tendency on average, but threshold classification can still be sensitive to baseline noise.

The analysis should account appropriately for baseline score. An analysis of covariance on the continuous follow-up outcome is often more precise than comparing raw change scores. For responder analyses, modeling the binary outcome with baseline adjustment may help, but the method and estimand should be prespecified.

Floor and ceiling effects matter too. Someone close to the best possible score may be unable to improve by the required amount. A single absolute-change threshold can structurally exclude people who have less room to change.

Relative measures can make effects look larger#

Responder reports often present a relative risk or odds ratio. Those measures need the raw proportions. A rise from 2% to 4% is a doubling, but the absolute difference is 2 percentage points. A rise from 40% to 55% has a smaller relative ratio but a 15-point absolute difference.

Odds ratios are not risk ratios and can look much farther from one when response is common. Report the event counts, percentages, absolute difference with confidence interval, and a relative measure if useful. A number needed to treat should include its time point and uncertainty.

When responder analysis adds real value#

Responder analysis is strongest as a prespecified complement to the continuous result. It translates a well-validated threshold into a frequency while the continuous analysis preserves information. Agreement across mean change, distribution plots, and credible thresholds makes your interpretation more stable. It can also clarify heterogeneity when the instrument has a real state transition, such as remission, and when the binary endpoint maps to care. It is weaker when the cutoff is arbitrary, data-derived, one of many tested, or detached from how the measure was validated.

A practical appraisal sequence#

Identify the original scale and who completed it. Find the exact responder rule and its justification, and confirm that it was set before anyone saw unblinded results. Then read the raw counts, the absolute difference, the confidence interval, and the missing-data handling, and set all of it beside the continuous estimate and the full distribution.

Ask one final question: would a slightly different defensible threshold tell the same story? If the conclusion depends on moving a line by one point, the result deserves caution.

References#

  1. Snapinn and Jiang, responder analyses and clinically relevant effects
  2. Altman and Royston, the cost of dichotomising continuous variables
  3. Collister and colleagues, continuous and responder analyses for patient-reported outcomes
  4. FDA guidance on patient-reported outcome measures
  5. ICH E9(R1) estimands and sensitivity analysis

Questions and answers

Is a responder analysis always inappropriate?

No. It can be clinically useful when the threshold is meaningful and prespecified. It is usually most informative beside, rather than instead of, the continuous analysis.

Is a minimal important difference automatically a responder threshold?

No. A group-level difference and meaningful within-person change are different quantities. The proposed threshold needs evidence for individual interpretation in the target population.

Why can a small mean difference produce a large responder-rate difference?

If many observations cluster near the cutoff, shifting the distribution slightly can move many people across it. The effect depends on the distribution and threshold.

Should missing participants be counted as nonresponders?

Not automatically. The rule must match the trial's estimand and plausible reasons for missingness. Sensitivity analyses should test the assumptions.

What graph is most helpful beside a responder percentage?

An empirical cumulative distribution or another full-distribution display shows results over many thresholds. A histogram, density plot, or box plot can add context, provided the group sizes and uncertainty are clear.