Evidence explainer

Evidence and research methods

When a Single Hazard Ratio Hides a Time-Varying Effect

A hazard ratio is a tidy summary only while the event-rate ratio stays reasonably stable. When an effect arrives late, fades, or reverses, one average number can hide the pattern that matters.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Hazard is not the same as risk
  3. What the proportional-hazards assumption says
  4. What violations look like
  5. Formal diagnostic checks
  6. Better summaries when effects vary
  7. A reader's audit of a survival result

A hazard ratio can compress years of follow-up into one number, but that compression may be misleading when the treatment effect changes with time. The proportional-hazards assumption says the hazard ratio is constant over time. Before you accept one ratio as the whole result, look at the survival curves, at the diagnostic checks the authors ran, and at an absolute time-based measure.

Key points#

Hazard is not the same as risk#

Risk is the probability that an event occurs over a specified period. Hazard is the instantaneous rate at which events occur at a particular time among those still event-free. A hazard ratio of 0.75 compares those instantaneous rates between groups under the model, but it does not mean that every participant's risk falls by 25 percent, nor does it directly state how many events were prevented.

The distinction becomes especially important because the group still under observation changes over time. People who experience the event leave the risk set. Competing events, loss to follow-up, and treatment changes can further alter who remains. A hazard ratio is therefore a conditional, time-indexed comparison rather than a simple probability ratio.

What the proportional-hazards assumption says#

In a Cox proportional-hazards model, the baseline hazard can rise and fall freely over time. The model assumes that a covariate multiplies that baseline hazard by a constant factor. If the treatment hazard is 0.7 times the control hazard near the start, it is assumed to remain 0.7 times the control hazard later.

This can be a useful approximation even when biology is not perfectly constant. No model is exact. The concern is material departure: a delayed treatment benefit, early procedural harm followed by later benefit, waning effectiveness, or an effect that reverses direction.

In those settings, the reported hazard ratio is a weighted average of different time periods. Its value can depend on follow-up length and censoring patterns. Two studies of the same treatment could report different average hazard ratios simply because one stopped before the late benefit emerged.

What violations look like#

Kaplan-Meier curves offer the first clue, although they display cumulative event-free probability rather than the instantaneous hazard itself.

Late separation#

The curves overlap for months and then diverge. This may happen when a treatment needs time to work. An average hazard ratio blends the early period of little effect with the later period of benefit.

Early harm, later benefit#

A procedure can carry an immediate risk followed by durable protection. The curves may first favor control and later favor intervention. A single ratio can obscure both clinically important phases.

Waning effect#

The curves separate early and then become parallel or converge. The average may imply a stable benefit that no longer exists late in follow-up.

Crossing curves#

The preferred group changes over time. A hazard ratio near 1 can arise from averaging meaningful early harm with meaningful later benefit. “No effect” would be a poor description.

Curves can be noisy when few participants remain, so do not overread a late wiggle. The risk table under the plot, the censoring, the event counts, and the confidence bands are what tell you whether the pattern is real or rests on a handful of remaining participants.

Formal diagnostic checks#

Schoenfeld residuals are commonly used to evaluate whether a covariate's effect varies systematically with time; a global test can assess the model as a whole, while covariate-specific tests examine individual terms. Time-by-treatment interactions and plots of scaled residuals provide more detail.

These tests have limits. A small trial may have little power to detect nonproportionality. A very large trial may flag a minor departure that barely changes interpretation. Testing after inspecting the curves can also invite selective modeling, and a paper that tells you only that “the test was nonsignificant” has not established that a constant ratio is scientifically adequate.

The 2020 JAMA commentary by Stensrud and Hernán went further, arguing that routine testing does not solve the deeper interpretive problems of hazard ratios, and their later work emphasized that proportional hazards is often implausible and not required for many clinically interpretable alternatives. Whether or not a reader accepts that broader critique, it reinforces a practical point: the estimand should be chosen for the clinical question, not merely because Cox software is familiar.

Better summaries when effects vary#

Restricted mean survival time#

Restricted mean survival time is the area under the survival curve up to a prespecified time. It can be read as the average event-free time accumulated during that window; the difference between groups is expressed in days, months, or years rather than as a relative instantaneous rate.

For example, a 45-day difference in event-free survival through 2 years answers a concrete question, though the time horizon must be chosen before viewing results, supported by adequate follow-up in both groups, and interpreted for the specific endpoint.

Milestone probabilities#

Reporting survival or cumulative incidence at 6 months, 1 year, and 3 years shows how absolute differences develop. These values should include intervals and numbers at risk. Multiple milestones should not be selected after the fact because one looks favorable.

Piecewise or time-varying effects#

Researchers can estimate separate hazard ratios for prespecified periods or model a treatment effect that changes smoothly with time. This makes the pattern visible but introduces additional choices and can become unstable when events are sparse. The time cut points and model form need justification.

Cumulative incidence with competing risks#

When another event prevents the outcome of interest, such as death before disease recurrence, ordinary survival methods may answer a different question from the one you have in mind. Cumulative-incidence functions and competing-risk estimands may be needed. Nonproportionality and competing risks are separate issues, but both can make a lone cause-specific hazard ratio incomplete.

A reader's audit of a survival result#

Before you repeat the headline hazard ratio, ask:

  1. What exact event and time origin were used?
  2. Is the effect measure a cause-specific hazard ratio, subdistribution hazard ratio, risk ratio, or something else?
  3. Were the survival or cumulative-incidence curves shown with numbers at risk?
  4. Do the curves separate immediately, late, temporarily, or in opposite directions?
  5. Did the protocol prespecify the Cox model and checks of proportionality?
  6. What diagnostic test and graphical assessment were reported?
  7. Were alternative summaries such as restricted mean survival time or milestone risk differences provided?
  8. Does differential censoring or treatment switching complicate the comparison?
  9. Is the follow-up window clinically meaningful and similar across groups?
  10. Does the conclusion describe the time pattern, or only the average ratio?

If the assumption appears reasonable and the ratio answers the chosen estimand, the Cox model can be useful. If the effect changes, the answer is not necessarily to discard the trial. It is to report the effect in a form that respects time.

Sources and further reading

  1. JAMA, why test for proportional hazards
  2. American Journal of Epidemiology, why use methods that require proportional hazards
  3. Journal of Clinical Oncology, moving beyond the hazard ratio
  4. BMC Medical Research Methodology, restricted mean survival time as an alternative

Questions and answers

Do crossing Kaplan-Meier curves automatically invalidate a Cox model?

They are a strong warning that a constant treatment hazard ratio may be a poor summary. The full judgment also considers uncertainty, event counts, censoring, prespecified analyses, and alternative models.

Is a nonsignificant proportional-hazards test proof that the assumption holds?

No. The test may lack power, and a minor statistical departure may not matter clinically. Curves, residual plots, design knowledge, and sensitivity analyses should be read together.

Is restricted mean survival time always superior?

No. It answers a different, time-bounded question and depends on a meaningful prespecified horizon with sufficient follow-up. Its advantage is direct interpretation without requiring proportional hazards.