Evidence explainer

Evidence and research methods

Hazard Ratios: What They Measure and Why Absolute Risk Still Matters

A hazard ratio compresses a whole time-to-event comparison into one relative rate. What it cannot give you is the absolute benefit.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Start with risk and rate as different quantities
  2. A numerical example separates the claims
  3. What the Cox model estimates
  4. Proportional hazards is a strong summary assumption
  5. Delayed effects are common in medicine
  6. Risk sets become selected over time
  7. Noncollapsibility complicates adjusted comparisons
  8. Censoring requires information assumptions
  9. Competing risks need event-specific methods
  10. Composite endpoints can hide component patterns
  11. Confidence intervals describe precision, not all uncertainty
  12. Kaplan-Meier curves need their own audit
  13. Restricted mean survival time is an interpretable alternative
  14. Observational hazard ratios need time-zero discipline
  15. A reporting set that supports decisions
  16. References

Time-to-event studies ask not only whether an event occurred, but when. Participants enter follow-up, experience the event, reach the study end, or become unavailable. A hazard ratio is one way to compare those evolving event processes, and it is widely reported because the Cox model can use incomplete follow-up efficiently and adjust for covariates without specifying the entire baseline hazard.

Its convenience has encouraged a misleading translation: “hazard ratio 0.70 means a 30 percent reduction in risk.” The first half may describe the modeled relative hazard. The second half sounds like 30 percent fewer events, which does not follow. If you want the clinical magnitude, the ratio needs time-specific absolute outcomes beside it.

Start with risk and rate as different quantities#

Risk is a probability over an interval. A five-year risk of 20 percent means that 20 of 100 comparable people would be expected to experience the event by five years under the stated conditions. It has a time horizon and a denominator defined at baseline.

The hazard at a moment is the instantaneous event rate among people who remain at risk immediately before that moment. It can rise, fall, or change shape over follow-up, and a hazard is not constrained between zero and one because it is a rate, though the probability over a very short interval is related to that rate.

The hazard ratio divides the hazard in one group by the hazard in another, so a ratio below one means the instantaneous rate is lower in the numerator group at that moment. Under a proportional-hazards model, one coefficient summarizes that ratio as constant over time.

A numerical example separates the claims#

Imagine a trial with a five-year event risk of 10 percent in control and 7 percent with treatment. The absolute risk reduction is 3 percentage points. The risk ratio is 0.70, and 33 people would need treatment for five years to prevent one event under a simple reciprocal calculation.

Another trial could have risks of 50 percent and 40 percent. Its absolute reduction is 10 points and risk ratio 0.80. Depending on when events occur, both trials might report a hazard ratio around 0.70. The identical hazard ratio would not mean identical benefit.

Conversely, a hazard ratio of 0.70 does not guarantee a five-year risk ratio of 0.70. The mathematical link depends on the baseline hazard and proportionality. Clinical interpretation therefore needs risks at prespecified times.

What the Cox model estimates#

David Cox's proportional-hazards model expresses each person's hazard as an unspecified baseline hazard multiplied by factors associated with treatment and covariates. The partial-likelihood method compares covariates among participants in the risk set whenever an event occurs.

Not specifying the baseline hazard is a strength for estimation, but it means the coefficient is a relative conditional parameter. To obtain survival probabilities, the baseline process must also be estimated. A table that gives you only regression coefficients has left out the starting event rate that determines absolute impact.

Covariate-adjusted hazard ratios compare participants with the same modeled covariate values. In a randomized trial, adjustment can improve precision when prespecified. In an observational study, it attempts to control measured confounding but does not repair unmeasured confounding, immortal-time bias, or inappropriate time zero.

Proportional hazards is a strong summary assumption#

The standard Cox coefficient is easiest to describe when the hazard ratio is constant across follow-up. The underlying hazards can both change; their ratio is assumed not to. If treatment has an early harm and late benefit, one number averages effects that point in opposite directions.

Kaplan-Meier curves can show gross nonproportionality. Curves may cross, separate after a delay, or separate early and later converge. Formal tests using Schoenfeld residuals can detect time trends, but they should not replace visual and clinical reasoning: a large study can flag a trivial deviation, while a small study may miss an important one. Stensrud and colleagues argue that testing proportionality as a gatekeeper can distract from the estimand: what has to be decided is whether a constant relative hazard is scientifically meaningful, with time-varying effects or alternative summaries reported when it is not.

Delayed effects are common in medicine#

Vaccines require time to generate immunity. Surgery can create early procedural risk followed by later benefit, cancer immunotherapy may produce delayed responses, and screening can shift diagnosis earlier before mortality changes at all; each of those creates a hazard pattern that one proportional ratio may compress poorly.

A landmark analysis can compare outcomes after a prespecified time among people event-free at that point, but it changes the target population and can introduce selection. Piecewise hazard ratios can describe early and late periods if chosen before seeing the curves. Flexible time-varying models offer another route. What a report should not do is pick the interval that looks favorable after inspecting the results; prespecified absolute risks and restricted mean survival time often stay understandable across complex patterns.

Risk sets become selected over time#

At baseline, randomization balances treatment groups on average. After follow-up begins, the people who remain event-free can differ. If treatment prevents early events among frailer participants, those participants remain in the treated risk set while similar control participants have already experienced the event and left it.

The hazard ratio at a later time compares these evolving survivors, not the original randomized populations. Their prognosis can differ because of treatment's earlier effects and individual susceptibility. Hernán called attention to this conditioning when discussing the causal interpretation of hazard ratios.

This does not make hazard models useless. It means the coefficient should not automatically be read as the treatment's effect on each person's instantaneous risk. Marginal survival curves preserve the original group assignment more directly for an intention-to-treat comparison.

Noncollapsibility complicates adjusted comparisons#

Hazard ratios are noncollapsible. Even without confounding, a covariate-adjusted conditional hazard ratio can differ from an unadjusted marginal hazard ratio. The change does not by itself show that adjustment removed bias.

This property makes comparisons across models difficult. A paper may present a treatment hazard ratio that moves after adding a strong prognostic variable, even if that variable was balanced. Readers should distinguish a change in estimand from evidence of confounding. Standardized survival curves or risks can translate a model into marginal outcomes for a target population, so an author should state whether an estimate is conditional or marginal, rather than leave you comparing coefficient size across differently adjusted models as if the scale were fixed.

Censoring requires information assumptions#

Right censoring occurs when follow-up ends before the event. Administrative censoring at a common study end is often straightforward. Loss to follow-up can be informative when it depends on prognosis, treatment effects, or care access.

Conventional survival estimates assume censoring is independent of future event time, unconditionally or after accounting for measured variables. If sicker participants leave one group more often, the remaining curve can look better. Reasons for loss and numbers by group should be reported. Inverse-probability-of-censoring weights can address measured predictors under assumptions, though extreme weights and unmeasured causes remain concerns, and treating a competing event as ordinary censoring creates a different estimand that can distort event probabilities.

Competing risks need event-specific methods#

If death prevents a nonfatal event, one minus a Kaplan-Meier curve that censors deaths generally overestimates real-world event probability. The cumulative incidence function accounts for competing events. Cause-specific and Fine-Gray subdistribution hazard ratios answer different questions.

A cause-specific hazard compares the event rate among people currently free of all events. A subdistribution hazard is linked to cumulative incidence through a constructed risk set. Neither should be called simply “risk.” A subdistribution hazard ratio is especially difficult to interpret without cumulative-incidence values. Good reports show you the event of interest and the competing outcomes together, because a treatment that increases early death can mechanically reduce later nonfatal events, and that lower cumulative incidence is not necessarily benefit.

Composite endpoints can hide component patterns#

Time to “cardiovascular death, heart attack, or stroke” ends at the first component. The composite hazard ratio can be driven by the most frequent or earliest event, which may not be the most serious. Components can move in opposite directions.

The exact hierarchy, adjudication, recurrent events, and competing death rules matter. A first-event analysis ignores later events and total burden. Recurrent-event models answer a different question and bring additional assumptions. So inspect the component event counts and curves yourself, remembering that trials are rarely powered for each component: a persuasive composite has clinically coherent components and an interpretation that a major harmful component does not contradict.

Confidence intervals describe precision, not all uncertainty#

A 95 percent confidence interval around a hazard ratio describes sampling uncertainty under the model and analysis. If it spans one, the data are compatible with no proportional-hazard difference and with the range shown. A P value does not state the probability that treatment works.

Narrow intervals do not cover bias from outcome misclassification, informative censoring, unblinded decisions, treatment switching, nonadherence, or model misspecification. A very precise adjusted hazard ratio from an observational database can still be confounded.

Clinical importance also cannot be inferred from statistical significance. A ratio near one can represent a meaningful absolute effect when baseline risk is high, while a dramatic ratio can translate into few prevented events when risk is tiny.

Kaplan-Meier curves need their own audit#

A survival curve estimates the probability of remaining event-free over time under censoring assumptions. Look for the numbers at risk printed beneath it. Late curve segments based on a handful of participants are unstable even when the line looks smooth to you.

Curves should include confidence bands and a clinically relevant time scale. Truncating the y-axis can magnify a small difference; using a full 0-to-1 axis can hide it. Both visualizations can be supplemented with numeric risks and differences.

Median survival is the time when the estimated survival curve reaches 50 percent. It cannot be estimated if fewer than half experience the event, and it ignores earlier and later curve differences. Reporting median follow-up is not the same as median survival.

Restricted mean survival time is an interpretable alternative#

Restricted mean survival time, or RMST, is the area under the survival curve up to a prespecified horizon. It represents average event-free or survival time accumulated during that period. A difference of 1.5 months over three years has a direct unit.

RMST does not require proportional hazards. Its interpretation depends on the chosen horizon, which must be clinically meaningful and supported by follow-up in both groups. Choosing the horizon after seeing where curves separate can bias presentation.

Other alternatives include risk difference or risk ratio at fixed times, percentile differences, and milestone survival. No single measure captures every curve shape. A small set of complementary summaries usually communicates better than replacing one talisman with another.

Observational hazard ratios need time-zero discipline#

In nonrandomized studies, treatment initiation, eligibility, covariate measurement, and follow-up should align. Classifying anyone who “ever” receives treatment as treated from cohort entry gives them a period they had to survive before initiation. That immortal time can create a protective hazard ratio.

Time-varying treatment models can allocate person-time correctly, but time-varying confounding affected by prior treatment may require g-methods. Confounding by indication remains because prognosis influenced treatment choice. The sophisticated survival model does not create randomization.

The paper should state the causal question, comparator, new-user criteria, censoring, treatment switching, and absolute outcomes. “Associated with” is usually the calibrated verb when identification assumptions remain uncertain.

A reporting set that supports decisions#

For each group, report participants, events, follow-up, censoring, and adherence. Show Kaplan-Meier or cumulative-incidence curves with numbers at risk, and give absolute risk and risk difference at prespecified, clinically meaningful times, with confidence intervals.

Report the hazard ratio, its model, covariates, whether it is cause-specific or subdistribution, and how proportionality was evaluated. If effects vary over time, present the chosen time-varying summary and an alternative such as RMST.

Finally, translate into the population actually studied. A relative estimate from a high-risk trial cannot tell you your absolute benefit if your own risk is lower, not without a valid baseline risk to apply it to. The hazard ratio is a compact statistical coordinate, not the whole clinical map.

References#

  1. Cox regression models and life-tables
  2. Understanding hazard ratios in clinical trials
  3. The hazards of hazard ratios
  4. Why test for proportional hazards?
  5. Alternatives to hazard ratios for comparing survival
  6. How hazard ratios can mislead

It does not provide medical advice or interpret an individual's prognosis.*

Questions and answers

Does a hazard ratio of 0.70 mean 30 percent fewer people had the event?

No. It means the modeled hazard, or instantaneous event rate among those still event-free, was 30 percent lower under the comparison and model. Absolute event reduction must be calculated at a stated time.

What is the proportional-hazards assumption?

It assumes the ratio of hazards between groups is constant through follow-up. Each group's hazard can change, but their ratio is modeled as stable. Crossing or delayed-separation curves can violate this summary.

Why can hazard ratios be hard to interpret causally?

At each time they compare people who remain event-free. Treatment and earlier events can change who remains, so later risk sets may be selected groups even after baseline randomization.

What should be reported with a hazard ratio?

Event counts, follow-up and censoring, survival or cumulative-incidence curves, numbers at risk, absolute risks and differences at useful times, confidence intervals, component outcomes, and assessment of time-varying effects.

What is restricted mean survival time?

It is average survival or event-free time accumulated up to a prespecified horizon. The group difference is expressed directly in days, months, or years and does not require proportional hazards.