Clinical trials often combine several outcomes into one primary endpoint. A heart-failure trial, for example, might combine death, hospitalization, and change in symptoms, and a conventional time-to-first-event analysis treats the first qualifying event as decisive, even when that event is less important than another component. A frequent hospitalization can therefore dominate a composite that also includes death.
The win ratio is one way to preserve a clinical hierarchy. It compares a participant assigned to treatment with a participant assigned to control, judges that pair first on the highest-priority outcome, and moves down to the next outcome only when the comparison is tied. Across all comparisons, treatment wins are divided by treatment losses.
That sounds intuitive, and sometimes it is. Yet a win ratio is not self-interpreting. To read one you need to know who was compared, what counted as a win, how follow-up and censoring were handled, how many comparisons tied, and which component produced the apparent advantage.
Why ordinary composites can mislead#
Composite endpoints increase event counts and can make a trial more efficient; they are most convincing when components have similar importance, occur with similar frequency, and move in a consistent direction. Those conditions often fail.
Consider a composite of cardiovascular death or hospital admission. If admissions are common and deaths are rare, the primary result may mostly reflect admissions, and a headline that says the composite fell can sound like a survival result even when mortality was unchanged. Time-to-first-event analysis also ignores events after the first one. A participant hospitalized early and then stable may be treated the same as one hospitalized repeatedly. The usual remedy is to inspect every component and its absolute event rate; the win ratio adds another option, by encoding clinical priority directly into the analysis.
A simple pairwise example#
Suppose one treatment participant and one control participant are compared on a hierarchy of death first and heart-failure hospitalization second.
If the control participant dies during follow-up and the treatment participant survives, treatment wins the pair. The hospitalization history is not consulted because the higher-priority outcome already resolves the comparison. If neither participant dies, the method moves to hospitalization. The one with a later first hospitalization, fewer recurrent hospitalizations, or a better prespecified hospitalization outcome may win, depending on the protocol. If neither level separates them, the pair ties.
Repeat that procedure across the selected comparisons. If treatment has 140 winning comparisons and 100 losing comparisons, the win ratio is 1.40. This means treatment generated 40 percent more wins than losses under that exact comparison rule. It does not mean mortality fell 40 percent or 40 of every 100 people benefited.
Matched and unmatched versions answer through different constructions#
The original proposal described matching treatment and control participants on baseline risk, then comparing each matched pair. Matching can make pairs clinically comparable, but it may leave some participants unmatched. Results can depend on the matching variables, tolerance, and order. A poorly chosen match can discard information or create imbalance.
The unmatched approach compares every treatment participant with every control participant. A trial with 500 participants in each group produces 250,000 pairwise comparisons. This uses everyone, but it does not create 250,000 independent observations. Each participant appears in many pairs, so valid variance and confidence-interval methods must account for dependence.
Reports sometimes say only "win ratio" without specifying which version was used. That omission matters. For the pairing rule, the covariate adjustment, the handling of stratification, and the analysis population, you have to go to the protocol or the statistical analysis plan.
The hierarchy is a clinical claim#
The order should be prespecified and justified before results are known. Death will often rank above hospitalization, but later levels can be debatable. Is a particular symptom-score improvement more important than an urgent outpatient visit? Should kidney failure rank above a nonfatal stroke? Do patients share the investigators' ordering?
An endpoint can also mix outcomes of different forms: time to death, count of admissions, and change in a quality-of-life scale, so the comparison rule must say exactly how each form creates a win, loss, or tie. Small changes on a lower-ranked continuous scale can resolve many pairs when serious events are uncommon. Patient input can improve the hierarchy, but preference is rarely uniform, so a transparent rationale is more defensible than claiming the ordering is universally obvious.
Why the component that drives the result still matters#
Hierarchical priority does not guarantee that the highest-ranked outcome drives the statistic. If few deaths occur, most pairs tie on death and move down to hospitalization, and if hospitalizations are also uncommon, a symptom or functional score may resolve the largest number of pairs.
This can be clinically reasonable if the lower-ranked outcome matters and is measured reliably. It can also make a dramatic-looking overall ratio rest almost entirely on a subjective component. In an unblinded trial, knowledge of assignment can affect symptom reporting, visit decisions, or decisions to admit.
The component win table is therefore essential. It should show how many pairs were won and lost at each level and how many continued to the next. You still need separate participant-level event rates, because pair counts can feel larger and less intuitive than the number of people affected.
Ties are information, not clutter#
A tie means the hierarchy did not distinguish a pair under the specified rules and observation window. Ties can be common when serious events are rare or follow-up is short: the win ratio's denominator includes wins and losses but excludes ties, so the ratio alone does not reveal how much information was unresolved.
Imagine 1,000 comparisons: 70 wins, 50 losses, and 880 ties. The win ratio is 1.40, the same as 700 wins and 500 losses with no ties. Those settings do not carry the same amount or pattern of information. Confidence intervals partly reflect precision, but you should still see the tie count or tie proportion.
Related measures can help. The win odds incorporates half of ties into each side under one convention. Net benefit expresses the difference between probabilities of a win and loss. None replaces component event rates and absolute effects.
Follow-up and censoring can change the question#
Participants are rarely observed for identical lengths of time. Some complete follow-up, some withdraw, some are lost, and some enter later before a common study end. If one member of a pair has shorter observation, it may be impossible to know who would have had the later event.
Different win-ratio procedures handle censoring in different ways. The estimand, meaning the treatment effect the analysis intends to estimate, must specify a meaningful time horizon and the conditions under which outcomes are compared. Recent methodologic work has emphasized that an unspecified horizon can make the result depend on trial-specific censoring patterns rather than a stable clinical question. So a strong report explains administrative censoring, loss to follow-up, competing events, terminal events, recurrent events, and sensitivity analyses, and it does not describe a win as though every pair had been observed for a lifetime.
Confidence intervals and statistical significance#
The estimated ratio should appear with a confidence interval. If the interval excludes 1 under the prespecified method, the result meets the usual criterion for statistical significance. That does not establish clinical importance or eliminate model assumptions.
A ratio can be statistically precise because the trial is large while representing a small absolute difference. It can also be unstable if a small number of losses makes the denominator small. Multiplicity matters when several hierarchies, cutoffs, time windows, or subgroups were tried, and the protocol and statistical analysis plan are what let you tell a confirmatory primary analysis from an exploratory re-analysis selected after the outcomes were seen.
What the win ratio cannot tell you by itself#
It does not tell you how many deaths were prevented, how many admissions were avoided, how much a symptom score changed, or how benefits and harms balance. It usually does not map cleanly to a number needed to treat. It can prioritize benefit components while leaving adverse events in a separate analysis.
It also does not fix a poor composite. An implausible hierarchy, inconsistently measured outcomes, heavy missingness, ascertainment bias, or a clinically trivial lower component remains a problem. Sophisticated arithmetic cannot make an unsuitable endpoint patient-important.
A practical reading checklist#
Start with the clinical question and population. Then find the complete hierarchy and ask whether it was prespecified. Identify matched or unmatched comparisons, the rule at each level, follow-up horizon, and handling of censoring.
Next, inspect wins, losses, and ties overall and by component. Compare absolute participant-level event rates and treatment effects for each outcome. Look for discordance, such as a favorable overall ratio with neutral or unfavorable mortality. Review adverse events separately.
Finally, check the confidence interval, sensitivity analyses, missing data, blinding, and whether the chosen lower-ranked outcomes could be influenced by treatment knowledge. The win ratio is most useful when it adds order without hiding the raw clinical events.
Sources and further reading
Questions and answers
Is a win ratio of 1.5 the same as a 50 percent risk reduction?
No. It means there were 50 percent more treatment wins than treatment losses across the specified pairwise hierarchy. It is not a direct estimate of an individual's event risk.
Why not just report each outcome separately?
Trials may need one prespecified primary endpoint for power and error control. A hierarchy can combine outcomes while respecting priority, but separate components still need to be reported for interpretation.
Are unmatched comparisons better than matched pairs?
Neither is automatically better. Unmatched analysis uses all cross-group pairs; matched analysis can improve baseline comparability but may omit participants or depend on matching choices. The method should fit the design and be fully specified.
What does a large number of ties mean?
It means many pair comparisons were unresolved under the hierarchy and observation window. The ratio can still favor treatment, but readers need the tie proportion and confidence interval to understand the result's information base.
Can a lower-ranked symptom score drive the whole result?
Yes. When serious events are rare or tied, lower levels resolve more pairs. Component-specific wins and absolute outcomes reveal whether the headline reflects mortality, hospitalization, symptoms, or another endpoint.