A quality-adjusted life year, or QALY, is one number that folds together two things a treatment can change: how long a person lives and how good that added time is judged to be. You multiply the years spent in a health state by a utility weight for that state, where 1 stands for full health and 0 stands for death, and you add the products across every state a person passes through. A year in full health is worth one QALY. Two years lived at a weight of 0.5 also come to one QALY. The design is deliberately plain so that a hip replacement, a cancer drug, and a diabetes program can all be scored on one ruler, and that same plainness is where almost every objection to the measure begins.
Key points#
- A QALY equals time in a health state multiplied by a utility weight from 1 (full health) to 0 (death); weights can go below 0 for states rated worse than death.
- Utility weights are not clinical measurements. They are elicited from people using methods such as the time trade-off, the standard gamble, or a questionnaire like the EQ-5D scored against a general-population value set.
- Paired with cost, QALYs produce an incremental cost-effectiveness ratio (ICER), a cost per QALY gained that agencies like NICE compare against a threshold range.
- The arithmetic treats every quality-adjusted year as identical regardless of who receives it, which is defended as fairness and criticised as blindness. Severity weighting is one attempt to soften that.
Where the two numbers actually come from#
The time input is the easier of the two. It usually comes from trial follow-up, survival data, or a model that projects how a condition unfolds over years.
The utility weight is the harder part, because it is not something you can read off an instrument. Blood pressure has a true value waiting to be measured; the desirability of living with, say, moderate hearing loss does not. It has to be drawn out of people, and three approaches dominate.
- The time trade-off asks how many years in full health a respondent would treat as equivalent to a longer stretch in a worse state. Trading eight of ten years to avoid the worse state implies a weight around 0.8.
- The standard gamble asks what risk of immediate death someone would accept for a chance at full health instead of staying in the impaired state.
- Multi-attribute questionnaires such as the EQ-5D ask people to rate their health across fixed dimensions (mobility, self-care, usual activities, pain or discomfort, and anxiety or depression), then convert that profile into a single index using a value set built from a population survey.
Two consequences follow that are easy to miss. Some states are rated as worse than being dead, so a weight can fall below zero and subtract from the total rather than add to it. And the weight pinned to a given condition depends heavily on which method was run and whom it was run on. A person living with a condition and a member of the public imagining it routinely give different answers, and choosing between those viewpoints is a methodological decision with real effects on the result.
Turning the weight into a decision#
On its own a QALY says nothing about value for money. It becomes a decision tool once cost is attached. Analysts compute an incremental cost-effectiveness ratio, or ICER: the extra cost of a new treatment divided by the extra QALYs it delivers, reported as a cost per QALY gained. A lower cost per QALY signals better value.
England's National Institute for Health and Care Excellence (NICE) shows how this plugs into appraisal. Its evaluations manual sets QALYs as the reference-case measure of health effect and, for adults, prefers the EQ-5D so that comparisons stay consistent from one appraisal to the next. Historically NICE has treated a most plausible ICER below roughly £20,000 per QALY as generally cost-effective, with acceptance up to about £30,000 needing extra justification such as greater uncertainty or wider benefits. NICE has said this range will shift to £25,000 to £35,000 once a regulatory change gives it the power to apply the higher figures, describing that as agreed but still contingent on the regulatory step rather than already in force. Both cost and health effects are discounted, currently at 3.5% a year, so a QALY far in the future counts for less than one now.
NICE also applies a severity modifier. Conditions that impose a larger QALY shortfall, judged both in absolute terms and as a share of the health a person would otherwise have expected, can have their QALYs weighted up by a factor of 1.2 or 1.7. A peer-reviewed account from authors at NICE and the University of Sheffield describes this modifier replacing an earlier end-of-life weighting, and it is a plain admission that a bare QALY count does not by itself reflect how much is at stake for the sickest patients.
What no weighting fixes#
The severity modifier points at the deeper issue: an unadjusted QALY treats every quality-adjusted year as interchangeable, no matter who gains it. A year won by a twenty-year-old and a year won by someone near the end of a long illness score the same. Some call that neutrality fair and some call it blind, and both descriptions fit the identical arithmetic.
Several specific gaps are well documented in the literature.
- Because the measure rewards gains in both length and quality of life, a treatment that mainly helps people who begin from a permanently lower utility, including some people living with disability, can produce fewer QALYs for the same clinical effort. That has driven longstanding concern that the metric can count against them.
- Averaged population value sets can misstate the lived reality of people who have adapted to a condition over years.
- The EQ-5D's five dimensions leave out much that people care about: cognition, relationships, dignity, and the value of reassurance or information.
- A narrow reference case sets aside carer burden, lost productivity, and effects that land outside the health system.
- The whole calculation is sensitive to modelling choices and to the discount rate, so two competent analyses of the same treatment can land on different ICERs.
None of this makes the QALY worthless. It makes it a structured, transparent input rather than a verdict handed down.
Reading a cost-per-QALY claim#
When a headline reports that a treatment costs a certain amount per QALY, the number is only as good as the assumptions folded inside it. Three questions do most of the work. Which utility method and which population produced the weights? What time horizon and discount rate were assumed? And what was left out of the model entirely? A QALY carries its assumptions inside it, so appraising a health technology claim really means appraising those assumptions, not the tidy single figure they generate.
Sources and further reading
Questions and answers
Can a QALY be negative?
Yes. If a health state is rated as worse than death, its utility weight is below zero, and time spent in that state subtracts from the QALY total rather than adding to it.
Is a QALY the same as a life year?
No. A life year counts only time. A QALY discounts that time by a utility weight for quality of life, so one year in poor health is worth less than one QALY, and one year in full health is worth exactly one.
Why do different studies report different costs per QALY for the same drug?
Because the result depends on modelling choices: which utility instrument and population set the weights, the time horizon, the discount rate, and which costs and benefits the model includes. Change those inputs and the ICER moves, even for identical clinical evidence.