Ask most people what a cost-per-QALY threshold represents and they will say it is the most a health system is willing to pay for a year of good health. That answer is intuitive, and it is wrong in an important way. When a health system runs on a fixed budget, a threshold does not measure what health is worth. It measures what health is lost. Funding a new treatment means moving money that was already producing benefit for other patients, and the size of that forgone benefit, the opportunity cost, is what a well-set threshold is trying to capture.
The distinction sounds academic until you notice how completely it changes the argument. Below I walk you through why the fixed budget is the load-bearing assumption, how a benchmark turns an incremental cost-effectiveness ratio into a decision, how the National Institute for Health and Care Excellence applies its stated range, and why economists still disagree about where the real number sits.
Key points#
- A threshold estimates health displaced elsewhere, not a valuation of health itself.
- The logic only holds when the budget is fixed, so new spending must come from existing spending.
- An incremental cost-effectiveness ratio (ICER) is meaningless until compared with what displaced care would have produced.
- NICE describes 20,000 to 30,000 pounds per QALY as a range applied with judgment, not a hard cutoff.
- A 2015 study put the empirical figure lower, near 13,000 pounds per QALY, and the debate is not settled.
The invisible patients behind every decision#
Start with the uncomfortable part, because it drives everything else. In a fixed budget the choice is never between a new treatment and nothing. It is between the patients who would receive the new treatment and a second group, usually anonymous and statistical, who lose the care that the same money was already buying. Those second patients rarely have names or advocates. They are the diagnostic clinic that does not expand, the nursing post that goes unfilled, the older, cheaper drug that reaches fewer people.
A threshold is the system's attempt to make that invisible group count. When it blocks a beneficial therapy, the natural objection is that the payer is putting a price on life. The opportunity-cost view answers that the payer is doing the opposite: it is refusing to ignore the lives on the other side of the ledger. Whether any particular threshold does that job accurately is a separate question, and a genuinely hard one.
Why the fixed budget is doing all the work#
None of this reasoning survives without the fixed-budget assumption, so it is worth stating plainly. A system with a set annual allocation cannot pay for something new by inventing extra money. It pays by displacing existing spending, and that displaced spending was itself producing health, through a medicine, a scan, a clinic, or a community service.
The quality-adjusted life year, or QALY, is what makes the comparison possible. One QALY is roughly a year lived in full health, with years in poorer health counted as fractions of one. Because a QALY puts a chemotherapy regimen, a hip replacement, and a smoking-cessation program on the same scale, it lets a payer ask a sharper question than "is this treatment worthwhile?" The sharper question is whether the treatment yields more health per pound than the care it will push aside.
Reading an ICER against the right yardstick#
The instrument for that comparison is the incremental cost-effectiveness ratio. The NICE manual defines the ICER as the extra cost of a technology divided by its extra QALYs, both measured against the relevant alternative. A therapy that costs an additional 300,000 pounds and delivers an additional 12 QALYs has an ICER of 25,000 pounds per QALY.
By itself that figure decides nothing. It only acquires meaning against a benchmark, and the benchmark is where opportunity cost enters. Suppose the care being displaced elsewhere generates one QALY for every 15,000 pounds. The new therapy, at 25,000 pounds per QALY, buys less health per pound than what it crowds out. The patients who receive it are helped, unambiguously, yet the population loses more QALYs from the displaced care than the therapy adds. The system ends up worse off in total health while every individual story about the new drug remains true. That paradox is the whole reason a threshold exists.
A range, not a single line#
NICE has long characterized technologies costing between 20,000 and 30,000 pounds per QALY as generally an acceptable use of NHS resources. Two features of that statement matter. First, it is a range rather than one number. Second, the manual frames it as a set of maximum acceptable ratios weighed with judgment, not an automatic pass-fail line.
Near the bottom of the range, an intervention is usually accepted as good value with little further debate. As the ratio climbs toward the top, the manual says other considerations carry more weight and the appraisal committee looks harder at the wider case. Structured modifiers can also apply, for example additional weight for treatments that address more severe conditions, which can lift the acceptable ratio in defined situations. The underlying discipline never changes: a higher ICER means more health displaced elsewhere, so it demands a stronger justification for accepting that loss.
The part that is genuinely unsettled#
Here the tidy logic meets a stubborn measurement problem. If a threshold is meant to reflect the health actually displaced at the margin, then its correct value is not something a committee gets to choose. It is a property of the system that has to be measured, namely how much the health system currently spends to produce one more QALY.
Claxton and colleagues attempted exactly that measurement in a 2015 study for the Health Technology Assessment programme. Using variation in NHS spending and outcomes across different programmes of care, they estimated the marginal productivity of NHS spending and arrived at a central figure of roughly 13,000 pounds per QALY, well below the stated 20,000 to 30,000 pound range. The implication was striking: if that lower number is right, approving therapies at the upper end of the range would strip more health from other NHS patients than it delivers.
That study did not end the argument. A later editorial surveying the field called opportunity cost in this context contentious and ambiguous across its theoretical, empirical, and public-opinion dimensions. It pointed to analysis by Zamora and Towse contending that once the structural uncertainty inside these models is taken seriously, the case for a true threshold far below the NICE range weakens, and that reasonable modelling choices can be consistent with 20,000 to 30,000 pounds being a defensible estimate. The disagreement is not about the economic logic, which both camps accept. It is about how confidently the productivity of an entire health system can be inferred from observational data, and how much weight you should give the resulting figures.
Sources and further reading
- NICE health technology evaluations manual (PMG36), economic evaluation
- NICE health technology evaluations manual (PMG36), committee recommendations
- Claxton et al., Methods for the estimation of the NICE cost-effectiveness threshold, HTA 2015;19(14), scientific summary
- Editorial, Opportunity costs in health care, cost-effectiveness thresholds and beyond
Questions and answers
Does a higher threshold mean a system values health more?
Not in the opportunity-cost view. A higher threshold means the system is willing to accept more health displaced elsewhere to fund new treatments. It signals tolerance for that trade-off, not a richer valuation of health.
Is the threshold a fixed rule?
No. NICE presents its range as guidance applied with committee judgment, and modifiers such as disease severity can shift the acceptable ratio. The range anchors the decision rather than automating it.
Why can't the correct threshold just be calculated once?
Because it depends on the marginal productivity of the whole health system, which must be estimated from imperfect observational data. Different reasonable modelling assumptions yield different numbers, which is why the empirical value remains debated.