Evidence explainer

Chronic disease in primary care

How Frailty Is Measured: Phenotype Versus Deficit Index

Frailty is scored two ways. One flags a physical syndrome from five signs; the other counts health problems as a fraction. A frailty number means little until you know which tool produced it.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. One word, two rulers
  3. The phenotype: a five-part checklist
  4. The deficit index: a running fraction
  5. Shrinking 40 items to a nine-point scale
  6. Where the two rulers disagree
  7. Matching the tool to the question

Frailty gets measured in two dominant ways, and they answer different questions. One asks whether a recognizable syndrome is present, scoring a person as frail when at least three of five physical signs appear. The other asks how much has gone wrong across the whole body, dividing the number of health problems a person has by the number checked. Both were built to predict falls, disability, institutional care, and death, and both do it well, yet they locate frailty in different places.

Key points#

One word, two rulers#

Ask whether an older adult is frail and you can mean two things. You can ask whether a distinct clinical syndrome is present, a state of low physiologic reserve with signs you can see and measure. Or you can ask how much damage has piled up, treating frailty as a running total spread across many body systems. Think of a smoke alarm versus a fuel gauge. The alarm either sounds or it does not; the gauge reads anywhere from full to empty. The phenotype is the alarm. The index is the gauge. Neither is wrong, and they were validated on different logic, which is exactly why they can disagree about the same patient.

The phenotype: a five-part checklist#

The physical frailty phenotype was defined by Fried and colleagues in the Journals of Gerontology, using the Cardiovascular Health Study, a cohort of more than 5,000 community-dwelling adults aged 65 and older. It names five components: unintentional weight loss (about 10 pounds over a year), self-reported exhaustion, measured grip weakness, slow walking speed over a short course, and low physical activity. Each has a defined cut-point, and two of the five, grip and gait speed, are physically measured rather than reported. Three or more criteria classify a person as frail, one or two mark an intermediate prefrail state, and none is robust.

Prediction was the design goal. Over roughly three years of follow-up, meeting the phenotype forecast new falls, worsening mobility, disability in daily activities, hospitalization, and death, and the raised risk held after adjustment for many other factors. The prefrail group also converted to frailty at higher rates, which supports the idea of frailty as a stage on a trajectory rather than a fixed tag.

The deficit index: a running fraction#

Mitnitski, Rockwood, and colleagues took the opposite approach in a 2001 paper in The Scientific World Journal. Rather than specify a syndrome, they counted deficits of any kind: symptoms, signs, laboratory abnormalities, diagnoses, and functional limitations. The frailty index is simply the count of deficits present divided by the count assessed, so a person with 12 problems out of 40 checked scores 0.30. Working from the Canadian Study of Health and Aging, they found deficits accumulate at a fairly steady average rate with age and that the index behaves like a state variable tracking the risk of death.

A 2008 methods paper by Searle and colleagues in BMC Geriatrics laid down rules for building one of these indices. Each candidate deficit should relate to health, grow more common with age, avoid becoming near-universal too early in life, and stay stable on repeat measurement, and the full list should span multiple systems, with at least 30 to 40 items recommended. Built this way the index is continuous and graded rather than a yes-or-no category, and in real populations it tends to level off well short of its theoretical maximum of 1.0, a submaximal ceiling that carries meaning of its own.

Shrinking 40 items to a nine-point scale#

Counting 40 items is impractical during a clinic visit, so Rockwood and colleagues published a brief clinical measure in 2005 in the Canadian Medical Association Journal. Their Clinical Frailty Scale ranks a person along an ordered spectrum from very fit to severely dependent, based on an overall read of function and comorbidity. In more than 2,000 older Canadians followed for five years, each single-step rise on the scale meaningfully increased the risk of death and of moving into institutional care, and the short scale predicted those outcomes about as well as a detailed multi-item index. The lesson was that a quick, judgment-based rating can carry much of the predictive signal of the long count.

Where the two rulers disagree#

The instruments pick out overlapping but not identical people. The phenotype is precise and reproducible, and it isolates a physical state that plausibly shares a biology of low energy and reserve; its cost is narrowness, since it says little about cognition, mood, or accumulated disease and needs standardized performance testing. The index is broad and sensitive across the whole range of health, and its continuous score catches small differences the three-of-five rule flattens; its cost is that the number depends on which deficits were chosen, and a high score tells you how much is wrong without naming any single mechanism.

Applied to the same population, they usually agree on who is clearly robust and who is clearly frail. The disagreement clusters in the large middle, where a person can clear the phenotype's threshold yet still carry a meaningful load of quieter deficits, or vice versa.

Matching the tool to the question#

Which measure fits depends on why you are asking. A trial testing whether exercise rebuilds physical reserve may want the phenotype's specificity, so its outcome maps onto the thing being treated. A health system stratifying surgical or hospital risk across an entire population may prefer the index's breadth and fine gradation. The practical takeaway is that a frailty statistic read without its source can mislead, because the same word points to two different measurements.

This kind of measurement question, how an instrument is defined and validated before its numbers are trusted, sits close to the evidence-appraisal work that informs primary care. The same discipline applies whether the outcome is frailty, a diabetes screening cut-point, or a childhood growth metric.

Sources and further reading

  1. Fried et al., Frailty in Older Adults: Evidence for a Phenotype, J Gerontol 2001
  2. Rockwood et al., A Global Clinical Measure of Fitness and Frailty, CMAJ 2005
  3. Mitnitski et al., Accumulation of Deficits as a Proxy Measure of Aging, ScientificWorldJournal 2001
  4. Searle et al., A Standard Procedure for Creating a Frailty Index, BMC Geriatrics 2008

Questions and answers

Is one frailty measure better than the other?

Not in general. They were validated for different purposes. The phenotype is stronger when you want a clean physical category; the index is stronger when you want a graded, whole-body score. Many groups now use both.

Can a person be frail on one measure and not the other?

Yes, most often in the middle range. Someone can miss the phenotype's three-of-five threshold while accumulating enough smaller deficits to score high on the index, or the reverse.

Do these tools give medical advice?

No. They are research and screening instruments, not diagnoses. Frailty assessment belongs in a conversation with a qualified clinician who can pick and interpret the right tool.