Evidence explainer

Evidence and research methods

How Depression Rating Scales Are Scored and Why the Cutoffs Are Debated

A depression rating scale turns an interview into a single number. How far that number must move before anyone would notice is unresolved.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The short answer
  2. Key points
  3. Turning an interview into a number
  4. "Significant" is not the same as "noticeable"
  5. What the anchoring study found
  6. Why three points draws fire
  7. The case for the other side
  8. A short checklist for reading the numbers

The short answer#

Depression rating scales work by giving a trained interviewer a fixed list of symptoms, scoring each one, and adding the scores into a single total. The two you meet most often in research are the Hamilton Rating Scale for Depression (HAM-D) and the Montgomery-Asberg Depression Rating Scale (MADRS). The hard part is not producing the number. It is deciding how much the number has to move before the change matters to a real person. A long-used convention treats a three-point gap between a drug and placebo on the HAM-D as clinically meaningful, yet studies that tie scale points to what clinicians can actually perceive put the smallest noticeable change closer to seven points. That mismatch is why the cutoff is genuinely disputed.

Key points#

Turning an interview into a number#

Both scales are administered through a structured conversation, and both lean on the interviewer's judgment rather than a self-report questionnaire. What differs is their design philosophy.

The HAM-D, published by Max Hamilton in 1960, is the older instrument. Its most common form, the 17-item version, rates depressed mood, guilt, suicidal thoughts, several kinds of insomnia, anxiety, agitation, appetite, and a cluster of bodily complaints. Some items run from 0 to 4, others only 0 to 2, and everything is summed. A notable share of the possible total comes from sleep, weight, and physical symptoms, categories that overlap with ordinary states of being unwell and with the side effects of medication.

The MADRS, introduced by Montgomery and Asberg in 1979, was deliberately built to fix what its authors saw as that weakness. It has 10 items, each scored 0 to 6, and it concentrates on the mood and cognitive core of depression while giving little weight to sleep and somatic complaints. Because the two scales carry different numbers of items and different point ranges, a two-point move on one is not the same size of change as a two-point move on the other. Any comparison has to keep that in mind.

"Significant" is not the same as "noticeable"#

A statistically significant result in a trial tells you one narrow thing: the difference between groups is unlikely to be a fluke of chance. It says nothing about whether a patient sitting in front of you would feel better or whether a clinician watching them would see any change. Those are different questions, and they need a different yardstick.

That yardstick is the minimal clinically important difference: the smallest change in score that lines up with a real, perceptible improvement. The reason it is hard to pin down is that a scale point has no fixed meaning on its own. Two points of improvement on the HAM-D might reflect slightly better sleep, or it might reflect a genuine lift in mood, depending entirely on which items moved. To give the raw numbers meaning, researchers anchor them to a judgment clinicians already make, most often the Clinical Global Impression rating, on which a rater simply calls a patient unchanged, minimally improved, much improved, and so on.

What the anchoring study found#

The most influential attempt to do this is by Leucht and colleagues in the Journal of Affective Disorders in 2013. Pooling data from dozens of drug trials, they matched HAM-D scores against clinicians' global impressions of change. The finding that drives the debate is this: a rating of "minimally improved," the faintest benefit a clinician would register at all, corresponded to a HAM-D drop of about seven points. "Much improved" did not map onto any fixed point count; it tracked a reduction of more than half the starting score.

They also showed the anchor drifts with severity. In milder cases, a minimal improvement lined up with roughly six points; in the sickest patients, roughly eight. So the same raw movement means different things depending on where a person began, which is exactly why a single universal threshold is awkward to defend.

Why three points draws fire#

For years, guidance from the UK's National Institute for Health and Care Excellence treated a three-point HAM-D difference between drug and placebo as a working line for clinical significance. The objection is easy to state. If the smallest change a clinician can even perceive sits around seven points, then a three-point rule stamps "clinically important" onto a gap too small for anyone to see at the bedside. On that reading, the conventional cutoff is far too generous, and trials that clear it may still be delivering changes below the threshold of noticeability.

This is not an abstract worry, because average drug-placebo gaps in antidepressant trials tend to land right in the disputed zone. A 2020 meta-analysis by Hengartner and colleagues in PLOS ONE pooled more than 130 placebo-controlled trials and looked at the HAM-D and MADRS separately. The standardized effect sizes were close, about 0.27 and 0.30, and the raw gaps were on the order of two points on the HAM-D and three on the MADRS. Both fell under the minimal-improvement benchmarks the authors used, leading them to call the average benefit of doubtful importance for a typical patient.

The case for the other side#

It would be dishonest to stop there, because the question is not settled. A 2020 paper by Hieronymus, Jauhar, Ostergaard, and Young in the Journal of Psychopharmacology makes the opposite case: judging antidepressants by the full HAM-D total actually understates them. The argument turns on the scale's own construction. By bundling core depressive symptoms together with sleep, somatic items, and complaints that overlap with side effects, the full sum adds noise. When they examine individual core items or a short core subscale such as the six-item HAM-D, the drug-placebo separation widens, and effects on features like depressed mood look larger than the headline number suggests. On this view, part of the small average is an artifact of the measuring stick, not the treatment.

Read together, the two camps are pointing at one problem from different angles. A single summed number squeezes a varied illness into one figure, and the meaning of any cutoff depends on how the scale was assembled and which symptoms are driving the total. That is precisely why a three-point line cannot be judged in a vacuum, apart from the anchoring evidence, the choice of instrument, and how ill the trial population was.

A short checklist for reading the numbers#

When a study reports an X-point improvement, three questions do most of the interpretive work:

None of this decides whether a particular treatment will help a particular person. What it does is let you read the reported figure honestly, which is the whole reason for understanding the scoring in the first place.

Sources and further reading

  1. Leucht et al., What does the HAMD mean? (J Affect Disord 2013)
  2. Hengartner et al., Efficacy of antidepressants assessed with the MADRS (PLOS ONE 2020)
  3. Hieronymus, Jauhar, Ostergaard & Young, One (effect) size does not fit at all (J Psychopharmacol 2020)

Questions and answers

Is a higher score better or worse?

Worse. On both the HAM-D and the MADRS, the score is a severity measure, so a higher total means more symptoms and a downward change over time signals improvement.

Why can't you just compare HAM-D and MADRS points directly?

Because they are built differently. The scales have different numbers of items and different point ranges, and the MADRS deliberately downplays sleep and physical symptoms that the HAM-D counts. A given point change on one is simply not the same quantity as the same change on the other.

Does the debate mean antidepressants do not work?

No. It is a debate about how to measure and interpret effects, not a verdict on treatment. One side argues the common cutoff is too lenient; the other argues the full scale hides real benefit in its core items. Both are questions about the ruler, and treatment decisions rest on far more than a single trial average.