The short answer#
Placebo response is large in depression trials because the placebo group is never just receiving a sugar pill. It is also living through the natural rise and fall of a mood episode, being pulled back toward its own average by the statistics of enrollment, and receiving weeks of structured, attentive contact with clinicians who ask careful questions. Across many antidepressant trials, placebo response has run near 35 to 40 percent, a level summarized in a 2020 analysis by Holper and Hengartner in BMC Psychiatry. Because that improvement is real, it raises the floor the drug arm has to beat, which is why the measured distance between drug and placebo often looks narrow. The aim here is to read that gap correctly, not to advise anyone on treatment.
Key points#
- The placebo arm measures much more than expectation: it also captures recovery over time, regression to the mean, and the therapeutic effect of frequent study visits.
- Depression is rated on subjective scales, so scores respond more to expectation than a laboratory value would.
- A large placebo response mathematically compresses the drug-placebo gap; a small gap is not the same as no effect.
- Whether placebo response has truly risen over the decades is contested once trial design changes are accounted for.
Start with the arithmetic#
The number that ends up in a results table is a subtraction: drug response minus placebo response. If the placebo arm improves by 35 to 40 percent, then the drug is not climbing from zero. It is climbing from an already high floor, and only the part that rises above that floor shows up as the treatment effect. This is the single most useful thing to hold in mind. A modest effect size can mean the drug does little, or it can mean the comparator is unusually strong. The arithmetic alone cannot tell you which, so the rest of the reading is about what filled the placebo arm.
What the placebo arm is actually capturing#
It helps to keep two ideas apart. The placebo effect, in the strict sense, is the improvement you can attribute to expectation and the ritual of being treated. The placebo response is broader: it is everything that happens in the inactive group over the course of the study. Most of that observed response is not the mind healing itself. It is the sum of several ordinary phenomena.
Depression tends to come in episodes, and people usually enroll when they feel at their worst. That is precisely the moment when the next measurement is most likely to read better, simply because extreme values tend to be followed by less extreme ones. That statistical pull toward the average is regression to the mean, and it acts on every arm regardless of what is in the capsule. Sitting on top of it is the natural history of the disorder, the tendency of a bad stretch to lift over weeks. Then add the trial itself: frequent visits, rating scales, and clinicians paying close, sustained attention. That setting is therapeutic in its own right, and it nudges people to feel and to report better.
Why expectation is unusually potent here#
Mood is measured through what patients and raters say, not through a blood value or a scan. Instruments such as the Hamilton and Montgomery-Asberg scales are built from judgment, and judgment-based endpoints bend more easily under expectation than a fixed laboratory number does. So the same hopeful expectation that barely moves a hard outcome can shift a depression score by a meaningful amount. When a participant knows there is a genuine chance of receiving an active antidepressant, that hope becomes part of what the scale records.
Why the gap looks small even when the drugs work#
The most rigorous synthesis of antidepressant efficacy is the 2018 network meta-analysis by Cipriani and colleagues in the Lancet, which pooled 522 trials and roughly 116,000 participants across 21 drugs. Two findings sit side by side. First, every antidepressant studied beat placebo for acute major depression. Second, and easier to overlook, the advantage was modest: the odds ratios for response clustered in a low range and the overall separation from placebo was small.
Those statements do not contradict each other. A drug can be reliably better than placebo and still separate from it by a narrow margin, because the placebo arm is nowhere near zero. Reading a small effect size as proof the drugs do nothing gets the causation backward. It is better read as evidence that the comparator is strong.
The rising-response story, and why it is contested#
For years a tidy narrative held: placebo response had been climbing decade over decade, eating into the drug-placebo difference and producing more so-called failed trials. Early reports described placebo response in published studies rising by several percentage points per decade through the 1980s and 1990s. Working from regulatory data spanning 1987 to 2013, Khan and colleagues reported in World Psychiatry that placebo response kept growing, up by roughly six percent since 2000, while the drug-placebo difference stayed about the same across the antidepressants approved in that window.
The picture is less settled than the headline. Reviewing 252 published and unpublished trials in Lancet Psychiatry in 2016, Furukawa and colleagues found placebo response looked stable since about 1991 once they adjusted for how trials had changed over time: their length, the number of study centers, and the use of fixed dosing. Part of the apparent rise, in other words, was a side effect of design drift rather than a real change in patients. Baseline severity and certain design features moved the placebo number; the calendar year, on its own, did far less than the story suggested.
Design features that push placebo response up#
A few recurring factors are worth watching when you read a single trial:
- More study sites tend to raise placebo response, partly because rating stays less consistent across many raters.
- Enrolling patients whose intake scores are inflated leaves more room for regression to the mean to work.
- A higher chance of receiving active drug, as in a design with several drug arms and one placebo arm, lifts expectation for everyone in the study.
- Weak blinding matters: older drugs with obvious side effects could tip off raters, which, as Holper and Hengartner discuss, can distort the comparison in either direction.
How to read the gap#
A large placebo response is not a flaw to be argued away. It is a signal that the outcome is soft, that the disorder resolves on its own for many people, and that the trial environment is itself a form of care. The practical discipline is to avoid two opposite mistakes. Do not treat a small drug-placebo difference as proof of no effect, and do not treat a strong placebo arm as proof the drug is pointless. Instead, ask what the placebo arm was measuring: how sick people were at intake, how many sites and arms the trial ran, and whether the blind likely held. Those questions do more to explain a result than the effect size alone ever will.
Sources and further reading
- Cipriani et al. 2018, Lancet network meta-analysis of 21 antidepressants
- Holper & Hengartner 2020, BMC Psychiatry, comparative efficacy of placebos in antidepressant trials
- Khan et al. 2017, World Psychiatry, rising placebo response, FDA data 1987-2013
- Furukawa et al. 2016, Lancet Psychiatry, placebo response rates in antidepressant trials
Questions and answers
Does a large placebo response mean antidepressants do not work?
No. It means the comparator is strong. The measured effect is the drug response minus a placebo response that already captures recovery over time, regression to the mean, and the care built into the trial. Pooled evidence still finds the drugs outperform placebo; the margin is just compressed.
Why is placebo response bigger in depression than in some other conditions?
Depression is scored on subjective, rater-dependent scales rather than a hard laboratory value, and many episodes improve on their own. Both features let expectation and natural recovery show up strongly in the placebo arm.
Has placebo response really been rising over the decades?
It is contested. Some analyses of regulatory data report a rise, while others find the trend largely disappears once you account for how trial design, length, and severity at intake have shifted over time.