Debates about antidepressants often force a false choice: “they work” or “they do not.” Randomized evidence supports a mean acute benefit over placebo for approved antidepressants in adults with major depressive disorder. That average is usually modest on symptom scales, varies across trials and drugs, and does not determine whether you, or any other particular person, will benefit, experience harm, or prefer another treatment.
Start with the randomized contrast#
Depression scores often improve substantially in both groups of a trial. Participants receive repeated assessment and clinical contact; symptoms fluctuate; enrollment often occurs during a severe period; and expectation, spontaneous improvement, and concomitant care contribute. Improvement from baseline in the treatment group cannot isolate the medicine's effect.
The causal estimate comes from comparing change or final scores between randomized groups under appropriate analysis, and if the antidepressant group improves by 12 points and placebo by 10, the randomized mean difference is 2 points, not 12. This does not mean the remaining improvement is “fake.” It means the trial cannot attribute it specifically to the drug. Clinical recovery can combine specific and nonspecific influences.
Raw points and standardized effects#
A raw mean difference states the group difference on a named scale, such as the 17-item Hamilton Depression Rating Scale or Montgomery-Asberg Depression Rating Scale. If you know that instrument, you can set the difference against the score range, the baseline severity, the measurement error, and published meaningful-change estimates.
A standardized mean difference divides the mean difference by a standard deviation. This puts different scales on a common unit. Rough labels such as small, medium, and large are not universal clinical rules. The denominator depends on population heterogeneity and measurement; the same raw benefit produces a smaller standardized effect in a more variable sample.
The 2018 network meta-analysis of 522 trials and 116,477 participants reported that all 21 included antidepressants were more efficacious than placebo for acute response, with drug-versus-placebo odds ratios ranging roughly from 1.37 to 2.13. That is evidence of average efficacy, but an odds ratio is not a response probability and should be translated using baseline risk.
What the FDA participant-level analysis found#
A 2022 analysis by FDA researchers used participant data from 73 randomized placebo-controlled trials submitted to the agency, covering 15 antidepressants and more than 17,000 participants. On the HAMD-17 scale, the average drug-placebo difference favored drug by 1.75 points, 95% confidence interval 1.63 to 1.86.
Both groups improved, with a greater mean improvement under antidepressant. The analysis also examined response distributions rather than relying only on a mean. Such results help describe the average magnitude in registration trials, but they inherit the trial populations, durations, scales, and medicine set submitted to FDA.
A 1.75-point mean difference should not be interpreted as every treated person feeling exactly 1.75 points better, and means can arise from broad small shifts, a subgroup with large benefit, different symptom-specific effects, or mixtures. Aggregate endpoint data cannot identify a reliable “responder type” without stronger interaction evidence.
Response rates depend on a cutoff#
Trials often define response as at least a 50% reduction from baseline and remission as a score below a threshold. These categories are easy to communicate but discard information.
Imagine two participants ending one point apart, with one just above and one just below the response cutoff. They are labeled nonresponder and responder despite near-identical outcomes, and a modest shift in the whole score distribution can move many people across a threshold and produce an apparently larger difference in response proportions.
Report both continuous and categorical outcomes, exact definitions, time point, denominator, and confidence intervals. A relative risk or odds ratio should be paired with absolute response rates and a risk difference. Number needed to treat varies with the control response, so do not carry one unchanged into your own setting.
Odds ratios are not risk ratios#
When response is common, an odds ratio looks farther from 1 than the corresponding risk ratio. If 40% respond on placebo and the odds ratio is 1.5, the treated response probability under simple assumptions is about 50%, not 60%.
Network meta-analysis can compare many drugs using direct and indirect trials, but transitivity matters. Dose, severity, and recruitment era may differ. So may setting, sponsorship, and comparator. Ranking one medicine first does not mean it is best for the person you are actually treating, or meaningfully better than the alternatives just behind it. Acceptability in the Cipriani analysis was measured mainly as all-cause discontinuation, a composite of benefit, adverse effects, expectations, and study conduct. It is not synonymous with tolerability or safety.
Statistical significance and clinical importance#
Large meta-analyses can estimate small effects precisely. A confidence interval excluding zero shows incompatibility with no mean difference under the model; it does not by itself show that the difference is important.
Clinical importance depends on symptom composition, baseline severity, and function. It depends on quality of life, patient priorities, adverse effects, and alternatives. A small mean difference can be worthwhile when depression is severe and the treatment is acceptable; the same difference may not justify burdens for someone else.
Minimal-important-difference thresholds for depression scales vary by method and context. Comparing a group mean difference directly with an individual-change threshold is a category error. The mean treatment effect and an individual's meaningful improvement answer different questions.
Baseline severity and effect modification#
It is plausible that drug-placebo differences vary with baseline severity, but subgroup claims have been inconsistent and method-sensitive; trials often restrict enrollment by severity, creating limited range, and baseline scores contain measurement error. Regression to the mean can create apparent relationships.
To claim effect modification, analyze a treatment-by-baseline interaction using participant-level data and an appropriate model. Showing significance in severe depression and nonsignificance in mild depression does not prove different effects. Severity is also not the only consideration. Depression subtype, comorbidity, and bipolar-spectrum features shape care. So do psychosis, anxiety, and substance use. So do previous response, age, pregnancy, medicines, and treatment preference, often without definitive randomized interaction evidence.
Individual response is not established by extra variance#
If some people have large drug-specific benefit and others none or harm, outcome variability might be greater in the antidepressant group than placebo, and meta-analyses comparing endpoint variances have generally not found strong evidence of substantially greater variability attributable to heterogeneous response, after accounting for mean-variance relationships.
That does not prove identical effects for everyone. Trial-level variances are an indirect test, adherence varies, outcomes are noisy, and individual causal effects cannot be observed because the same person cannot simultaneously take drug and placebo in the same episode. Clinical trial-and-monitor practice can still be reasonable. It should be described as learning an individual's observed course, not as evidence that baseline characteristics reliably predicted a drug-specific effect.
Publication and reporting bias changed the apparent literature#
The 2008 comparison of FDA trial records with publications found that most studies judged positive by FDA were published as positive, while many negative or questionable studies were unpublished or presented in ways that conveyed a positive outcome. The published literature overstated apparent effect size.
The 2018 network meta-analysis sought unpublished information from registries, regulators, companies, and authors, an important strength. Yet risk of bias, sponsorship, incomplete outcome data, selective analysis, and short follow-up remain relevant. If you search only journal articles, you are not reviewing the evidence on a drug.
Acute trials do not answer long-term questions#
Many efficacy trials last six to eight weeks and use selected adults without some common comorbidities or suicide risk. They establish acute symptom effects under those conditions. Maintenance and relapse-prevention trials ask whether continuing treatment after response reduces relapse, often using enriched randomized-withdrawal designs.
Enrichment means only people who tolerated and responded to treatment enter randomization. That is useful for a continuation question but overstates general tolerability if applied to everyone starting treatment. Abrupt or rapid discontinuation in a control arm can also complicate relapse interpretation through withdrawal symptoms. Long-term observational studies face confounding by indication, adherence, access, and illness course. Evidence should distinguish continuation benefit, chronic adverse effects, withdrawal, recurrence, and functional recovery.
Harms belong beside efficacy#
Common adverse effects vary by medicine. They can include gastrointestinal symptoms, sleep change, and activation or sedation. They can include sexual dysfunction, weight change, sweating, and blood-pressure effects. Some serious risks are age-, drug-, dose-, and interaction-specific. Discontinuation symptoms can occur and should not be confused automatically with relapse.
All-cause dropout does not capture these outcomes. Read adverse events, discontinuations due to adverse events, and serious events separately. Read suicidality assessment, sexual function, and withdrawal methods separately too. Short trials are poorly suited to rare or delayed harms.
Treatment decisions and any dose changes should be made with a qualified clinician. Abruptly stopping an antidepressant can cause symptoms and clinical deterioration.
A practical appraisal sequence#
Identify diagnosis, severity, and setting. Identify medicine, dose, comparator, duration, and outcome scale. Separate within-group change from between-group effect. Translate standardized or odds-ratio results into raw points and absolute probabilities when possible.
Check response and remission definitions, missing-data handling, and blinding. Check adverse effects, discontinuation, and sponsorship. Check registration and unpublished evidence. For network results, examine direct evidence and transitivity. Then decide whether the effect is large enough to matter, given function, quality of life, safety, and the alternatives you have.
Finally, do not predict one person's result from an average. Evidence informs a monitored shared decision; it does not replace diagnosis, preference, or follow-up.
Sources and further reading
- Cipriani and colleagues, Comparative Efficacy and Acceptability of 21 Antidepressants, Lancet (2018)
- Stone and colleagues, Response to Acute Antidepressant Monotherapy in Trials Submitted to FDA, BMJ (2022)
- Turner and colleagues, Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy, New England Journal of Medicine (2008)
- Munkholm and colleagues, Individual Response to Antidepressants, Meta-Analysis and Simulation Study, PLOS One (2020)
- National Institute for Health and Care Excellence, Depression in Adults, Treatment and Management, NG222
Questions and answers
Do antidepressants work better than placebo on average?
Yes, randomized acute adult trials show an average advantage. Its magnitude is generally modest on symptom scales and varies by medicine, study, and outcome.
Does a small mean difference mean nobody benefits greatly?
No. A mean cannot reveal every individual effect. It also cannot prove that a distinct large-responder subgroup exists or can be identified in advance.
Why can a response odds ratio sound larger than the score difference?
Categorizing a continuous score at a cutoff can move many people near the boundary, and odds ratios differ numerically from risk ratios when outcomes are common.
Is improvement in the placebo group imaginary?
No. It combines natural course, regression to the mean, expectation, clinical contact, measurement, and other care. Randomization estimates the added average effect of assignment to the drug.
Can acute trial results establish long-term benefit and safety?
No. Maintenance, withdrawal, recurrence, function, rare harms, and long-term tolerability require different and longer evidence.