Evidence explainer

Mental and behavioral health

How a Network Meta-Analysis Ranks 21 Antidepressants

A network meta-analysis links hundreds of trials so any two antidepressants can be compared, including pairs never tested head to head. It maps where reasonable options cluster, not which one wins.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The short answer
  2. Key points
  3. The problem it was built to solve
  4. Two yardsticks, not one
  5. Why the order is softer than it looks
  6. What the study genuinely supports

The short answer#

A network meta-analysis does not name one best antidepressant, and the most-cited study of its kind was never designed to. It stitches hundreds of randomized trials into a single connected web, then scores every drug on two separate measures: efficacy (how often depressive symptoms respond) and acceptability (how often people stay on the medicine rather than stopping it). When Cipriani and colleagues did this in the Lancet in 2018 across 21 antidepressants, 522 trials, and 116,477 adults, the finding beneath the headlines was softer than the coverage suggested. Every drug beat placebo, the gaps between reasonable choices were small, and the rankings came wrapped in enough uncertainty that no medicine could be called the correct answer for a given person.

Key points#

The problem it was built to solve#

Picture the evidence base as a map with most cities unconnected by any road. Drug A was tested against placebo. Drug B was tested against placebo. A trial pitting Drug A directly against Drug C may simply never have been run. A conventional meta-analysis can only pool trials that compared the same two things, so it stalls the moment you ask which of two never-compared drugs works better.

Network meta-analysis restores the missing roads. If two drugs were each tested against a shared comparator, their performance can be estimated relative to each other through that common anchor. Add in whatever direct comparisons do exist, link enough of these routes together, and every drug becomes reachable from every other, including pairs no trial ever placed side by side. The Cipriani team published the full underlying dataset openly, which means other researchers can rebuild the network and stress-test its assumptions rather than take the results on trust.

The method leans on one unspoken assumption called transitivity: the trials being linked must resemble each other closely enough, in patient severity, age, and design, that borrowing strength across them is fair. Chain together trials that studied very different populations and the indirect comparisons start to drift. That fragility is exactly why the analysis rewards careful reading over a glance at whatever sits on top.

Two yardsticks, not one#

The study measured two things that everyday conversation tends to fuse.

Efficacy was defined as response, roughly a halving of depression-scale scores over about eight weeks of acute treatment. On this axis all 21 drugs separated from placebo, with odds ratios running from about 2.13 for amitriptyline down to 1.37 for reboxetine. In the head-to-head data a higher-scoring cluster emerged that included agomelatine, amitriptyline, escitalopram, mirtazapine, paroxetine, venlafaxine, and vortioxetine.

Acceptability looked through the opposite lens: how many participants stopped the drug for any reason, a rough stand-in for tolerability and staying power in real use. Here a different cluster looked better, including agomelatine, citalopram, escitalopram, fluoxetine, sertraline, and vortioxetine, while several potent drugs such as amitriptyline and venlafaxine carried heavier dropout.

The important detail is the gap between the two lists. A medicine can press harder on symptoms and still be harder to live with. Only a few drugs land in a favorable region on both axes at once, which is why the two-yardstick view is more honest than any single ranked column.

Why the order is softer than it looks#

The tempting move is to treat the output as a league table and reach for whatever sits on top. The authors argued against precisely that, and the reasons sit inside the statistics.

Each rank is a single point surrounded by a wide credible interval. When the uncertainty band of the third-placed drug overlaps the band of the tenth, the apparent order is mostly noise dressed as precision. Summary ranking scores such as SUCRA compound the illusion by turning fuzzy, overlapping distributions into a crisp sequence, flattering small gaps into decisive ones.

Evidence quality caps how hard the numbers can be pushed. Certainty was graded moderate to very low, and most trials carried moderate or high risk of bias. Many were short, industry-funded, and enrolled tidier patients than a clinic typically sees, so the estimates describe an average acute response in trial populations rather than a forecast for you, with your comorbidities and your medication history.

The scope is narrow too. The analysis speaks only to acute treatment. It says nothing about maintenance over months, individual side-effect profiles, drug interactions, or your own preference, all of which shape the decision that actually gets made.

What the study genuinely supports#

Read for what it is, the analysis is valuable. It confirms that antidepressants as a class separate from placebo for acute major depression. It shows that the distances between sensible first choices are small. And it hands clinicians a defensible short list rather than a lone champion. A handful of drugs, escitalopram and vortioxetine among them, sit in a favorable zone on both distributions in this particular dataset. That placement marks where the two curves happen to overlap in these trials. It is a useful prompt for a conversation with a clinician, never proof that either drug is objectively best or a reason to prefer it over an equally reasonable alternative.

That is the fair use of a ranking: a map of where the workable options cluster, with the size of the gaps taken seriously and the uncertainty left visible. Treating one antidepressant as definitively the best converts a probability distribution into a verdict, and the underlying data was never strong enough to carry that weight.

Sources and further reading

  1. Cipriani et al., 2018, Lancet (PubMed)
  2. Comparison of 21 Antidepressants, Focus (APA reprint)
  3. GRISELDA open dataset (Mendeley, CC BY 4.0)

Questions and answers

Does this study tell me which antidepressant to take?

No. It compares average results across trial populations on two measures, efficacy and acceptability, and the differences among reasonable options are small. The choice for an individual depends on side-effect profile, other conditions, past responses, and preference, none of which a population ranking can settle. Use it as a starting point for a discussion with a clinician.

If every drug beat placebo, are the differences meaningful?

They are real but modest. All 21 drugs outperformed placebo for acute major depression, yet the spread between most active drugs is small and the rankings overlap heavily. The clearest signal is at the extremes of the list rather than in the crowded middle.

Why compare drugs that were never tested against each other?

Because direct trials cover only a fraction of the possible pairings. By linking drugs through shared comparators, network meta-analysis fills in the missing comparisons. This works only when the linked trials are similar enough for the indirect estimates to be trusted, an assumption every reader should check before leaning on the results.