Evidence explainer

Evidence and research methods

What the STAR*D Reanalysis Debate Teaches About Trial Fidelity

STAR*D remains influential because it studied sequential depression treatment in usual-care settings. Its reanalysis debate shows why a protocol, dataset, and headline must be compared line by line.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What STAR*D was designed to do
  2. Remission is an operational definition
  3. How the cumulative headline was built
  4. What the 2023 reanalysis changed
  5. Missing data are not blank space
  6. Protocol fidelity is necessary but not mechanical
  7. What STAR*D does not establish
  8. How to audit a trial with a disputed headline
  9. The larger lesson for evidence communication
  10. References

The Sequenced Treatment Alternatives to Relieve Depression study, known as STAR*D, asked a practical question. What happens when a person with major depression does not reach remission with the first treatment and proceeds through additional medication or psychotherapy options? More than four thousand adults entered the first treatment level, many through primary-care settings, and common comorbidities were allowed.

The study became known for a cumulative message: after as many as four acute treatment steps, roughly two-thirds of participants who remained in the reported analysis achieved remission at some point. A 2023 reanalysis using patient-level data and the authors' reading of the original protocol reported a much lower cumulative rate, and the gap was not caused by a new blood test or newly discovered participants. It came from analytic choices.

What STAR*D was designed to do#

STAR*D was funded by the National Institute of Mental Health and conducted across psychiatry and primary-care sites. Everyone began with open-label citalopram in level 1; participants who did not achieve an acceptable result or could not tolerate treatment could move to later levels that offered switches, additions, and psychotherapy options. Some comparisons were randomized, while treatment preferences and clinical eligibility shaped which options a participant could enter.

The design favored applicability. Many conventional antidepressant trials exclude people with other psychiatric or medical conditions, use brief fixed dosing, and compare one drug with placebo. STAR*D used measurement-based care and dose adjustment over a longer acute phase, then offered sequential choices resembling clinical practice.

That advantage brought complexity. People left between levels. Not every participant was eligible or willing for every comparison. Later-level groups were selected by earlier response, tolerance, and preference. The study was open label, and different symptom scales served different purposes. A single cumulative percentage compresses all of that structure.

The NIMH STAR*D archive links the study questions and major results. Reading the archive makes clear that STAR*D was not one simple four-arm randomized experiment. It was a program of connected treatment steps.

Remission is an operational definition#

Remission in depression research means symptom scores fell below a prespecified threshold. It does not guarantee permanent recovery, absence of every symptom, restored function, or no relapse. The instrument and assessor matter.

STAR*D collected the Hamilton Rating Scale for Depression, or HRSD, through research outcome assessors who were not part of the treatment visit. It also used the Quick Inventory of Depressive Symptomatology, including a self-report form, in clinical measurement-based care, and the original protocol named the blinded HRSD threshold below 8 as the primary research definition of remission.

Published STAR*D reports often emphasized remission based on the clinic-administered or self-reported QIDS threshold, in part because exit HRSD data were missing for many participants; these scales are correlated but not identical. Their items, administration, and thresholds classify some people differently. The level 1 report presented outcomes from citalopram treatment and showed that remission estimates varied by measure. That is already a warning against saying “the remission rate” without telling you the instrument and the population.

How the cumulative headline was built#

The 2006 cumulative report summarized acute and longer-term outcomes across as many as four steps. It estimated that about 67 percent achieved remission when results were accumulated over sequential treatments. Later steps had lower acute remission, and relapse was more frequent among people who required more steps.

A cumulative remission calculation needs rules. Who enters the denominator? Is a person counted at a level if already below the remission threshold before starting it? What happens when the designated exit scale is absent but another scale is available? Is a remission counted if the participant later relapses? How are people who leave the study handled?

Different defensible estimands can answer different questions. “Proportion of all initial entrants ever observed in remission” differs from “proportion eligible for analysis at each level who met the protocol's primary outcome.” The problem arises when a publication presents one as though it were self-evident, and you carry it into a broader claim.

What the 2023 reanalysis changed#

The BMJ Open reanalysis was conducted under the Restoring Invisible and Abandoned Trials approach using STAR*D patient-level data. It sought to apply the authors' interpretation of the original protocol: use blinded HRSD remission, exclude participants who did not satisfy specified analysis criteria, and avoid counting people who were already in remission at entry to a treatment level.

Under those rules, the authors reported cumulative HRSD remission of 35.0 percent. When they supplemented missing exit HRSD assessments with QIDS remission, the estimate rose to 41.3 percent. Both were substantially below the widely repeated 67 percent.

The reanalysis argued that original publications departed from the protocol by using a nonblinded clinic measure for key outcomes and including participants who should not have entered certain analyses. It also highlighted the large number of missing exit HRSD assessments.

Those findings are important because they identify specific, testable sources of divergence. The debate is not simply one group trusting medicines and another distrusting them. It is about which rows, outcomes, and rules produce the estimate.

Missing data are not blank space#

Participants without an exit assessment may differ systematically from those with one. Some leave because they improve, some because they worsen, some because of adverse effects, logistics, preference, or unrelated events. Treating every missing result as nonremission can be conservative for one purpose but biased if successful participants disproportionately leave. Substituting another instrument adds people but changes the outcome definition.

Complete-case analysis assumes that the observed subset can answer the question without selection bias, an assumption often implausible in a long sequential study. Last observation carried forward, multiple imputation, inverse-probability weighting, worst-case assumptions, and composite failure definitions each make different claims.

The right response is not to hide one method behind a single number. Report how many assessments are missing at each level, why when known, and how estimates move across plausible assumptions. The difference between 35.0 and 41.3 percent in the reanalysis itself shows the effect of one missing-data choice.

Protocol fidelity is necessary but not mechanical#

Prespecification limits the freedom to select a favorable outcome after seeing results. A dated protocol and statistical analysis plan identify the intended primary endpoint, exclusions, comparisons, and missing-data strategy. Deviations can increase bias, especially when they improve the headline.

Still, a protocol is not sacred text. It can contain contradictions, omit necessary details, or specify a method later shown to be unsuitable. Real studies encounter unexpected problems. A change can be scientifically justified if it is timed, documented, explained, and distinguished from the prespecified result.

The RIAT initiative promotes correction of abandoned or misreported trials using clinical study reports and participant-level information. Its value rests on transparency. A restoration should identify source documents, reproduce prior results where possible, specify every changed rule, release code when permitted, and invite replication.

Reanalysts also make choices. Their classification of protocol fidelity, treatment of missingness, and preferred estimand can be debated. “Reanalysis” does not mean free from researcher degrees of freedom.

What STAR*D does not establish#

The disagreement does not prove that antidepressants never help. STAR*D level 1 lacked a placebo group, so it cannot separate medication effects from natural history, expectation, clinical attention, regression to the mean, and other influences. Conversely, response during an open-label course cannot be dismissed as meaningless simply because the exact cause is uncertain.

STAR*D also does not show that every later treatment is equivalent for every person. Group averages and underpowered comparisons can miss meaningful heterogeneity. Treatment choice includes prior history, symptoms, bipolar screening, safety, interactions, patient priorities, psychotherapy access, and monitoring.

The reanalysis does not replace the wider evidence from placebo-controlled trials, comparative trials, psychotherapy studies, safety surveillance, and guidelines. It revises how one influential program should be described.

Do not start, stop, or switch an antidepressant on the strength of a cumulative statistic you read in an article. Abrupt changes can cause withdrawal symptoms and clinical deterioration. Individual decisions require a treating clinician and a current assessment, especially when suicidal thoughts, severe functional decline, mania, psychosis, or acute safety concerns are present.

How to audit a trial with a disputed headline#

Begin with the study question and estimand. Identify the population, treatment strategies, outcome, time horizon, and event summary. Then compare the protocol, registry, analysis plan, publications, and dataset documentation.

For each outcome, record the instrument, assessor, threshold, timing, and hierarchy. Rebuild the participant flow from screening through every analysis. List exclusions made before and after treatment. Quantify missingness by group and reason. Check whether amendments were made before outcomes were known.

Reproduce the published estimate before you change it. Then change one rule at a time, which will show whether the difference comes from the outcome scale, the denominator, the missing data, the eligibility, or the coding. Sensitivity analyses should correspond to plausible clinical and statistical assumptions, not only favorable alternatives. Finally, separate acute remission from sustained remission, relapse, function, quality of life, adverse effects, and treatment burden: a person who meets a symptom threshold for one visit has a different outcome from someone well and functioning a year later.

The larger lesson for evidence communication#

A memorable percentage can outlive the details that produced it. Guideline summaries, educational materials, and media reports repeat the number until it seems like a natural fact. STAR*D shows that the denominator and measurement system must travel with the headline.

The responsible summary is longer but clearer: in this open-label sequential effectiveness program, cumulative remission varied substantially according to outcome measure, eligibility, and missing-data rules. The original summary and a protocol-focused reanalysis produced materially different estimates. That sentence preserves the study's importance and the uncertainty.

The best repair is prospective. Register the analysis, preserve versions, publish protocols and amendments, make deidentified data and code available under appropriate governance, and label exploratory work. Those practices do more than settle an old controversy. They reduce the chance that the next trial needs restoration.

References#

  1. NIMH STAR*D study archive
  2. STAR*D level 1 citalopram outcomes
  3. STAR*D cumulative outcomes report
  4. 2023 RIAT reanalysis
  5. NIMH Data Archive STAR*D collection
  6. RIAT proposal

Questions and answers

Did the reanalysis prove that antidepressants never work?

No. It reevaluated outcomes in one complex study and challenged a cumulative remission headline; it did not compare antidepressants with placebo across all evidence or determine one person's response.

Why did the reported cumulative remission percentages differ?

The analyses used different outcome measures, eligibility rules for analysis, treatment-step denominators, and approaches to missing exit assessments.

Was STARD a conventional blinded randomized trial?

No. It was an open-label multistep effectiveness study. Some later comparisons were randomized, but preferences and prior response shaped later participation.

Does following a protocol always produce the only valid analysis?

No. Prespecified analyses deserve priority, but protocols can be ambiguous or flawed. Justified changes should be transparent and reported beside the planned result.

What should a reader take from the debate?

Read the protocol, participant flow, outcome instrument, analysis population, missing-data rules, and sensitivity analyses. Treat a cumulative percentage as the end of an analytic chain, not a self-contained fact.

Was STAR*D a conventional blinded randomized trial?

No. It was an open-label, multistep effectiveness study, with some later treatment comparisons randomized and participant preference affecting available options.