SURMOUNT-1 was a phase 3 randomized trial of tirzepatide for chronic weight management. The result was not simply that participants “lost 20%.” There were three dose groups, two prespecified ways to handle treatment discontinuation and rescue events, a distribution of individual responses, and a placebo group that received the same lifestyle program.
The 72-week findings established substantial weight reduction in the studied population. A later 176-week analysis in the subgroup with prediabetes added evidence about durability during continued treatment and progression to type 2 diabetes. Neither analysis shows that everyone will reach the average, that benefit continues unchanged after stopping, or that the original trial established fewer heart attacks, longer life, or safety beyond its follow-up.
Start with who entered the trial#
Participants were adults with body-mass index at least 30, or at least 27 plus one or more weight-related complications, and at least one unsuccessful dietary effort to lose weight. People with diabetes were excluded. The average baseline weight was about 105 kg, average BMI about 38, and roughly 68% were women. About 41% had prediabetes.
Those details set the target population. Results cannot be assumed to be identical in adolescents, people with type 2 diabetes, people at lower BMI without complications, pregnancy, severe gastrointestinal disease, or populations underrepresented in the study. The racial and geographic composition also affects transportability.
Randomization was in a 1:1:1:1 ratio to three dose strategies or placebo. All groups received a reduced-calorie diet and physical-activity counseling. The placebo comparison therefore estimates adding the medicine to that trial program, not medicine versus no care.
The trial was sponsored by the manufacturer. Sponsorship does not invalidate a randomized trial, but you should inspect protocol control, analysis, missing data, author roles, and whether independent studies reproduce the result.
The two co-primary outcomes#
One co-primary outcome was percent change in body weight from baseline to week 72. The other was the proportion of participants who lost at least 5% of baseline weight. Both were prespecified and tested within a multiplicity plan.
The continuous outcome shows average magnitude. The threshold outcome shows how many crossed a clinically recognizable boundary. Neither reveals the whole distribution. Two groups can have the same mean while one has tightly clustered responses and another has a mix of large loss, little change, and discontinuation.
In the primary report, estimated mean changes under the treatment-regimen estimand were -15.0%, -19.5%, and -20.9% in the active groups and -3.1% with placebo. At least 5% loss was estimated in 85%, 89%, and 91% versus 35%. At least 20% loss occurred in 50% and 57% of the two higher-dose groups versus 3% with placebo; the lowest-dose group was not included in that particular multiplicity-tested comparison.
These are model-based group estimates with confidence intervals, not promises. A 20% mean reduction does not mean that each participant lost 20% or that every kilogram was fat rather than lean tissue.
Why the estimand changes the sentence#
An estimand defines the treatment effect. ICH E9(R1) asks analysts to specify population, treatment condition, outcome variable, how intercurrent events are handled, and the population-level summary.
SURMOUNT-1 reported a treatment-regimen estimand that addressed effect regardless of discontinuation or initiation of another weight-management therapy, similar to a policy question after assignment; it also reported an efficacy estimand focused on what would be expected if participants remained on assigned treatment without selected intercurrent events.
The efficacy question generally produces a larger effect because it targets continued use. The treatment-regimen question incorporates some consequences of stopping or adding another therapy. Neither is automatically more truthful. The useful choice depends on whether a reader wants biological efficacy under continued treatment, expected effect of initiating a strategy, or another clearly defined question.
Missing weights still require assumptions. The model uses observed data to impute what was not measured. Sensitivity analyses ask whether plausible departures from those assumptions change the conclusion, and a high completion rate helps, but does not make missing outcomes irrelevant when dropout differs by group or relates to response.
Completion and treatment persistence are different#
About 86% of randomized participants completed the 72-week trial, while about 82% completed assigned treatment. A participant could remain in outcome follow-up after stopping study drug, which strengthens a treatment-regimen analysis.
Treatment discontinuation can reflect adverse events, perceived lack of benefit, personal preference, access, pregnancy, protocol rules, or loss to follow-up. Lumping all causes together hides whether stopping is related to outcome. A per-protocol analysis that simply removes discontinuers can break the balance created by randomization.
Real-world persistence can be lower than in a trial with regular visits, supplied medicine, counseling, and selection for participation. Effectiveness after prescribing therefore depends on affordability, supply, tolerance, follow-up, and the ability to sustain the plan.
Harms belong beside the weight curve#
The most common adverse events were gastrointestinal, including nausea, diarrhea, and constipation, and occurred mainly during dose escalation. Most were reported as mild or moderate. Treatment discontinuation because of adverse events occurred in 4.3%, 7.1%, and 6.2% of active groups versus 2.6% with placebo in the primary trial.
A 72-week trial can characterize common events better than rare or delayed harms. Product labeling and postmarket evidence add risks, contraindications, interaction concerns, and monitoring beyond the article's abstract. Trial eligibility can exclude people at higher risk of certain harms.
The correct comparison is not “weight loss versus side effects” as if each were a single outcome. It includes metabolic, functional, psychological, nutritional, gallbladder, pancreatic, gastrointestinal, reproductive, and treatment-burden considerations relevant to the individual. Rapid or substantial loss can also reduce lean mass, so nutrition and resistance activity may matter within professional care.
What the 176-week extension added#
Participants with prediabetes at baseline were eligible to continue randomized treatment to week 176, followed by 17 weeks off treatment, and this was a selected subgroup of the original trial, not a new randomization of everyone with obesity.
The 2025 peer-reviewed report found mean weight changes at week 176 of -12.3%, -18.7%, and -19.7% across active groups versus -1.3% with placebo. Type 2 diabetes was diagnosed in 1.3% of pooled active participants and 13.3% of placebo participants during the 176-week treatment period, corresponding to a hazard ratio of 0.07. After the 17-week off-treatment period, cumulative proportions were 2.4% and 13.7%.
That is strong evidence that continued treatment delayed or prevented diagnostic progression during the observed period in this prediabetes subgroup. It does not establish lifetime prevention. Weight loss itself mediates part of the benefit, diabetes ascertainment can be affected by ongoing pharmacology, and risk can reappear after treatment stops.
Long follow-up improves durability evidence but introduces attrition and survivor selection. Compare who entered the extension, who remained, and how outcomes were adjudicated before you take the durability figure at face value.
Weight regain tests the chronic-treatment model#
SURMOUNT-1 did not randomize continued treatment against withdrawal after initial loss. SURMOUNT-4 addressed that question: after an open-label lead-in, participants were randomized to continue tirzepatide or switch to placebo. Continued therapy maintained and extended weight reduction, while withdrawal led to substantial regain on average.
That finding supports obesity as a chronic, relapsing condition for many people. It does not mean every person needs the same medicine forever, nor that stopping always returns weight to baseline. It means a 72-week endpoint should not be read as a permanent treatment-free effect.
Maintenance decisions include benefit, adverse effects, goals, pregnancy plans, cost, access, alternative treatments, and what happened during dose changes or interruption. Abrupt changes without clinical guidance can create avoidable problems.
What outcomes were not established by the original trial#
Body weight is directly meaningful but remains an intermediate outcome for some health claims. SURMOUNT-1 measured cardiometabolic risk factors, yet it was not powered to show fewer cardiovascular deaths or major cardiovascular events. Other indication-specific trials are needed for sleep apnea, heart failure, diabetes, and cardiovascular outcomes.
Quality of life and physical function matter, but questionnaire changes should be interpreted by scale, missing data, and whether participants knew or guessed their assignment because of obvious effects. Cancer, fracture, fertility, and very long-term nutritional outcomes require different follow-up.
The trial establishes a large average weight effect under studied conditions. It does not make a universal claim about health, appearance, worth, or the correct goal for every body.
A six-question reading method#
For any weight-management trial, ask:
- Who was enrolled and excluded?
- What did the comparator receive?
- Was the endpoint mean change, a responder threshold, health outcome, or all three?
- Which estimand handled discontinuation, rescue therapy, and missing values?
- How many completed follow-up and treatment, and why did people stop?
- What happened after the randomized period or after withdrawal?
This method keeps a striking average in context without minimizing it; it also makes comparisons across medicines safer because trials can differ in population, duration, estimand, lifestyle program, dose strategy, and missing-data assumptions.
Look beyond the average response#
An average combines people who lost much more, much less, or no weight. Responder thresholds help describe that spread, but a row of 5%, 10%, 15%, and 20% cutoffs still does not show the full distribution. A cumulative distribution plot can reveal whether one group shifts broadly or whether a subgroup produces most of the mean difference.
Prespecified subgroup analyses can explore whether effects differ by sex, baseline BMI, prediabetes, age, or geography, but they are usually underpowered for interaction, and a result that is significant in one subgroup but not another does not itself prove a subgroup difference. The relevant test is the interaction, interpreted with multiplicity, biological plausibility, and replication.
Absolute kilograms and percent change answer related but distinct questions. Percent change makes different starting weights more comparable, while kilograms can be easier to picture. Waist circumference, blood pressure, laboratory measures, physical function, and patient-reported outcomes add context, but improvement in one marker does not guarantee improvement in every outcome.
The most useful individual forecast remains a monitored treatment trial within appropriate care. Early response, tolerance, nutritional status, function, and your goals can support reassessment. Population averages inform that conversation; they do not replace it.
References#
- SURMOUNT-1 72-week primary report
- SURMOUNT-1 176-week obesity and diabetes-prevention report
- ClinicalTrials.gov NCT04184622
- ICH E9(R1) estimands framework
- FDA clinical review of tirzepatide for chronic weight management
- SURMOUNT-4 randomized withdrawal trial
For your own health, talk with your clinician.*
Questions and answers
Did the average participant in SURMOUNT-1 lose more than 20% of body weight?
The estimated mean reached about 20% in the two higher-dose groups under the main analysis, while the lowest group averaged about 15%. Individual responses varied, and not everyone remained on treatment.
Why did SURMOUNT-1 exclude diabetes?
The primary trial targeted obesity or overweight without diabetes. People with type 2 diabetes can have different weight responses, medicines, and safety context and were studied in a separate program.
Does the 94% relative diabetes-risk reduction mean no one developed diabetes?
No. During 176 weeks, 1.3% in pooled active groups and 13.3% with placebo developed diabetes. The relative comparison is large, but it was not zero risk and does not prove lifetime prevention.
What happens when treatment stops?
Randomized withdrawal evidence shows average weight regain after stopping, though amount varies. Continued support and an individualized maintenance plan are important; medication changes should be discussed with the treating clinician.
Can SURMOUNT-1 tell me which weight-loss medicine is best?
Not by itself. It was placebo-controlled, not a head-to-head comparison with every alternative. Cross-trial averages are confounded by different populations, durations, estimands, and care programs.