A statistician can calculate that a randomized trial needs 700 participants to test its primary efficacy hypothesis. The final phase 3 program may enroll 3,000, span dozens of countries, and include several trials. That does not necessarily mean the calculation was wrong. It means the primary hypothesis was only one of the questions the program must answer.
ICH E9 states the principle directly: a trial should be large enough to provide a reliable answer to its questions, and a study sized for safety or an important secondary objective may need more participants than one sized for primary efficacy. Global development adds questions about whether results apply across regions, whether enough people with relevant characteristics were studied, and whether regulators can characterize common harms and longer use.
There is no universal multiplier between a power calculation and a final program. The binding constraint changes by disease, endpoint, expected effect, treatment duration, background event rate, product novelty, safety signals, and intended population. Good design identifies each constraint explicitly instead of hiding the final number behind "regulatory requirements." If you want to know why a program is the size it is, that is the list to ask for.
What the usual power calculation does#
For a primary efficacy endpoint, designers specify an effect to detect or a noninferiority margin, variability or event rate, type I error, desired power, allocation ratio, and analysis method. The calculation estimates how many evaluable participants are needed under those assumptions.
Each input is conditional. If the true event rate is lower, variability higher, treatment difference smaller, or withdrawal rate greater than expected, the study may have less information than planned. Designers therefore add an allowance for missing primary outcomes and may use blinded checks of nuisance parameters when appropriate.
The output is not "the number needed to approve a drug." It is the number needed for a specified chance of answering one prespecified statistical question. A program can meet that number while remaining too small or too narrow for safe and credible use.
A trial and a program are different units#
One confirmatory trial may test one dose against placebo for one endpoint. A regulatory program integrates dose-finding studies, confirmatory trials, active-comparator work, long-term follow-up, clinical pharmacology, special-population studies, and accumulated safety data.
Some indications can rely on one pivotal trial plus confirmatory evidence under a justified framework. Others use two adequate and well-controlled trials. Global submissions may need the same core evidence to support several regions, each with questions about standard of care, dose, comparator, and applicability.
Add the enrollment across every study and a program will look far larger than its pivotal power calculation. That arithmetic only means something if you are comparing like with like. Participants contributing to safety, dose selection, or a regional bridging question do not all contribute equally to the primary confirmatory estimate.
Safety does not use the same denominator#
An efficacy trial is commonly powered for a treatment difference that is frequent enough to observe. A serious adverse reaction occurring in 1 of 1,000 treated people is a different detection problem. Even several thousand participants may provide limited precision for rare events.
A useful approximation is the rule of three. If an event is not observed among N treated participants, the upper end of a rough 95% confidence bound for its true rate is about 3 divided by N. So when zero events occur among 300 people, that is still compatible with a rate near 1%, and zero among 3,000 is compatible with a rate near 0.1%. The absence of an event is not proof of zero risk.
ICH E1 gives general duration and participant-count principles for medicines intended for long-term treatment of non-life-threatening conditions. It is a baseline framework, not a universal ceiling. Expected widespread use, a new mechanism, concerning nonclinical findings, class effects, or a signal in early studies can justify a larger or longer safety database.
Duration can be more limiting than head count#
Three thousand people treated for four weeks do not answer the same safety question as several hundred treated for a year. Some harms emerge after cumulative use, delayed physiology, dose escalation, or interaction with changing illness. Sustained efficacy may also matter when the intended treatment is chronic.
Programs therefore track both number of participants and treatment duration. They may need enough people at the proposed dose for six or 12 months, enough withdrawals followed after stopping treatment, and enough observation to understand whether benefit persists. Those time requirements can enlarge the program, because the fast efficacy trials finish long before the long-term cohort matures, and a sponsor will often run an extension or a dedicated long-term study in parallel rather than hold every efficacy result back until it does.
Global trials must support regional interpretation#
ICH E17 addresses multiregional clinical trials conducted under one protocol across geographic regions. The aim is to produce evidence that can support regulatory decisions in multiple regions while planning for factors that may change treatment response.
Intrinsic factors include genetics, age, body size, organ function, and disease phenotype. Extrinsic factors include diet, medical practice, background therapy, adherence patterns, diagnostic criteria, and health-system access. Region itself is usually a proxy for these variables, not a biological mechanism.
The total sample may need to accommodate a scientifically justified allocation across regions. A country with only a handful of participants contributes little to understanding local use. On the other hand, requiring each region to reproduce the overall p-value would multiply enrollment dramatically and defeat the purpose of a shared trial.
ICH E17 instead emphasizes evaluation of consistency, clinically relevant differences, and explanatory factors. Allocation can consider disease prevalence, feasibility, regional population size, and the information regulators need. Statistical significance in every region is not the default definition of consistency.
Local significance is one option, not the only rule#
Regional planning strategies include proportional allocation, equal allocation, a fixed minimum, preservation of a fraction of the overall effect, or enough enrollment for a local hypothesis test. Each answers a different policy question and has different costs.
A local-significance rule is demanding because regional subgroups are smaller and interaction is noisy. A fixed minimum can secure useful descriptive information without pretending it offers a definitive local estimate. Proportional allocation may reflect future use but leave small regions with sparse data. Whichever strategy a program picks, it should be agreed before the results are known, because selecting a favorable regional rule after observing the data converts an applicability assessment into outcome-driven storytelling.
Subgroups consume information quickly#
A trial that is well powered overall may be unable to estimate effects precisely in older adults, people with kidney impairment, a minority disease subtype, or users of an important concomitant medicine. Splitting 1,000 participants into ten groups does not create ten reliable trials.
Subgroup goals should be distinguished. Representation asks whether relevant people were included. Precision asks whether their effect and safety estimates are informative. Interaction testing asks whether response truly differs. These goals need different sample sizes.
FDA's June 2024 diversity action plan document was still draft guidance at the research date, but the underlying principle is broader than that draft: the studied population should reflect those likely to use the product, and enrollment goals should be planned rather than explained after the fact. Numerical diversity alone is insufficient if recruitment concentrates participants in sites or settings unlike future care.
Several doses and comparators divide the sample#
Suppose a program needs 600 participants on the proposed dose for the primary comparison. Adding a second dose, placebo, and an active comparator changes allocation. If the goal is to compare each dose with control while controlling multiplicity, total enrollment rises.
Active comparators may be needed because placebo alone does not answer how a new treatment fits current care. Add-on designs may be necessary when withholding standard treatment is unethical. Different regions may also use different background therapies, creating stratification and interaction questions.
Dose selection is not merely an early-phase task. Phase 3 data may need to show whether a lower dose retains benefit with fewer harms, whether titration works, and whether a fixed combination adds value beyond its components.
Missing data and treatment discontinuation add uncertainty#
Designers commonly inflate enrollment for expected loss to follow-up, but simple inflation does not solve biased missingness. People may stop treatment because it failed, caused adverse effects, or became burdensome. Their missing outcomes are informative.
ICH E9(R1) asks trials to define the treatment effect of interest through an estimand, including how events such as discontinuation, rescue medicine, or death are handled. Follow-up after treatment stopping may be critical to the clinical question, and a global program can have different withdrawal patterns across regions because visit burden, transport, background care, and study expectations all differ. Extra enrollment restores some precision; retention, outcome collection, and appropriate sensitivity analyses are what protect validity.
Operational uncertainty encourages contingency#
Enrollment estimates rely on site projections that are often optimistic. Disease prevalence does not equal eligible prevalence. Competing trials, laboratory criteria, vaccination or treatment changes, seasonal patterns, and a changing standard of care can slow recruitment or reduce event rates.
Programs sometimes open more sites or countries than the efficacy calculation appears to require so that enrollment finishes on time. This can be reasonable, but too many low-enrolling sites create training, monitoring, and consistency problems. Site count is not a substitute for realistic feasibility. Adaptive sample-size reassessment or event-driven designs can respond to some uncertainty when they are prespecified and protected against bias, whereas unplanned expansion after somebody has seen the comparative outcomes can inflate false-positive risk and undermine credibility.
Manufacturing and formulation changes can add studies#
The product used early in development may differ from the commercial formulation, device, manufacturing process, or dosing schedule. Analytical comparability and pharmacokinetic bridging may be sufficient for some changes. Others require clinical data.
A device that changes administration, a long-acting formulation, or a combination product can introduce usability and error questions not answered by the original efficacy trial. Global programs may also need human-factors work, immunogenicity assessment, or data across lots. Those participants are part of the development program even when they do not affect the pivotal p-value, and counting them without saying which question they answer is what makes a program look inefficient when it is in fact meeting a separate requirement.
Larger programs can still leave evidence gaps#
A study of 20,000 people can be narrow if eligibility excludes older adults, organ impairment, polypharmacy, or the settings where treatment will be used. Conversely, a smaller pragmatic trial with broad participation may answer applicability better while offering less precision for rare harms.
Enrollment size does not repair a poor comparator, biased outcome, unblinded assessment, weak adherence measurement, or missing follow-up. Nor can preapproval trials reliably detect every rare or delayed harm, so postmarketing surveillance and further trials remain necessary. Read a program by information per participant rather than by head count alone: the added participants are justified when they answer a defined decision question that cannot be answered more efficiently with better design or with valid existing data.
An audit trail for the final number#
A transparent program can show each constraint in a table: primary efficacy, key secondary endpoints, safety frequency and duration, regional allocation, important subgroups, dose arms, withdrawal allowance, and special studies. The largest constraint sets the minimum for a trial, while other questions may be answered across the integrated program.
Assumptions should be labeled and stress-tested. What happens if event rates fall by one third? If withdrawal doubles? If one region enrolls slowly? If safety requires a longer period? Scenario planning is what lets somebody explain the number to you rather than assert it.
It also exposes unjustified excess. "More data" is not an ethical blank check, because every additional participant is a person accepting inconvenience and some risk to answer a question somebody chose to ask. Scientific necessity, feasibility, and participant protection have to stay aligned.
Sources and further reading
- ICH E9, Statistical Principles for Clinical Trials
- ICH E17, General Principles for Multiregional Clinical Trials
- ICH E1, Clinical Safety Population for Long-Term Treatment
- ICH E5 R1, Ethnic Factors in Acceptability of Foreign Clinical Data
- ICH E8 R1, General Considerations for Clinical Studies
- FDA draft guidance on diversity action plans, June 2024
Questions and answers
Is the power calculation the minimum number needed for approval?
No. It estimates the number needed for a defined statistical question under assumptions. Safety, duration, regional relevance, subgroups, doses, and other regulatory questions may require more.
Must every country show a significant treatment effect?
No. ICH E17 supports evaluating regional consistency and clinically meaningful differences. Local statistical significance is one allocation strategy, not a universal requirement.
Why can safety require more participants than efficacy?
Clinically important harms may be much less common than the efficacy endpoint. Larger and longer treated cohorts improve the chance of observing them and narrow uncertainty about their rate.
Does adding more sites always speed a global trial?
No. More capable sites can improve recruitment, but many low-volume sites add training, monitoring, and variation. Feasibility and data quality matter as much as site count.
Is a larger trial automatically more trustworthy?
No. Size improves precision but cannot fix biased design, inappropriate comparators, poor follow-up, unreliable outcomes, or an unrepresentative population.