Multiplicity arises when a trial has more than one opportunity to make a confirmatory claim. Testing several endpoints, doses, groups, time points, subgroups, or interim data cuts can turn ordinary random variation into an apparently persuasive result unless the analysis controls the combined false-positive risk.
Key points#
- A 5% type I error rate applies to one prespecified test, not automatically to a collection of tests.
- Familywise error control limits the chance of at least one false rejection across a defined family of hypotheses.
- Bonferroni, Holm, fixed-sequence, gatekeeping, and graphical procedures allocate or recycle alpha in different ways.
- Interim efficacy analyses need boundaries or alpha-spending rules because every look adds a chance to stop for an exaggerated result.
- A small p value is confirmatory only if the hypothesis had a valid place in the prespecified testing plan.
Why repeated opportunities change the probability#
For one test conducted at alpha 0.05, the long-run chance of a false positive is 5% when the null hypothesis and model assumptions hold. If 20 null hypotheses are tested separately at 0.05 and the tests are independent, the chance of at least one false positive is:
1 - 0.95^20, which is about 64%.
Real trial tests are often correlated, so that calculation is only an illustration. The principle remains: more opportunities to select a favorable result raise the chance of doing so by luck.
The relevant error rate for a confirmatory family is often the familywise error rate, the probability of at least one false claim in that family. Controlling it at 5% preserves the intended evidentiary standard across the collection rather than pretending each test exists alone.
Defining the family is a scientific decision#
Not every analysis in a paper must belong to one enormous family. The family should correspond to the set of hypotheses that can support a particular confirmatory claim.
For example, a trial may have two primary endpoints, three treatment doses compared with one control, and several key secondary endpoints. The protocol is where you find out which findings can establish efficacy, which can support additional claims, and how alpha moves among them. Exploratory analyses can be presented without confirmatory error control if they are labeled honestly and not promoted as settled evidence. Ambiguity about the family creates room for selective interpretation: if authors say there was only one primary endpoint but later treat a favorable secondary endpoint as an equal route to success, the effective claim family was larger than advertised.
Several ways to control the budget#
Alpha is often described as a budget. The metaphor is useful as long as it is not taken literally: procedures work through mathematical rules that control error under stated conditions.
Bonferroni and Holm procedures#
Bonferroni divides alpha across hypotheses. With five equally weighted tests and an overall 0.05 target, each test would use 0.01; it is simple and valid under broad dependence structures, but can be conservative when tests are correlated.
Holm's procedure orders p values and compares them with progressively less stringent thresholds. It controls familywise error and is at least as powerful as ordinary Bonferroni. Both methods can make sense when hypotheses have no natural hierarchy.
Fixed-sequence testing#
A fixed-sequence procedure ranks hypotheses before the data are examined. The first can be tested at the full alpha, and if it succeeds the next can use the full alpha too, continuing until a test fails. Hypotheses after that first failure cannot support confirmatory claims through the sequence.
This approach can be efficient when the order reflects clinical importance and expected power. Its consequence must be respected. A favorable secondary endpoint below a failed primary endpoint remains descriptive unless another prespecified path allocated it alpha.
Gatekeeping and graphical procedures#
Gatekeeping methods place hypotheses into families, often requiring success in a primary family before alpha reaches secondary outcomes, and graphical procedures display each hypothesis as a node with an initial alpha allocation and prespecified rules for transferring alpha after rejection. These approaches can handle complex programs, including multiple doses and endpoints, while keeping the logic auditable, but the complexity is justified only if the protocol and report explain it clearly enough for you to reconstruct which claims passed.
Co-primary endpoints can mean two different things#
Sometimes all co-primary endpoints must succeed to establish benefit. Requiring success on every endpoint does not inflate the type I error in the same way as allowing any one endpoint to establish success, although power may fall because one failure defeats the claim.
In other trials, success on any of several primary endpoints is sufficient. That creates multiple routes to a positive conclusion and requires adjustment. The phrase “co-primary” does not tell you which rule applies. The success criterion must be stated.
Multiple groups share information and error#
A multi-arm trial may compare several treatments or doses with one control group. The comparisons are correlated because they share control participants. Testing each at 0.05 without adjustment increases the chance of at least one false claim.
Bonferroni remains available, but methods that account for correlation can be more efficient. The appropriate procedure depends on whether every comparison can support a separate claim, whether doses are ordered, whether arms may be dropped, and whether the hypotheses were specified before enrollment.
Adding an arm after a trial begins requires special care. Preserving error control depends on timing, information already observed, and the adaptation rules. A new comparison cannot simply be inserted into the original plan as if it had always been there.
Interim looks spend alpha across time#
An interim analysis examines accumulating trial data before the planned final analysis. It may support early stopping for clear benefit, harm, or futility. Repeated efficacy looks at the ordinary 0.05 threshold inflate the false-positive rate.
Group-sequential boundaries solve this by setting stricter early thresholds and an adjusted final threshold. O'Brien-Fleming-type designs require very strong early evidence and leave much of the error budget for the final analysis. Pocock-type designs use more similar thresholds across looks.
Alpha-spending functions express the amount of type I error that may be used by a given information fraction; they allow some flexibility if analyses occur at different calendar times, provided information time and the total spending rule are respected.
The exact schedule belongs in the protocol or statistical analysis plan. A statement that a trial stopped “because the result was significant” is not enough. You need the planned boundary, the actual information fraction, the stopping recommendation, and the final adjusted estimate.
Early stopping can also exaggerate effect size, especially when few events have accumulated. Crossing a valid boundary controls false-positive probability, but it does not guarantee that the observed magnitude is stable. Follow-up and uncertainty remain important.
Adjusted p values and uncertainty intervals#
A multiplicity procedure should be visible in the results. Authors may report the local alpha threshold, multiplicity-adjusted p values, simultaneous confidence intervals, or the decision path through a testing graph.
An unadjusted 95% confidence interval for each of many endpoints does not provide 95% simultaneous coverage across the family, and if a paper emphasizes several confirmatory effects, intervals aligned with the multiplicity plan can help you see which magnitudes remain plausible after adjustment.
Failure to reject a hypothesis within a hierarchy is not proof of no effect. It means the prespecified procedure did not support a confirmatory claim. Descriptive estimates can still guide future research if they are reported with appropriate restraint.
A reader's multiplicity audit#
Ask these questions before accepting the highlighted result:
- How many primary endpoints, treatment comparisons, time points, and interim looks were planned?
- What exact collection of hypotheses could establish success?
- Was the testing strategy documented before unblinded results were known?
- Which procedure controlled the familywise error rate?
- If testing was sequential, did an earlier failure close the path?
- Were treatment groups added, dropped, or pooled, and under what rules?
- Did subgroup analyses receive alpha or remain exploratory?
- Were adjusted p values or compatible intervals reported?
- If the trial stopped early, was the prespecified boundary crossed?
- Does the abstract distinguish confirmatory findings from descriptive ones?
Multiplicity does not make secondary endpoints or subgroup analyses worthless. It determines what level of claim the evidence can support. Transparent planning preserves the difference between a result that happened to be noticed and a hypothesis that passed a fair test.
Sources and further reading
Questions and answers
Is Bonferroni always required?
No. It is one valid option. Fixed sequences, Holm procedures, gatekeeping, graphical methods, and designs that use correlation can control error more efficiently when their assumptions and rules fit the trial.
Must every secondary endpoint be adjusted?
Only endpoints intended to support confirmatory claims need a valid place in the relevant error-control plan. Other endpoints may be reported as exploratory, with that status made clear.
Does a p value below 0.05 remain significant after an interim look?
Not necessarily. The result must cross the boundary assigned to that look. Early efficacy thresholds are often more stringent than 0.05.