Evidence explainer

Evidence and research methods

The Fragility Index, Explained

Changing one participant's outcome can move some trial results across P equals 0.05. The fragility index counts how many changes it would take. That is all it counts.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The calculation in plain language
  2. Why the index became popular
  3. A worked thought experiment
  4. A threshold-sensitivity measure, not a truth meter
  5. Why sample size and event frequency matter
  6. The loss-to-follow-up comparison
  7. The classic index has a narrow domain
  8. Significant and nonsignificant results
  9. The index does not measure clinical importance
  10. It also does not measure risk of bias
  11. Better questions to ask alongside the index
  12. The appropriately fragile conclusion
  13. References

Two randomized trials can report the same P value and still feel very different. One result may depend on a few outcome events among 120 participants. Another may rest on hundreds of events among thousands. The fragility index translates one aspect of that difference into a count.

For a conventional two-group randomized trial with a binary outcome, the fragility index is the smallest number of participant outcome statuses that must change for a statistically significant result to become nonsignificant, usually using a two-sided Fisher exact test and a 0.05 threshold.

If the fragility index is 1, one hypothetical event-status change crosses that line. If it is 20, twenty changes are required under the specified calculation. The count is easy to communicate. Its meaning is much narrower than words such as “robust,” “reliable,” or “true.”

The calculation in plain language#

Start with a two-by-two table. Its rows are treatment groups. Its columns are participants with and without the binary outcome. The published comparison is statistically significant according to a chosen test.

The classic procedure then:

  1. changes one participant's status in the direction that reduces the between-group contrast;
  2. recalculates the two-sided Fisher exact test;
  3. repeats the change until the P value no longer falls below 0.05;
  4. reports the minimum number of changes required.

The procedure is hypothetical. It does not claim that researchers miscoded those participants or that their outcomes actually differed; it asks how far the observed event table lies from a specified statistical boundary.

Different algorithms can change outcomes in different arms or use different tests. A report should state the definition, direction of modification, test, and threshold. An unlabeled number is not fully reproducible.

The P value is not expressed in participants. A statement such as P equals 0.047 can be mistaken for strong evidence because it sits on one side of a conventional line, though the fragility index reveals that the same finding may be one outcome change away from P equals 0.05.

In 2014, Michael Walsh and colleagues calculated the index for 399 randomized trials with statistically significant binary outcomes: the median sample size was 682, the median number of events was 112, and the median fragility index was 8. One quarter of results had an index of 3 or less. Among trials that clearly reported follow-up losses, 53% had an index smaller than the number lost. Those findings did not prove that half the trials were wrong. They showed that the conventional significant-versus-nonsignificant label could hinge on fewer outcomes than many readers expected.

A worked thought experiment#

Imagine a trial comparing a new intervention with usual care for hospitalization by 90 days, and the report gives event counts for both groups and a two-sided P value below 0.05.

An analyst changes one non-event in the better-performing group into an event, then recalculates Fisher's exact test. The P value remains below 0.05. The analyst repeats the process. After the fourth total change, the P value rises above the threshold. Under that defined procedure, the fragility index is 4.

What can you conclude?

The binary significance classification is four outcome changes from its boundary. That is all the index establishes. It does not establish that four records are wrong, that four missing participants had events, or that the treatment effect is unimportant; the risk difference and confidence interval might still support a meaningful effect, or they might show wide uncertainty. Those quantities must be read directly.

A threshold-sensitivity measure, not a truth meter#

The fragility index inherits the binary logic it is designed to illustrate. It asks when P crosses alpha. The American Statistical Association warns that scientific conclusions should not rest only on whether a P value passes a particular threshold.

A result at P equals 0.049 and another at P equals 0.051 are not scientifically opposites. Their data can be nearly identical. An index that counts the path between those categories makes the discontinuity visible, but it also keeps the discontinuity at the center.

The index does not provide:

A trial with a large index can still have biased outcome measurement or selective reporting. A trial with a small index can be rigorously conducted and appropriately interpreted as uncertain.

Why sample size and event frequency matter#

Large trials generally have more opportunity to accumulate outcome differences; all else equal, changing a fixed number of events has less influence on a large event table than on a small one. Rare outcomes often produce small counts even in moderately large samples.

The fragility quotient divides the index by a denominator, usually total sample size. It expresses the changes as a proportion. An index of 5 in 100 participants and an index of 5 in 10,000 participants then no longer appear equivalent.

The quotient does not solve the central limitations. It still depends on the significance test, event rate, and threshold. There is no accepted cutoff that separates fragile from robust. Both index and quotient should be treated as descriptive sensitivity measures. Some authors also divide by total events rather than sample size. Because competing denominators answer different questions, a paper must name which quotient it reports.

The loss-to-follow-up comparison#

You will often see the fragility index set beside the number of participants lost to follow-up, and if 12 outcomes are unknown and the index is 3, it is natural to worry that the missing data alone could cross the threshold.

That comparison is a warning signal, not a reanalysis, and it assumes only a count, while bias depends on which group lost participants, why outcomes are missing, and what those outcomes would have been. All missing participants need not move in the unfavorable direction, and changing an observed non-event to an event is not the same process as resolving an unknown outcome.

A better assessment examines:

The index can motivate this audit. It cannot replace it.

The classic index has a narrow domain#

The original calculation is best suited to a two-group randomized trial, a dichotomous outcome, and a conventional hypothesis test on a simple event table.

Continuous outcomes#

Pain scores, blood pressure, and quality-of-life scales contain more information than a yes-or-no endpoint. Dichotomizing them discards information and makes the index depend on an extra cutoff. Proposed continuous fragility measures exist, but they are not the classic index and require additional assumptions.

Time-to-event outcomes#

Survival analyses use whether and when events occur, account for censoring, and may estimate hazard ratios over follow-up. A simple two-by-two table ignores timing. Extensions for time-to-event data exist, but applying the classic count to final event totals can contradict the primary analysis.

Adjusted analyses#

Many trials prespecify regression models that adjust for stratification factors or strong prognostic variables, and replacing that analysis with an unadjusted Fisher exact test can produce a fragility index of zero even when the reported adjusted result is below 0.05. The zero may reflect a different test rather than a changed outcome.

Cluster, crossover, and noninferiority designs#

These trials have dependence structures or hypotheses that a simple event table does not preserve. A conventional fragility calculation may be invalid or answer a different question.

Observational studies#

Changing event counts does not address confounding, selection bias, or model specification. A small index adds little, while a large one cannot make a nonrandomized association causal.

Significant and nonsignificant results#

The classic index was defined for significant results. A “reverse fragility index” or fragility index for nonsignificant findings counts outcome changes required to cross into significance.

That extension may be useful for describing threshold proximity, but it can encourage a harmful interpretation: that a nonsignificant result is nearly positive. The primary question should be whether the confidence interval excludes effects that matter, not how many changes would produce a favored label.

For a noninferiority trial, the relevant boundary is a prespecified noninferiority margin, not necessarily a null effect. Any fragility method must align with the actual hypothesis and estimator. A generic calculator may not.

The index does not measure clinical importance#

Statistical sensitivity and clinical importance are different axes.

A huge trial can produce a large fragility index for a tiny absolute risk difference that few patients would value. A smaller trial can have a modest index for a large, important effect that remains imprecisely estimated. Neither count communicates benefits, harms, burden, cost, or patient preferences.

For a binary outcome, look for both the relative and the absolute effect. If hospitalization falls from 20% to 15%, the relative risk, the absolute risk reduction of 5 percentage points, and the number needed to treat of 20 each tell you something different about the same finding. The confidence interval shows you the range compatible with the data and the analysis.

CONSORT 2025 recommends reporting group results, effect estimates, precision, and both absolute and relative effects for binary outcomes. Those quantities remain primary. A fragility index can be supplemental.

It also does not measure risk of bias#

Risk of bias asks whether trial design, conduct, analysis, or selective reporting could systematically move the estimate.

Cochrane's RoB 2 framework examines the randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result, and none of those domains is encoded in the fragility index.

For example:

Better questions to ask alongside the index#

When an article reports a fragility index, ask:

  1. Does the classic method fit the trial design and outcome?
  2. Was the same statistical method used as in the primary analysis?
  3. What event-status changes were allowed, and in which group?
  4. What were the total sample size and total number of events?
  5. How does the count compare with missing outcomes in each group?
  6. What are the absolute effect and confidence interval?
  7. Was the endpoint clinically important and prespecified?
  8. Do other trials and a systematic review point in the same direction?

Replication across credible studies is a stronger form of robustness than distance from one trial's P-value boundary.

The appropriately fragile conclusion#

The fragility index has one excellent teaching function: it turns an abstract P-value boundary into a participant count. That can puncture false certainty created by the word “significant.”

Its strength is also its limit. It measures how many outcome changes cross one statistical line under one calculation. It does not assess truth, bias, causal validity, clinical importance, or the durability of an evidence base.

Use it as a supplemental stress test. Then go back to the quantities and judgments that actually carry your decision: absolute and relative effects, confidence intervals, missing-data sensitivity, prespecification, risk of bias, outcome relevance, and consistency across studies.

References#

Questions and answers

Is a fragility index of 1 bad?

It means one outcome-status change crosses the chosen significance boundary, which is extreme threshold sensitivity; whether the evidence is useful still depends on effect size, precision, trial quality, outcome importance, and other studies.

What fragility index is considered robust?

There is no accepted universal cutoff. The count depends on sample size, event rate, test, allocation, and threshold, so context matters more than a fixed label.

Is the fragility index better than a confidence interval?

No. A confidence interval directly describes uncertainty around an effect estimate; the index adds an intuitive sensitivity count for a narrow class of analyses but should not replace effect estimation.

Should the index be compared with loss to follow-up?

Yes, as a prompt for closer missing-data review. It does not justify assuming that all unknown outcomes would move against the result.

Can I calculate it for a survival outcome?

The classic two-by-two method ignores event timing and censoring, so it should not replace a time-to-event analysis. Specialized extensions need their own definitions and validation.