Evidence explainer

Diabetes and metabolic health

Bioequivalence for Narrow Therapeutic Index Drugs: How Tight Is Tight Enough

Most generics must land the 90 percent confidence interval for their test-to-reference blood-level ratio inside 80.00 to 125.00 percent. Narrow therapeutic index drugs are held tighter still.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. The 80 to 125 window, and what it is not
  3. Where a one-fifth swing turns dangerous
  4. One scaling method, running in both directions
  5. The second hurdle: matching consistency, not just the average
  6. Reading a substitution with this in mind

Not every generic is judged against the same yardstick, and knowing that changes how you read a substitution. Most oral generics must prove that the 90 percent confidence interval for the test-to-reference ratio of blood levels lands entirely inside 80.00 to 125.00 percent. Narrow therapeutic index drugs, where a small change in blood concentration can flip a patient between working therapy and harm, are held to a stricter, reference-scaled target. Highly variable drugs, whose levels swing widely on their own, are allowed a controlled wider one. The FDA finalized all three of these paths in its May 2026 guidance.

Key points#

The 80 to 125 window, and what it is not#

You will often see the familiar range misread as a promise that a generic can run 20 to 25 percent off the brand in any given person. It says nothing of the kind. The rule governs the 90 percent confidence interval around the ratio of average blood levels across a study group, usually measured as peak concentration and area under the concentration-time curve. The whole interval, not just its center, has to fall inside the limits. Because of that, marketed generics tend to cluster far nearer the reference than the boundaries imply: a product whose true average drifts toward an edge produces an interval that pokes past it and fails.

The window also carries a built-in judgment, that a difference of roughly a fifth in average blood level does not change outcomes for most drugs. The lopsided look of the numbers (80 on the low side, 125 on the high side) is an illusion of scale. On a logarithmic axis the two limits are exact mirror images, which is why blood-level data are log-transformed before the statistics are run.

Where a one-fifth swing turns dangerous#

For some drugs, that same one-fifth difference is the margin between control and crisis. A narrow therapeutic index drug is one in which small shifts in dose or blood concentration can tip a patient toward treatment failure or serious toxicity. The classic teaching set includes warfarin, levothyroxine, lithium, tacrolimus, and several antiseizure medications, and the same worry surfaces with some agents used in chronic metabolic and endocrine care, which is one reason the topic matters in everyday primary care. For these drugs the assumption behind the standard window collapses. A 20 percent change is no longer a rounding error; it can be the difference between a therapeutic and a toxic level.

The FDA's answer, described in detail by Donnelly and colleagues in Clinical Pharmacology and Therapeutics in 2025, is a fully replicated, four-period crossover study. Each healthy volunteer takes the reference product twice and the test product twice. That repetition buys two things an ordinary two-period study cannot. It reads the reference drug's own dose-to-dose variability directly, and it lets the same subjects show whether the test product is any more variable than the brand.

One scaling method, running in both directions#

Reference-scaled average bioequivalence pins the acceptance limits to how consistent the reference product is, captured as its within-subject standard deviation. The logic then cuts two ways depending on the drug.

When the reference is very steady from dose to dose, the limits pull inward from 80 to 125 and demand a closer match. This is the narrow index case. The FDA also sets a firm ceiling here: the scaled limits are never allowed to exceed the conventional 80.00 to 125.00 percent, and that cap takes over once the reference within-subject standard deviation climbs to about 0.21, the point where scaling would otherwise loosen the window. Below that, a low-variability narrow index drug faces a genuinely tighter target. The worked analysis by Jiang and colleagues in the AAPS Journal shows the effective band contracting toward roughly 90 to 111 percent when the reference is highly reproducible.

Highly variable drugs pose the mirror-image problem. Defined by a within-subject variability around 30 percent or more, their levels bounce so much that the reference product compared against itself can miss the standard window by chance, which would leave some truly equivalent generics unapprovable. Here scaling widens the limits in step with the reference's variability, up to a capped maximum, so the criterion tracks the drug's built-in noise instead of punishing it. The 2026 final guidance folds the statistics for both the narrow index and highly variable cases into a single document, retiring the 2001 guidance and finalizing the December 2022 draft.

The second hurdle: matching consistency, not just the average#

Scaling by itself does not finish the job for a narrow index drug, because a test product could hit the average dead-on yet swing more wildly within individual patients. So the FDA adds a variability comparison. It is an explicit statistical test that the generic is not meaningfully less consistent than the brand: the upper bound of the 90 percent confidence interval for the ratio of test-to-reference within-subject standard deviations must stay under a fixed constant. A generic narrow index drug therefore has to clear two bars in one study, matching both the average blood level and the steadiness of that level. An ordinary generic is never asked to prove the second.

Reading a substitution with this in mind#

The practical message is reassuring and slightly counterintuitive: generics are not all held to one number, and the standard is strictest precisely where the stakes are highest. A generic narrow index drug that reaches your pharmacy has been studied in a replicated design and has met both a scaled average limit and a consistency limit that no ordinary generic faces. That is a stronger body of evidence than the headline 80 to 125 figure suggests. It does not settle every clinical question, since formulation excipients, your own response, and the setting of a switch still matter, but the regulatory framework already builds in the principle that tight drugs deserve tight limits.

Sources and further reading

  1. FDA final guidance, Statistical Approaches to Establishing Bioequivalence (Federal Register, May 29, 2026)
  2. Donnelly et al., Narrow Therapeutic Index Drugs: FDA Experience, Views, and Operations, Clin Pharmacol Ther 2025
  3. Jiang et al., A Bioequivalence Approach for Generic Narrow Therapeutic Index Drugs (PMC4476992)
  4. GovInfo record, 91 FR 32056, FDA guidance availability notice 2026-10705

Questions and answers

Does a generic being bioequivalent mean patients will not notice a switch?

For most drugs, bioequivalence is strong reassurance that average blood levels match closely. For narrow index drugs the added scaled and consistency tests raise that confidence further. Even so, individual factors and formulation differences can matter, so any concern about a specific switch is worth raising with the prescriber and pharmacist.

Why are the limits widened for highly variable drugs rather than kept the same?

Because their blood levels swing so much on their own that even the brand compared against itself could fail the fixed window by chance. Scaling the limits to that intrinsic variability keeps the test fair without lowering the real standard of equivalence.

Are levothyroxine and warfarin really held to a tighter standard?

Yes. As narrow therapeutic index drugs, their generics are evaluated with the replicated, reference-scaled approach and the added consistency test, a stricter path than the conventional 80.00 to 125.00 percent criterion used for most medicines.