Evidence explainer

Evidence and research methods

Two Ways to Count a Trial: Intention-to-Treat and Per-Protocol

The same randomized trial can be counted two ways, and the two counts answer different questions. Knowing which one you are reading, and why, settles many arguments about what a trial actually showed.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. The rule that feels wrong at first
  3. What per-protocol is reaching for
  4. A lesson from the placebo arm
  5. Where the caution flips
  6. Reading the gap between the two numbers

Picture a trial that has already finished. The randomization worked, the follow-up is complete, and now someone has to decide who counts. That single choice, made after the data are in, produces two very different numbers from the same study. Intention-to-treat counts every participant in the group they were assigned to, no matter what they did afterward. Per-protocol counts only the people who took the treatment the way the protocol intended. Both are legitimate. They just answer different questions, and reading a trial well starts with knowing which question you are looking at. For anything about your own care, talk with a clinician who knows your history.

Key points#

The rule that feels wrong at first#

Intention-to-treat has one instruction: once randomized, always analyzed. If someone was assigned to the drug and never swallowed a tablet, they stay in the drug group. If they crossed over to the other arm halfway through, they stay where they started. Newcomers to trial methods almost always push back here. Why give a treatment credit for a person who refused it?

The answer is that randomization is the only step in a trial that makes the two groups genuinely comparable, across everything measured and everything nobody thought to measure, before treatment begins. That balance is easy to lose. The instant you drop or move participants based on something that happened after randomization, you are sorting people by their own later behavior, and the two groups you compare are no longer the two groups randomization built. Intention-to-treat guards the single feature that separates a randomized trial from a carefully dressed observational one.

There is a practical reason too. In real clinics, patients miss doses, stop early, and change their minds. A treatment with excellent biology that most people cannot stay on is worth less than its mechanism suggests. By folding actual adherence into the estimate, intention-to-treat answers the question a health system truly faces: if we offer this to a population, what happens on average to the people we offer it to? That is an effectiveness question, and it is usually the one that matters for a decision.

What per-protocol is reaching for#

Per-protocol asks something narrower and, at first glance, sensible. Among the people who actually took the treatment as designed, how well did it work? This edges toward a pure efficacy question, the effect of a drug when the drug is truly in the body. Clinicians care about this. To learn whether a molecule does what its mechanism promises, the noise of missed doses is exactly what you might want to remove.

The problem is the method of removal. To build a per-protocol population you exclude the people who did not adhere, and adherence is never random. People who stay on a treatment tend to differ from people who stop, and they differ in ways that also shape outcomes. They may be healthier, better supported, or simply spared the side effects that drive others to quit. Keep the adherent and discard the rest, and you are no longer comparing like with like. You have rebuilt a selected, observational-style comparison inside a randomized trial, and the protection is gone.

A frequent misreading is to treat non-adherence as random noise sprinkled evenly across both arms. It rarely is. If the active drug causes more side effects, the people who leave the treatment arm are systematically unlike the people who leave the placebo arm, and the remainers you compare are mismatched in a direction you cannot observe.

A lesson from the placebo arm#

The clearest demonstration of this trap comes from a large cardiovascular trial reported in the New England Journal of Medicine in 1980. Among participants assigned to placebo, those who took their placebo faithfully had markedly lower mortality than those who did not. No one believes the placebo saved lives. Adherence was standing in for a whole cluster of healthier behaviors and circumstances that had nothing to do with the pill.

The lesson generalizes. Adherent people tend to do better for reasons unrelated to the treatment, so any analysis that conditions on adherence inherits that head start. Apply this to a genuine drug and per-protocol will usually push the estimate away from the null, making the treatment look stronger than the policy effect implies. Sometimes that larger figure is closer to the true biological efficacy. Often it is closer to what we hoped to find. The per-protocol estimate alone cannot tell you which, and that ambiguity is the whole problem.

Where the caution flips#

There is one important reversal, and appraisers who miss it get the direction of skepticism backward. In a non-inferiority trial, where the goal is to show a new treatment is not meaningfully worse than an established one, intention-to-treat stops being the strict analysis and becomes the lenient one. Non-adherence blurs the arms together and drags both toward each other, which makes any two treatments look more alike. Since looking alike is what a non-inferiority trial is trying to demonstrate, that blurring makes it easier to declare success. A systematic review of antibiotic non-inferiority trials makes the point concrete: careful trialists report both analyses and lean on per-protocol as the conservative check, precisely because intention-to-treat can flatter the weaker conclusion here. The safer analysis depends on the claim being tested.

Reading the gap between the two numbers#

A well-reported trial hands you both figures: intention-to-treat as the primary result and per-protocol as a sensitivity analysis. Treat the distance between them as information rather than an inconvenience.

When they agree, confidence rises, because the result does not hinge on who happened to stay in the study. When per-protocol runs well above intention-to-treat, resist the pull of the bigger number. That gap is usually telling you that dropout carried information, or that the treatment helps the people who can stay on it more than it helps everyone offered it. That is a real and useful message, but it is a message about selection, not permission to upgrade the effect.

Three habits keep a reader honest. Check how non-adherence and crossover were handled and how many participants were affected, since a per-protocol analysis that throws out a large share of people deserves suspicion. Ask whether dropout differed between the arms, because asymmetry is where the bias hides. And match the analysis to the claim: for the practical question of whether to adopt a treatment, intention-to-treat is the anchor, and per-protocol is a lens you hold beside it, never one you swap it for.

Sources and further reading

  1. Coronary Drug Project, adherence and mortality (NEJM 1980)
  2. Intention-to-treat, as-treated, and per-protocol approaches to analysis (PMC)
  3. ITT versus per-protocol in antibiotic non-inferiority trials, systematic review (PMC)

Questions and answers

Is intention-to-treat always the right primary analysis?

For superiority trials, it is usually the primary result, because it preserves randomization and estimates the real-world effect of offering a treatment. For non-inferiority trials the picture is more nuanced, and both analyses should be reported and should agree.

Does per-protocol overestimate or underestimate the effect?

In a superiority setting it typically pushes the estimate away from the null, making a treatment look stronger, because adherent participants tend to fare better for reasons unrelated to the drug. That is why a large gap above intention-to-treat should raise questions rather than excitement.

What should I do when the two analyses disagree a lot?

Read the disagreement as a signal. Look at how much dropout and crossover occurred, whether it was balanced across arms, and how the participants were excluded. A big divergence usually points to informative non-adherence rather than a hidden, larger true effect.