A crossover trial compares treatments within the same person. Participants receive two or more interventions in a randomized sequence, such as treatment A followed by B or B followed by A. This removes much between-person variation and can improve precision, but only if time, carryover, and missing periods do not undermine the comparison.
Reconstruct the sequence first#
The common two-period, two-treatment design has at least two sequences: one group receives A in period one and B in period two, and the other receives B first and A second. Randomizing sequence helps separate treatment from calendar order.
Each participant should contribute an outcome under both treatments. The central contrast is the difference between that person's outcome on A and their outcome on B. Those within-person differences are then combined across participants. This differs from a parallel trial, where the main contrast is between separate groups of people.
More complex designs can repeat sequences, include more treatments, or use incomplete blocks. Before you interpret a result, draw the periods, the treatment order, the washouts, the measurement windows, and the number of participants contributing to each cell. A crossover label without a clear sequence diagram leaves too much hidden.
Why the design can be efficient#
People differ in baseline symptom burden, physiology, genetics, behavior, and many other features. In a parallel trial, those differences add variability to the treatment comparison. A crossover design controls stable personal characteristics by comparing each person with themselves.
The relevant sample-size input is therefore within-person variability, not only the variability between people, and when repeated measurements in one person are highly correlated, a crossover trial may estimate a treatment difference with fewer participants than a parallel design.
That efficiency is conditional. If outcomes fluctuate substantially within a person, measurement is unreliable, or the disease changes across periods, the expected advantage shrinks. A small headcount is not automatically excused because you saw the word “crossover” in the methods.
The condition and intervention must fit#
The target condition should be sufficiently stable across the study. Stable does not mean symptom-free or perfectly constant. It means that any background change can be separated from the treatment contrast and does not make the first and second periods fundamentally different.
The intervention should act within each period and have effects that reverse after it stops. Short-acting symptom treatments can fit this structure. A curative operation, vaccine, psychotherapy that teaches durable skills, or medicine that permanently changes disease course usually does not.
Stopping and restarting must also be ethically and clinically acceptable. A design that requires withdrawal of an essential therapy may be inappropriate even if its pharmacology looks convenient. The methods has to say why the condition, treatment, period length, and outcome together support crossover.
Carryover is the defining threat#
Carryover occurs when the effect of treatment in one period persists into the next. Residual drug can be present, but carryover can also reflect longer biological changes, withdrawal, behavior learned during treatment, or delayed adverse effects.
A washout interval aims to let the prior effect fade before the next measurement period. Its duration should be justified using pharmacokinetics and pharmacodynamics, not chosen as a round number. Several half-lives may reduce drug concentration while a downstream biological effect persists longer. Conversely, a short-acting intervention may not need a long gap.
The analysis should not rely on a statistical test that asks whether carryover is “significant.” Such tests have low power, and selecting an analysis based on their result can introduce bias. In the standard two-period design, treatment-by-period interactions are difficult to distinguish cleanly from carryover. Prevention, clinical justification, and sensitivity analysis are more credible than a claim that carryover was proved absent. And if meaningful carryover cannot be ruled out, analyzing only first-period data turns the study into a small parallel trial and discards the reason for choosing crossover; that fallback may be informative as a sensitivity check, but it is not a routine repair.
Period and sequence can distort the contrast#
A period effect means outcomes differ systematically between earlier and later periods regardless of treatment. Seasonal symptoms, disease progression, adaptation to study procedures, or improved measurement technique can create this pattern. Because both treatments occur in both periods across randomized sequences, a suitable model can estimate an average treatment effect while accounting for period.
A sequence effect means outcomes differ according to whether A or B came first; it may reflect carryover, dropout, baseline imbalance by chance, or an interaction between treatment order and behavior. Sequence should be reported and examined, but a small trial may have little power to interpret it. What protects the comparison is the randomized order plus a prespecified model: pooling every observation on A against every observation on B, without period and participant structure, can hand you misleading precision.
Paired analysis is necessary#
The simplest analysis calculates each participant's difference between treatments and then analyzes those differences. Regression or mixed-effects models can handle period, treatment, sequence, repeated observations, and additional design features. Whatever the method, it should recognize that measurements from the same person are correlated.
Treating period observations as if they came from unrelated groups usually wastes the efficiency of pairing and may misstate uncertainty, and a report should give you the within-person treatment contrast with a confidence interval, not only separate means within each treatment period.
Baseline deserves care. A single baseline before the entire trial does not capture changes before later periods. Period-specific baseline measurements can help assess stability, but adjusting for post-randomization measurements requires a prespecified rationale. The report should distinguish baseline from outcomes and avoid using an affected measurement as if it were unaffected.
Missing periods can be especially damaging#
Dropout is not merely a smaller sample, and if a participant stops after adverse effects on the first treatment, the outcome under the second treatment is missing for a reason related to treatment and sequence. Complete-case analysis then keeps only people who tolerated and completed both periods.
Reports should show participant flow by sequence and period, with reasons for discontinuation. The estimand should state how intercurrent events, rescue treatment, nonadherence, and missing outcomes are handled. Sensitivity analyses should examine plausible departures from the missing-data assumptions. Repeated crossover designs can face a further problem, which is participant fatigue and declining adherence across periods; more cycles increase information only if the later measurements remain reliable and representative.
Blinding and outcomes still matter#
Using each participant as their own control does not remove expectation effects. If treatments look or feel different, participants may infer the period assignment. Assessors can also be influenced. Blinding, matching placebos, and objective outcomes remain valuable where feasible.
Outcome timing should align with onset and offset of treatment. Measuring too early may miss benefit; measuring too late can mix treatment with carryover or background change. Multiple daily measurements may provide rich data, but their correlation and prespecified summary still need a coherent analysis. Harms require full reporting too, because an average symptom improvement can hide the fact that one sequence caused early withdrawals, or that adverse effects accumulated over repeated periods.
Read the report against current guidance#
The crossover extension to CONSORT specifies information unique to this design, while the general CONSORT statement was updated in 2025. The title should identify the crossover design. The report should describe sequence generation and concealment, period and washout durations, eligibility rationale, sample-size assumptions, participant flow by period and sequence, and the paired analysis. The results should state how many participants received and completed each treatment in each period, with treatment, period, and sequence effects reported according to the prespecified plan and any carryover concern discussed using clinical and design evidence.
A six-question appraisal#
Ask whether the condition is stable, whether treatment effects reverse, whether the washout is biologically adequate, whether sequence was truly randomized and concealed, whether the analysis is paired, and whether missing periods could depend on earlier treatment.
Then interpret the target. A crossover trial often estimates a short-term average effect among people able to complete repeated periods. It may not show durability, long-term harms, or feasibility in a broader population. Its strength is a precise within-person comparison, not universal scope.
Sources and further reading
Questions and answers
Does every crossover trial need a washout period?
It needs a credible strategy to prevent prior treatment from affecting the next measurement. That may be a formal washout, a run-in interval, or evidence that the effect ends before the next period. The rationale should be explicit.
Can researchers test for carryover and proceed if the test is negative?
A negative carryover test does not prove absence, especially in a small two-period trial. Design and biological reasoning should address carryover before the data are analyzed.
Is a crossover result automatically personalized medicine?
No. It estimates an average within-person effect across enrolled participants. It uses personal comparisons, but the reported treatment effect still summarizes a study population unless results are analyzed as individual N-of-1 trials.