A cluster randomized trial randomizes intact groups, such as clinics, practices, schools, hospital wards, or communities. People within a group usually resemble one another and share staff or local conditions, so they do not contribute fully separate pieces of information. The central appraisal task is to trace the cluster structure from randomization through recruitment, analysis, and interpretation.
Draw the trial before reading its p-value#
Start by identifying three units. The unit of randomization is what enters the allocation process. The unit of intervention is what receives the program or policy. The unit of observation is where outcomes are measured. In a clinic trial, practices may be randomized, clinicians may deliver the intervention, and outcomes may be recorded for patients.
Those units can differ without creating a flaw. The problem arises when a report describes thousands of patient observations but gives little information about the much smaller number of randomized clinics. A ten-clinic trial remains a ten-cluster randomized experiment even if each clinic contributes a thousand records.
A simple diagram helps. List the clusters by arm, show when individuals were identified, and mark where outcomes were measured. That will tell you whether the unit named in the title is the unit that actually generated the comparison.
Why randomize clusters at all?#
Some interventions exist only at group level, and a new electronic reminder, staffing model, school meal program, or public-health campaign cannot be assigned cleanly to one person while withheld from a neighbor in the same system.
Contamination offers a second reason. If clinicians are taught a communication method, they may use it with all patients, and individual randomization within the same practice would blur the contrast, whereas group allocation can preserve a meaningful difference between arms.
Logistics can also make cluster delivery more feasible, but convenience alone is not enough. Cluster randomization sacrifices statistical efficiency and introduces recruitment and analysis challenges. Authors should explain why those costs were justified for the intervention and question.
Similar participants mean less information#
Patients in one clinic may share geography, referral patterns, socioeconomic conditions, staff, and local practice style. Their outcomes are correlated. The intracluster correlation coefficient, or ICC, summarizes how much of the outcome variation is associated with cluster membership.
Even a small ICC can matter when clusters are large. For equal cluster size m, a common approximation for the design effect is:
1 + (m - 1) × ICC
If each cluster has 100 participants and the ICC is 0.02, the design effect is about 2.98, and a raw sample of 3,000 observations then carries roughly the precision of about 1,000 unrelated observations under the simplified assumptions. This is an illustration, not a substitute for the trial's actual calculation.
Unequal cluster sizes usually reduce efficiency further. Adding many more people to a few large clusters may contribute less information than adding new clusters; sample-size planning should state the assumed ICC, expected cluster sizes and variation, attrition at both levels, and the number of clusters needed.
Few clusters can leave important imbalance#
Randomization balances prognostic features on average over repeated trials, but a small number of clusters can produce noticeable differences by chance. One arm may contain larger hospitals, more rural practices, or populations with different baseline risk.
Stratification, restricted randomization, or matching can improve balance on selected cluster features. Those design choices must then be respected in analysis. Baseline tables should show both cluster-level and participant-level characteristics by arm. A claim that “baseline variables were not statistically different” is not enough, because a small cluster count has little power to detect meaningful imbalance. Adjustment for prespecified prognostic factors can improve precision and address chance differences, but it cannot repair every unmeasured distinction between a handful of sites, and that should temper how far you generalize.
Recruitment after allocation is a distinctive bias#
In many cluster trials, the clinic or school is randomized before individual participants are identified or consented. Recruiters may therefore know the cluster assignment. That knowledge can influence who is invited, who accepts, or how eligibility is interpreted.
Suppose staff in intervention clinics recruit patients who are especially motivated, while control clinics enroll a broader group. The resulting participant difference is created after randomization. An apparently beneficial intervention effect may partly reflect who entered each arm.
The report will tell you when participants were identified relative to randomization, who applied the eligibility criteria, whether that person knew the assignment, and whether recruitment rates differed. Using preexisting registries, objective eligibility rules, or masked recruitment can reduce the risk. An intention-to-treat label does not erase selection that happened before a person entered the analyzed cohort.
The analysis must preserve correlation#
Treating every patient as an unrelated randomized observation generally produces standard errors that are too small and confidence intervals that are too narrow. Suitable methods include cluster-level analyses, mixed-effects models, and generalized estimating equations with an appropriate correlation structure and small-sample handling.
The method should match the number and size of clusters, outcome type, design features, and target effect. With few clusters, large-sample approximations may be unreliable. Degrees-of-freedom corrections, randomization-based methods, or cluster-level summaries may be needed. The report should explain rather than merely name the software command.
Missingness can occur at two levels. An individual may lack an outcome, or an entire cluster may withdraw. Losing one of twelve clinics is not equivalent to losing one of several thousand patient records. Flow diagrams and sensitivity analyses should keep those levels separate.
Decide what effect the trial estimates#
A cluster trial may estimate the effect of offering a system strategy to a population, not the biological effect among people who fully use it. That policy effect can be exactly the relevant question, but it differs from a treatment effect under perfect adherence.
The analysis should define its estimand: the population, intervention condition, comparator, outcome, summary measure, and handling of intercurrent events. It should also state whether effects are averaged across individuals or clusters. Weighting every person equally lets large clusters dominate; weighting every cluster equally answers a different question.
Informative cluster size creates another complication. A larger clinic may differ systematically in staffing, case mix, or quality. If cluster size is related to outcome or treatment effect, the weighting approach can change the result, and you want a rationale and a sensitivity analysis here, not an assumption that there is one natural average.
Contamination does not disappear automatically#
Cluster allocation reduces spillover but may not eliminate it. Clinicians may work at multiple sites, patients may cross clinic boundaries, or a community campaign may reach control areas. Reports should describe these pathways and, when possible, measure crossover or implementation fidelity.
Contamination usually pulls arms toward each other, but its direction is not guaranteed. Awareness of the intervention can alter control behavior. Conversely, uneven implementation can make an effective program appear weak. The intervention description should make clear what each arm actually received, not only what the protocol intended.
Reporting guidance is a map, not a validity certificate#
The cluster extension to CONSORT adds design-specific items to randomized-trial reporting, while the general CONSORT statement was updated in 2025. A complete report should identify the cluster design in its title, explain the rationale, describe cluster and participant eligibility, show flow at both levels, report the ICC or other correlation information, and state how clustering entered sample-size planning and analysis.
Following a checklist improves transparency. It does not prove that the design choice, recruitment process, or model was sound. Use the report to reconstruct the trial, then judge the decisions yourself.
A compact appraisal sequence#
First count clusters, not only people. Then ask why cluster allocation was necessary, check when and how individuals were recruited, find the ICC assumptions and whether unequal cluster size was considered, and confirm that the analysis and the confidence intervals account for clustering. That leaves you cluster loss, baseline imbalance, contamination, and the estimand.
Finally, consider transportability. A precise average across twenty highly selected health systems may not apply to a small practice with different infrastructure; cluster trials often test delivery as much as treatment, so local context is part of the intervention rather than background noise.
Sources and further reading
Questions and answers
Is a cluster trial with thousands of patients necessarily large?
No. Statistical information depends strongly on the number of clusters, their sizes, and within-cluster correlation. A huge patient count across a few sites can still yield imprecise estimates.
Can ordinary regression fix a cluster design?
Only if the model properly accounts for the correlation and design. Treating all patient observations as unrelated generally overstates precision.
Does cluster randomization prevent selection bias?
It protects the assignment of clusters. It does not prevent biased individual recruitment after staff know the assigned arm. The timing and masking of recruitment require separate appraisal.