A control group gives a study its “compared with what?” Without that comparison, you cannot separate improvement after an intervention from improvement caused by time, natural recovery, regression toward average symptoms, other care, measurement, or the experience of being studied. The control is an estimate of the outcome that would have occurred under a different assigned strategy.
It is not a literal copy of what the same person would have experienced without treatment. That individual counterfactual can never be observed at the same time, so a well-designed control group makes a group-level counterfactual credible enough to support a causal comparison, with uncertainty and assumptions stated.
Key takeaways#
- The comparator defines the claim: placebo, usual care, an active treatment, another dose, and historical data answer different questions.
- A concurrent control alone does not guarantee fairness; allocation, blinding, follow-up, outcome measurement, and co-interventions also matter.
- Randomization makes groups comparable on average before treatment, while a control supplies the outcome comparison. These are related but distinct features.
- Placebo designs can isolate specific treatment effects, but withholding effective care requires ethical justification and risk limits.
- Uncontrolled evidence can establish feasibility, describe rare cases, or generate hypotheses, but it usually cannot estimate a treatment effect reliably.
A control group estimates the missing counterfactual#
For each participant, two potential outcomes are of interest: the outcome under the tested intervention and the outcome under the comparison strategy. Only one can be observed. If a participant receives the intervention, the untreated or alternative-treatment outcome is missing by design.
A comparison group addresses that missing outcome at the population level, and if groups are sufficiently comparable before treatment and are followed in the same way afterward, their average outcome difference can estimate the effect of assignment. Randomization is powerful because it gives every eligible participant a known chance of each assignment, balancing both measured and unmeasured prognostic factors on average.
Balance is expected, not guaranteed in every finite sample. Chance differences can remain. More importantly, comparability can be damaged after randomization by unequal loss to follow-up, crossovers, different additional care, selective outcome measurement, or analysis that excludes participants according to post-assignment events. The control therefore does more than sit in the second column of a table: its credibility depends on how it was selected, treated, measured, retained, and analyzed.
Why improvement alone cannot identify a treatment effect#
Many health outcomes change without the tested intervention. Acute infections resolve. Pain fluctuates. Blood pressure varies between visits. Depression scores can move with life events, concurrent therapy, or repeated measurement. Chronic diseases may progress at different rates.
Regression to the mean is especially deceptive. People often seek care or enter a study when symptoms are unusually severe. Even if nothing effective happens, an extreme measurement tends to be followed by a less extreme one because temporary influences do not repeat perfectly; a before-and-after study can show that expected movement and call it treatment benefit.
Expectations and context can also affect outcomes. Attention, reassurance, a treatment ritual, and belief about assignment may change symptom reports, adherence, or behavior. Outcome assessors can interpret borderline findings differently if they know treatment. A placebo or sham can help balance some of those effects, while blinding can reduce differential measurement.
Natural history and contextual effects are not interchangeable: an untreated control, wait-list group, or observational history may improve differently from a blinded placebo group because study participation and treatment ritual differ. Research comparing placebo groups with untreated natural-history groups illustrates why the “placebo effect” should not be inferred simply from improvement in a placebo arm.[5]
Five broad control choices and the questions they answer#
ICH E10 organizes major control types and the design issues that follow from them.[1] The categories can overlap in multi-arm studies, but each has a distinct purpose.
Placebo or sham control#
A placebo resembles the tested medicine without its active component. A sham procedure imitates aspects of a device or procedure without the hypothesized active element, and when participants, clinicians, and assessors cannot distinguish assignment, expectations and assessment conditions may be better balanced.
The primary question is whether the intervention outperforms the matched control under the trial conditions. The difference does not equal every biological effect in isolation. Adherence, rescue treatment, unblinding, and interactions between context and treatment can still influence the estimate.
Matching can be difficult. Side effects may reveal assignment. A sham operation may carry risk without therapeutic prospect. Behavioral interventions may be impossible to conceal from the person delivering them; reports should explain what was matched, who was blinded, whether blinding was assessed appropriately, and what happened when assignment became apparent.
No-added-treatment or wait-list control#
A no-added-treatment group may continue baseline care without the new intervention. A wait-list group receives the intervention later. These designs can estimate the effect of offering the intervention compared with not offering it during the study interval.
They usually do not balance expectation, contact time, or treatment ritual. Participants know whether they are receiving the program. Outcome reporting and retention may differ, but such a control can still be appropriate when outcomes are objective, placebo effects are unlikely to dominate, or blinding is impractical, but the interpretation must match the design.
Active or usual-care control#
An active-control trial compares the new intervention with an established treatment. This answers a clinically relevant question: how does the new option compare with what would otherwise be used? Trials may seek superiority, noninferiority, or sometimes equivalence within defined margins.
“Usual care” is not self-defining. It can vary between clinicians, sites, countries, and time periods. CONSORT guidance asks investigators to describe both comparator and intervention as actually delivered, including adherence and fidelity.[4] A label without dose, access, timing, co-interventions, and delivery conditions makes replication and transport difficult.
Fairness matters. An active control used at a suboptimal dose, in an unsuitable population, or with poor adherence can make the new intervention look favorable; conversely, unusually intensive delivery of the comparator may not represent ordinary care.
Dose-comparison control#
Different doses or regimens of the same intervention can reveal dose-response, identify a minimally effective dose, or compare efficacy and harms, though a dose-comparison trial does not necessarily establish that any dose works if every arm could be ineffective. Including placebo or an established active control can provide an anchor when feasible. This design also needs adequate separation between doses. If achieved concentrations overlap substantially because of adherence or metabolism, the nominal labels may not represent meaningfully different interventions.
External or historical control#
An external control comes from outside the concurrent randomized trial, such as a registry, prior trial, chart cohort, or well-described natural history. It can be valuable when randomization is infeasible, a disease is rare, outcomes are dramatic and predictable, or early evidence is being developed.
The difficulty is exchangeability. Changes in diagnostic criteria, supportive care, calendar time, geography, eligibility, outcome ascertainment, follow-up, and data quality can create differences unrelated to treatment. Statistical adjustment addresses measured factors under assumptions; it cannot repair every unmeasured difference. External controls are most persuasive when the disease course is stable, eligibility and time zero align, outcomes are objective, and data are collected comparably.
A control group and randomization are not the same thing#
A study can have a control group selected by patient preference, clinician choice, hospital, calendar period, or existing records. That makes it controlled but not randomized. Baseline differences may then explain some outcome difference.
Randomization can create two or more assignment groups, one of which functions as the comparator, and allocation concealment protects the sequence before assignment so recruiters cannot predict or influence the next group. Blinding protects later behavior and measurement. These methods solve different problems:
- Randomization addresses baseline comparability on average.
- Allocation concealment protects the assignment process.
- Blinding reduces differential behavior and assessment after assignment.
- Complete follow-up limits attrition bias.
- Intention-to-treat analysis preserves the randomized comparison for the effect of assignment.
Calling a trial “controlled” does not certify the other features. Each needs separate evaluation.
The ethical boundary around placebo#
Placebo use is not ethical or unethical by label alone. The 2024 revision of the World Medical Association Declaration of Helsinki states that a new intervention should generally be tested against the established effective intervention.[2] Placebo or no intervention may be acceptable when no proven intervention exists, or when there is a compelling and scientifically sound methodological reason and participants will not face added risk of serious or irreversible harm from not receiving the established intervention.
The ethical analysis considers disease severity, duration of withholding, rescue criteria, monitoring, informed consent, available treatment, and the scientific need for a placebo comparison. Add-on designs can preserve established background care while randomizing the new intervention against placebo. Early escape rules can limit time without effective control. An active comparator may be necessary when withholding care would create unacceptable risk.
Scientific weakness is itself an ethical problem. A control that cannot answer the question may ask participants to accept burden without producing reliable knowledge. Ethical review and scientific design therefore meet at the comparator.
Why active-control trials need assay sensitivity#
Suppose a new treatment and an established therapy produce the same outcome. Several explanations fit: both worked; neither worked in this particular trial; the outcome was measured poorly; adherence was low; participants had little room to improve; or the study could not detect a difference.
Assay sensitivity is the ability of a trial to distinguish effective from ineffective treatment. In a superiority trial, a clear difference can demonstrate that capacity directly, while in a noninferiority trial with no placebo arm, investigators often rely on historical evidence that the active control has a predictable effect under conditions similar to the current study.
That reliance creates the constancy assumption. The active control's effect now must be sufficiently similar to its effect in earlier trials. Changes in population, endpoints, adherence, rescue therapy, or background care can weaken that assumption. The noninferiority margin must also preserve a clinically important portion of the established effect. A poorly chosen margin can declare success while allowing meaningful loss of benefit.
The worked example on reading a noninferiority trial shows how margins and active controls shape inference. Clinical-trial protocol design explains how the comparison and estimand are set before data are seen.
How controls can become unfair after assignment#
Even a sensible comparator can be undermined during conduct. Watch for:
- Different visit schedules or contact time between arms.
- Unequal access to rescue treatment or additional services.
- Differential encouragement, adherence support, or outcome prompting.
- More missing outcomes in one group.
- Switching that is handled differently by arm.
- Outcome assessors who know assignment for subjective endpoints.
- A comparator delivered below its accepted standard.
- Analysis windows that favor one treatment's timing.
These features create performance or measurement differences beyond the intended intervention contrast. The report should describe care as delivered, not only as planned.
Contamination can move groups toward each other when control participants receive parts of the tested intervention. Co-intervention can move them apart when one group receives additional care. Both affect what the trial actually estimates.
What uncontrolled studies can still contribute#
Uncontrolled evidence is not useless. A case report can reveal an unexpected adverse event. A case series can describe a phenotype. A single-arm feasibility study can test recruitment, delivery, adherence, measurement, and short-term tolerability. In rare diseases with dramatic outcomes and well-characterized natural history, an external comparison may sometimes support stronger inference.
The claim must remain proportional. “Participants improved after treatment” is descriptive. “Treatment caused the improvement” requires a credible counterfactual or an unusually compelling alternative design. FDA regulations on adequate and well-controlled studies explicitly emphasize distinguishing drug effects from spontaneous change, placebo effects, and biased observation.[3]
Historical comparisons that use “before treatment” as the control face extra problems. Calendar time, diagnostic intensity, supportive care, and eligibility may change. Survivorship and immortal-time biases can enter if time zero differs. A striking before-and-after chart should make you ask where the concurrent group went.
A control-group checklist for readers#
Before you interpret the result, ask:
- What exactly did the control group receive?
- Does that comparator answer the decision people actually face?
- How were participants assigned, and was allocation concealed?
- Who knew the assignment?
- Were follow-up, measurement, and co-interventions similar?
- Did adherence, switching, and missing data differ?
- For an active control, could the study have detected a difference?
- For an external control, do eligibility, time zero, outcomes, and calendar time align?
- Does the conclusion compare only what the trial actually compared?
Related articles explain allocation concealment versus blinding, protocol design, and outcome switching. The site's research overview connects these design checks to evidence appraisal.
The sentence every result needs#
Finish every treatment claim you read with the comparator: the intervention improved an outcome relative to what? Placebo, active treatment, usual care, no added treatment, another dose, and historical data are not interchangeable. Once the comparator is named and its credibility examined, the study's legitimate claim becomes much clearer.
References#
- ICH E10: Choice of Control Group and Related Issues in Clinical Trials
- World Medical Association Declaration of Helsinki, 2024 Revision
- 21 CFR 314.126: Adequate and Well-Controlled Studies
- CONSORT 2025: Intervention and Comparator as Actually Administered
- Separating Placebo Effects From Natural History in Epilepsy Trials
Questions and answers
Is a control group always given a placebo?
No. Controls may receive placebo or sham, no added intervention, usual care, an established active treatment, another dose, or external historical care, and the choice determines the question and the biases that require attention.
Does having a control group make a study randomized?
No. A controlled study can allocate by preference, clinician choice, site, or time. Randomization is a separate procedure that balances prognostic factors on average and supports causal inference when conduct and analysis preserve the comparison.
Why can an uncontrolled study not prove that a treatment worked?
Because the observed change combines treatment with natural history, regression toward average symptoms, expectation, measurement, other care, and selection. Without a credible comparison, those components cannot usually be separated.
Is placebo use ethical when an effective treatment exists?
It depends on scientific necessity and participant risk. Withholding effective care is not acceptable when it adds serious or irreversible harm. Add-on designs, active comparators, rescue criteria, and limited placebo duration can sometimes answer the question while protecting participants.[2]
What is wrong with a poorly chosen active control?
It can bias the comparison or remove assay sensitivity: an apparent tie might mean both treatments work, neither worked under the study conditions, or the trial could not reveal a difference. Dose, delivery, adherence, population, endpoint, and noninferiority margin all matter.