Evidence explainer

Evidence and research methods

External Control Arms: Reading a Trial That Never Randomized Its Comparator

An external control arm judges a new drug against patients from outside the study. Groups that were never randomized can differ in ways that fake a benefit, or hide one.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Start with what randomization buys you
  3. Why this is treated as a fallback
  4. Three ways the comparison can mislead
  5. A short checklist for reading a single-arm claim

An external control arm asks a blunt question in a difficult way: how would these treated patients have fared if they had not received the drug? Instead of answering it with a randomized comparator built into the same trial, an externally controlled study borrows that answer from outside, from a historical cohort, a disease registry, or a natural history dataset. The design is sometimes acceptable to regulators, but only in a narrow band of situations, and you should read it with more suspicion than almost any other kind of clinical evidence. The reason is easy to state and stubbornly hard to solve: two groups that were never randomized can differ in ways that either invent the look of a benefit or conceal a real one, and no amount of statistical polishing fully repairs a comparison that time and biology have already tilted.

Key points#

Start with what randomization buys you#

In an ordinary randomized controlled trial, chance decides who gets the new treatment and who gets the comparator. That coin toss does something no spreadsheet can reproduce: it tends to balance the traits investigators measured and, just as importantly, the ones they never thought to record. When two arms start out alike on average, a later difference in outcomes is more plausibly the drug.

A single-arm trial gives that up. Everyone enrolled receives the investigational drug, so there is no internal comparator at all. To make any claim about benefit, the outcome has to be measured against an external reference for how similar patients would otherwise have done. That reference is the external control arm.

The International Council for Harmonisation guideline ICH E10 sorts these references into two families. A historical control is a group treated at an earlier time, before the new drug existed. A concurrent external control is treated during the same period but somewhere else, such as another hospital or a registry. Both live outside the trial, which is what makes them external, and both carry the same central weakness.

Why this is treated as a fallback#

ICH E10 does not soften the point: externally controlled designs raise serious concerns about whether the treated and control groups are comparable and about keeping bias in check, so the guideline reserves them for unusual circumstances. That caution has held for decades, and it explains why single-arm submissions cluster in one particular corner of medicine.

Three conditions define that corner, and they tend to appear together. The first is a serious or life-threatening disease with real unmet need, where a randomized trial may be impractical or ethically uncomfortable. The second is a disease whose natural course is well mapped and highly predictable, so the untreated trajectory is known rather than assumed. The third is an expected effect large enough to be self-evident and tightly linked in time to the treatment, so the signal towers over the noise that bias adds. A 2024 framework published in Therapeutic Innovation and Regulatory Science examined non-oncology approvals decided by the FDA and the European Medicines Agency between 2019 and 2022 and reported that the single-arm cases that succeeded overwhelmingly involved rare diseases with an established natural history and little or no standard of care, with a large effect size appearing again and again among the applications that got through.

Three ways the comparison can mislead#

The groups may not match#

Without randomization, treated and external patients can differ systematically from the start. People who enroll in a live trial are often healthier, watched more closely, and treated at more capable centers than the historical patients they are measured against. The 2024 analysis found that a failure to balance baseline characteristics between the two groups was a recurring regulatory criticism, and reviewers noted that gaps could survive even after adjustment. Matching and propensity methods can line up the traits someone remembered to record, but they cannot balance the unrecorded ones the way chance does. That leftover gap is where a false signal hides.

The clocks may not agree#

Every fair comparison needs a shared starting line, an index point from which outcomes are counted on both sides. Lining that moment up across a running trial and an outside dataset is harder than it sounds, because the two rarely flag the same clinical event in the same way. Get it wrong and immortal time bias can creep in, where the treated group only looks longer-lived because a patient had to survive long enough to receive the drug and be recorded as receiving it. The FDA draft guidance on externally controlled trials, issued on February 1, 2023 and still a draft, spends considerable attention on defining and aligning that start time, and the 2024 framework found that non-contemporaneous external controls drew explicit criticism in a large share of the FDA applications it reviewed.

The rulers may not read the same#

The two groups also have to measure the same outcome the same way. When an endpoint is subjective, or rests on an imaging read or a biomarker assessed differently across eras and institutions, an apparent difference can reflect the ruler rather than the medicine. These comparisons are usually unblinded, so knowing who received the drug can color how an outcome is judged. The 2024 framework reported that a meaningful share of FDA applications were faulted for subjective, imaging, or biomarker endpoints without solid reliability studies. It is one reason a hard, unambiguous outcome such as death is far easier to defend than a clinician's judgment call.

A short checklist for reading a single-arm claim#

When a headline rests on a single-arm result, four questions do most of the work. What was the control, and how recent is it? Did the two groups look alike at the start? How was the start time defined on each side? Was the endpoint measured identically and objectively on both? Then one more that outranks the rest: is the effect big enough to survive all of those doubts? A small difference from an external comparator is precisely the setting in which bias, not the drug, is the likeliest explanation. The FDA document remains a draft open to exactly these questions, not a finished endorsement of the design, and reading it in that spirit is the honest stance.

Sources and further reading

  1. FDA Draft Guidance: Considerations for the Design and Conduct of Externally Controlled Trials (2023)
  2. Framework for Regulatory Acceptance of Single-Arm Trials (PMC, 2024)
  3. ICH E10: Choice of Control Group and Related Issues in Clinical Trials
  4. Federal Register notice, Externally Controlled Trials Draft Guidance (Feb 1, 2023)

Questions and answers

Is an externally controlled trial the same as a real-world evidence study?

They overlap but are not identical. An external control arm is one way to build evidence, and it often draws its comparator from real-world sources such as registries or electronic records. Real-world evidence is the broader category of data collected outside a controlled trial, which can be used for external controls but also for many other purposes.

Why not just adjust for the differences statistically?

Adjustment, matching, and propensity methods can align the characteristics that were actually recorded. What they cannot do is balance the factors no one measured, which is the very thing randomization handles automatically. That is why regulators still treat a large, unmistakable effect as the real safeguard rather than the statistics.

When is a single-arm design most convincing?

When the disease is rare and serious, its untreated course is well understood and predictable, there is little or no effective standard of care, and the treatment produces a change so large and so closely timed that ordinary bias cannot plausibly account for it.