Observational health data can answer questions that a randomized trial has not answered. They can include large populations, long follow-up, uncommon outcomes, and groups often missing from trials. Yet a database does not assign treatment at random. People who start, continue, avoid, or stop a treatment differ for reasons that may also affect their outcome.
Target trial emulation is a framework for making the causal question and study design explicit before analyzing those data, and researchers describe the hypothetical randomized trial that would answer the question, then show how each part is represented in the available records. This discipline can prevent avoidable design errors. It cannot turn nonrandomized data into a randomized experiment or remove unmeasured confounding by declaration.
Begin with the trial that would answer the question#
Suppose the question is whether starting treatment A rather than treatment B reduces a five-year risk among people newly eligible for either. A randomized trial would define who qualifies, assign eligible participants to one strategy at a specified moment, begin follow-up then, measure the same outcomes in both groups, and analyze a defined contrast.
An ordinary database analysis can drift away from that question. It might compare current users with nonusers, include people who survived long enough to become users, use future adherence to define groups, or start follow-up at different moments. A sophisticated model cannot fully repair a causal question that was encoded incorrectly. The target-trial framework puts design before estimation: it asks researchers to write a protocol sufficiently clear that another team could reconstruct the intended experiment and see where observational data fall short.
The protocol components#
Eligibility criteria define the target population and the moment at which each person can enter, and they may include age, diagnosis, prior treatment, laboratory values, contraindications, and enough data history to measure the confounders you will later need to adjust for. In an emulation, every one of those criteria needs an operational definition built from fields that already exist at or before time zero. Nothing later counts.
Treatment strategies must be specific enough to distinguish. "Treated" and "untreated" can hide initiation, continuation, switching, adherence, and co-interventions. A strategy could compare initiation of A with initiation of B, or sustained strategies over time. The data need to support the distinction being claimed.
Assignment is random in the hypothetical trial. In observational data, investigators classify people according to the strategy their records show at baseline. Because classification is not random, causal interpretation requires assumptions about measured confounding and the process that produced treatment choice.
Follow-up states when observation begins and ends. Reasons for ending can include outcome, death, loss to follow-up, administrative end, or deviation from a strategy, depending on the causal contrast. The start must align with eligibility and assignment.
Outcomes require definitions, timing, and ascertainment rules. A code in claims or records may represent a clinical event only imperfectly, so validation, equal ascertainment across groups, competing events, and missing outcomes all matter.
The causal contrast says which effect is sought. An assignment-style contrast compares strategies as classified at baseline regardless of later adherence, analogous to an intention-to-treat effect, and a per-protocol contrast asks about following the strategies as specified and usually needs additional methods for time-varying adherence and confounding.
The analysis plan connects the estimand to estimation. It describes adjustment, censoring, missing data, models, effect measures, subgroup analyses, and sensitivity analyses. Absolute risks and risk differences often make a result more interpretable than a relative measure alone.
The causal estimand is the question in mathematical form#
An estimand identifies the population, treatment strategies, outcome, time horizon, and contrast. Saying "the effect of treatment" is incomplete. The effect of starting a treatment can differ from the effect of remaining on it. Risk at one year can differ from risk at ten years. A population eligible for first-line therapy differs from people with several prior treatments.
Clear estimands prevent a common slide between questions. Researchers may describe initiation in the methods but interpret the result as sustained adherence. They may estimate a hazard ratio while discussing an absolute event reduction. They may censor at treatment change and treat the remaining groups as comparable without adjusting for reasons for change.
The target protocol keeps those choices visible. It also reveals when the available data cannot answer the desired question. That is a useful result of design work, even if no estimate follows.
Why time zero matters#
Time zero is the moment when eligibility is assessed, treatment is assigned or classified, and follow-up begins. In a randomized trial, these events are normally aligned. Observational studies can separate them accidentally.
Suppose you classify someone as treated because a treatment appears at any point in the first six months, while counting outcomes from the first day. To enter the treated group, that person must remain alive and event-free until treatment begins. The period before initiation is "immortal" with respect to the group definition: an outcome during that period would prevent classification as treated. If you credit that pre-initiation time to treatment, the treated group receives an artificial survival advantage.
Starting treated follow-up at initiation but untreated follow-up at an earlier eligibility date creates a different mismatch. Using a future event to decide eligibility or group membership can introduce selection. The 2026 BMJ guidance on "starting right" emphasizes aligning eligibility and assignment when people can qualify at several times.
One solution is to emulate a sequence of trials at repeated eligible times, classifying each person based on the strategy at that time and applying methods that account for repeated participation. Other designs use cloning and censoring when strategies cannot be distinguished at baseline. These are specialized choices that need clear reporting and suitable assumptions.
The hormone-therapy lesson#
Postmenopausal hormone therapy and coronary heart disease became a classic example because earlier observational analyses appeared more favorable than randomized evidence; it was tempting to attribute the entire discrepancy to the absence of randomization.
Hernán and colleagues reanalyzed Nurses' Health Study data as a sequence of trials comparing eligible initiators with noninitiators and aligned design and analysis more closely with the randomized Women's Health Initiative question. The estimated coronary effects for early follow-up and overall follow-up became more similar to the randomized findings than earlier observational summaries had been.
The example does not show that design removes all confounding or settle every hormone-therapy question. Differences in time since menopause, treatment formulation, adherence, eligibility, and outcome period still matter. It shows that observational-randomized disagreement can arise from several sources at once: confounding, who is compared, when follow-up starts, and which effect is estimated. So a reconstructable protocol tells you more than the label "real-world evidence."
Mapping the protocol to data#
After specifying the target trial, researchers should name the observational analog for each component. If eligibility requires a clinical feature recorded only intermittently, how is it measured? If pharmacy records show dispensing but not ingestion, what treatment definition is used? If an outcome algorithm has known sensitivity and positive predictive value, how could misclassification affect groups?
Baseline confounders must be measured before or at time zero. They should be causes of treatment choice and outcome, not variables affected by treatment. Standardization, weighting, matching, stratification, regression, or some combination of them can do the adjusting, and which one you reach for should follow from the estimand and the data structure.
For sustained strategies, confounders can change over time and be affected by earlier treatment while also influencing later treatment. Conventional adjustment can then create bias. G-methods, including inverse-probability weighting and the parametric g-formula, were developed for such settings, but they demand additional modeling and data quality.
Missing data need description across treatment, outcome, and confounders. Complete-case analysis can select a nonrepresentative group when missingness relates to prognosis, and loss to follow-up or health-plan exit can be informative rather than random, so the records that stop may not be a random sample of the records that continue. Who is missing is part of the result.
Assumptions that remain#
Conditional exchangeability means that, after accounting for measured variables as specified, treatment groups can be compared as though assignment were random within those covariate patterns. It cannot be tested completely from observed data. An unrecorded cause of both treatment choice and outcome can still bias the estimate.
Positivity requires that people within relevant covariate patterns have a nonzero chance of each strategy, and if everyone with a certain contraindication receives only one treatment, the data cannot support comparison for that group without extrapolation. Extreme weights can signal practical positivity problems.
Consistency requires that the observed treatment correspond to the strategy being analyzed and that treatment versions are sufficiently well defined. A vague category covering materially different care can violate that link.
Measurement validity, correct model specification, appropriate handling of competing events, and no important selection bias are additional concerns. Negative controls, quantitative bias analysis, alternate definitions, and benchmarking against a trial can probe robustness. They do not prove that all assumptions hold.
What the TARGET Statement adds#
The 2025 TARGET Statement provides reporting guidance for observational studies that emulate a target trial. It asks authors to identify the design, explain why emulation is needed, specify the target protocol and identifying assumptions, map each component to the data, report estimates with precision, and show sensitivity to design and analysis choices.
Reporting guidance makes appraisal possible. A study can follow every reporting item and still have unmeasured confounding, weak outcome measurement, or a causal question the data cannot identify. Conversely, poor reporting can make a strong analysis impossible to verify. Use TARGET as a map, then judge the terrain yourself.
A reader checklist#
First, can the hypothetical trial be reconstructed? Look for explicit eligibility, strategies, assignment, time zero, follow-up, outcomes, causal contrasts, and analysis. If one component is absent, ask yourself what question the number actually answers.
Second, are eligibility, classification, and follow-up aligned? Watch for future information used to define baseline groups, prevalent-user comparisons, and different start dates.
Third, how were protocol elements mapped to records? Check code definitions, validation, prescription versus use, outcome ascertainment, data history, and missingness.
Fourth, which confounders were measured before treatment, and which important determinants were unavailable? A very large sample reduces random error but does not erase systematic bias.
Fifth, does the estimate match the stated contrast and include absolute effects, precision, and time? Review sensitivity analyses, negative controls, alternate definitions, and any benchmark against randomized evidence.
Finally, check whether conclusions stay within the population, strategies, data era, and assumptions studied. A target trial emulation can make causal reasoning more disciplined. It does not make causal language automatic.
References#
- Hernán and Robins, Using Big Data to Emulate a Target Trial
- Hernán and colleagues, hormone therapy observational emulation
- JAMA, Target Trial Emulation framework, 2022
- BMJ, applying randomized-trial principles to observational studies
- TARGET Statement for transparent reporting, 2025
- BMJ, aligning eligibility and treatment assignment at time zero, 2026
- TARGET reporting guideline site
Questions and answers
Is target trial emulation a randomized trial?
No. It uses a hypothetical randomized protocol to structure an observational analysis. Treatment was not randomized, so causal interpretation still depends on assumptions about confounding, selection, measurement, and positivity.
What is time zero?
It is the aligned moment when eligibility is established, treatment strategy is assigned or classified, and follow-up begins. Misalignment can give one group guaranteed event-free time or introduce other design bias.
Does a large database solve confounding?
No. More records can improve precision and permit subgroup study, but systematic differences between treatment groups remain if important causes are unmeasured, mismeasured, or modeled poorly.
Why write the hypothetical trial if it cannot be run?
The protocol clarifies the causal question and exposes design choices. It helps researchers see whether the available data can represent eligibility, strategies, time, outcomes, and needed confounders before fitting a model.
Does compliance with TARGET prove that an estimate is valid?
No. TARGET supports transparent reporting. Validity still depends on study design, data quality, assumptions, analysis, precision, and how well the emulation matches the protocol.