Evidence explainer

Evidence and research methods

Immortal Time Bias: How Misassigned Follow-Up Creates False Benefit

If patients had to stay alive long enough to qualify for a group, that time is not the treatment's doing. A wrong time zero manufactures benefit before therapy starts.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. A simple timeline reveals the error
  2. Future information creates the “immortality”
  3. Treatment-status misclassification is one mechanism
  4. Selection can create a related immortal period
  5. A fixed Cox model does not know the dates are wrong
  6. Time zero must align three decisions
  7. Grace periods need explicit assignment rules
  8. Landmark analysis trades bias for a narrower question
  9. Minimum-duration definitions are especially risky
  10. Transplant and surgery studies illustrate the stakes
  11. Prescription databases have an additional invisible interval
  12. Competing events can disguise the pattern
  13. Signals readers can detect without raw data
  14. Correct timing does not remove confounding
  15. Reporting should make the timeline auditable
  16. References

An observational study classifies patients as “treated” if they receive a therapy at any point after diagnosis. Follow-up, however, begins on the diagnosis date. Anyone eventually treated had to remain alive and eligible until the treatment date. If an outcome had occurred earlier, that person could never have entered the treated group.

Crediting that guaranteed survival interval to treatment creates immortal time bias. The treatment seems protective partly because its group was defined using future survival. No regression adjustment can undo a timeline that was encoded incorrectly.

A simple timeline reveals the error#

Consider 1,000 people discharged after a serious illness. A study asks whether starting rehabilitation within 60 days improves one-year survival, and researchers divide participants into those who ever started rehabilitation and those who never did, then count survival from discharge.

A person who began on day 45 contributes 45 days before receiving rehabilitation. To be labeled treated, that person had to survive those 45 days. Someone who died on day 20 could only enter the never-treated group; if all 45 days are assigned to rehabilitation, the treated group receives death-free time that occurred before treatment.

The error can be large when treatment is delayed, early events are common, or the qualifying window is long, and it can produce a dramatic protective association even if the therapy has no effect. The apparent benefit begins before the first dose or session, a clue that the design is granting impossible credit.

Future information creates the “immortality”#

The word does not mean participants literally could not die. It means that, under the study's classification rule, an event during the interval would prevent future membership in the group. Group assignment depends on surviving the future.

Examples include receiving a transplant, filling a prescription, completing a procedure, accumulating 90 treatment days, responding at six months, or becoming adherent during follow-up. Each label requires time to pass. If follow-up is counted from an earlier point and the entire interval is assigned to the future label, the outcome is structurally excluded from that group during the qualifying period.

The same logic applies to nonfatal outcomes. A patient must remain stroke-free until a medication starts to be an eventual starter. “Immortal” is shorthand for event-free with respect to the endpoint.

Treatment-status misclassification is one mechanism#

Suissa described immortal time in pharmacoepidemiology as person-time misclassified with respect to treatment. An ever-treated indicator assigns all follow-up to the treated category, including days before initiation. Those days belong to untreated status if the question concerns current treatment.

A time-dependent variable can classify the person as untreated until the first dose and treated afterward. That repairs this particular allocation. The analysis then compares hazards under current status, subject to confounding and other assumptions.

Treatment discontinuation complicates the timeline further. An as-treated analysis needs a definition for grace periods, stockpiling, gaps, switching, and lingering biological effects. Time-varying coding is not automatically correct merely because it changes; it must represent the causal strategy.

Some designs require survival to enter the cohort itself. Suppose eligibility includes completing a 90-day treatment course, and time zero is placed at the first dose: everyone in the cohort has survived and tolerated 90 days, but early discontinuers and deaths may be excluded.

This can create selection bias even if person-time is not literally mislabeled. Conditioning on future adherence or response selects people based on post-baseline factors affected by prognosis. Zhou and colleagues distinguish structural pathways that produce immortal time through misclassification and selection. The lesson is broader than any one software setting: no eligibility criterion, group label, or analysis population should require information from after the intended start of follow-up unless you have explicitly changed the target question.

A fixed Cox model does not know the dates are wrong#

Including an ever-treated indicator in a Cox regression does not solve the problem. The model assumes that the supplied covariate correctly describes status from time zero. It can estimate the wrong comparison with impressive precision.

Adjusting for age, severity, and comorbidity also does not fix the guaranteed survival. Those variables address confounding, not the allocation of pre-treatment time; a propensity score defined using future treatment inherits the same timeline problem.

A time-dependent Cox model can represent initiation correctly for some descriptive or causal questions. If time-varying confounders both influence treatment and are affected by prior treatment, standard adjustment can introduce bias; marginal structural models or other g-methods may be required.

Time zero must align three decisions#

Target-trial emulation makes the alignment explicit. At time zero, a participant should meet eligibility, be assigned or classified to a treatment strategy, and begin outcome follow-up. Randomized trials naturally synchronize these steps at assignment.

Observational data often scatter them. Eligibility may be assessed at diagnosis, treatment starts weeks later, and covariates are measured during the gap, so the treated strategy benefits from surviving until initiation while the untreated strategy starts immediately. One way out is to compare strategies available at a common decision point, such as initiating drug A versus drug B on the prescription date among new users eligible for either; another is to emulate a sequence of daily or weekly trials in which eligible untreated people can initiate or remain uninitiated at each time.

Grace periods need explicit assignment rules#

Clinical strategies often allow treatment within a window rather than at one instant. “Start within 30 days” cannot be known on day zero. Simply labeling everyone who starts within 30 days as treated creates 30 potential immortal days.

Clone-censor-weight methods can assign each eligible person initially to copies of all strategies compatible with their observed data. A copy is censored when the person's behavior becomes inconsistent with that strategy, and inverse-probability weights address informative artificial censoring under measured assumptions.

Alternatively, researchers can define baseline assignment using information available at time zero or use sequential trials at each possible initiation day. These approaches estimate different effects. The paper owes you the estimand, the grace period, the adherence rule, and the weighting diagnostics, not just a sentence saying immortal time was “controlled.”

Landmark analysis trades bias for a narrower question#

A landmark analysis chooses a fixed post-baseline time, includes people alive and event-free at that time, classifies treatment using information up to the landmark, and follows outcomes afterward. No participant receives survival credit before the landmark in the analyzed outcome period.

The cost is exclusion of all early events and people who do not reach the landmark, and the estimand becomes the effect or association among landmark survivors, not among everyone eligible at baseline. Treatment before the landmark can also affect who survives to be included.

The landmark should be clinically justified and prespecified. Trying several landmarks and highlighting the most favorable result reintroduces data-driven selection, and early outcomes should still be reported so that you can see what was left out of your view.

Minimum-duration definitions are especially risky#

Studies may define use as two prescriptions, 90 covered days, or a cumulative dose. These definitions improve certainty that medicine was used, but they require future survival and persistence; assigning follow-up from the first prescription to a group defined by later persistence selects tolerant survivors.

An intention-to-treat-like new-user analysis can classify at first treatment and ignore later duration for the primary contrast. A per-protocol analysis can model adherence over time with appropriate methods. Neither should exclude early discontinuers as if they never started.

Dose-response analyses can also create immortal categories. A high cumulative dose requires more survival time than a low dose, and cumulative dose should be treated as a time-varying history or analyzed through strategies that align eligibility and accumulation.

Transplant and surgery studies illustrate the stakes#

The historical heart-transplant example made the problem visible: patients receiving a transplant had to survive long enough for an organ to become available. Counting survival from referral and classifying recipients as transplanted from day one makes transplantation appear to protect the preoperative waiting period.

The correct analysis can treat transplant as a time-varying state or emulate strategies around listing and receipt, depending on the question. Waiting-list changes, contraindications, disease progression, and organ availability create time-varying confounding and competing events.

Procedures after hospitalization create similar structures. People healthy enough to undergo surgery weeks later are selected survivors: a comparison with everyone who never undergoes the procedure can combine immortal time, confounding by health status, and access differences.

Prescription databases have an additional invisible interval#

Inpatient medicine use may not appear in outpatient pharmacy claims. A hospitalized patient can survive while receiving unmeasured therapy, then fill a prescription after discharge. Classifying outpatient fills without accounting for hospital days creates immeasurable time and differential gaps.

Data availability, not biology, determines when treatment becomes visible. Researchers should map enrollment, hospitalization, dispensing, administration, and outcome capture. Restricting to observable periods can help but may change the population. A claims study should spell out whether a prescription represents dispensing or ingestion, how inpatient stays are handled, and whether death data are complete, because the database clock and the clinical clock have to be reconciled before either can be trusted.

Competing events can disguise the pattern#

If treatment initiation competes with death, early deaths accumulate in the never-treated group. Reporting only a treatment hazard ratio hides that pathway. Plots of time to treatment, event timing, and risk-set composition can expose it.

Competing-risk methods alone do not correct immortal assignment. They appropriately handle mutually exclusive outcomes after the study design defines states and time zero. A flawed future-defined treatment group remains flawed inside a Fine-Gray model. Multistate models can represent untreated, treated, event, and death transitions. Their clarity depends on correctly timed state changes and a scientific question for each transition.

Signals readers can detect without raw data#

Watch for “ever received,” “completed,” “responded,” “adherent,” “at least two prescriptions,” or “survived to” when follow-up begins earlier. Compare the cohort-entry date with the first possible treatment date. Ask where events occurring in between were counted.

Look at the treated group's event curve before treatment could plausibly act. A separation that begins immediately despite delayed initiation is suspicious. Check whether early deaths were excluded, whether the treated group had guaranteed follow-up, and whether treatment timing is reported. The methods section will tell you whether treatment is baseline or time-varying, how the risk set updates, and what was done about grace periods, switching, censoring, and covariate timing, because a vague sentence about time-dependent analysis does not establish that any of it was implemented correctly.

Correct timing does not remove confounding#

After immortal time is repaired, treated and untreated patients can still differ. People who initiate therapy may be healthier, sicker, wealthier, more connected to care, or eligible by clinical criteria absent from the data. Confounding by indication and healthy-user bias remain.

Time-varying confounding can be severe because worsening disease prompts treatment and predicts outcome. A correctly timed but naively adjusted estimate can still be biased. Negative controls, active comparators, measured severity, target-trial design, and sensitivity analyses add evidence. Repairing the timeline can move an estimate dramatically toward no association, as shown in time-related bias examples, but the corrected number is not automatically causal. Each bias needs its own design response.

Reporting should make the timeline auditable#

A useful diagram shows eligibility assessment, treatment assignment or initiation, covariate windows, grace period, follow-up start, and event ascertainment. A table reports how many events occurred before treatment and how much person-time preceded initiation.

Authors should present absolute risk curves under defined strategies, not only a hazard ratio. Sensitivity analyses can vary grace periods, treatment definitions, induction windows, and analytic approaches. Results that hang on crediting one particular interval deserve your caution.

The core audit fits in one sentence: could a participant have experienced the outcome during time that the analysis counted as belonging to a group they had not yet qualified to join? If yes, the claimed treatment benefit may contain survival that happened before treatment.

References#

  1. Immortal time bias in pharmacoepidemiology
  2. Immortal time bias in cohort studies
  3. Using observational data to emulate a target trial
  4. Structural description of biases generating immortal time
  5. Time-related biases in pharmacoepidemiology
  6. ENCePP guide to methods addressing bias and confounding

It does not provide medical advice or determine whether a treatment is effective.*

Questions and answers

What is immortal time bias?

It is bias created when follow-up during which a participant had to remain event-free to qualify for a future group is credited to that group. The usual result is an artificial survival advantage.

Why is the time called immortal?

To enter the future-defined group, a participant must survive or remain event-free through that interval, so an earlier event would have made group membership impossible.

Does a Cox model automatically fix immortal time bias?

No. A fixed ever-treated indicator preserves the wrong allocation. Treatment must be timed correctly, or the design must assign strategies at a synchronized baseline with appropriate methods.

Is landmark analysis always the solution?

No. It can avoid counting pre-landmark outcomes, but it excludes early events and changes the target population to landmark survivors. It should be prespecified and interpreted as a narrower question.

What should readers compare in the methods?

Align eligibility, treatment assignment, and start of follow-up. Then inspect how pre-treatment person-time, early events, grace periods, switching, and post-baseline eligibility were handled.