Evidence explainer

Evidence and research methods

When Real-World Evidence Helps and When It Misleads

Routine-care data can answer questions trials never measured, but their scale does not neutralize bias. The strongest studies define the decision first and make every assumption inspectable.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Match the data to the job
  3. Limits begin with confounding
  4. Time zero is where many studies fail
  5. Selection can enter at every data boundary
  6. Measurement error is rarely neutral
  7. Pre-specification reduces analytical freedom
  8. Use triangulation instead of one heroic model
  9. A reader's five-line summary

Real-world evidence is most useful when it measures actual care that a conventional trial could not efficiently capture, and it can show treatment patterns, reach populations excluded from trials, follow uncommon harms, and estimate longer-term outcomes. It can also produce a very precise answer to the wrong question.

The central risk is not that routine data are “messy.” It is that care decisions, data collection, and follow-up are connected to health; those connections can create an apparent treatment effect even when treatment made no difference.

Key points#

Match the data to the job#

Routine data are particularly strong for describing what happened. Claims can count prescriptions, procedures, and hospitalizations across a large insured population; electronic records can add laboratory values, vital signs, and clinical notes; registries can standardize disease-specific outcomes; and wearables and devices can measure behavior or physiology at high frequency.

Description answers questions such as: Who receives this treatment? How quickly is a test adopted? Which regions have access? How often is a coded outcome recorded? These are valuable without claiming causation.

Prediction asks what is likely to happen to a person with a given information set. A model can predict hospitalization well even if none of its variables causes hospitalization. It needs calibration, discrimination, and external validation in the intended setting.

Causal comparison asks what would happen if similar people followed different strategies. That requires a counterfactual design. A correlation between treatment and outcome does not become causal because the dataset is large or the model contains many covariates.

Limits begin with confounding#

Suppose clinicians tend to prescribe a newer drug to people whose disease is worsening. If those people later have more hospitalizations, the drug may look harmful even if it reduced their risk, which is confounding by indication: the reasons for treatment also predict the outcome.

The reverse can happen when a treatment requires specialist access, health literacy, or ability to attend follow-up. Treated people may start healthier or better supported. Routine data capture some of these factors and miss others.

An active-comparator, new-user design can help: instead of comparing users with nonusers who may be at entirely different clinical stages, it compares people initiating two reasonable options at the same decision point. Restriction, matching, standardization, propensity scores, and weighting can further balance observed variables.

Balance diagnostics should show whether measured covariates became comparable. A high c-statistic for the propensity model is not the goal. The goal is covariate balance and adequate overlap. People whose treatment choice was nearly certain may have no credible counterpart in the other group.

Residual confounding remains possible. A quantitative bias analysis, negative-control outcome, negative-control factor, or comparison with a randomized benchmark can probe it. These tests are informative precisely because no adjustment method proves that unmeasured confounding is absent.

Time zero is where many studies fail#

Eligibility, treatment classification, and follow-up should begin at a common time. If people are classified as treated only after filling several prescriptions, they had to remain alive and observable long enough to qualify. Counting that earlier “immortal” period as treated time favors the treatment group.

Other time problems include using future information to define baseline, giving one group a longer opportunity to have an outcome detected, and censoring people when they change treatment even though the reasons for change predict prognosis.

The target-trial framework helps by writing the observational study as though it were emulating a hypothetical randomized trial. Specify eligibility, strategies, assignment, start of follow-up, outcome, causal contrast, and analysis. Then map each element to the data you actually hold. A gap becomes visible before any code is run.

For repeated treatment decisions, time-varying confounding can require methods such as marginal structural models. Ordinary regression may adjust for a variable that is both affected by earlier treatment and predicts later treatment, blocking part of the effect while introducing bias.

Selection can enter at every data boundary#

Health-system records include people who use that system and the care recorded there. If treatment changes the probability of returning, loss to follow-up can depend on both treatment and outcome risk. Claims lose people when insurance changes. A voluntary registry may enroll people with more resources, stronger symptoms, or special interest in the condition.

Selection also occurs through analytic requirements. Requiring a complete laboratory panel may exclude people with less access or milder disease. Conditioning on a post-treatment event, such as hospitalization, can create collider bias by selecting a group through pathways affected by both treatment and health. So look for a cohort flow diagram, counts at each inclusion rule, the characteristics of the people kept and the people dropped, and an explanation of how anyone enters and leaves the data system. If a paper calls its sample “representative,” make it show why.

Measurement error is rarely neutral#

Billing codes reflect reimbursement and documentation as well as biology. A coded diagnosis can mean confirmed disease, rule-out evaluation, historical disease, or a copied problem-list entry. Prescription records show dispensing, not ingestion. Notes may contain richer context but natural-language extraction adds model error.

Outcome validation should match the use you intend. Positive predictive value matters when identifying cases; sensitivity matters when counting all cases; agreement on timing matters in survival analysis. Validation in a different database or decade may not transfer.

Treatment can change observation. People taking a medicine that requires laboratory monitoring may have more opportunities for an abnormality to be detected; a new device may initially be used at centers with more complete documentation. This surveillance bias can create group differences even if underlying event rates are identical.

Missing data need a causal explanation. “Missing at random” is an assumption, not a property confirmed by an imputation command. Report how much is missing by group and over time, why it may be missing, which variables inform imputation, and how results change under plausible departures.

Pre-specification reduces analytical freedom#

Routine databases offer many ways to define factor, outcome, washout, grace period, follow-up, and adjustment. If analysts try several and publish the most favorable, confidence intervals no longer reflect the true selection process.

A dated protocol and statistical analysis plan should identify the primary design and outcome. STaRT-RWE provides a structured template for recording implementation details. RECORD extends reporting guidance for research using routinely collected health data. Code lists, algorithms, and reusable code should be shared when governance permits. Changes after seeing results are not forbidden, but they should be labeled, dated, and justified, so that an exploratory analysis can generate a hypothesis without masquerading as confirmation.

Use triangulation instead of one heroic model#

Confidence rises when different designs with different likely biases point in the same direction, and a pragmatic randomized trial, a claims analysis, a disease registry, and a self-controlled design may each answer a related question. Agreement is more persuasive when their errors are not shared.

Within one study, useful sensitivity analyses include alternative outcome definitions, grace periods, negative controls, high-dimensional adjustment, quantitative bias analysis, different missing-data assumptions, and analyses restricted to areas of treatment overlap. A long sensitivity appendix is not automatically reassuring. The tests should target the assumptions most capable of overturning the conclusion you are being asked to accept.

FDA's Sentinel Initiative illustrates a mature use of distributed routine data for active safety surveillance. It combines a common data approach, prespecified queries, and governance across data partners. Even there, a signal is a reason for evaluation, not automatic proof of causation.

A reader's five-line summary#

For any real-world study, write down:

  1. The exact decision and whether it is descriptive, predictive, or causal.
  2. The target population, strategies, time zero, outcome, and follow-up.
  3. The important reasons treatment and outcome could be related before treatment.
  4. What the database measures well, poorly, or not at all.
  5. Which analysis most directly tests the weakest assumption.

If you cannot fill one of those lines from the report and its supplement, your certainty should fall. More rows make random error smaller. They do not make structural bias smaller.

Sources and further reading

  1. FDA, Real-World Evidence resources and guidance
  2. FDA, assessing electronic health records and claims data
  3. BMJ, target trial framework for causal inference from observational data
  4. RECORD statement for studies using routinely collected health data
  5. STaRT-RWE structured template for real-world evidence plans
  6. FDA Sentinel Initiative

Questions and answers

Is real-world evidence the same as observational research?

No. Routine-care data can support observational studies or pragmatic randomized trials. “Real world” describes the data context, not treatment assignment.

Can machine learning remove confounding?

It can model complex relationships among measured variables. It cannot adjust for a factor that is missing, mistimed, or measured too poorly.

Is a statistically precise estimate more trustworthy?

Only regarding random error under the model. A narrow confidence interval can surround a biased estimate.

What is the clearest warning sign in a causal RWE study?

Misaligned time zero is a major one. If treatment status requires future survival or future information, the comparison may be biased before adjustment begins.