Evidence explainer

Evidence and research methods

Real-World Data After Approval: What It Can Add

Approval answers one evidence question at one point in time. Real-world data can extend the record, but routine records do not turn into evidence without a design.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Approval is a beginning, not a finish line
  2. Data source and evidence claim are different objects
  3. Postmarket safety uses several complementary systems
  4. Confounding by indication is the central trap
  5. Time can create bias before analysis begins
  6. Outcomes need validation
  7. Missing data are produced by care
  8. Real-world evidence can support more than safety
  9. A protocol should exist before the result
  10. Read a real-world claim in seven lines
  11. References

The evidence record for a medicine does not close on the approval date. Premarket trials answer specific questions in selected participants over a defined period. After approval, use spreads across larger, older, younger, more medically complex, and more diverse populations: treatment lasts longer, prescribing patterns change, rare events accumulate, and the medicine interacts with routine care.

Real-world data can help examine that wider experience. The term includes electronic health records, insurance claims, and product and disease registries. It includes pharmacy data, patient-generated information, and digital health measurements. Real-world evidence is not the database itself. It is the clinical evidence produced from a study whose data and methods can support the inference.

Approval is a beginning, not a finish line#

Randomized trials remain powerful because assignment can balance measured and unmeasured prognostic factors on average. They often use standardized outcomes, scheduled visits, and adjudication, though those strengths can come with narrower eligibility, limited duration, modest sample size for rare harms, and care processes unlike routine practice.

Suppose a serious adverse event occurs in one person per 20,000 treated. Even a trial with several thousand participants may observe none. A database containing millions of treatment episodes may identify an imbalance. Longer observation can reveal latency, cumulative dose effects, pregnancy outcomes, or interactions that were not estimable before approval.

Routine data can also study effectiveness rather than efficacy: how treatment performs amid missed doses, switching, comorbidity, variable monitoring, and the access barriers your patients actually face, and the answer may differ from a protocol-controlled trial without contradicting it.

Data source and evidence claim are different objects#

An electronic record contains clinical notes, diagnoses, laboratory values, prescriptions, and encounters. Claims record billable services and dispensings. A registry follows a defined population or product. A wearable produces repeated measurements from people willing and able to use it.

Each source sees part of the clinical story. A prescription order does not prove the medicine was filled or taken. A dispensing does not prove adherence. A billing diagnosis may be a rule-out, a historical condition, or a code chosen for reimbursement. Death can disappear when a person leaves an insurance plan unless records are linked. Symptoms and function are often absent.

So fit for purpose depends on the claim you are making. Claims may measure hospitalization well but miss blood pressure. An oncology registry may capture stage and tumor characteristics but not every comorbidity. A health system record may contain rich labs while losing outcomes that occur elsewhere.

FDA guidance asks whether data are relevant and reliable for the question. Relevance means the data contain the population, intervention, comparator, outcomes, and follow-up you need. Reliability includes how data were accrued, curated, transformed, and quality controlled.

Postmarket safety uses several complementary systems#

Spontaneous adverse-event reports can surface unusual patterns rapidly. They are valuable for signal detection, especially when an event is rare, distinctive, and temporally related, but they cannot usually estimate incidence because the number treated and the completeness of reporting are uncertain.

Active surveillance queries defined data networks rather than waiting only for voluntary reports. FDA's Sentinel Initiative uses distributed data from multiple partners to examine medical-product safety while source data remain under partner control. A common data model supports standardized analyses, but important differences in coding and care remain.

Product registries can follow pregnancy, devices, rare diseases, or long-term outcomes. Formal postauthorization studies may test a regulator's specific concern. Case-control, cohort, and self-controlled case series designs each handle time and confounding differently. So do interrupted time series and other designs. No single stream is complete. Convergence across pharmacology, trials, and spontaneous reports makes a causal judgment stronger. So does convergence across active surveillance and independent datasets.

Confounding by indication is the central trap#

In routine care, treatment is chosen for a reason. People receiving one medicine may be sicker, frailer, wealthier, more closely monitored, or more able to adhere than people receiving another; the outcome difference can reflect those starting conditions rather than the treatment.

New-user, active-comparator designs can improve comparability. Instead of contrasting current users with everyone else, the study follows people when they start one of two treatments used for the same decision. Baseline covariates are measured before treatment, and follow-up begins from a shared time origin.

Propensity scores, standardization, inverse-probability weighting, and outcome regression can adjust measured factors. They do not repair an important confounder that was never recorded or measured badly. Negative controls and quantitative bias analyses can probe whether remaining bias could explain the result. The largest database is not automatically the most credible. More observations shrink your confidence interval around a biased estimate just as neatly as around a valid one.

Time can create bias before analysis begins#

Immortal-time bias occurs when a person must survive long enough to be classified as treated, but that guaranteed survival time is credited to treatment. Time-varying treatment can be mishandled when assignment changes after baseline. Conditioning on future adherence or survival can select a distorted group.

Hernán and Robins propose specifying the hypothetical randomized trial the observational study seeks to emulate. The target trial defines eligibility, treatment strategies, and assignment. It defines follow-up, outcomes, causal contrast, and analysis. Aligning time zero, eligibility, and treatment assignment prevents several avoidable biases.

The framework does not make observational data randomized. It forces the question and time structure into the open. When emulation is impossible because key variables are missing, that limitation becomes visible to you before a polished estimate is produced.

Outcomes need validation#

A code for myocardial infarction may perform well for hospitalization but poorly in outpatient problem lists. Kidney injury can be defined from diagnosis codes, laboratory changes, or both. Bleeding definitions vary by site, severity, and transfusion. Small changes can alter event counts and apparent treatment effects.

Validation compares the algorithm with a credible reference such as chart review or adjudication in a sample. Positive predictive value alone is not enough if sensitivity is very low. Transport also matters: a code set validated in one country, insurer, or year may not perform the same way in yours. Researchers should publish executable phenotype definitions where possible, state whether assessors were blinded, and report missingness. Changing an outcome algorithm after seeing group differences risks selective analysis.

Missing data are produced by care#

In a clinical record, a missing lab is rarely random: it may mean a clinician saw no need to test, a person missed a visit, testing occurred elsewhere, insurance ended, or illness prevented follow-up. Complete-case analysis can therefore select people with a particular care pathway.

Informative censoring occurs when departure from the dataset relates to prognosis or treatment. Methods can adjust for measured predictors of loss, but assumptions remain. Linkage to mortality, registries, or other systems can improve completeness while introducing matching errors and governance obligations. Digital data add device wear, battery failure, software changes, and selective participation. A stream with millions of measurements from a narrow group is still narrow evidence.

Real-world evidence can support more than safety#

Regulators may consider real-world evidence for label expansions, postapproval requirements, external controls in rare diseases, or confirmation of effectiveness when randomized trials are impractical. FDA's framework evaluates whether the evidence is fit for regulatory use, whether the design provides an adequate scientific basis, and whether conduct meets regulatory requirements.

Pragmatic randomized trials can use routine-care systems for enrollment, treatment delivery, or outcome collection. Randomization and real-world data are not opposites. A randomized design embedded in practice can retain protection against confounding while improving operational relevance. European regulators are developing DARWIN EU to generate evidence from health-care databases across Europe. Common methods and federated analyses can increase scale, but governance, data harmonization, and local validation remain necessary.

A protocol should exist before the result#

Postmarket analyses are vulnerable to trying many definitions, lags, subgroups, and models. A prespecified protocol and statistical analysis plan reduce the opportunity to select the most favorable answer. Public registration, versioned code, data provenance, and transparent deviations strengthen auditability.

Sensitivity analyses should test plausible choices: outcome definitions, grace periods, and unmeasured confounding. They should test competing events, negative controls, and alternative comparators. Replication in a genuinely independent system asks whether the result depends on one coding culture or care network.

The related article on real-world evidence in regulation explains the regulatory pathway, while when real-world evidence helps and misleads focuses on causal design.

Read a real-world claim in seven lines#

Identify the target population, treatment strategy, and comparator. Identify the outcome, time zero, follow-up, and causal contrast. Then ask whether the data measure each element reliably. Check why treatment was chosen, what could be missing, how losses were handled, and whether the analysis was specified before results.

Real-world evidence is valuable precisely because routine care is complex. Credibility comes from modeling that complexity honestly, not from claiming that a dataset is naturally unbiased because it is large or familiar.

Governance belongs in that appraisal. You should know who can access identifiable data, how transformations are logged, which software version produced the result, and whether the sponsor can suppress an unfavorable analysis. Independent replication is difficult when contracts prevent methods or code from being shared. Privacy protection and reproducibility are not opposing goals: controlled environments, synthetic examples, versioned definitions, and auditable query code can reveal the analytical path without releasing person-level records. Evidence is stronger when another qualified team can understand how the raw observations became the reported estimate.

References#

  1. FDA real-world evidence program
  2. FDA framework for the real-world evidence program
  3. FDA considerations for using RWD and RWE
  4. FDA Sentinel Initiative
  5. EMA DARWIN EU
  6. Using big data to emulate a target trial

Questions and answers

Are real-world data the same as real-world evidence?

No. Real-world data are observations collected during care or daily life. Real-world evidence is a clinical inference produced by analyzing fit-for-purpose data with an appropriate design.

Can routine records detect adverse effects missed before approval?

Yes. Very large populations and longer follow-up can reveal rare, delayed, or subgroup-specific patterns, although signals require evaluation before causation is claimed.

Does a larger database remove confounding?

No. More records can narrow random error while leaving systematic bias untouched or making a biased estimate look more precise.

Can real-world evidence replace randomized trials?

Sometimes it can answer questions that trials cannot or support regulatory decisions, but suitability depends on the causal question, data, design, and remaining uncertainty.

Why are data collected for billing or care hard to analyze?

Codes and records serve operational purposes. Diagnoses, outcomes, adherence, reasons for treatment, and care received elsewhere may be missing or measured inconsistently.