Evidence explainer

Evidence and research methods

What Real-World Data Sources Actually Look Like

Real-world data are produced by care, payment, registries, and devices rather than by a traditional trial protocol. Their value depends on whether the source can validly answer a specific question.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Claims data: a map of billable care
  2. Electronic health records: detail inside a care system
  3. Registries: depth within a defined frame
  4. Pharmacy, laboratory, and imaging systems
  5. Digital health technologies and patient-generated data
  6. Public-health and mortality sources
  7. Linkage: complementary data with additional uncertainty
  8. Common data models do not make sources identical
  9. From source data to an analytic variable
  10. Descriptive and causal questions need different designs
  11. A fit-for-purpose checklist
  12. References

The phrase real-world data can make a database sound like an unfiltered recording of reality. It is not. Every source is built for a purpose. A claim supports payment. An electronic health record supports care and operations. A registry follows a defined population or product. A wearable records what its sensors and software can detect. The research dataset you receive is a transformed remainder of those processes.

The FDA definition is broad: real-world data are routinely collected information about patient health status or delivery of health care. Real-world evidence is the clinical evidence about use and potential benefits or risks of a medical product derived from analysis of those data. The distinction matters. A billion rows are data. They become evidence only when you supply a question, a design, an analysis, and a credible interpretation.

Claims data: a map of billable care#

Administrative claims usually contain enrollment dates, diagnosis and procedure codes, places and dates of service, clinician or facility identifiers, and amounts billed or paid, and pharmacy claims can record that a prescription was submitted and adjudicated. Their strength is longitudinal capture across many participating clinicians while a person remains within the covered system.

Their limitation is equally fundamental: claims are created to obtain or document payment, so a diagnosis code may represent a confirmed disease, a suspected condition used to justify a test, a rule-out, or a coding convention. Clinical severity, symptoms, and laboratory values are often absent. So are imaging findings, social context, and reasons for choosing a treatment.

A filled prescription is not proof that a person took it. A medication purchased without insurance may be missing. Services delivered outside the plan may disappear. Enrollment churn can make an apparent end of follow-up look like freedom from events. Claims lag can also change the most recent months as submissions are corrected or adjudicated.

Claims can be strong for questions with well-recorded billable events, such as many hospitalizations or procedures, and they are weaker when the outcome is a subtle symptom, a disease stage not tied to coding, or a clinical judgment that never produces a separate claim.

Electronic health records: detail inside a care system#

Electronic health records can contain vital signs, laboratory results, medication orders, imaging reports, problem lists, clinician notes, allergies, referrals, and patient messages; they may reveal why a decision was made and how the person responded. Their granularity makes them attractive for research.

The record is still an operational artifact. A medication order is not necessarily dispensed, and a copied problem-list item can persist long after it stopped being true; measurements are taken more often in people who are sicker, worried, insured, or already receiving care. Free text contains context but requires natural-language processing or review, both of which can make errors.

Care outside the health system may be absent. A person who stops appearing can be healthy, hospitalized elsewhere, uninsured, relocated, or dead. Different sites use different laboratory methods, units, note templates, and coding habits. A software migration can create an apparent clinical trend that is actually a documentation change.

The FDA guidance on EHR and claims data therefore asks whether sources are appropriate for the specific question, whether key variables have valid operational definitions, and whether data accrual and transformation are controlled. “From the EHR” is a provenance label, not a validity argument.

Registries: depth within a defined frame#

A registry is an organized system that uses observational methods to collect uniform information about a population defined by a disease, condition, procedure, or product. Examples include cancer registries, transplant registries, pregnancy registries, and medical-device registries.

Registries can capture disease stage, product identifiers, procedures, outcomes, and follow-up not reliably available in claims. A well-run registry may use standardized definitions, adjudication, active follow-up, and quality checks. It can support natural-history research, safety surveillance, comparative effectiveness, and postmarket obligations.

The denominator needs scrutiny. Is participation mandatory, population-based, hospital-based, voluntary, or limited to selected centers? A registry from specialty centers may overrepresent severe or unusual cases. Loss to follow-up can depend on prognosis. Data may be richer for participants who return often and sparse for those who do not.

FDA's registry guidance focuses on relevance and reliability. Relevance includes whether the population, intervention, comparator, outcomes, and follow-up fit the question. Reliability includes data accrual, quality assurance, completeness, and verification. A famous registry does not automatically satisfy either requirement for a new use.

Pharmacy, laboratory, and imaging systems#

Dispensing systems can clarify when a medicine left a pharmacy, quantity supplied, refill timing, and sometimes discontinuation, and they improve on prescribing orders but still do not show ingestion, correct technique, sharing, or stockpiling. Cash purchases and pharmacies outside a network can remain invisible.

Laboratory databases offer numeric results with timestamps and reference information. They can support phenotyping and outcome ascertainment when methods are stable. Selection is a central problem: only tested people have values. A missing test is not a normal test. Changes in assays, reference ranges, specimen types, and reporting limits need harmonization.

Imaging archives hold images and reports. Images can support new measurement or machine-learning research, while reports encode clinical interpretations. Protocol, equipment, reconstruction, contrast, and reader behavior vary. Training a model on images obtained only after clinicians suspected disease can produce a dataset unlike a screening population.

Digital health technologies and patient-generated data#

Wearables, home monitors, apps, and implanted devices can record heart rate, rhythm, and movement. They can record sleep proxies, glucose, and blood pressure. They can record symptoms or medication events. Frequent measurement can reveal patterns missed during clinic visits. It also produces device-specific data shaped by wear time, firmware, sensor placement, calibration, and proprietary algorithms.

Missingness is rarely random. People remove devices, lose connectivity, stop using an app, or charge it at particular times. A software update can alter an endpoint without any change in health. Consumer devices may not preserve raw signals or version history. A high-frequency stream can contain millions of correlated measurements from relatively few people, so row count can badly exaggerate the effective sample size you actually have.

Patient-reported data contribute symptoms, function, quality of life, treatment burden, and priorities unavailable in billing records, though the instrument, language, recall window, prompting schedule, accessibility, and completion pattern define what the data mean. A validated questionnaire is not interchangeable with a single app rating.

Public-health and mortality sources#

Vital records, disease surveillance, immunization information systems, and public-health reporting add events beyond one care network, and death indexes can reduce outcome loss, although cause-of-death coding is imperfect and updates may lag. Surveillance data can identify outbreaks or population trends but depend on case definitions, testing, reporting rules, and jurisdiction.

Policy changes can create discontinuities. When testing becomes more available, reported incidence may rise even if true incidence does not. A new mandatory-reporting rule can look like a disease surge. Calendar time, coding transitions, and health-system strain belong in the model, not merely in the limitations paragraph.

Linkage: complementary data with additional uncertainty#

Researchers link claims to records, registries to mortality, mothers to infants, or devices to clinical outcomes because one source rarely contains every needed variable. Deterministic linkage uses shared identifiers. Probabilistic linkage combines imperfect fields such as name, date of birth, address, and service dates.

Linkage success can differ by housing stability, name changes, and data-entry error. It can differ by language, age, and other characteristics. False matches assign another person's event; missed matches erase a true event. If either error differs between comparison groups, linkage can bias the effect rather than merely add noise.

Privacy-preserving linkage can reduce direct exchange of identifiers but does not eliminate governance duties. A paper should give you the match method, the quality metrics, the proportion linked, the duplicate handling, and the sensitivity analyses. “Linked dataset” is not sufficient detail.

Common data models do not make sources identical#

Distributed networks such as the FDA Sentinel Initiative transform local data into a common structure and run standardized queries while data partners retain control, which enables large surveillance studies and consistent code execution.

A common data model harmonizes names, formats, and vocabularies. It cannot create a laboratory result that was never measured, repair an inaccurate diagnosis, or make different care pathways equivalent, and two sites can populate the same field from different workflows. Semantic consistency and clinical comparability remain empirical questions you have to answer.

From source data to an analytic variable#

Most studies do not analyze raw records. They construct a cohort and derive variables. “Diabetes” might mean one diagnosis code, two codes separated in time, or a medication. It might mean an HbA1c threshold, registry confirmation, or a machine-learning phenotype. Each definition changes sensitivity, specificity, and the population included.

Treatment start might be an order date, dispensing date, administration record, or first day covered after a washout period. An outcome might use a code list, a laboratory rule, text extraction, chart review, or adjudication, and these choices should be prespecified and validated against a credible reference in a relevant sample.

The data journey also includes extraction, cleaning, and mapping. It includes deduplication, linkage, date shifting, and analysis. Reproducibility requires versioned code, source refresh dates, data-quality reports, and an audit trail; provenance is part of scientific evidence, because you cannot evaluate an estimate without knowing how it was produced.

Descriptive and causal questions need different designs#

Real-world data can describe utilization, incidence, natural history, and safety signals. A causal question, such as whether treatment A prevents more strokes than treatment B, imposes stronger demands. Treatment choices reflect disease severity, contraindications, clinician preference, access, and patient priorities. Those same factors may affect outcome, creating confounding by indication.

Design tools include an active comparator, new-user cohort, and explicit time zero. They include a baseline covariate window, propensity methods, and negative controls. They include quantitative bias analysis and sensitivity analyses. None guarantees exchangeability. Unmeasured disease severity or treatment rationale can remain. The FDA non-interventional study guidance therefore emphasizes early specification of design and analysis, data relevance and reliability, and transparent handling of confounding, missing data, and misclassification, and a credible emulation of a target trial states eligibility, treatment strategies, assignment, follow-up, outcomes, causal contrast, and analysis before anyone looks at a result.

A fit-for-purpose checklist#

Start with the decision you have to make, not the database. Define the target population, treatment or condition, comparator, outcome, and time horizon. Then ask whether the source observes each element with adequate completeness, timing, and validity.

Inspect who enters and leaves the data. Identify why each field exists operationally. Validate important variables. Check data lag, software changes, code-set revisions, and care outside the system. Describe transformation and linkage. Examine missingness and measurement frequency. For causal work, draw the time line and identify confounders before modeling.

The All of Us data overview illustrates a deliberately multimodal research resource combining surveys, electronic records, physical measurements, wearable data, and genomics. Even a purpose-built program still requires you to understand participation, available releases, harmonization, and which participants contributed each modality.

References#

  1. FDA real-world evidence overview
  2. FDA guidance on EHR and claims data
  3. FDA guidance on registry data
  4. FDA guidance on non-interventional studies
  5. FDA Sentinel Initiative
  6. All of Us data sources

Questions and answers

Are real-world data the same as real-world evidence?

No. Data are the records and measurements. Evidence is a conclusion generated by applying a defensible design and analysis to data suitable for the stated question.

Are electronic health records more accurate than insurance claims?

Neither source wins universally. Records often have richer clinical detail; claims may have broader continuity for billed services. Accuracy depends on the variable and purpose.

Does a very large database remove bias?

No. It can narrow confidence intervals around a confounded or misclassified estimate. Validity depends on design and measurement, not only size.

Why link more than one data source?

Linkage can combine complementary variables, such as clinical measurements, dispensing, and death. False and missed matches can also introduce unequal error and need validation.

Can real-world evidence support an FDA decision?

Yes, for specified purposes. FDA evaluates relevance, reliability, traceability, design, analysis, and how the result fits the total evidence rather than accepting a source label alone.