“Registry” describes several information systems that serve different jobs. A cancer registry records tumors and outcomes. A device registry follows implants. A quality registry compares care processes and results. A clinical-trial registry creates a public record of planned research. A transplant waiting list is operational as well as informational.
The shared idea is structured, purposeful collection for a defined population. The differences are just as important. A trial registry does not contain the same person-level longitudinal detail as a disease registry, and a voluntary product registry does not have the population coverage of a statutory national registry.
Start with the registry's purpose#
The AHRQ user's guide defines a patient registry as an organized system using observational methods to collect uniform clinical and other data for a specified population, with a scientific, clinical, or policy purpose. That purpose drives every later choice.
A natural-history registry may follow people from early disease through treatment and outcomes. A product registry may monitor long-term safety after a device or medicine reaches wider use. A procedure registry may assess complications and practice variation. A quality registry may feed timely performance information back to participating sites.
Trying to make one registry serve every purpose can weaken all of them. Minimal fields improve completion but may omit important confounders. Detailed forms capture nuance but increase cost and missingness. Real-time quality feedback needs faster data than a carefully adjudicated research endpoint.
The protocol should specify the target population, inclusion criteria, and setting. It should specify outcomes, schedule, and the analysis plan. It should specify governance and intended users. A registry assembled first and explained later invites data-driven questions that its structure was never built to answer.
Coverage determines the denominator#
A population-based registry seeks to capture all eligible people in a geographic or administrative population; coverage can support incidence, prevalence, and survival estimates when case ascertainment and migration are understood. A center-based registry represents patients who reached participating centers. A voluntary online registry represents people who learned about it, had access, and chose to enroll.
These designs can all be useful. Problems arise when a narrow denominator is described as if it were universal. Referral centers may see more severe or unusual disease. Insurance databases omit people without that coverage. Procedure registries exclude patients who never received the procedure.
Report how participants enter, how many eligible people are captured, why records are excluded, and which groups are underrepresented. Compare registry distributions with an external population when possible. The number of rows does not tell you who is missing.
A common data dictionary creates comparability#
Uniform collection requires definitions. What counts as disease onset: first symptom, first test, or diagnosis date? Is death from any cause or disease-specific? Is a complication self-reported, coded, adjudicated, or confirmed by laboratory criteria?
A data dictionary defines variable names, permitted values, and units. It defines timing, source, and handling of unknown values. Version control records when a definition changes. Without it, a registry can mix old and new criteria and still hand you a single trend line.
Core outcome sets can improve comparison across centers and studies. They should include outcomes important to patients, not only values convenient to extract. Burden matters too. Repeated long questionnaires can increase dropout and selectively lose people with worse health or fewer resources. Quality assurance can include automated range checks, logic checks, and duplicate detection. It can include source verification, training, audit samples, and feedback to sites. Perfect data are impossible; visible error processes are more credible than a claim of completeness.
Follow-up loss changes the story#
A registry can start with broad enrollment and still become selective over time. People may move, change insurance, or receive care elsewhere. They may withdraw, die, or stop responding. Loss is especially biased when it relates to treatment response, disability, income, language, or dissatisfaction.
Retention should be reported by time and subgroup. Passive linkage to vital records, claims, or electronic records can recover outcomes, subject to authority and data quality. Active follow-up can collect symptoms and function that administrative sources miss. Combining methods often provides a fuller record.
Missingness is not a single category. “Not measured,” “not applicable,” “unknown,” and “lost to follow-up” carry different information. Converting all of them to a blank cell destroys that distinction.
Registry comparisons are vulnerable to confounding#
If people receiving treatment A have better outcomes than people receiving treatment B, treatment may be responsible. They may also differ in disease severity, age, or access. They may differ in clinician selection, contraindications, calendar time, or follow-up.
Statistical adjustment helps only for measured factors recorded accurately. A registry designed for quality reporting may lack the reason a treatment was chosen. Center effects can combine expertise, volume, referral patterns, and local resources. Missing severity can dominate the result.
Credible comparative studies specify a target trial, align eligibility and treatment start, use appropriate comparators, validate outcomes, and test sensitivity to unmeasured confounding; the related registry-study strengths and limits guide shows how to appraise these choices. Registries are particularly strong for description, uncommon conditions, rare outcomes, long-term trajectories, and hypothesis generation. Causal claims require additional design, not a larger font on the sample size.
Product registries can extend safety follow-up#
Premarket studies may be too small or short to identify rare or delayed harms. A product registry can collect implantation details, operator factors, and revisions. It can collect failures, pregnancy outcomes, or other events over time.
FDA's 2023 guidance on registry data asks whether the population, timing, variables, and follow-up are relevant to the regulatory question and whether data accrual and quality are reliable. It also addresses governance, common definitions, missing data, linkage, and access for verification.
A product registry can suffer channeling bias if higher-risk patients receive one device, and surveillance bias if one group is watched more intensively; voluntary enrollment may miss failures treated elsewhere. Linkage to claims, mortality, or device identifiers can strengthen ascertainment when governance permits.
Trial registries solve a different problem#
Clinical-trial registries record protocols rather than routine care. Prospective registration places the title, sponsor, and design in a public system before participant outcomes are known. It also places interventions, eligibility, outcomes, and timing there.
This matters because a completed study can disappear if results are unfavorable. A primary outcome can be replaced by a more favorable secondary one. Analyses can be presented as planned even when chosen after data review. Registration creates a dated comparison point.
ClinicalTrials.gov includes registration and results-reporting requirements under United States law and policy for specified studies. WHO's International Clinical Trials Registry Platform links data from recognized primary registries. ICMJE requires public trial registration at or before first participant enrollment as a condition of publication in member journals.
Registration is not full transparency by itself. Records can be late, incomplete, vague, or not updated. Results can remain unreported. Compare the registration, the protocol, and the statistical plan. Compare the publication and the results record before you trust any one of them.
A public record improves accountability only if maintained#
A trial record should update recruitment status, protocol changes, completion, and results. Changes need dates and reasons. A registry that says “recruiting” years after a study ended misleads participants and reviewers.
Outcome fields should be specific enough to compare with the publication. “Safety and efficacy” is not a usable prespecification. Name the measure, metric, aggregation, and time point. Secondary and exploratory outcomes should remain distinguishable.
Identifiers should connect the registry record, protocol, publications, data-sharing statement, and regulatory reports. This chain reduces ambiguity when titles and author lists change, and the article on trial registration and results reporting examines those duties in detail, while preregistration and registered reports explains related methods beyond clinical trials.
Privacy and participation need governance#
Longitudinal registries can contain genetic data, rare diagnoses, and locations. They can contain procedures and dates that make reidentification possible even after direct identifiers are removed. Governance should define access, permitted uses, and retention. It should define linkage, security, breach response, and oversight.
Consent or another lawful basis should match the registry's function and jurisdiction. Broad future-use language is not unlimited permission. Community and participant representatives can help set priorities, acceptable uses, return-of-results policies, and communication.
Data sharing can improve reproducibility while protecting participants through controlled access, data-use agreements, and minimization. Secure environments and disclosure review help too. A claim that data are “anonymous” needs technical and legal scrutiny.
A registry appraisal in eight questions#
Ask what decision the registry was built to support. Define the target population and enrollment route. Check coverage and retention. Inspect variables, sources, and versioned definitions. Determine how outcomes were validated. Identify treatment-selection mechanisms and missing confounders. Review linkage quality and governance. Finally, compare the analysis with a dated protocol.
A registry is neither a weak trial nor a perfect mirror of practice. It is a designed observation system. Its credibility comes from knowing which people and events it can see, which it cannot, and how that boundary shapes the claim you are being asked to believe.
Useful registries also return value to the people and sites supplying data. Feedback can show missing fields, delayed follow-up, outcome variation, or an emerging safety question, and public reports should protect confidentiality and explain case-mix adjustment so that ranking does not punish centers caring for more complex populations. If you have contributed data, you can reasonably ask what the registry learned, whether its purpose changed, and how to withdraw from future contact where withdrawal is available; a system that only extracts data and never communicates its use may satisfy a technical protocol while weakening the trust needed for long-term completeness.
References#
- AHRQ registries user's guide
- FDA guidance on assessing registries for regulatory use
- WHO International Clinical Trials Registry Platform
- ClinicalTrials.gov reporting requirements
- ICMJE clinical-trial registration policy
- STROBE observational-study reporting guideline
Questions and answers
Is a patient registry the same as a clinical trial registry?
No. A patient registry collects longitudinal data about defined people or care. A trial registry publicly records planned studies and, in many cases, their results.
Can registry data prove that one treatment caused a better outcome?
Usually not by simple comparison. Treatment selection, missing variables, follow-up, and outcome definitions can confound observational estimates.
Why register a clinical trial before enrollment?
Prospective registration creates a dated public record of the question, outcomes, design, and sponsor before results can influence what is disclosed.
What makes a registry representative?
Broad and known coverage, clear eligibility, high capture, transparent exclusions, and follow-up across relevant settings support representativeness. Large size alone does not.
Can registry information be linked to other datasets?
Yes, where law, consent or another legal basis, governance, and secure methods permit it. Linkage can fill outcome gaps while introducing matching error and privacy risk.