Evidence explainer

Chronic disease in primary care

How Common Is Diagnostic Error, and What Do the Estimates Count?

There is no single rate of diagnostic error, because definitions, denominators, and detection methods differ. Burden models show scale and concentration; they do not count cases.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. First decide what the word means
  2. Separate error from harm
  3. The denominator changes the headline
  4. Six windows onto the problem
  5. What the 795,000 estimate means
  6. Why the Big Three finding matters
  7. Why estimates legitimately disagree
  8. A five-question reading check

Diagnostic error is common enough to be a major patient-safety problem, but no single percentage describes it across all care, because a chart review, an autopsy series, a malpractice database, and a national disease model measure different events in different populations. Their numbers become useful only after you ask what counted as an error, what counted as harm, and what denominator was used.

First decide what the word means#

The National Academies defined diagnostic error as failure to establish an accurate and timely explanation of a patient's health problem, or failure to communicate that explanation to the patient, and this definition is broader than choosing the wrong disease name. A diagnosis can eventually be correct but unacceptably delayed. A sound conclusion can also fail the patient if it is never communicated or acted upon.

The definition is deliberately patient centered, but measurement still requires judgment. “Timely” depends on the condition and the available information. A one-day delay has different meaning for a rapidly evolving infection than for a slowly changing chronic problem. Accuracy can also evolve as a disease declares itself. Uncertainty at an early visit is not automatically an error if the uncertainty is recognized, dangerous possibilities are addressed, and follow-up is appropriate.

Researchers therefore often use a related idea: a missed opportunity to make a correct or timely diagnosis based on the information reasonably available at the time. This approach focuses review on an actionable process rather than treating every later change in diagnosis as proof that the first clinician failed.

Separate error from harm#

An error does not always injure a patient. A wrong label may be corrected before treatment changes. Conversely, diagnostic harm can be profound when a time-sensitive condition is missed, even if the disease was unusually difficult to recognize.

At least four outcomes may appear in the literature:

  1. Diagnostic disagreement, such as a later interpretation differing from the original one.
  2. A diagnostic error under a defined review standard.
  3. A missed opportunity that reviewers judge preventable or actionable.
  4. Harm attributable to the diagnostic process, sometimes restricted to death or permanent disability.

A study counting any disagreement will produce a different rate from one counting only severe, attributable harm. Neither is inherently wrong. They answer different questions.

The denominator changes the headline#

Even with one definition, the rate depends on what you put below the fraction. Possible denominators include all outpatient visits, all diagnoses, all patients with a particular disease, all closed malpractice claims, all hospital deaths, or an entire national population for one year.

Consider a study that begins with people later proven to have stroke and asks how many were initially missed. That produces a disease-based error rate. Another study may begin with all dizziness visits and ask how often a serious cause was missed. That produces an encounter-based rate. The two figures cannot be substituted for each other because the populations and clinical questions differ.

Lifetime statements introduce another denominator. The National Academies concluded that most people are likely to experience at least one diagnostic error during their lives, though a lifetime probability accumulates many encounters, and you should not read it as a per-visit error rate.

Six windows onto the problem#

Retrospective chart review#

Reviewers reconstruct a diagnostic timeline and judge whether earlier information offered a missed opportunity, and this can reveal process failures across visits, but documentation is incomplete and hindsight can make the eventual diagnosis seem more obvious than it was.

Electronic triggers#

Health systems can search for patterns such as an unexpected hospital admission soon after an outpatient visit, or a new cancer diagnosis after a prior abnormal result. A trigger efficiently finds records worth reviewing. It does not prove an error by itself, and its yield depends on the rule and data quality.

Autopsy and diagnostic comparison studies#

Autopsy studies compare clinical conclusions with findings after death. Second readings of images or pathology specimens can identify important discrepancies. These methods can reveal missed disease, but they sample selected populations and may not represent everyday ambulatory care.

Malpractice claims and incident reports#

Claims contain rich accounts of severe events, while voluntary reports can identify hazards that routine data miss. Both are filtered. Many errors never lead to a claim or report, and the events that do may differ systematically from those that do not.

Standardized patients and vignettes#

Researchers can present the same scripted problem to multiple clinicians and compare decisions against a defined answer; this improves control and allows direct comparison, but performance in a simulation may differ from performance with real patients, teams, and time constraints.

Disease-burden models#

Models combine disease incidence, estimated diagnostic-error rates, and estimated probabilities of serious harm. They can describe national scale even when no reporting system counts every event. Their output inherits uncertainty from every input and from assumptions used to combine them.

What the 795,000 estimate means#

Newman-Toker and colleagues built a national, disease-based model of serious harm from dangerous conditions being misdiagnosed, and their central estimate was about 795,000 deaths or permanent disabilities per year in the United States. The probabilistic plausible range was approximately 598,000 to 1,023,000.

That figure is not a registry tally. It is a modeled estimate created by combining disease frequency with error and harm rates drawn from clinical literature. The authors also compared the result with setting-based estimates and performed analyses using more conservative assumptions. The appropriate interpretation is an estimated order of magnitude with substantial uncertainty, not an exact annual count.

The model also should not be converted casually into a percentage of all medical encounters. It starts from dangerous diseases and estimates serious outcomes. It does not count every minor, rapidly corrected, or communication-only error covered by the broader National Academies definition.

Why the Big Three finding matters#

The same analysis found that vascular events, infections, and cancers accounted for about 75.8 percent of serious misdiagnosis-related harms, and within those categories a smaller group of diseases contributed a large share. Stroke, sepsis, pneumonia, venous thromboembolism, and lung cancer were prominent contributors.

Concentration makes prevention more tractable. Health systems can build disease-specific detection triggers, reliable test-result follow-up, escalation pathways, and feedback around high-harm presentations. That does not make uncommon diagnoses unimportant. It means a focused program can potentially address a meaningful fraction of severe harm without pretending that one generic checklist solves every diagnostic problem.

Why estimates legitimately disagree#

Two careful studies may report very different numbers because they differ in setting, disease mix, severity threshold, follow-up period, reviewer standard, or denominator. A study can also count a process failure even when the final diagnosis is correct, while another counts only wrong final diagnoses. Some ask whether an error occurred; others ask whether it was preventable or caused harm.

The most useful response to disagreement is not to average the percentages. It is to map each result to its question. How often is a serious disease missed at the first visit? How many patients suffer permanent harm nationally? How many closed claims involve diagnosis? How many abnormal tests lack timely follow-up? These are complementary measures of a system, not rival estimates of one universal rate.

A five-question reading check#

Before you repeat a diagnostic-error statistic, ask:

These questions do not weaken the case for diagnostic safety. They make the evidence more actionable. A credible measure identifies the part of the diagnostic process being observed and supports an intervention that can change it.

Sources and further reading

  1. National Academies, Improving Diagnosis in Health Care (2015)
  2. AHRQ, Foundational Terminology for Diagnosis Relating to Suboptimal Processes and Outcomes
  3. Newman-Toker and colleagues, Burden of Serious Harms From Diagnostic Error in the United States, BMJ Quality and Safety (2024)
  4. AHRQ, Measure Dx Resource To Identify and Learn From Diagnostic Safety Events
  5. AHRQ, Diagnostic Error in the Testing Process

Questions and answers

Is every delayed diagnosis a diagnostic error?

No. Some diseases cannot be identified with reasonable confidence at the first encounter. Delay becomes concerning when information available at the time supported a more accurate or timely explanation, when necessary follow-up failed, or when the uncertainty and safety plan were not communicated.

Does a second opinion that changes a diagnosis prove the first diagnosis was negligent?

No. Disagreement can flag a case for review, but it does not determine whether the earlier conclusion was unreasonable. The available information, disease evolution, uncertainty, and care process all matter.

Should the 795,000 estimate be quoted as a precise count?

No. It is a model-based central estimate with a broad plausible range. It supports the conclusion that serious diagnostic harm is large and concentrated, while leaving uncertainty about the exact total.