Medicine has always used evidence in the ordinary sense of observation. A symptom changed, a wound healed, a patient deteriorated. Evidence-based medicine, or EBM, added a harder question: what methods make an observation reliable enough to guide another patient?
The movement that acquired its name in the early 1990s did not invent experiments, statistics, clinical judgment, or respect for patients. It organized them into a repeatable practice: formulate an answerable question, find the best available research, appraise its validity and relevance, combine it with clinical expertise and the person's values, and evaluate the result.
That history is not a straight march from anecdote to algorithm. It is a continuing effort to make comparisons fair, knowledge cumulative, decisions transparent, and uncertainty visible.
Fair comparisons came before the name#
The basic problem is ancient. Patients can improve because of natural recovery, fluctuating disease, placebo and context effects, concurrent care, or selective memory, and a treatment given only to people expected to recover can look effective even when it is not.
Early clinicians attempted concurrent comparisons to separate treatment from background change. James Lind's 1747 scurvy experiment is a famous example. He placed 12 sailors into six pairs, kept their basic circumstances similar, and compared proposed remedies. Citrus produced the clearest recovery. Lind did not report random allocation, and the sample was tiny, but the comparison was prospective and concurrent.
In nineteenth-century France, Pierre Charles Alexandre Louis promoted a “numerical method” for comparing outcomes across groups, including his analysis of bloodletting. His data and methods would not meet modern standards, yet the project challenged therapeutic authority with counted outcomes. Between them, Lind and Louis established the habit EBM later formalized: plausible theory and senior opinion have to face comparative observation.
Randomization changed causal confidence#
Alternating treatments can create balanced groups but allows recruiters to predict the next assignment. Coin tosses and random-number methods remove that predictable pattern. Concealment keeps recruiters from using foreknowledge to alter who enters.
The UK Medical Research Council's 1948 trial of streptomycin for pulmonary tuberculosis became a landmark: a central schedule based on random sampling numbers assigned 107 young patients to streptomycin plus bed rest or bed rest alone. Assignments stayed concealed until eligibility was confirmed. Radiographs were assessed without treatment identifiers or chronology.
The trial showed substantial early benefit and also revealed emerging drug resistance. Its influence came from both the result and the clear design. It gave clinical investigators a model for creating comparable groups and documenting how bias was limited.
Randomized trials then spread through therapeutics. Statistical methods for survival, repeated measurements, and cluster allocation broadened what trials could answer. So did methods for noninferiority and adaptive designs. Ethical oversight and informed consent developed through a separate but intersecting history.
Clinical epidemiology moved methods to the bedside#
Population epidemiology had studied patterns and causes of disease. Clinical epidemiology focused those tools on questions arising in patient care: How accurate is this test? What is the prognosis? Does treatment help? What predicts harm?
John R. Paul used the term clinical epidemiology in 1938. Alvan Feinstein later developed rigorous approaches to clinical measurement, prognosis, comorbidity, and reasoning. At McMaster University, David Sackett and colleagues built clinical epidemiology into medical education from the late 1960s onward.
The practical insight was that clinicians did not need to become full-time statisticians to use research methods, though they did need enough skill to recognize selection bias, confounding, chance, weak measurement, and an inapplicable population.
In 1981, McMaster authors published a series of guides in the Canadian Medical Association Journal to help clinicians read research critically. The method became known as critical appraisal. It converted abstract epidemiology into questions a reader could use during care.
Archie Cochrane made effectiveness a health-system question#
Archie Cochrane's 1972 book Effectiveness and Efficiency argued that limited health resources should support interventions shown to work through reliable evaluation, with a strong emphasis on randomized trials.
Cochrane's concern extended beyond one clinician reading one paper. Health systems needed organized, current summaries of all relevant trials. Otherwise, effective care could remain unused while ineffective or harmful care persisted.
During the 1980s, researchers assembled systematic reviews of pregnancy and childbirth interventions. Cochrane praised that cumulative approach and urged other specialties to follow it. In 1993, the Cochrane Collaboration was established to prepare and maintain systematic reviews across health care, and that was the institutional shift: evidence synthesis became planned infrastructure rather than an occasional narrative review by an expert choosing the studies they already knew.
The term evidence-based medicine#
At McMaster around 1990, Gordon Guyatt used “scientific medicine” for a residency curriculum centered on critical appraisal and bedside use of research, but the phrase provoked resistance because it implied that existing medicine was unscientific. “Evidence-based medicine” became the more durable name.
The Evidence-Based Medicine Working Group's 1992 JAMA article announced a new approach to teaching practice. It challenged intuition, unsystematic experience, and pathophysiologic rationale as sufficient grounds for decisions. It emphasized efficient searching and formal appraisal of clinical research.
The rhetoric of a “new paradigm” was intentionally sharp. It helped distinguish the curriculum, but it also created misunderstandings. Critics heard that experience and patient individuality no longer mattered.
The 1996 clarification#
David Sackett and colleagues responded in a short 1996 BMJ article on what EBM was and was not. Their formulation joined three elements:
- the best relevant research evidence;
- the clinician's expertise in diagnosis, context, and care;
- the individual patient's values, circumstances, and preferences.
No element could replace the others. Research without clinical judgment could be applied to the wrong person. Expertise without external evidence could become outdated. A technically favorable option could conflict with what the person values or can sustain. The article also rejected “cookbook medicine.” Population evidence provides estimates and options; it does not eliminate uncertainty or make every patient interchangeable.
The Users' Guides created a shared language#
From 1993 through 2000, the EBM Working Group published a long series of JAMA Users' Guides to the Medical Literature; the guides taught readers to appraise therapy, diagnosis, prognosis, harm, economic analysis, systematic reviews, and clinical decision rules.
Concepts such as absolute risk reduction, number needed to treat, and likelihood ratios moved into routine teaching. So did confidence intervals and applicability. Structured clinical questions came to be organized around population, intervention, comparator, and outcome, often called PICO.
The guides did not make appraisal mechanical. They made hidden judgments discussable. You could say why a paper was credible, or why it did not answer the question at the bedside.
Information retrieval was part of the revolution#
Critical appraisal is useless if you cannot find the relevant evidence in time. Printed indexes, library searches, and journal clubs once made retrieval slow. MEDLINE, PubMed, the Cochrane Library, online journals, and point-of-care summaries changed the practical scale.
Better access also created overload. You cannot fully appraise millions of articles during a consultation. Evidence services began organizing information into alerts, critically appraised topics, and systematic reviews. They organized it into guidelines and clinical decision support. This produced a new responsibility: appraise the evidence product, not just its citations. A convenient summary can be outdated, commercially influenced, or disconnected from its source.
Reporting standards made missing methods visible#
You cannot judge a method the authors did not report. The CONSORT statement, first developed in the 1990s and updated through CONSORT 2025, specifies essential information for randomized trial reports. Related standards address diagnostic accuracy, observational studies, systematic reviews, prediction models, and health-economic evaluations.
Reporting guidelines do not guarantee good methods. They make omissions easier to see and replication more feasible. Trial registration, public protocols, and statistical analysis plans add a time-stamped record before results are known. That is the real evolution: the job grew from “read the paper” to “compare the paper with the protocol, the registry, and the full evidence record.”
GRADE separated evidence certainty from recommendations#
By the late 1990s, organizations used many incompatible grading systems. A letter grade could refer to study design in one guideline and recommendation strength in another.
The GRADE Working Group began in 2000 and developed a transparent framework. For each important outcome, certainty can be rated high, moderate, low, or very low. Randomized evidence can be rated down for risk of bias, inconsistency, indirectness, imprecision, or publication bias. Observational evidence can sometimes be rated up under defined circumstances.
GRADE also separates certainty from recommendation strength. A recommendation considers the balance of benefits and harms, values, and resources. It considers equity, acceptability, and feasibility. Strong evidence for a small effect does not automatically create a strong recommendation. Evidence-to-Decision frameworks put those additional judgments on the page instead of hiding them behind a show of hands.
Shared decision-making completed the bedside loop#
EBM's early writing focused heavily on clinician skills. Shared decision-making made the patient's role operational.
The clinician explains options, likely benefits and harms, and uncertainty in understandable terms. The patient contributes goals, past experience, tolerance for burden, and preferences. Together they choose a course and revisit it as outcomes or priorities change.
Decision aids, natural frequencies, absolute risks, and teach-back can improve the conversation. Shared decision-making is not merely asking the patient to choose without guidance. It is a joint process informed by credible evidence and professional interpretation.
Critiques improved the movement#
Several criticisms identify real failures in the evidence ecosystem:
- trials can exclude the people who most need care;
- industry funding and academic incentives can shape questions and reporting;
- statistically significant changes may be clinically trivial;
- guidelines can multiply into burdensome single-disease rules;
- averages can obscure meaningful variation;
- publication bias hides unfavorable results;
- access barriers and social conditions limit feasible choices;
- overreliance on hierarchies can devalue qualitative evidence and mechanism.
These are not reasons to return to intuition alone. They are reasons to improve transparency, representation, and outcome selection. They are reasons to improve data access, synthesis, conflict management, and patient involvement. EBM also became broader evidence-based health care, nursing, public health, policy, and practice. Each field adapted the core methods to different interventions and decision contexts.
Reproducibility and open science#
The twenty-first-century reproducibility movement emphasized public protocols, data and code availability, computational checking, replication, and correction. Registered reports separate publication decisions from study results. Living systematic reviews update conclusions as new evidence arrives.
These practices treat the published article as one view of a research process rather than the whole record. They also acknowledge that error is normal in science and that correction should be designed into the system.
Artificial intelligence is a retrieval tool, not an evidence grade#
Modern language and search systems can help formulate questions, screen citations, extract structured data, and summarize large literatures. They can also invent citations, flatten uncertainty, mix current with outdated guidance, and obscure why a source was selected.
An AI-generated answer does not become evidence because it sounds synthesized. Evidence provenance, source date, and study design still require checking. So do risk of bias, effect size, and applicability. Human accountability remains essential for clinical interpretation. Used well, automation shortens the distance between your question and the relevant research; it does not do the appraising for you.
The five recurring actions#
Many teaching models summarize EBM as a cycle:
- Ask an answerable question.
- Acquire the best available evidence.
- Appraise validity, size, and relevance.
- Apply it with expertise and patient priorities.
- Assess the decision and improve the process.
The cycle is more important than memorizing a hierarchy. A strong decision may use randomized benefit evidence, observational safety data, and diagnostic studies. It may use qualitative research and the patient's own response over time.
The cumulative conclusion#
Evidence-based medicine began before its name and continues beyond its original curriculum. Its history joins fair comparison, randomization, clinical measurement, literature retrieval, systematic review, transparent grading, and shared decisions.
Its deepest commitment is not to a pyramid or a P value. It is to disciplined uncertainty. Claims should be testable, evidence should be findable, methods should be appraisable, conclusions should be proportional, and decisions should fit the person and context.
The history remains unfinished because the evidence system is never finished. New studies revise estimates, new methods reveal bias, and patient priorities change what counts as a good outcome.
References#
- Evidence-Based Medicine Working Group. Evidence-based medicine: a new approach to teaching the practice of medicine. JAMA. 1992.
- Sackett DL, et al. Evidence based medicine: what it is and what it is not. BMJ. 1996.
- Medical Research Council. Streptomycin treatment of pulmonary tuberculosis. BMJ. 1948.
- Cochrane. Archie Cochrane and the history behind the organization.
- Sackett DL. Clinical epidemiology: what, who, and whither. Journal of Clinical Epidemiology. 2002.
- Schunemann HJ, et al. The history and evolution of GRADE. GRADE Book. 2026.
- Schunemann HJ, et al. Grading the certainty of evidence. Cochrane Handbook.
- Hopewell S, et al. CONSORT 2025 statement. BMJ. 2025.
Questions and answers
Who invented evidence-based medicine?
No single person invented its methods. Gordon Guyatt and colleagues established the name and teaching movement, building on centuries of trials, statistics, epidemiology, critical appraisal, and evidence synthesis.
Did EBM begin in 1992?
The influential JAMA article appeared in 1992. Controlled trials, clinical epidemiology, critical appraisal, and systematic review had developed earlier.
Does EBM replace clinical expertise?
No. Expertise is needed to diagnose, judge applicability, explain options, perform care, and respond to the individual. EBM asks expertise to work with external research rather than substitute for it.
Is EBM only about randomized trials?
No. Randomized trials are powerful for treatment effects. Diagnosis, prognosis, harms, prevalence, mechanism, implementation, and experience require other designs as well.
What is the biggest modern change in EBM?
Evidence is increasingly treated as a living system: preregistered studies, public protocols, structured reporting, continuous synthesis, explicit certainty ratings, and decisions made with patients.