Evidence explainer

Evidence and research methods

What Good Clinical Evidence Looks Like

Good clinical evidence is not one prestigious design or a small p-value. It is a precise question, a design that can answer it, and a report you are able to audit.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Begin with the question, not the evidence pyramid
  2. Make the causal contrast explicit
  3. What randomization can and cannot do
  4. Risk of bias belongs to a result
  5. Precision is not the same as lack of bias
  6. Read magnitude before the p-value
  7. Patient-important outcomes and surrogate endpoints
  8. Harms need their own evidence strategy
  9. Transparency makes appraisal possible
  10. Replication and robustness
  11. Systematic reviews are methods, not prestige labels
  12. Certainty belongs to a body of evidence
  13. Applicability and equity
  14. A rapid but disciplined reading sequence
  15. Sources

Good clinical evidence earns your confidence in a specific conclusion. It begins with a clear question and uses a design capable of answering it, and the study then needs credible conduct, relevant outcomes, enough information to distinguish important benefit from uncertainty, and transparent reporting that allows verification.

No single feature settles the appraisal. Randomization does not correct a badly measured outcome. A large sample does not remove confounding. Peer review does not reveal missing studies. A meta-analysis cannot make biased inputs reliable. Evidence quality is the alignment of question, design, conduct, analysis, reporting, and application.

Begin with the question, not the evidence pyramid#

A therapy question asks whether assigning an intervention changes outcomes compared with an alternative. A diagnostic question asks how a test classifies or predicts in the intended pathway. A prognosis question asks what happens over time. A harm question may require very large cohorts or pharmacovigilance because rare events are not captured well in ordinary trials.

Population, intervention, comparator, outcome, and time, often shortened to PICO plus time, make the claim testable. “Drug A works” is vague. “In adults with condition X, does adding Drug A to current care reduce symptomatic relapse over 12 months compared with current care alone, and at what risk of serious harm?” can guide design and interpretation.

The best design follows the question. Randomized controlled trials are powerful for treatment effects because random assignment can balance measured and unmeasured baseline causes on average. Cohort studies can estimate prognosis and uncommon or delayed harms. Cross-sectional designs can estimate prevalence. Diagnostic-accuracy studies need representative participants and an appropriate reference standard.

Evidence hierarchies are useful reminders, not automatic verdicts, and a randomized trial with major loss to follow-up and selective outcome reporting can be less credible than a carefully designed observational study for the same question. The design creates opportunities for validity; execution determines whether they were realized.

Make the causal contrast explicit#

Every effect is a comparison; improvement from baseline does not show that an intervention caused the change because symptoms can fluctuate, regression to the mean can occur, and other care may change. A concurrent control group estimates what would have happened under an alternative strategy.

The comparator should be clinically relevant. Placebo can isolate efficacy where ethical and appropriate, but it may not establish advantage over the best current treatment. “Usual care” must be described because it differs by setting and time. Noninferiority trials need a justified margin and evidence that the active control would have worked in the trial context.

The treatment effect being estimated also needs definition. Assignment to a strategy despite nonadherence answers a policy-like question. Effect while adhering answers a different question and usually needs stronger assumptions; ICH E9(R1) calls this precise target an estimand, including how events such as discontinuation, rescue treatment, or death are handled. Changing the causal contrast after seeing the data invites biased storytelling, which is why protocols and statistical analysis plans should specify the primary question before outcomes are unmasked, with any exploratory analyses labeled honestly.

What randomization can and cannot do#

An unpredictable allocation sequence prevents clinicians and participants from choosing the next assignment. Concealment before assignment protects against selection bias. Masking after assignment can reduce differences in care and outcome assessment, although it is not feasible for every intervention.

Randomization balances baseline factors in expectation, not perfectly in every finite trial. Chance imbalances can occur. Prespecified covariate adjustment can improve precision, but post hoc adjustment should not be used to manufacture a preferred result.

Randomization does not prevent bias from missing outcomes, differential co-interventions, crossover, unmasked subjective assessment, protocol deviations, or selective reporting. It also does not make an irrelevant endpoint patient-important or a narrow sample broadly applicable.

Cluster trials, crossover trials, and adaptive trials have additional design-specific concerns. Analysis must respect the unit of randomization, period and carryover, or adaptation rules. The word randomized is the beginning of appraisal, not the end.

Risk of bias belongs to a result#

Cochrane's RoB 2 framework evaluates bias in a specific randomized-trial result across five domains: the randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result.

This result-level focus matters. Mortality may be measured completely and objectively while a patient-reported symptom in the same study has high attrition and unmasked assessment. Calling the entire study simply “good” or “bad” loses that distinction.

Missingness depends on why data are absent and whether absence relates to the unobserved value, and ten percent missing at random may be less concerning than five percent missing mainly among participants whose condition worsened. Sensitivity analyses can ask how strong departures from assumptions would need to be to change the conclusion.

Selective reporting occurs when investigators measure many outcomes, time points, definitions, or models and publish the favorable subset. Comparing the article with the registry, protocol, and analysis plan is how you find the switches and omissions. A polished manuscript cannot substitute for that audit.

Precision is not the same as lack of bias#

Random error produces uncertainty even in an unbiased study. Confidence intervals show a range of effect estimates compatible with the data and model. A narrow interval can exclude important effects; a wide interval may include substantial benefit, no meaningful difference, and harm.

Large samples improve precision but do not correct systematic error. A million biased records can yield a very precise wrong answer. Conversely, a small well-designed trial may be unbiased but too imprecise for the decision you face.

Stopping a trial early for apparent benefit can exaggerate effects, especially with few events. Repeated interim looks require prespecified statistical control. Very low event counts make estimates fragile even when participant counts sound large.

“No statistically significant difference” is not evidence of equivalence. Equivalence and noninferiority require prespecified margins, appropriate design, and intervals that exclude unacceptable differences. Absence of evidence and evidence of absence are different.

Read magnitude before the p-value#

A relative risk of 0.75 means a 25% relative reduction. If baseline risk is 40%, risk falls to 30%, an absolute reduction of 10 percentage points. If baseline risk is 0.4%, it falls to 0.3%, an absolute reduction of 0.1 percentage point.

Both can be statistically convincing in a large trial, but their practical implications differ. Absolute effects depend on baseline risk and often vary across settings or patient groups. Number needed to treat is the inverse of an absolute risk difference over a stated time, so it should never be reported without outcome and horizon.

For continuous outcomes, a mean difference can conceal response distribution. A two-point average change might be noticeable or trivial depending on the scale, variability, measurement reliability, and patient priorities. Dichotomizing a continuous outcome can make interpretation easier but loses information and can create arbitrary thresholds. Clinical importance should be prespecified where possible. A minimal important difference is context-dependent and uncertain, not a universal line between meaningful and meaningless.

Patient-important outcomes and surrogate endpoints#

People usually care about living longer, feeling or functioning better, avoiding disability, and reducing treatment burden. Biomarkers and surrogate endpoints can shorten trials, but a change in a marker does not always translate into those outcomes.

A valid surrogate must reliably capture the effect of treatment on the clinical outcome in the relevant disease and intervention class. Correlation between marker and outcome is not enough. A treatment can improve a marker while causing offsetting harm through another pathway.

Composite outcomes can increase event counts but may be driven by frequent, less important components while the most serious component is unchanged; each component, definition, and competing event should be examined.

Outcome ascertainment should be valid, reliable, and similar across groups. Central adjudication can reduce some bias, but the clinical relevance of the definition remains. A technically clean endpoint can still answer the wrong question.

Harms need their own evidence strategy#

Trials are often powered for common benefits, not rare serious harms. Eligibility restrictions and short follow-up can further limit safety inference. “No difference in adverse events” may mean too few events or incomplete collection rather than proven safety.

Harms should be defined, actively collected where appropriate, graded, timed, and reported by randomized group. Discontinuation, dose modification, hospitalization, and quality of life can reveal burden beyond a list of event names.

Observational studies, registries, spontaneous reports, and linked records can detect rare or delayed signals after broader use. They add scale and duration while introducing confounding, selection, and measurement problems. Consistency across designs can be more informative than insisting on one perfect source.

Benefit and harm also compete within individuals. A small average benefit may be worthwhile for some risk profiles and not others. Good evidence reports both sides without turning a population estimate into a personal instruction.

Transparency makes appraisal possible#

Prospective trial registration records the question, outcomes, and analysis before results are known. ClinicalTrials.gov includes study records and, for many trials, structured summary results. The site itself cautions that listing is not government approval of the study's safety or science.

Protocols and statistical analysis plans reveal what was intended. CONSORT 2025 provides a 30-item minimum reporting checklist and participant-flow diagram for randomized trials. It improves completeness of reporting but is not a risk-of-bias certificate.

Data and code sharing can allow result reproduction, error detection, and new analyses, subject to consent, privacy, governance, and legitimate participant protections. A statement that data are “available on request” is less transparent when conditions and decision authority are unspecified.

Funding and conflicts should be disclosed. A financial tie does not automatically make a result false, and absence of a commercial tie does not guarantee rigor; it can influence design, comparator, interpretation, and publication, so you need to be told.

Replication and robustness#

Reproducibility can mean rerunning the same code on the same data. Replicability can mean obtaining a consistent result in new data. Both matter, but neither is a mechanical yes-or-no label.

Analytic robustness asks whether reasonable decisions about exclusions, missing data, covariates, outcome definitions, and models lead to similar conclusions. A finding that survives only one unexplained specification is not a finding you can lean on.

External replication can fail because the original was wrong, the new study was underpowered, the populations differ, or the intervention and outcome were not truly the same. The cause should be investigated rather than choosing the preferred result. Repeated small studies do not automatically equal one definitive study. If they share bias, selective publication, or overlapping data, a pooled estimate can be precisely misleading.

Systematic reviews are methods, not prestige labels#

A systematic review begins with a protocol, explicit eligibility criteria, a comprehensive search, duplicate study selection or verification, structured data collection, risk-of-bias assessment, and appropriate synthesis. PRISMA 2020 guides transparent reporting of these steps.

A meta-analysis is the statistical combination, not the review itself. Pooling can improve precision when studies address sufficiently similar questions. It can mislead when interventions, populations, outcomes, or biases differ in ways a single average hides.

Heterogeneity should be explored through prespecified, credible explanations rather than many data-driven subgroups, and a random-effects model acknowledges variation in effects but does not explain it or make every pooled result applicable.

The review can only synthesize accessible evidence. Unpublished or selectively reported results can distort the body. Trial registries, regulatory documents, preprints, and contact with authors may reduce but not eliminate dissemination bias.

Certainty belongs to a body of evidence#

GRADE rates confidence in an effect estimate for a particular outcome and decision. Randomized evidence may begin at high certainty, then be rated down for risk of bias, inconsistency, indirectness, imprecision, or publication bias. Nonrandomized evidence can sometimes support higher confidence when appropriate criteria are met.

Certainty is outcome-specific. Evidence for short-term symptom improvement can be high certainty while evidence for mortality or rare harm is low. It is also context-specific: direct evidence for one population may be indirect for another.

A recommendation adds further considerations such as benefit-harm balance, values, resources, equity, acceptability, and feasibility. Strong evidence does not always mandate one choice, and a strong recommendation can sometimes be made despite limited evidence in exceptional circumstances. Guidelines should keep evidence certainty and recommendation strength visibly apart, and you should not read a recommendation label as if it were the underlying study result.

Applicability and equity#

Eligibility, recruitment, setting, clinician skill, adherence support, comparator, outcome, and follow-up determine how far a result travels. A trial can be internally valid yet indirect for older adults with multimorbidity or for a system unable to deliver the same support.

Average benefit can hide unequal performance or access. Subgroup analysis should focus on credible effect modifiers and avoid stereotyping broad categories. Small subgroup samples often provide uncertainty, not proof of no difference.

Equity appraisal asks who was included, who could use the intervention, who bears burdens, and whether implementation could widen disparities. Digital access, language, cost, transportation, and time can be part of the causal pathway from intervention to outcome.

Applicability is not a license to ignore randomized evidence. It is a structured statement about the target and the evidence gaps, with uncertainty carried into the decision.

A rapid but disciplined reading sequence#

First rewrite the headline as PICO plus time. Locate the registry and protocol, then identify the prespecified primary outcome and analysis. Check participant flow, missing data, crossover, masking, and whether groups received different care beyond the intended contrast.

Read the effect estimate and confidence interval in absolute terms, then compare the range with the threshold that would change your decision. Review harms with the same attention as benefit, and ask whether follow-up was long enough and whether a surrogate stood in for the outcome that matters.

Then move outward: has the result been replicated, does it fit the broader evidence, are important studies missing, and does your target setting differ? Note funding and conflicts without using them as a shortcut.

Your conclusion should be calibrated. “This trial proves” is rarely needed. Better language states what design showed, how large and precise the result was, what biases remain, and where it is likely to apply.

Sources#

  1. Official GRADE Book
  2. Cochrane Handbook for Systematic Reviews of Interventions
  3. Cochrane guidance on Risk of Bias 2
  4. CONSORT 2025 statement for randomized-trial reporting
  5. PRISMA 2020 statement for systematic-review reporting
  6. ClinicalTrials.gov registry and results-database overview

Questions and answers

Are randomized trials always the best evidence?

They are often the strongest design for treatment causality, but not for every question. Conduct, outcome quality, missing data, precision, and relevance can make a particular randomized result weak, while other designs may be necessary for prognosis or rare harms.

Does statistical significance mean an effect is clinically important?

No. Statistical significance does not state the absolute benefit, patient relevance, certainty, burden, or harm. Effect size and interval should be compared with what would change a decision.

Does peer review prove a study is reliable?

No. Reviewers can identify problems, but they rarely audit all source data or rerun every analysis. Registration, protocols, complete reporting, reproducible code, replication, and post-publication correction remain important.

Is a systematic review automatically high-quality evidence?

No. A review can have an incomplete search, biased eligibility, invalid included studies, inappropriate pooling, or outdated evidence. Methods and certainty must be appraised.

What is the fastest way to appraise a clinical claim?

Define the exact question, verify design and comparator, find the prespecified outcome, read effect and interval, inspect attrition and bias, review harms and applicability, then compare with the full evidence body.