Evidence explainer

Evidence and research methods

Bayesian and Frequentist Trial Results Answer Different Questions

Frequentist methods describe how data behave over repeated sampling. Bayesian methods give a probability about the effect itself. They answer different questions.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. One hypothetical trial, two summaries
  3. The frequentist question
  4. The Bayesian question
  5. Priors are inputs to inspect, not defects to hide
  6. Probability of any benefit is not probability of enough benefit
  7. What a reanalysis can and cannot do
  8. A reader's checklist

Bayesian and frequentist analyses can use the same trial data and report different-looking answers because they are designed to answer different questions. Neither framework rescues a weak design. Your job when you read one is to work out which question it answered, look at the assumptions it needed, and decide whether the number it reports is the number your decision turns on.

Key points#

One hypothetical trial, two summaries#

Consider a randomized trial of a treatment intended to reduce 30-day mortality: the estimated risk ratio is 0.85, with a 95 percent confidence interval from 0.70 to 1.03 and a two-sided p value of 0.09.

A conventional frequentist summary is that the trial did not cross the prespecified 0.05 threshold. The interval does not exclude no average difference, and it also remains compatible with a clinically meaningful benefit. It would be wrong to translate p equals 0.09 into a 91 percent probability that the treatment works.

A Bayesian analysis could use the same outcome data, specify a prior distribution for the risk ratio, and calculate a posterior probability that the ratio is below 1. It might report, for example, a high probability of any mortality reduction but a lower probability of a reduction large enough to change practice. Those probabilities are conditional on the chosen prior, likelihood, outcome definition, and model. The two summaries are not contradictory; they answer different questions and may emphasize different decision thresholds.

The frequentist question#

In frequentist inference, the treatment effect is treated as fixed and unknown. Probability describes what would happen to data or procedures over hypothetical repetitions.

A p value begins with a model that includes a null hypothesis, such as no average treatment effect; it calculates how often the specified analysis would produce data at least as incompatible with that hypothesis as the observed data, assuming the model and null are correct. It does not compare the probability of the null with the probability of a benefit.

A 95 percent confidence interval comes from a procedure designed so that, under its assumptions, 95 percent of intervals generated across repeated samples contain the true parameter. For the one interval in hand, standard frequentist language is that the procedure produced a range of values compatible with the data at that confidence level. The framework does not assign a 95 percent probability to the fixed parameter lying within this particular interval.

These long-run guarantees are valuable. They can control false-positive rates in confirmatory trials and allow a design to be audited against a prespecified analysis plan, and their interpretation becomes misleading only when you convert them into posterior probabilities they were not built to provide.

The Bayesian question#

Bayesian inference treats uncertainty about the effect as a probability distribution. The analysis has three main components:

  1. The prior distribution describes uncertainty before the current trial data are incorporated.
  2. The likelihood describes how compatible different effect sizes are with the observed data under the model.
  3. The posterior distribution updates the prior with the likelihood.

From the posterior, researchers can calculate the probability of any benefit, the probability of exceeding a clinically important threshold, or the probability of harm, and a 95 percent credible interval contains 95 percent of the posterior probability for the effect under the stated model and prior. That direct language is attractive when you have a decision to make, but it does not remove assumptions; it makes one of them, the prior, especially visible.

Priors are inputs to inspect, not defects to hide#

Every analysis contains choices: outcome definitions, covariates, missing-data rules, functional forms, and decision thresholds. A Bayesian prior adds another explicit choice about pretrial information.

Priors can serve different purposes:

Borrowing is not automatically valid. Earlier studies may use different populations, endpoints, comparators, or standards of care. If those differences make the historical information nonexchangeable with the current trial, a strong prior can produce false precision.

A transparent report gives the prior's shape and parameters, explains its evidence base, and shows sensitivity analyses. If a neutral prior gives a positive conclusion and a plausible skeptical prior reverses it, the data alone did not settle the decision. If reasonable priors converge, the finding is more robust.

Probability of any benefit is not probability of enough benefit#

Bayesian reports often highlight a posterior probability that the effect is on the favorable side of zero. That can sound decisive even when the expected gain is tiny.

Suppose the posterior assigns 96 percent probability to any reduction in symptom score but only 42 percent probability to improving by the minimum amount patients consider worthwhile. Both statements can be correct. The first answers whether the direction is favorable. The second addresses clinical value. The same distinction exists in frequentist reading, where a very precise, statistically significant effect can be too small to matter, and both frameworks improve when the clinically important threshold is prespecified and reported on an absolute scale.

What a reanalysis can and cannot do#

A Bayesian reanalysis of a completed frequentist trial can reveal information hidden by a binary significant or nonsignificant label. The ANDROMEDA-SHOCK reanalysis is a useful example: it translated the trial into posterior probabilities under several priors, allowing readers to see how conclusions changed with prior assumptions.

Such a reanalysis does not change randomization, outcome quality, missing data, treatment crossover, or external validity. It cannot turn a biased estimate into an unbiased one. It also should not be chosen after inspecting the result merely because it sounds more favorable. Prospective specification remains important, particularly when an analysis will support a regulatory or clinical decision.

Regulatory guidance reflects this distinction. FDA has long provided guidance for Bayesian medical-device trials; in January 2026 it issued draft guidance for drug and biological-product trials, including use in primary inference, trial monitoring, subgroup analyses, and incorporation of external information. The document is draft guidance, not a binding final rule, and that status belongs beside any claim drawn from it.

A reader's checklist#

When you have a frequentist analysis in hand, ask:

For a Bayesian analysis, add:

For either framework, return to the design. Randomization, allocation concealment, outcome measurement, missingness, adherence, and representativeness usually matter more than the philosophical label on the analysis.

Sources and further reading

  1. FDA 2026 draft guidance on Bayesian methodology in drug and biological-product trials
  2. FDA guidance on Bayesian statistics in medical-device clinical trials
  3. Therapeutic Innovation and Regulatory Science, tutorial on modern Bayesian methods in clinical trials
  4. American Journal of Respiratory and Critical Care Medicine, Bayesian reanalysis of ANDROMEDA-SHOCK

Questions and answers

Is Bayesian analysis more subjective?

It makes the prior explicit, which can look more subjective. Frequentist analyses also contain consequential judgments. The appropriate response is transparency, justification, and sensitivity analysis rather than pretending choices are absent.

Can a Bayesian trial control false positives?

Yes. Designers can simulate repeated trials under specified scenarios and choose posterior decision thresholds with acceptable operating characteristics. Bayesian probability statements and frequentist design evaluations can coexist.

Which framework should a clinician prefer?

Prefer the analysis that answers the decision-relevant question with assumptions that are visible and credible. Often the most informative report presents the effect size, uncertainty, absolute outcomes, and sensitivity analyses clearly enough that readers do not have to choose a statistical identity before understanding the evidence.