Evidence explainer

Evidence and research methods

Prespecified Versus Post Hoc Analyses: How to Judge the Difference Fairly

Prespecification protects a confirmatory claim by fixing key decisions before results can influence them. Post hoc analysis is not worthless, but it answers a more exploratory question.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What prespecification is trying to protect
  2. Registration, protocol, and analysis plan
  3. An outcome name is not enough
  4. Primary and secondary do not mean important and unimportant
  5. Multiplicity exists even when only one result is printed
  6. Subgroups need more than separate P values
  7. When changing the plan is legitimate
  8. Post hoc is a timing description, not an insult
  9. How to compare plans with publications
  10. Registration does not solve every bias
  11. A practical hierarchy for reading claims

Research data can support many reasonable analyses. Outcomes can be defined in several ways, measured at several times, and adjusted with different covariates. They can be split into subgroups and transformed using different models. If investigators choose among those options after seeing the results, conventional confidence intervals and P values no longer have their usual confirmatory meaning.

Prespecification reduces that problem by recording important decisions before results can influence them, but it does not guarantee that a study is unbiased, and it does not make every later analysis invalid. It changes how confidently you can treat a result as a test of a prior hypothesis rather than a pattern discovered in the same data.

What prespecification is trying to protect#

A confirmatory test assumes that the hypothesis and analysis rule were not selected because they happened to fit the observed noise. Without constraints, a researcher might inspect several time points, outcomes, subgroup cutoffs, covariate sets, and missing-data methods, then present the most favorable result as though it were the only test considered.

No misconduct is required. Humans naturally notice coherent patterns and forget unproductive paths. Prespecification creates a record of the intended test and makes analytical flexibility visible, and that protection is strongest when the plan is written before enrollment or, at minimum, before allocation is unblinded and outcome data are examined. Timing depends on design, but the principle is consistent: the decision must precede access to information that could steer it.

Registration, protocol, and analysis plan#

A trial registry provides a public, timestamped record of core design elements. Prospective registration helps readers identify unpublished studies, compare planned and reported outcomes, and see version history. Registry entries are necessarily concise and may not contain enough analytical detail.

The protocol describes objectives, design, participants, and interventions. It describes outcomes, sample size, oversight, and broad analysis. SPIRIT guidance supports complete trial protocols, while CONSORT guidance supports transparent reports.

The statistical analysis plan, or SAP, translates the protocol into operational rules. It can define analysis populations, estimands, covariates, and model forms. It can define multiplicity procedures, missing-data methods, and intercurrent events. It can define subgroup tests, sensitivity analyses, and output conventions. The three documents should align, but none of them automatically overrides another. So when details conflict you need the dates, the versions, the amendment history, and a stated hierarchy.

An outcome name is not enough#

“Pain at three months” sounds specific until the choices are listed. Which instrument? Which score or subscale? Change from baseline or final value? One visit or a time window? Mean score, responder proportion, or time to improvement? How are deaths and missed visits handled?

The same issue applies to cardiovascular events, remission, hospitalization, or quality of life. Complete prespecification identifies the outcome domain, exact measure, and metric. It identifies the aggregation method, assessment time, and analysis population. Modern estimand thinking also asks what effect is targeted in the presence of intercurrent events such as treatment discontinuation, rescue therapy, or death. Vague registration can preserve almost as much flexibility as no plan at all: a timestamp proves that words existed, not that those words constrained the analysis.

Primary and secondary do not mean important and unimportant#

The primary outcome is the main confirmatory target used to anchor sample size and control false-positive risk; secondary outcomes can be clinically important and rigorously prespecified, but their interpretation depends on multiplicity and the trial's testing strategy.

Exploratory outcomes can identify safety signals, mechanisms, and future hypotheses. Labeling them honestly does not make them scientifically worthless. Problems arise when a secondary or newly defined outcome is promoted to primary after results are known without transparent explanation. A paper may also report the registered primary outcome faithfully while its abstract and conclusion lean on a more favorable secondary result, so outcome switching includes hierarchy and emphasis, not only omission.

Multiplicity exists even when only one result is printed#

If twenty analyses were tried and one is reported, the article can still look like a single test, and the unreported analytical search still affects the probability that the displayed result arose by chance.

Multiplicity can come from several endpoints, repeated time points, or doses. It can come from subgroups, models, interim looks, and alternative definitions. Prespecified procedures can control a familywise error rate, false-discovery rate, or hierarchical sequence, depending on purpose. Not every analysis needs a formal adjustment, but confirmatory claims should state which family of tests was protected and how. Confidence intervals need the same context: a nominal 95 percent interval attached to a selected result does not automatically account for the selection process.

Subgroups need more than separate P values#

A common error is to report that treatment was significant in one subgroup and not significant in another, then conclude the effects differ: the correct question is whether there is evidence of an interaction.

Credible subgroup evidence has a prespecified variable and cutoff, biological or clinical rationale, limited number of tests, adequate information, a formal interaction analysis, and consistency across related outcomes or studies. Even then, subgroup effects are often estimated imprecisely.

Post hoc subgroups can be valuable when an unexpected safety pattern or mechanistic clue emerges. They should be described as exploratory and tested in new data rather than converted immediately into a treatment rule.

When changing the plan is legitimate#

Rigid adherence to a flawed plan can be worse than a justified amendment. A statistical assumption may prove untenable, a measurement instrument may become unavailable, an external event may disrupt visits, or a newer method may address missing data more appropriately.

The strongest amendment is made before unblinding, documented with date and rationale, and approved through the proper governance process. It is reflected consistently across registry, protocol, SAP, and report. If outcome data were already available, the report should say who had access and what was known.

Changes after results are examined can still answer useful questions, but they should be labeled post hoc, accompanied by the original analysis where possible, and treated as sensitivity or exploratory work. Transparency lets you decide whether the reason is compelling.

Post hoc is a timing description, not an insult#

Many scientific advances begin with unexpected observations. Post hoc analysis can investigate why an intervention failed, identify a measurement problem, explore treatment heterogeneity, or generate a mechanism for later testing.

Its weakness is using the same data to discover and confirm a pattern; the estimated effect can be exaggerated, subgroup boundaries can be optimized, and uncertainty usually ignores the search. Replication in new data, preferably with the new hypothesis and analysis prespecified, converts discovery into a more credible test. Exploratory work is strongest when the authors show the full analytical context, avoid causal language the design cannot support, and make code and specifications available where possible.

How to compare plans with publications#

Do not look only at the current registry page. Registries preserve version histories, and the entry you are looking at may have been edited after the study finished. Identify the first entry, dates of enrollment, and each amendment. Identify the final protocol, SAP date, and publication.

Build a table of planned and reported primary and secondary outcomes. Compare definitions, time points, metrics, analysis populations, and hierarchy. Note omitted outcomes, newly introduced outcomes, and changes in emphasis.

Then read explanations. A discrepancy is a finding about reporting, not proof of intent. Some registry fields are entered incorrectly, protocols evolve, and journals impose space limits. The appropriate response is transparent reconciliation, not automatic accusation.

Empirical studies continue to find discrepancies between prespecified and published outcomes in trials and other intervention studies. A 2026 BMJ cohort of registered cohort-intervention studies found that primary outcomes were often changed, newly introduced, or omitted. The exact frequency in one sample should not be generalized to all research, but it reinforces the need to inspect source documents.

Registration does not solve every bias#

A perfectly prespecified analysis can still suffer from poor randomization, missing data, or measurement error. It can suffer from nonadherence, biased outcome assessment, an inappropriate model, or selective publication of the entire study. A registered hypothesis can also be based on earlier undisclosed analyses of related data. Prespecification narrows one family of problems, data-driven analytical choice and selective reporting, and it should be combined with sound design, data quality, blinded procedures where feasible, complete reporting, code review, and replication.

A practical hierarchy for reading claims#

The greatest confirmatory weight usually belongs to a clearly defined primary analysis in a prospectively registered study, supported by a dated protocol and SAP, with appropriate multiplicity control and minimal unexplained deviation.

Prespecified secondary and sensitivity analyses can strengthen interpretation, transparent deviations may be reasonable but deserve scrutiny, and post hoc analyses can explain a result and generate hypotheses. Undisclosed changes are the different case. You find them only by comparing documents, and they weaken confidence because you cannot reconstruct the analytical path.

This hierarchy is not a mechanical score. The scientific question, design, and effect size still matter. So do consistency, risk of bias, and external evidence.

Sources and further reading

  1. ICMJE, Clinical Trial Registration Recommendations
  2. COMPare Project, Goldacre and colleagues, A Cohort Study of Outcome Reporting in Trials, Trials (2019)
  3. Mayo-Wilson and colleagues, What Are the Odds That a Trial Is Prespecified?, BMC Medicine (2018)
  4. Consistency of Prespecified Outcomes in Registered Cohort Intervention Studies, BMJ (2026)
  5. CONSORT and SPIRIT, Reporting Guidance and Resources

Questions and answers

Is every unregistered analysis post hoc?

Not necessarily. It may have been specified in a dated protocol or SAP that was not placed in the registry. Readers need access to the document and its timing to assess the claim.

Can investigators improve an analysis after the trial starts?

Yes. Amendments made before relevant results are known can improve rigor. Their date, rationale, and information access should be documented.

Should post hoc findings be ignored?

No. They can be informative and generate strong hypotheses. Their uncertainty is greater because the data helped select the question, so new-data confirmation is usually needed. Prespecification earns trust through constraint and traceability. Post hoc analysis earns value through candor, complete context, and a clear path to confirmation.