Evidence explainer

Evidence and research methods

How to Audit Difference-in-Differences and Interrupted Time Series Studies

Both designs estimate a policy effect without randomization. Their credibility rests on one thing, which is whether the untreated future was reconstructed plausibly.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. One policy question, two designs
  2. Reading an interrupted time series
  3. The stable-trend assumption
  4. Reading difference-in-differences
  5. Parallel trends are not certified by one p-value
  6. Timing can contaminate the estimate
  7. Other threats shared by both designs
  8. A compact audit for readers

Quasi-experiments try to recover a causal comparison from a change that was not randomized: a city introduces a smoke-free law, a health system changes a coverage rule, or a clinic adopts a new protocol on a known date. Researchers cannot observe the same population both with and without that change, so they construct the missing outcome from time trends, comparison groups, or both. Reading the study well means finding that constructed counterfactual and asking yourself how easily it could be wrong.

One policy question, two designs#

Suppose a hospital introduces an electronic reminder intended to increase vaccination. Monthly vaccination rates are available for two years before and two years after launch.

An interrupted time series analysis uses the hospital's own pre-launch pattern as the counterfactual. It estimates whether the rate changed immediately at launch, whether the post-launch slope differs from the earlier slope, or both.

A difference-in-differences analysis adds a hospital that did not introduce the reminder. It compares the change at the intervention hospital with the change at the comparison hospital. If both were affected by a national campaign, that shared movement should subtract out. The two designs answer the same broad question through different assumptions, and combining them in a controlled interrupted time series can be stronger than relying on either alone, provided the comparison series is genuinely informative.

Reading an interrupted time series#

A useful interrupted time series needs many observations on both sides of a clearly defined intervention date, and a simple before-and-after average is not enough because it cannot distinguish a policy effect from a trend already underway. Segmented regression instead models the underlying time pattern and estimates two main quantities:

Those effects have different substantive stories. A new prescribing restriction might produce an immediate drop. A training program may take months to alter practice, making a gradual slope change more plausible. Credible analyses state the expected pattern before selecting a model. Trying many breakpoints and shapes after viewing the data increases the chance of mistaking noise for an effect.

Time observations also violate the assumption that each point is unrelated to its neighbors. This serial correlation can make uncertainty look smaller than it is. Seasonality matters when respiratory illness, staffing, school terms, or holiday patterns affect the outcome. Models should account for these structures, and graphs should show enough pre-intervention time to reveal them.

The stable-trend assumption#

The interrupted time series estimate depends on a claim about the unobserved future: without the intervention, the earlier trajectory would have continued in a predictable way. This cannot be observed directly. A long, stable pre-period makes the claim more plausible, but it never proves it.

Ask what else changed at the same time. A new coding definition, supply shortage, public campaign, staffing reorganization, or unrelated policy can create the same visual break. Check whether the population entering the measure changed. If the denominator shifts from all eligible patients to only active patients, an apparent improvement may be measurement rather than care. Control outcomes and control series help here. If a reminder should affect vaccination but not an unrelated laboratory test, the unrelated measure can reveal a system-wide documentation change, and if a similar hospital without the reminder shows the same break, the local intervention is a weak explanation.

Reading difference-in-differences#

Difference-in-differences uses a double subtraction. Let the intervention group's outcome rise from 40 to 55, a change of 15 points. Let the comparison group rise from 35 to 43, a change of 8. The difference-in-differences estimate is 7 points. It attributes the common 8-point rise to background forces and the remaining 7 to the intervention.

This removes fixed differences between groups and shocks shared at the same time. It does not remove every difference. Everything rests on one requirement: without the intervention, the groups' outcomes would have changed in parallel. Their levels do not need to match, but their untreated trajectories must be comparable.

Why was the comparison group chosen? A nearby hospital may share labor markets and public-health campaigns, but it may also treat a different population or react to the policy indirectly, while a distant hospital may avoid spillover while experiencing different regional conditions. Selection should follow a causal argument, not merely which dataset was convenient.

Authors often report that pre-intervention trends were not statistically different and describe the assumption as satisfied. That is too strong. A low-powered test can miss a meaningful divergence, and a few pre-period points may reveal little about future behavior. Conversely, a tiny detectable difference may matter less than the overall trajectory.

Inspect the graph yourself. Are the groups moving similarly for a meaningful interval? Are estimates from an event-study plot near zero before treatment, with uncertainty shown? Were comparison groups selected before outcomes were examined? Would alternative comparison groups tell the same story?

Placebo dates are useful: if the model reports an effect before the intervention occurred, the design or trend model is suspect. Placebo outcomes that the intervention should not affect test a different route to spurious findings. Neither check proves the counterfactual, but each makes an implausible story easier to detect.

Timing can contaminate the estimate#

Policy adoption is often staggered. Some states begin in January, others in July, and some never adopt; a once-standard two-way fixed-effects model can then compare newly treated groups with untreated groups and with groups treated earlier. If effects grow, fade, or differ across groups, these comparisons can receive unexpected weights and produce a summary that does not represent any clear policy contrast.

Modern difference-in-differences methods address this by defining effects for each adoption cohort and time period, using never-treated or not-yet-treated units as appropriate comparison groups, then aggregating transparently. You do not need to reproduce the estimator to ask the key questions: who serves as the control at each time, are already-treated units used as controls, and are effects allowed to vary by cohort and time since adoption?

Other threats shared by both designs#

Several problems cross design labels:

Standard errors also need to reflect clustering when outcomes within a state, hospital, or person are related, because a precise-looking estimate built on only a few policy units may overstate what the design can tell you.

A compact audit for readers#

Begin with the graph, not the final coefficient. Mark the intervention date, trace the pre-period, and identify every group being used to represent the untreated future. Then ask yourself:

  1. Was the intervention timing determined without reference to the outcome?
  2. Is the expected immediate or delayed effect scientifically plausible?
  3. Are pre-trends long and comparable enough to support projection?
  4. Could another event, spillover, or measurement change explain the break?
  5. Are seasonality, serial correlation, clustering, and staggered timing handled?
  6. Do placebo checks, alternative windows, and alternative controls preserve the conclusion?

The final estimate earns causal weight only as far as those questions have convincing answers. Quasi-experiments are powerful because they turn real-world changes into structured comparisons. Their weakness is also visible: the missing counterfactual remains an argument, not an observed group created by randomization.

Sources and further reading

  1. Bernal, Cummins, and Gasparrini, interrupted time series regression tutorial, International Journal of Epidemiology (2017)
  2. Wing, Simon, and Bello-Gomez, designing difference-in-differences studies, Annual Review of Public Health (2018)
  3. Callaway and Sant'Anna, difference-in-differences with multiple time periods, Journal of Econometrics (2021)
  4. Goodman-Bacon, difference-in-differences with variation in treatment timing, Journal of Econometrics (2021)

Questions and answers

Are quasi-experiments the same as randomized trials?

No. Both seek causal effects, but quasi-experiments rely on naturally occurring timing or comparison groups. Their assumptions about trends and assignment require more substantive defense.

Does a flat pre-period prove parallel trends?

No. It supports the assumption within the observed window but cannot show what would have happened after the intervention. Limited data may also hide meaningful divergence.

Which design is stronger?

Neither label guarantees quality. A well-controlled time series with a clear intervention can be stronger than a difference-in-differences analysis with a poor control group. Design credibility comes from the counterfactual and its checks.