Regression to the mean is the tendency for an extreme measurement to be followed by one closer to the average when repeated measurements are imperfectly correlated. It requires no treatment effect, no biological recovery, and no faulty instrument. It follows from selecting an observation that contains both a persistent signal and a temporary fluctuation.
This matters whenever care or research begins at a peak: very high blood pressure, severe pain, a low mood score, a poor laboratory result, or an unusually successful clinic month. Improvement afterward may be real, but the before-and-after change alone cannot tell you how much of it the intervention caused.
Key points#
- A measurement combines a more stable underlying state with short-term variation and measurement error.
- Selecting people because the observed value is extreme also selects people whose temporary fluctuation was extreme.
- The return toward average can mimic treatment benefit in uncontrolled before-and-after comparisons.
- Randomized concurrent controls experience the same selection process and reveal the change beyond it.
- Repeating measurements before selection, prespecifying analyses, and using appropriate baseline adjustment reduce misinterpretation.
Signal plus fluctuation#
Consider an observed systolic blood pressure as:
observed value = underlying level + temporary variation + measurement error
The underlying level differs among people and can change over time. Temporary variation comes from pain, stress, sleep, recent activity, medication timing, and ordinary physiology. Measurement error comes from cuff fit, posture, device precision, and reading conditions.
Suppose a program enrolls only people with a reading above 180 mm Hg. Some have a persistently high underlying level. Others crossed 180 partly because temporary variation pushed the observed value upward. At the next measurement, that temporary component is unlikely to be equally high. The group average falls even without an effective program.
The same logic works in the other direction. Students selected for exceptionally high test scores may score lower next time. Hospitals selected for an unusually low complication rate may later look worse. Regression refers to movement toward the relevant population mean, not inevitable worsening or improvement.
Why imperfect correlation is enough#
If two repeated measurements were identical, the first would predict the second perfectly and no regression would occur, and if they are related but not identical, an extreme first value predicts a second value that is extreme in the same direction but usually less so.
This is not a claim that every individual moves toward average. Some become more extreme, some stay similar, and some cross to the other side. It is a pattern in the conditional average among people selected by the first measurement.
The amount depends on reliability. Highly reliable measures show less regression. Noisy measures show more. Restricting eligibility to more extreme values strengthens selection on the temporary component and often increases the apparent change.
How treatment gets undeserved credit#
Many people seek care when symptoms are at their worst. Treatment begins, and symptoms improve. If that has happened to you, several processes may have contributed:
- a true treatment effect;
- the natural course of an episodic condition;
- changes in other care or behavior;
- expectancy and reporting effects;
- regression to the mean.
An uncontrolled before-and-after study combines them. Subtracting the baseline mean from the follow-up mean does not isolate treatment, and a p value testing whether change differs from zero only shows that the average changed under the combined processes. Patient testimonials are especially vulnerable because the story often begins at a crisis point. The person's improvement can be genuine while the inferred cause remains uncertain.
A concurrent control reveals the excess change#
The most direct protection is to apply the same eligibility rule and timing to a comparison group; in a randomized trial, both groups are selected at the same extreme baseline and should experience similar regression, natural history, and measurement schedules on average. The difference between randomized groups estimates the assigned treatment effect under the trial assumptions.
Imagine both groups begin with an average pain score of 9. After four weeks, the treatment group averages 5 and the control group 7. The four-point within-group fall is not the treatment effect. The randomized between-group contrast is two points, subject to uncertainty and missing-data considerations.
A nonrandomized control can help but may differ in why or when it was selected. Historical controls are vulnerable to calendar changes. A wait-list group can differ in expectations and co-interventions. Design remains more important than a sophisticated adjustment after the fact.
Repeat measurements can improve eligibility#
If treatment decisions permit, confirm an extreme value with standardized repeat measurements. Home or ambulatory blood-pressure monitoring, repeat laboratory tests, or symptom diaries can estimate the underlying level more reliably than one peak value.
This is not always safe or appropriate. Some extreme findings require immediate action, and waiting solely to improve a research estimate would be wrong. The principle applies when repetition is clinically acceptable.
Using the average of several pre-intervention measurements reduces random noise; it also changes the target population, because eligibility now reflects persistent elevation rather than one transient peak, and a report should state which of the two it means. Run-in periods can stabilize measurement and adherence before randomization in a similar way, but they may exclude people unable to complete the run-in and reduce generalizability.
Baseline analysis choices matter#
In randomized trials with a continuous follow-up outcome, analysis of covariance commonly models follow-up while adjusting for baseline, and under appropriate assumptions it is usually more precise than comparing change scores and handles chance baseline imbalance more appropriately.
If you see a paper test each group's change from baseline and declare efficacy because one is significant while the other is not, that is invalid. The relevant test compares the groups directly. Percentage change can be unstable when baseline values are small and can create distribution problems.
In observational data, adjusting for baseline does not automatically solve selection or confounding. If treatment choice depends on the extreme measurement, treated and untreated people may differ in severity, prognosis, and care-seeking. The design must account for why each entered treatment.
Analyzing change against the initial value is also mechanically biased because baseline appears in both variables, and if change is follow-up minus baseline, a high baseline tends to correlate with a more negative change even without a true relationship. Specialized methods and repeated data are needed to study whether initial level modifies change.
Quality improvement and performance dashboards#
Regression to the mean also affects institutions. A quality program may target the worst-performing clinics this quarter. Their rates often improve next quarter partly because extreme rankings contained chance variation. Rewarding the program for the full change overstates its effect.
The opposite happens when high performers are selected for an award and later decline. The decline does not prove complacency. Funnel plots, shrinkage estimates, risk adjustment, repeated periods, and controlled interrupted time series can help separate signal from noise.
Control groups for policies can be difficult, but several pre-intervention time points are far better than one, and they show the underlying trend, the seasonality, and whether the baseline you were handed was an unusual spike.
Limits: regression is not the same as several related ideas#
Natural history is a real change in the condition over time. Measurement error is one source of imperfect repeatability. Regression to the mean is the statistical consequence of conditioning on an extreme observation under imperfect correlation.
The placebo effect refers to changes related to expectations and treatment context, although “placebo response” in an uncontrolled group also contains natural history and regression. Mean reversion in financial time series can involve dynamic processes and should not automatically be equated with the measurement-selection phenomenon.
Regression also does not mean the population average changes. The overall mean can stay constant while the initially high subgroup moves down and the initially low subgroup moves up.
A reader's checklist#
Ask:
- Were participants, clinics, or time periods chosen because a value was extreme?
- How reliable is that measurement?
- Was the baseline a single reading or a standardized average?
- Was there a concurrent group selected by the same rule?
- Was treatment assigned randomly?
- Is the reported effect a between-group contrast or only within-group change?
- Could natural course, co-interventions, or selective follow-up explain the pattern?
- Were multiple pre-intervention measurements available?
- Did analysts correlate change with the baseline value in a way that creates mathematical coupling?
- Does the conclusion claim more than the design can separate?
Regression to the mean is not an argument that improvement is unreal. It is a reason to require a fair comparator before deciding what caused it.
Sources and further reading
Questions and answers
Does everyone with an extreme result move toward average?
No. It is an average tendency across repeated measurements, not a prediction that applies to every individual.
Can repeating a test eliminate regression to the mean?
Repeating under standardized conditions can reduce noise and better estimate the underlying state. It does not eliminate biological variation or replace urgent action when needed.
Does randomization prevent regression to the mean?
It does not prevent the pattern. It distributes it across treatment groups so the randomized between-group contrast is not credited with the common change.
Is improvement in a placebo group all regression to the mean?
No. It can also include natural history, concomitant care, expectancy, reporting effects, and measurement change.