Evidence explainer

Evidence and research methods

Understanding Standardized Mean Differences

A standardized mean difference is a group difference expressed in standard deviations. It lets you pool studies that used different scales, and it hides the units while doing it.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Why standardize a mean difference
  2. Cohen's d and Hedges' g
  3. The construct must really be shared
  4. Align the direction first
  5. The denominator changes the story
  6. Rough labels need restraint
  7. Translate back to a familiar scale
  8. Change scores, final scores, and baseline adjustment
  9. Small studies and selective reporting
  10. Heterogeneity is more than I-squared
  11. SMD is not a response rate
  12. A checklist for reading a forest plot
  13. Keep the unitless number anchored
  14. References

A standardized mean difference, or SMD, reports the distance between two group means in standard-deviation units, and it is often used in meta-analysis when studies assess the same construct but use different instruments. Depression severity, pain, or physical function may each be measured with several valid scales that cannot be pooled in their original units.

The SMD solves a unit problem. It does not solve differences in population, measurement quality, timing, bias, or clinical meaning. So read a pooled value as a constructed effect measure carrying assumptions, not as a universal ruler.

Why standardize a mean difference#

If every study measures systolic blood pressure in millimeters of mercury, the mean difference preserves a familiar unit. A pooled reduction of 5 mm Hg can be discussed directly.

Suppose instead that six trials measure fatigue using four questionnaires with different ranges and scoring directions. A five-point change on one scale is not the same numerical quantity as five points on another, so the SMD rescales each trial's treatment-control difference by a study standard deviation, placing effects on a common dimensionless scale.

In simplified form:

SMD = (mean in the intervention group minus mean in the comparison group) divided by a pooled standard deviation.

The sign shows direction after scale coding. The magnitude shows how large the mean difference is relative to variability among participants. Standardization permits pooling, but the original units disappear.

Cohen's d and Hedges' g#

Cohen's d commonly uses the pooled within-group standard deviation. In small samples, its absolute magnitude tends to be biased upward. Hedges' g multiplies d by a correction factor that approaches one as sample size increases.

Many meta-analysis programs report Hedges' g while describing the effect generically as an SMD. You have to check the methods, because the label will not tell you which estimator produced the number.

Other denominators exist. Some designs use a control-group standard deviation, baseline standard deviation, change-score standard deviation, or another reference. These choices answer slightly different questions and affect magnitude. They should not be mixed silently.

The sampling variance of the SMD also matters. Study weights in an inverse-variance meta-analysis depend partly on that uncertainty, and a large point estimate from a small study can receive less weight than a precise estimate from a larger study.

The construct must really be shared#

Standardizing numbers does not prove they measure the same thing. Two scales labeled “quality of life” may emphasize physical ability, symptoms, relationships, or emotional well-being differently. Pooling them assumes that each captures a sufficiently common construct.

Time point also matters. Immediate pain after a procedure and pain at six months are not interchangeable because both are continuous; nor should symptom severity and functional ability be combined merely because both concern the same condition.

Eligibility, intervention, comparator, and context need clinical coherence before statistical synthesis. The PICO framework remains relevant. An SMD is a summary measure, not permission to pool unlike questions.

Scale validity should be considered. A poorly validated instrument can add measurement error or respond differently across language and culture. Standardization does not repair those limitations.

Align the direction first#

On some scales, a higher score means worse symptoms. On others, it means better function. If directions are not aligned, genuine benefits can cancel one another in a pooled estimate.

Review authors may multiply one scale's mean values by minus one so all positive effects point in the same clinical direction. That operation should be documented. Forest plots and text must state whether positive or negative favors the intervention.

The sign is a convention, not a quality judgment. An SMD of -0.4 and +0.4 can represent the same magnitude with opposite coding. Find the “favors” labels before you interpret the direction. Errors also arise when change scores and final scores use inconsistent directions. Data extraction should preserve the scale definition, range, and direction for every study.

The denominator changes the story#

Consider two trials in which the intervention improves a familiar score by 5 points more than comparison care. If the pooled standard deviation is 10 in one trial, the SMD is 0.5. If the standard deviation is 20 in another, the SMD is 0.25.

The raw difference is identical. The standardized effect differs because participants in the second study vary more; that variability may reflect a broader population, measurement noise, baseline severity, or genuine response heterogeneity.

Cochrane notes that SMD pooling assumes differences in standard deviations mainly reflect differences in measurement scale rather than real differences in variability between study populations, and when that assumption is doubtful, an SMD can entangle treatment effect with population heterogeneity.

This is one reason a larger SMD does not necessarily mean a biologically stronger intervention: a narrowly selected sample with low variability can yield a larger standardized value than a diverse routine-care population with the same raw benefit.

Rough labels need restraint#

Rules of thumb often label absolute values near 0.2 as small, 0.5 as medium, and 0.8 as large. These conventions came from broad behavioral-science contexts. They are not clinical thresholds.

A small average change can matter for a common, serious outcome or a low-cost intervention, while a large standardized change can be unimportant if the scale captures a transient surrogate or if harms outweigh benefit. Baseline risk, burden, duration, adverse effects, and alternatives remain part of the decision.

The confidence interval is essential. An SMD of 0.4 with an interval from 0.05 to 0.75 is compatible with effects ranging from slight to substantial under the model; a narrow interval around 0.1 can exclude a conventionally medium effect while still leaving you to judge what 0.1 means. Avoid describing the intervention as effective solely because an interval excludes zero. Magnitude, precision, bias, and clinical relevance should be reported separately.

Translate back to a familiar scale#

One approach multiplies the SMD by a representative standard deviation from a familiar instrument, and if SMD is 0.4 and a representative standard deviation is 10 points, the translated mean difference is 4 points.

The choice of standard deviation controls the translation. Using a small trial's value can produce a different answer from using a large observational cohort. Authors should justify the source and present sensitivity analyses when several values are plausible.

Another approach expresses the effect in units of a minimally important difference, if a credible threshold exists. This can show whether the average difference approaches a change considered meaningful. Individual-change thresholds and between-group differences are not always interchangeable, so the interpretation should be explicit.

The SMD can also be translated into an approximate probability that a randomly selected person from the intervention group has a better score than a randomly selected person from the comparison group. Such conversions assume distributional forms and can sound more intuitive to you than the data warrant. The safest presentation usually gives you all three: the SMD, a familiar-scale translation, and the assumptions behind that translation.

Change scores, final scores, and baseline adjustment#

Trials may report final values, changes from baseline, or regression-adjusted differences. Mean differences on a common scale can often combine final and change values under appropriate methods. SMDs are more complicated because their standard deviations differ.

The standard deviation of a change score depends on the correlation between baseline and follow-up. It can be smaller than the standard deviation of final scores. Mixing SMDs standardized by change variability with those standardized by final-score variability can produce inconsistent effect metrics.

Baseline-adjusted estimates from analysis of covariance are often efficient, but reconstructing an SMD requires the appropriate standard error or standardization approach; review authors should follow a prespecified method and consider sensitivity analyses rather than converting every reported statistic mechanically. Cluster-randomized and crossover trials also require correct unit-of-analysis handling. Treating observations as if they were all unrelated understates uncertainty, regardless of which effect-size formula is used.

Small studies and selective reporting#

The small-sample correction in Hedges' g addresses one mathematical bias. It does not address publication bias, selective outcome reporting, weak randomization, missing data, or flexible analysis.

Small studies often produce variable effect estimates. If only the most favorable are published, a meta-analysis can overstate benefit. Funnel plots and statistical tests have limitations, especially with few studies or heterogeneity. Protocols, registries, and searches for unpublished results remain important.

Choosing among multiple scales within a study after seeing results is another risk. A protocol should define the preferred instrument or a hierarchy. Otherwise, the selected SMD may reflect the most favorable measure rather than the planned outcome. Risk-of-bias assessment applies before pooling. A precise pooled SMD cannot cancel systematic bias shared across studies.

Heterogeneity is more than I-squared#

A random-effects meta-analysis estimates an average effect across a distribution of underlying effects. The pooled SMD may not describe any specific setting when effects vary widely.

I-squared summarizes the proportion of observed variation associated with between-study heterogeneity under assumptions, but it does not state how widely true effects vary in clinical units. Tau-squared is on the SMD scale. A prediction interval can show a range in which the effect of a future similar study might lie, although it is imprecise with few studies.

Potential sources include instrument type, follow-up, baseline severity, intervention intensity, comparator, risk of bias, and population variability. Subgroup analysis needs prespecification and enough studies. A difference in significance between subgroups is not evidence of a significant subgroup difference. When constructs or populations are too dissimilar, the correct response may be not to pool. The guide to deciding when not to meta-analyze explains that boundary.

SMD is not a response rate#

An SMD describes a difference between group means. It does not say that every participant improved by that amount. Two groups can have the same mean difference with very different individual response patterns.

Converting continuous scores into “responders” above a threshold may help you interpret them, but it discards information and depends on the cutoff. If a threshold is selected after results are known, bias increases.

Likewise, a mean near zero can hide benefit for some and harm for others, but subgroup claims require evidence. Distribution plots, variance, and prespecified treatment-effect analyses can add context. An SMD alone cannot identify who benefits.

A checklist for reading a forest plot#

Confirm that every scale measures the same construct and that score directions are aligned, and identify whether the effect is Cohen's d, Hedges' g, or another estimator and which standard deviation forms the denominator.

Read the confidence interval for each study and the pooled estimate. Check risk of bias, missing data, and whether the same time point is used. Examine heterogeneity and the prediction interval when available.

Look for a translation into a familiar scale or meaningful-change framework. Demand the source of the standard deviation used for back-translation. Consider harms and burden separately, because a symptom SMD does not encode them.

The forest-plot guide covers weights and confidence intervals, while the guide to clinical importance separates mathematical compatibility from decision relevance. The site's research overview connects both to broader evidence appraisal.

Keep the unitless number anchored#

The SMD is valuable when several sound instruments measure one construct. Its strength is comparability. Its weakness is abstraction.

Anchor the result to what was measured, when, in whom, with which variability, and against which comparator. Then show uncertainty and translate carefully. A unitless summary becomes useful to you only when the clinical units, the assumptions, and the consequences all stay visible.

References#

  1. Cochrane Handbook Chapter 6: choosing effect measures
  2. Cochrane Handbook Chapter 10: meta-analysis
  3. Cochrane Handbook Chapter 15: interpreting results
  4. Lakens: calculating and reporting effect sizes
  5. Small-sample properties of meta-analysis tests
  6. GRADE guidance on continuous outcomes

Questions and answers

What does a standardized mean difference of 0.5 mean?

It means the group means differ by half of the standard deviation used for standardization. It does not by itself say whether that difference is important to patients.

Are Cohen's d and Hedges' g the same?

They use closely related standardization, but Hedges' g applies a small-sample correction to reduce upward bias in the absolute effect estimate.

Can standardized mean differences combine any continuous outcomes?

No. The measures should assess the same underlying construct, use compatible populations and time points, and have directions aligned. Standardizing unrelated outcomes does not make them comparable.

Why can two studies have different standardized effects with the same raw mean difference?

The denominator is a standard deviation. A more variable study produces a smaller standardized effect for the same raw difference, even when the clinical difference is identical.

How should an SMD be made clinically understandable?

Present the confidence interval and certainty, translate it to a familiar scale when justified, compare it with a meaningful-change threshold, or show the probability of benefit without hiding assumptions.