Evidence explainer

Evidence and research methods

A Clinically Important Difference Is Not Just a Significant Difference

A minimal clinically important difference estimates the smallest change patients consider meaningful. It helps you read a scale, but it is context dependent and is often misapplied to group averages.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Why scales need an interpretive bridge
  3. Statistical significance answers a different question
  4. How an important difference is estimated
  5. Within-person change and between-group difference are not interchangeable
  6. There is rarely one permanent value
  7. Common interpretation errors
  8. A reader's checklist

A statistically detectable change is not necessarily a change that patients notice or value, and the minimal clinically important difference, often abbreviated MCID or MID, is an attempt to connect movement on a measurement scale with lived importance.

Key points#

Why scales need an interpretive bridge#

Symptoms, function, and quality of life are often measured with questionnaires. A trial may report that one group improved 4.2 points more than another on a 100-point scale. The p value may be small, but nothing in it tells you whether 4.2 points is an improvement anyone would notice.

An MID provides an interpretive bridge. In the foundational formulation, it is the smallest difference in a score that patients perceive as important and that could justify considering a change in management, in the absence of excessive burden. Later work has refined both the terminology and the methods. “Minimal” does not mean trivial, universal, or precisely known. It means the lower boundary of change considered important for a specified measure and context.

Statistical significance answers a different question#

A p value is influenced by effect size, variability, sample size, and the statistical model. With a very large sample, a tiny mean difference can be statistically significant. With a small sample, an important difference can miss the conventional threshold because the estimate is imprecise.

The uncertainty interval is therefore essential. Suppose a trial estimates a 5-point benefit and a credible MID is 8 points. If the interval ranges from 1 to 9, the average effect is below the benchmark but a clinically important average benefit remains possible. If the interval ranges from 7 to 13, the estimate is compatible mainly with effects near or above the benchmark. A p value alone cannot distinguish these situations.

Clinical importance also includes outcomes beyond the target scale: an improvement may be worthwhile when burden and harms are small, yet unattractive when the intervention is difficult or carries serious risk. The MID informs that judgment; it does not complete it.

How an important difference is estimated#

Anchor-based methods#

An anchor is an external measure intended to give the score change meaning. A common anchor asks participants whether they feel much better, a little better, unchanged, or worse. Investigators then estimate the average score change among people who report a small but important improvement.

Other anchors include a global rating by a patient or clinician, a functional milestone, or another validated measure. Patient-reported anchors are particularly relevant when the outcome is a symptom or quality-of-life construct.

The anchor must be understandable, sufficiently related to the target score, and able to distinguish change. A retrospective global rating may be influenced by current status and imperfect recall. A clinician rating may not match the patient's priorities. If the anchor is weak, the resulting threshold inherits that weakness.

Distribution-based methods#

Distribution-based approaches express change relative to variation or measurement error. Examples include a fraction of a standard deviation, the standard error of measurement, or a reliable change index.

These methods help answer whether a change is large relative to noise in the instrument. They do not by themselves show that patients consider the change important; they are useful as supporting information, especially when checking whether an anchor-based estimate exceeds measurement error, but they should not replace a meaningful anchor.

Combining evidence#

A credible estimate often draws on several anchors and statistical checks rather than one calculation. Reviews may identify a range of plausible MIDs and judge credibility using criteria such as the anchor's relevance, correlation with the score, precision, and similarity between the MID study and the target trial population.

Within-person change and between-group difference are not interchangeable#

Many MIDs are derived by observing how much an individual's score changed when that person reported feeling a little better, while a randomized trial commonly reports a difference between the average changes of two groups. Those are different quantities.

A group mean can be smaller than a within-person MID while still representing a meaningful shift in the distribution. For example, if the treatment causes an additional subset of participants to improve substantially, averaging responders and nonresponders may yield a mean difference below the individual threshold. Conversely, a mean at the threshold does not mean every participant improved by that amount.

This does not make group means useless. It means you should avoid the rigid rule that a mean difference one fraction below the MID is unimportant while one fraction above is important; the estimate and its interval should be interpreted alongside responder proportions and the full outcome distribution when available.

Responder analyses have their own tradeoffs. Defining response as improvement of at least one threshold is intuitive, but dichotomizing a continuous score discards information. Results can change with the chosen cut point, and participants close to either side of it may be clinically similar. Presenting continuous and responder results together is more informative than either alone.

There is rarely one permanent value#

An MID can vary because the meaning of change varies. Relevant features include:

People with severe symptoms may value a different absolute change from those with mild symptoms. The threshold for recognizing deterioration may differ from the threshold for improvement. Translation can change how items and response choices are understood. Using a value borrowed from a distant population may create false precision, so a paper should say why its chosen MID applies to the instrument, population, direction, and time horizon under study.

Common interpretation errors#

One error is treating the MID as a biological constant. Another is choosing from several published values after seeing the trial result. Selecting the smallest available threshold can make an effect appear important; selecting a larger one can make the same effect appear inadequate.

A third error is using a distribution-based number and calling it patient important without a suitable anchor. A fourth is comparing only the point estimate with the threshold and ignoring the uncertainty interval. A fifth is declaring equivalence because a nonsignificant result did not cross the MID. Demonstrating that effects larger than a threshold are unlikely requires a sufficiently precise interval and an appropriate design, not a null p value.

The MID also should not be confused with a minimal detectable change. Minimal detectable change concerns whether movement exceeds expected measurement error. A change can be reliably measured yet too small to matter, or meaningful but difficult to distinguish from noise with an imprecise instrument.

A reader's checklist#

When an article invokes an MCID or MID, ask yourself:

  1. Is the threshold tied to the exact instrument and scoring version?
  2. Was it derived in a similar population and setting?
  3. Does it concern improvement, deterioration, or both?
  4. Was the method anchor based, distribution based, or combined?
  5. Was the anchor relevant, understandable, and correlated with score change?
  6. Is the value an individual-change threshold or intended for group comparisons?
  7. Was the threshold selected before trial results were known?
  8. Are alternative credible values shown in sensitivity analyses?
  9. Does the uncertainty interval include effects both below and above the threshold?
  10. Are responder proportions, harms, burden, and other outcomes also reported?

A clinically important difference adds patient meaning to measurement. It works best as a transparent range supported by credible anchors, not as a decorative line that turns a complex result into a binary verdict.

Sources and further reading

  1. Controlled Clinical Trials, foundational definition of a clinically important difference
  2. JAMA, defining changes that matter to patients
  3. The Spine Journal, review of concepts and estimation methods
  4. BMJ, credibility assessment for minimal important differences
  5. BMJ, selecting credible minimal important differences for trial interpretation

Questions and answers

Are MCID and MID the same?

They are often used interchangeably. Some methodologists prefer “minimal important difference” because “clinical” may be too narrow for social, functional, or quality-of-life outcomes.

Can an effect be statistically significant but smaller than the MID?

Yes. A large sample can estimate a small average difference precisely. Whether that difference matters depends on the threshold's credibility, uncertainty, response distribution, burdens, and other outcomes.

Should every patient improve by at least the MID?

No. An individual threshold and a group average describe different things. Group results should be supplemented with distributions or responder proportions when those analyses are well planned.