Imprecision is not a synonym for "statistically nonsignificant." In GRADE, the question is whether random uncertainty leaves more than one materially different interpretation of the effect. A confidence interval that supports important benefit at one end and no worthwhile benefit or harm at the other cannot guide a stable conclusion, even if its pooled estimate looks favorable.
What is being rated#
GRADE assigns certainty to a specific outcome and effect estimate within a defined question. Randomized evidence generally begins at high certainty, while nonrandomized evidence generally begins lower, then domains such as risk of bias, inconsistency, indirectness, imprecision, and publication bias can change the rating.
Imprecision concerns random error. If another plausible sample from the same underlying process could lead to a different conclusion, your confidence in the estimate falls. This is separate from bias, which concerns systematic error, and inconsistency, which concerns unexplained differences across studies.
The rating target must be clear. A systematic review may ask whether an intervention has any nontrivial effect. A guideline panel may ask whether benefits outweigh harms enough to recommend one option. The same confidence interval can receive different judgments because the relevant thresholds differ.
Start with absolute effects#
Relative measures can look stable while the number of people affected varies greatly with baseline risk. A risk ratio of 0.80 could mean 20 fewer events per 1,000 when baseline risk is 100 per 1,000, or only 2 fewer when baseline risk is 10 per 1,000.
GRADE guidance therefore emphasizes confidence intervals around absolute effects when possible. The panel or review defines thresholds that separate conclusions such as:
- important benefit;
- trivial or no important effect;
- important harm.
If the entire interval lies within one region, imprecision may not lower certainty. If it crosses from important benefit into trivial effect, the data leave two conclusions open. If it spans important benefit, trivial effect, and important harm, uncertainty is more serious. Whichever way it falls, the thresholds should have been declared, justified, and ideally chosen without tailoring them to the observed result, because a hidden threshold makes the judgment impossible for you to audit.
Minimally and fully contextualized approaches#
GRADE distinguishes levels of context.
A minimally contextualized approach may compare effects with a threshold separating trivial from nontrivial benefit or harm. It is useful when a review rates whether an intervention produces an important effect without making a full recommendation.
A partially or fully contextualized approach incorporates the decision: baseline risks, benefits, harms, burdens, and the values placed on outcomes. A guideline panel may care whether the net effect crosses a threshold at which it would choose treatment A over treatment B.
This is not inconsistency in the method. Certainty is confidence that the true effect lies on the correct side of the threshold relevant to the question. Different questions can require different thresholds.
How many levels should certainty fall?#
GRADE commonly describes imprecision as not serious, serious, or very serious, corresponding to no rating down, one level, or two levels. Updated guidance recognizes circumstances where exceptionally broad uncertainty can justify three levels.
The number of thresholds crossed is informative. An interval that narrowly crosses a trivial-effect boundary may lead to one-level reduction. An interval spanning important benefit and important harm leaves sharply different decisions plausible and may warrant a larger reduction. The judgment also considers how far the interval extends into each region, not merely whether it touches a line. Panels should explain that reasoning in plain terms, so that you can see which conclusions were still open: "The interval includes a clinically important reduction and no important effect," or "The interval includes benefit, no effect, and harm." Either sentence tells you more than an unexplained footnote saying "downgraded for imprecision."
Optimal information size#
The optimal information size, OIS, adapts conventional trial sample-size reasoning to a body of evidence. It asks how many participants or events a single adequately powered trial would require to detect the target effect at chosen error rates and baseline risk.
If a meta-analysis falls far below that amount, a seemingly reassuring result may be unstable. Sparse data are especially concerning when a few events moving between groups would materially change the estimate.
Current GRADE guidance does not treat OIS as a mechanical hurdle in every case. When a confidence interval already crosses the relevant threshold, the interval supplies the primary reason for rating down. When the interval does not cross a threshold but the relative effect is large or the evidence is sparse, OIS can reveal that apparent precision rests on too little information. The calculation itself depends on assumptions about target effect, baseline risk, statistical power, and type I error, and changing any of those inputs changes OIS, so a report should state every one of those assumptions rather than present the number as universal.
Event counts can matter more than participant totals#
Ten thousand participants sound reassuring, but if only 12 outcomes occur, the effect estimate may remain fragile. For binary outcomes, information comes heavily from the number and distribution of events. Clustered designs, unequal follow-up, and losses can reduce effective information further.
Conversely, a smaller study of a common outcome may produce a narrower and more decision-relevant interval. Raw sample size should never replace inspection of event counts and uncertainty.
Rules of thumb can support consistency, but they are not a substitute for the actual decision threshold. The CDC ACIP handbook discusses OIS and alternative information-size approaches while retaining judgment about clinical thresholds and event rates.
Why significance can mislead#
A confidence interval may exclude the null while still cross a clinical-importance threshold. Suppose a treatment reduces pain by an estimated 1.2 points on a 10-point scale, with an interval from 0.1 to 2.3. The result is statistically different from zero. If patients regard a 1-point change as the minimum important difference, however, the interval includes both trivial and important benefit. GRADE can rate down for imprecision.
The reverse can also occur. An interval may include the null yet remain entirely within a range of effects that would not change a decision. If a panel's concern is ruling out a large harm and the entire interval does so, the result may be sufficiently precise for that purpose even without conventional significance. P-values answer how incompatible data are with a chosen null under a model. They do not identify which clinical decisions remain plausible.
Baseline risk and multiple outcomes#
Absolute effects depend on baseline risk, which may vary across populations. A panel may need to assess imprecision at several plausible baseline risks. If the decision changes across those risks, certainty or applicability may differ by group.
Each critical outcome also receives its own certainty rating. Mortality may be precise while a rare serious adverse event remains very imprecise. A recommendation should not collapse that difference into one unexplained study-quality label. When benefit and harm rest on different evidence bases, a fully contextualized decision may be uncertain even if the primary efficacy outcome is precise, and the evidence-to-decision process has to keep those uncertainties apart rather than average them.
A reader's audit#
- What exact outcome and effect are being rated?
- Are relative results translated into absolute effects using a defensible baseline risk?
- Which thresholds separate important, trivial, and harmful effects?
- How many decision regions does the confidence interval cross?
- Are participant and event counts adequate for the assumed target effect?
- If OIS was calculated, are its assumptions shown?
- Is the degree of rating down explained rather than merely labeled?
- Would another reasonable threshold or baseline risk change the judgment?
Work through those questions and imprecision looks contextual rather than arbitrary. The interval is observed; the thresholds are declared; the rating follows from how the two interact.
Sources and further reading
- Zeng and colleagues, GRADE Guidance 34 on minimally contextualized imprecision, Journal of Clinical Epidemiology (2022)
- Zeng and colleagues, GRADE Guidance 35 on contextualized imprecision and decisions, Journal of Clinical Epidemiology (2022)
- U.S. CDC Advisory Committee on Immunization Practices, GRADE handbook chapter on certainty
- Core GRADE 2, choosing the certainty target and assessing imprecision, BMJ
Questions and answers
Can a statistically significant result be imprecise?
Yes. Its interval may exclude no effect but still include both trivial and clinically important effects, or it may rest on sparse events below the relevant information size.
Is OIS just the total number of people in all studies?
No. It is a target derived from assumptions about the outcome, baseline risk, effect size, power, and error rate. Event counts and effective sample size matter.
Why can two panels rate the same interval differently?
They may be answering different decisions or using different justified thresholds and baseline risks. Transparent panels state those choices so readers can assess them.