Evidence explainer

Evidence and research methods

What a Funnel Plot Shows, and What It Cannot

A funnel plot compares study effect estimates with their precision. Asymmetry can flag small-study effects, but it cannot identify their cause on its own.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Read the axes before the shape
  2. A constructed pair shows what omission looks like
  3. Publication bias is only one mechanism
  4. Other causes of asymmetry
  5. When an Egger-type test helps
  6. Contour-enhanced plots add one clue
  7. Adjustment methods are sensitivity analyses
  8. A better missing-evidence investigation
  9. References

A funnel plot places each study's effect estimate on one axis and a measure of its precision or size on the other. Larger, more precise studies usually cluster near the pooled estimate. Smaller, less precise studies scatter more widely. When that scatter is roughly balanced, the points may resemble an inverted funnel. When one side looks sparse, the plot shows asymmetry.

The key interpretation is deliberately narrow. Funnel plot asymmetry is evidence of a pattern called small-study effects, not a diagnosis of publication bias. Missing results are one explanation. Differences in populations, methods, effect measures, and chance are others. The graph should trigger an investigation, not end one.

Read the axes before the shape#

The horizontal axis commonly shows an intervention effect, such as a log odds ratio, log risk ratio, or mean difference, and the vertical axis may show standard error, its inverse, or sample size. The chosen axes affect appearance, so the labels and the scale are part of what you are being shown.

If standard error is on the vertical axis, precise studies with smaller standard errors sit toward the top and imprecise studies sit lower. Dashed lines are often drawn around a center to show the region where about 95% of study estimates would be expected under a simplified sampling model, though those lines are not confidence limits for a pooled answer, and points outside them are not automatically invalid.

The center may be the fixed-effect estimate, random-effects estimate, or another reference. A random-effects synthesis assumes true effects can differ across studies, which complicates the simple funnel expectation. Always identify what the center represents before you call a plot balanced.

A constructed pair shows what omission looks like#

The left panel below contains all 14 fictional studies. The invented estimates widen as standard error rises. The right panel deliberately leaves out three imprecise positive estimates. The gap makes the selected set look asymmetric, but the construction already tells us why. A real plot will not tell you which points never appeared, or why.

Constructed paired funnel plotsPaired funnel plots show 14 fictional studies centered near zero. The complete panel includes imprecise estimates on both sides. The selected panel omits Iris at effect 0.20 and standard error 0.31, Juniper at 0.44 and 0.34, and Mesa at 0.64 and 0.41. Their deliberate absence creates lower-right asymmetry. Dashed guides show the constructed center plus or minus 1.96 standard errors; a comparable shape in real data would not establish its cause.Complete constructed set-0.900.9Constructed selected set-0.900.9Effect estimateEffect estimateStandard error (smaller is more precise)retainedintentionally omitted in right panel
Constructed example only, not meta-analysis results. The complete panel contains 14 invented studies; the selected panel deliberately omits three imprecise positive estimates to create a lower-right gap without implying that any real evidence is missing.
View the constructed data table
Constructed funnel-plot values
StudyEffectStandard errorSelected-panel status
Atlas0.010.06Retained
Birch-0.050.09Retained
Cedar0.070.1Retained
Delta-0.140.15Retained
Elm0.150.16Retained
Fjord-0.250.22Retained
Grove0.240.23Retained
Harbor-0.430.3Retained
Iris0.20.31Omitted in right panel
Juniper0.440.34Omitted in right panel
Keel-0.580.38Retained
Linden0.030.39Retained
Mesa0.640.41Omitted in right panel
North-0.680.44Retained

The paired construction also shows why the phrase “missing corner” needs care. The complete panel is not perfectly symmetric. Sampling variability does not arrange points neatly, even when every result is present. The right panel's gap is easier to see because the omitted values were chosen for teaching. Real evidence is messier.

Publication bias is only one mechanism#

Publication bias occurs when whether a study becomes available is associated with its results. Statistically detectable or favorable findings may be submitted, accepted, or published faster than null or unfavorable findings; selective outcome reporting creates a related problem within a published study when only some measured outcomes or analyses are available.

Either process can distort a meta-analysis if the available results differ systematically from those that are missing, and a funnel plot may show the resulting small-study pattern because small studies are more vulnerable to remaining unpublished when their findings appear unremarkable. It cannot tell you whether the decision occurred with investigators, sponsors, journals, or elsewhere, and it cannot recover the unobserved data.

Registration and regulatory records can provide stronger evidence. Comparing a protocol or registry entry with the report can reveal a prespecified outcome that disappeared or changed; an inception cohort of all registered studies can identify completed work with no results. Those designs investigate known records rather than inferring absence from geometry.

Other causes of asymmetry#

True heterogeneity is a central alternative. Smaller studies may enroll participants at higher baseline risk, use more intensive delivery, occur earlier in a technology's development, or take place in specialist settings, and if the intervention genuinely works differently there, effect size will be associated with study size even when reporting is complete.

Bias within studies can create a similar pattern. Small trials may have weaker allocation methods, less complete follow-up, more flexible analyses, or noisier outcome measurement. The resulting estimates can be exaggerated without any study being withheld. Risk-of-bias assessment and design-stratified analyses are needed to explore that possibility.

Some effect measures have a built-in mathematical relation with their standard errors. Odds ratios and standardized mean differences can produce apparent asymmetry under certain conditions. Poorly chosen axes, zero-event corrections, and a bounded outcome scale can add artifacts. The appropriate graph and test depend on the effect measure.

Chance remains an explanation. A handful of points can look strikingly lopsided even when generated by a balanced process. Conversely, a symmetric plot can occur despite missing evidence. The absence of visible asymmetry is not proof of complete reporting.

When an Egger-type test helps#

Regression tests ask whether effect estimates are associated with their standard errors more strongly than expected by sampling variation; the widely known Egger test was introduced in 1997, but it is not a universal add-on for every meta-analysis.

Current Cochrane guidance uses at least 10 studies as a practical minimum for considering asymmetry tests, while noting that even more may be needed, and you also need a useful range of standard errors across those studies. Ten nearly identical large trials provide little information about small-study patterns. Heterogeneity and the chosen effect measure can invalidate a simple test.

A nonsignificant result does not clear a review of missing-evidence concerns because power can be low. A significant result does not prove publication bias because the test detects association, not cause; report the test used, why it fits the effect measure, and how its result aligns with visual, clinical, and documentary evidence.

Contour-enhanced plots add one clue#

A contour-enhanced funnel plot shades regions by statistical significance. If the apparently missing area corresponds mainly to nonsignificant, unfavorable results, selective reporting becomes more plausible. If the gap lies in a region of highly significant results, another mechanism may fit better.

This is still circumstantial. Statistical significance depends on thresholds and standard errors, and reporting decisions can be influenced by direction, novelty, subgroup findings, or sponsor priorities rather than one P value. Contours refine a question; they do not certify an answer.

Adjustment methods are sensitivity analyses#

Methods such as trim-and-fill, selection models, and regression-based adjustments estimate how a pooled result might change under assumptions about missingness. None observes the missing studies. Each embeds a model about why results are absent and how their effects would be distributed.

Present adjusted estimates as sensitivity analyses with their assumptions. If conclusions change across plausible models, that instability is the result, and if they remain similar, the finding may be less sensitive to those particular assumptions, but that does not prove the evidence set is complete.

Restricting a synthesis to larger or lower-risk studies can also be informative when those studies better match routine use. It answers a different question and should be prespecified or clearly labeled. Do not try several restrictions and report only the one that preserves your preferred conclusion.

A better missing-evidence investigation#

Begin before drawing the funnel. Check whether the search covered trial registries, regulatory documents, dissertations, preprints, conference records, and non-English sources where relevant. Compare reports with protocols and analysis plans. Contact investigators for missing outcomes when feasible. Record reasons for unavailable data.

Then examine whether a funnel plot is suitable for the effect measure and number of studies. Interpret any asymmetry you find alongside population differences, intervention versions, risk of bias, and study size. Use a formal test only when its assumptions fit. Add sensitivity analyses that represent plausible missingness mechanisms.

ROB-ME provides a structured way to bring these strands together for a specific meta-analysis result. The final judgment can be low risk, some concerns, or high risk of bias due to missing evidence; that conclusion should explain the evidence, not simply repeat that a plot looked symmetric or asymmetric.

The guides to reading a forest plot and spotting outcome switching cover two adjacent checks. The site's research overview places them in the broader evidence-appraisal program.

References#

  1. Cochrane Handbook Chapter 13 on missing evidence
  2. Recommendations for examining funnel plot asymmetry
  3. Original graphical and regression test paper
  4. ROB-ME tool for bias due to missing evidence
  5. Updated analysis of selective antidepressant-trial publication

Questions and answers

Does a symmetric funnel plot prove there is no publication bias?

No. The plot may have low power, missing studies may not create asymmetry, or different mechanisms may cancel one another. Registration and protocol comparisons remain important.

How many studies are needed for a funnel plot?

There is no magic number. Cochrane commonly uses at least 10 studies before applying asymmetry tests, and some settings require more. Similar study sizes or substantial heterogeneity can make the method uninformative even above that count.

Is the Egger test a test for publication bias?

It is better described as a test for a relation between effect size and precision. Publication bias is one possible cause of that relation, not the only cause.

Can trim-and-fill recover the missing studies?

No. It imputes hypothetical values under a model. Its estimate is useful for sensitivity analysis, not as a reconstructed historical record.

Should funnel plots be used for diagnostic accuracy reviews?

Simple intervention-effect funnel methods may not fit diagnostic accuracy data, where sensitivity and specificity are paired and thresholds vary. Use methods designed for the data and follow current diagnostic-review guidance.