A funnel plot is often treated as a visual detector for missing negative studies. That shorthand is too confident. The plot compares study effect estimates with a measure of their size or precision. If smaller, less precise studies show systematically different effects from larger studies, the plot may be asymmetric; selective publication is one possible cause, but clinical differences, methods, chance, and effect-measure artifacts can produce the same picture.
Egger's regression and trim-and-fill answer different questions about that asymmetry. Egger's test assesses a statistical association between standardized effect and precision. Trim-and-fill imputes hypothetical studies to create a more symmetric plot and recalculates a pooled estimate; neither observes the studies that were never found, determines why results differ by study size, or guarantees a corrected treatment effect.
What a funnel plot represents#
Each point is a study. The horizontal position is usually its effect estimate. The vertical position is a measure such as standard error, inverse standard error, or sample size. Larger, more precise studies tend to cluster near the top. Smaller studies scatter more widely near the bottom because their estimates have more sampling variation.
If studies are estimating one common underlying effect, standard errors are appropriate, and there are no size-related biases, the scatter may resemble an inverted funnel, though symmetry is an expectation over repeated samples, not a geometric rule that every real plot must satisfy. A plot with six or eight points can look lopsided by chance. A plot with many points may be asymmetric because the studies truly address different clinical questions.
Contour-enhanced funnel plots add regions of statistical significance. If apparently missing studies fall mainly in nonsignificant regions, selective publication becomes more plausible. If gaps occur in highly significant regions, another mechanism may fit better. This remains diagnostic reasoning, not proof.
The outcome scale matters. For binary outcomes, some effect measures are mathematically related to their standard errors, creating asymmetry even without selection. Tests designed for odds ratios, risk ratios, or risk differences are not interchangeable. For continuous outcomes, standardized mean differences can also be coupled with their precision. Choose a plot and test that match your effect measure, and follow current methodological guidance.
Why smaller studies may show larger effects#
Publication and dissemination bias are familiar explanations. A small study with an exciting result may be submitted, accepted, and cited, while a small null study remains unpublished, and outcomes within a published study can also be selected after investigators see the results. Time-lag bias may make positive results appear earlier. Language and database restrictions can make some reports easier to find.
Several non-selection explanations remain:
- Smaller trials may enroll people with more severe disease, higher baseline risk, or a setting where an intervention works better.
- They may deliver a more intensive intervention or use more experienced centers than large pragmatic trials.
- Shorter follow-up can capture an early benefit while missing later loss of effect or harms.
- Weak allocation concealment, lack of blinded outcome assessment, attrition, and flexible analysis can inflate effects, and these problems may be more common in small studies.
- Dose, comparator, diagnostic threshold, or outcome definition can be related to study size.
- A few unusual results can create asymmetry by chance.
These mechanisms can coexist. Calling the pattern “publication bias” too early can prevent the more useful question: what feature changes with study size, and does it also change the true effect or the amount of bias?
What Egger's regression tests#
The original Egger approach regresses each study's standardized effect estimate on its precision. In a common formulation, the effect divided by its standard error is the outcome, inverse standard error is the predictor, and studies are weighted. Under the model, an intercept that differs from zero indicates asymmetry.
The resulting P value tests a statistical null within that model. It does not estimate the number of missing studies or identify the cause. A nonsignificant result can reflect low power rather than symmetry. A significant result can reflect genuine effect modification or an unsuitable test rather than selective publication.
Cochrane guidance commonly advises that asymmetry tests usually should not be used with fewer than ten studies because power is low. Ten is not a magic threshold. With ten heterogeneous or similarly sized studies, inference can remain weak. When there are many small studies and only one large study, the apparent pattern may depend heavily on that single point.
Prespecify the test, preferably in your review protocol. Trying several and reporting the one that crosses 0.05 creates another selection problem. Interpretation should include the intercept or regression estimate, confidence interval, number of studies, range of standard errors, heterogeneity, and visual plot, not the P value alone.
Why heterogeneity complicates the regression#
Suppose smaller studies took place in specialist clinics with intensive support, while larger studies tested the same named intervention in general practice. If support modifies effectiveness, effect size will be related to study precision even if every result is available. Egger's test detects the relation but cannot label it as bias.
Subgroup analysis or meta-regression can examine whether setting, dose, baseline risk, follow-up, or risk-of-bias features explain the gradient, but those analyses are observational across studies and can suffer ecological bias, multiple testing, and low power. Meta-regression and ecological bias explains those limits.
Statistical heterogeneity alone does not make a funnel plot useless, but it changes the question. Prediction intervals and careful study-level comparison may be more informative than forcing one symmetric funnel. See prediction intervals in meta-analysis for how between-study variation affects expected results in a new setting.
What trim-and-fill does#
Trim-and-fill is an algorithm built around funnel symmetry. In simplified terms, it identifies the more extreme small studies contributing to asymmetry, temporarily trims them to estimate a revised center, then fills the plot with imputed mirror-image studies. The pooled effect is recalculated using the observed and imputed points.
The method is attractive because it provides a visible plot and an adjusted number. The imputed points are not recovered experiments. Their outcomes, populations, methods, and even existence are unknown. They are mathematical placeholders generated by the selected estimator and assumptions.
Different estimators can infer different numbers of missing studies. Results can change with fixed-effect versus random-effects models, outcome scale, outliers, heterogeneity, and which side of the funnel is assumed to be missing. If true effects vary with setting or baseline risk, forcing symmetry around one center may create a misleading adjustment. In some situations trim-and-fill performs poorly or adjusts in the wrong direction.
For these reasons, an adjusted estimate should be presented as one sensitivity scenario. It can show how much the pooled estimate would move under a particular symmetry-based model. It should not replace the primary synthesis or be described as the effect “after correcting publication bias.”
A better missing-evidence assessment#
The first protection is review design. Search multiple bibliographic databases, trial registries, regulatory material, conference records, references, and relevant gray literature. Avoid unnecessary language restrictions. Record which investigators and sponsors were contacted and what was obtained. Compare protocols, registrations, statistical plans, and publications for prespecified outcomes and time points.
Next ask whether missingness could depend on the result. A registered completed study without results, a published article that omits a prespecified outcome, and an unavailable subgroup may each create different risks. Direction matters: would likely missing results favor the intervention, the comparator, or no clear direction?
ROB-ME provides a structured approach to risk of bias due to missing evidence in meta-analysis, and it separates known missing results from the possibility of unknown missing studies and asks reviewers to judge how the selection process could affect the synthesis. The tool does not turn incomplete evidence into complete evidence. It makes assumptions visible and links them to a risk judgment.
Sensitivity analyses can then test plausible departures. Selection models directly model the probability that a result is observed, but depend on strong assumptions and often need many studies. Worst-case or bounded scenarios may be more transparent. Restricting to prospectively registered or lower-risk studies can be informative if prespecified, though it may reduce precision and alter the clinical mix.
How to report these analyses#
A clear report distinguishes four layers, and your reader should be able to see all four:
- Observation: describe the plot, including number of studies, effect measure, axis, and any isolated points.
- Test: state the chosen asymmetry method, why it fits the outcome, estimate, confidence interval, and limitations.
- Explanation: compare selection with clinical, methodological, and chance explanations.
- Consequence: show sensitivity analyses and say whether the direction, magnitude, certainty, or decision changes.
Avoid “no publication bias was found” after a low-powered test. A safer statement, and the one you should write, is that the test did not detect asymmetry, followed by the study count and limitations. After a significant test, say that small-study effects were detected and explain why their cause remains uncertain.
The certainty of evidence can be downgraded for suspected publication bias when the whole evidence pattern supports that judgment. A funnel plot is one input. A small set of sponsored trials with unavailable outcomes may justify concern even when no funnel analysis is possible. Conversely, an asymmetric plot with a strong, prespecified clinical explanation need not be labeled undisclosed suppression without supporting evidence.
Worked interpretation without false correction#
Imagine 14 trials with a favorable pooled effect. The six smallest studies show the largest benefits, Egger's test is significant, and trim-and-fill moves the pooled estimate toward no effect. Inspection then shows that small trials used intensive supervised treatment for high-risk participants, while large trials used low-intensity delivery in a broad population. Several small studies also have unclear allocation concealment, and two registered studies have no posted results.
The proper conclusion is not that trim-and-fill revealed the true effect. There are at least three plausible contributors: effect modification by intervention intensity and baseline risk, methodological inflation, and selective nonreporting. Present stratified or meta-regression results cautiously, assess missing records through ROB-ME, run explicit sensitivity analyses, and lower certainty if your decision remains vulnerable.
That explanation is longer than a funnel-plot label, but it is more useful. It tells you what is observed, what is inferred, and which unanswered facts could change the conclusion.
References#
- Cochrane Handbook chapter 13: bias due to missing results
- Egger and colleagues, BMJ 1997
- Sterne and colleagues, BMJ 2011
- Duval and Tweedie trim-and-fill method
- ROB-ME tool, BMJ 2023
- Cochrane Handbook chapter 10: meta-analysis
Questions and answers
Does a symmetric funnel plot prove that there is no publication bias?
No. Selection can affect studies of all sizes or operate in a way that preserves apparent symmetry. With few studies, the plot may have little ability to reveal a pattern. Searches, registrations, and outcome-level comparisons remain necessary.
Is ten studies enough to run Egger's test?
Ten is a common rough minimum, not a guarantee. Power and validity also depend on the range of study sizes, heterogeneity, outcome scale, outliers, and the selected test. Interpretation should remain cautious near that threshold.
What does a significant Egger test mean?
It indicates statistical evidence of the modeled association between effect and precision, often called small-study effects. It does not determine whether the cause is publication bias, methodological weakness, clinical heterogeneity, an effect-measure artifact, or chance.
Should the trim-and-fill estimate replace the main meta-analysis?
No. The imputed studies are hypothetical and the adjustment depends on assumptions. Report it as a sensitivity analysis alongside the primary model, alternative explanations, and other missing-evidence assessments.
What is more informative than either method alone?
Prospective registrations, protocol-publication comparisons, broad searches, outcome availability, study-level risk of bias, subject-matter explanations for size-related differences, and structured ROB-ME judgments provide a stronger combined assessment.