Standard pairwise meta-analysis combines studies of the same two options. Network meta-analysis extends the idea to a connected set of three or more interventions. If trials compare A with C and B with C, the network can estimate A versus B indirectly through C. When head-to-head A-versus-B trials also exist, direct and indirect information can be combined. The extra reach depends on an extra assumption: transitivity.
How an indirect comparison works#
Imagine randomized trials comparing drug A with placebo C and separate trials comparing drug B with placebo C. If the log risk ratio for A versus C is -0.4 and for B versus C is -0.1, an indirect estimate for A versus B is their difference, -0.3, with uncertainty from both sets of trials.
That subtraction is valid only if C links comparable clinical questions, and if A trials enrolled severe disease after previous treatment while B trials enrolled mild untreated disease, baseline severity may modify relative response. The difference you compute could mix treatment effect with trial-population differences.
A network model generalizes this logic across many nodes, studies, and multi-arm trials while preserving randomization within each trial. It can improve precision and provide a coherent set of relative estimates. It cannot create randomization between people in different trials.
Define the nodes before admiring the network#
Each node represents an intervention category. Combining doses, routes, treatment durations, combinations, or versions into one node assumes they share the same relative effect for the question. Splitting every variant creates a sparse network; lumping distinct interventions creates a misleadingly connected one.
The protocol should justify node definitions clinically and methodologically before results are known. Placebo, usual care, waiting list, and active control may not be equivalent. “Usual care” can change by country and era. Sham procedures may produce different expectations and co-interventions than no treatment. Outcome nodes require comparable definitions and time points too; a network that treats 6-week remission, 12-week response, and final-visit symptom change as one endpoint may gain power by losing meaning.
Transitivity is the core causal assumption#
Transitivity asks whether any participant in the network could, in principle, have been randomized to any of the interventions being compared, subject to the review's eligibility criteria. It also requires important effect modifiers to be similarly distributed across comparisons.
Potential effect modifiers include baseline severity or risk, age, prior treatment, disease subtype, dose, follow-up, outcome definition, co-interventions, setting, calendar time, and risk of bias. The list depends on mechanism and clinical context; it cannot be generated solely by software.
Review authors should prespecify plausible modifiers, tabulate them by direct comparison, and explain important imbalances. Similar overall means do not guarantee similar distributions, and unreported variables cannot be declared balanced. Transitivity is not directly testable as one null hypothesis. It is a clinical and methodological judgment, supported or challenged by the structure and the data.
Heterogeneity and incoherence are different#
Heterogeneity is variation among studies that estimate the same direct comparison, such as A versus C, and incoherence is disagreement between different paths estimating the same contrast, such as direct A-versus-B trials disagreeing with the indirect estimate through C.
Local methods examine a particular loop or split direct from indirect evidence for one comparison. Global methods test the network more broadly. These diagnostics can identify tension but have limited power in sparse networks, and a nonsignificant test does not prove coherence, just as a nonsignificant heterogeneity test does not prove identical effects.
Statistical coherence can also coexist with violated transitivity if biases line up in similar directions. Conversely, a chance discrepancy can appear even in a clinically well-planned network. Investigators should examine effect modifiers, data errors, node definitions, outliers, and risk of bias rather than treating one p-value as a verdict.
Read the network diagram as a data map#
Nodes usually show interventions and edges show direct comparisons. Their size or thickness may represent participant or study count, but conventions vary. The diagram should set out disconnected components, dominant comparators, sparse links, and interventions supported by only one small trial.
A star-shaped network centered on placebo may have many estimates but little direct evidence between active options. An old comparator can act as the bridge between different treatment eras. A closed loop permits direct-indirect comparison; a tree without loops may rely entirely on the transitivity assumption without an empirical incoherence check. Multi-arm trials require models that account for correlated comparisons. Counting each pair as independent double-counts the shared group and understates uncertainty.
The model does not erase ordinary meta-analysis problems#
Randomization protects comparisons within a trial, not across missing studies or selectively reported outcomes. Allocation problems, deviations from assigned treatment, missing data, outcome measurement, and selective reporting still affect every direct estimate. Bias can propagate through the network and influence many indirect comparisons.
Between-study heterogeneity also remains. Fixed-effect and random-effects network models answer different questions and make different assumptions, and a common heterogeneity variance across all comparisons may stabilize estimation but can be implausible when outcomes or interventions differ. Sparse evidence makes comparison-specific heterogeneity difficult to estimate.
Small-study effects and publication bias can distort the network. Comparison-adjusted funnel plots may help only under assumptions about which treatments are “newer” and the comparability of contrasts. Searching trial registries, regulatory sources, and unpublished results remains essential.
Rankings are seductive and often fragile#
Network meta-analysis can estimate the probability each intervention occupies each rank. SUCRA, the surface under the cumulative ranking curve, summarizes the rank distribution. A value near 100% means an intervention tends to occupy higher ranks under the model; it does not mean a 100% response rate, a 100% probability of clinically important benefit, or high certainty.
Ranks discard magnitude. If five drugs differ by trivial amounts, the model must still place one first and one last, and a poorly studied intervention with a favorable but imprecise estimate can have a substantial chance of being best and also a substantial chance of being worst. Rankings may shift with a different outcome, time point, node definition, heterogeneity model, or inclusion decision.
Present pairwise effect estimates, absolute effects, confidence or credible intervals, and certainty before ranks. Benefits and harms need separate outcomes. The “best” symptom rank may come with the worst discontinuation or serious-adverse-event profile.
Direct, indirect, and network estimates should be visible#
For the comparison you actually care about, ask how much of the information is direct and how much arrives through other nodes; a network estimate that looks precise may be driven by an indirect path through a comparator used in a dissimilar era.
Node-splitting displays direct and indirect estimates separately when both exist. Agreement supports, but does not prove, the assumptions. Disagreement should not be solved automatically by choosing whichever estimate is more favorable. Investigators may need to revisit eligibility, modifiers, extraction, node definitions, or present evidence separately. Prediction intervals can show the range of underlying effects across future comparable settings when heterogeneity is estimable. League tables should distinguish effect direction and avoid duplicating reciprocal estimates in a confusing way.
Confidence is comparison-specific#
CINeMA evaluates confidence across within-study bias, reporting bias, indirectness, imprecision, heterogeneity, and incoherence. GRADE-based approaches similarly begin with the evidence supporting each contrast. A network does not receive one global badge of “high quality.”
One comparison may have multiple large head-to-head trials and high confidence. Another may rest on a long indirect path with serious imprecision. A treatment ranking that combines the two can hide that difference from you. Imprecision should be judged against a clinically meaningful decision threshold, not merely whether the interval crosses no effect. Heterogeneity matters most when it crosses boundaries that would change action.
A practical appraisal sequence#
First, verify a registered protocol and a systematic search that includes all relevant interventions and outcomes. Second, inspect how nodes were defined and whether the network is connected for defensible reasons. Third, map direct evidence, study dates, sample sizes, and important effect modifiers by comparison.
Fourth, appraise risk of bias and reporting bias in the direct evidence. Fifth, examine heterogeneity and both local and global incoherence diagnostics with their low power in mind. Sixth, read direct, indirect, and combined estimates for the comparisons that matter.
Then translate results into absolute benefits and harms for a relevant baseline risk. Read certainty assessments before rankings. Finally, test whether reasonable changes to node definitions, studies, models, and assumptions alter the decision you would make.
Sources and further reading
- Cochrane Handbook, Chapter 11, Undertaking Network Meta-Analyses
- Hutton and colleagues, PRISMA Extension Statement for Network Meta-Analyses, Annals of Internal Medicine (2015)
- Salanti, Indirect and Mixed-Treatment Comparison, Network, or Multiple-Treatments Meta-Analysis, Research Synthesis Methods (2012)
- Nikolakopoulou and colleagues, CINeMA Framework for Evaluating Confidence in Network Meta-Analysis, PLOS Medicine (2020)
- Chaimani and colleagues, Graphical Tools for Network Meta-Analysis in STATA, PLOS One (2013)
Questions and answers
Can a network compare two drugs that have never been tested head to head?
Yes, if they connect through one or more common comparators and the transitivity assumption is credible. The result remains indirect and should be labeled and graded accordingly.
Does a nonsignificant inconsistency test prove transitivity?
No. Sparse networks often have little power to detect disagreement, and some transitivity violations do not produce visible statistical incoherence.
Is the highest SUCRA score the best treatment?
Not necessarily. SUCRA reflects relative ranking under the model, not effect size, clinical importance, safety, certainty, cost, or patient preference.
Why include placebo if it is not a treatment choice?
Placebo can connect active interventions and preserve information from placebo-controlled trials. Its role as a linking comparator still requires similar settings and effect modifiers.
Is network meta-analysis stronger than head-to-head evidence?
It can synthesize more information, but a credible, adequately powered, low-bias direct trial may be more persuasive for its exact comparison. Strength depends on the evidence and assumptions, not the method's complexity.