A PRISMA flow diagram is the receipt for a systematic review. It shows, count by count, how the authors went from thousands of database hits to the handful of studies they actually analyzed, and it lets you check their subtraction. Read it from the top down through four levels (what the search returned, what survived a title and abstract pass, what survived full-text reading, and what entered the final synthesis), and confirm one thing at each step: every record that left the funnel is accounted for with a stated number and, where it matters, a reason. When those tallies balance and the exclusion reasons are specific, you are looking at a search you can audit. When they do not, you are being asked to trust a selection process you cannot see.
Key points#
- The flow diagram reports how evidence was found and filtered, not whether the included studies were any good.
- The arithmetic has to close: records entering a box, minus exclusions, must equal records leaving it.
- The exclusion reasons at full-text review are the most informative numbers on the page.
- A tiny included set drawn from a large search, with no searchable strategy, deserves a second look.
- Read the flow diagram before the forest plot.
Why this box matters before any result#
PRISMA stands for Preferred Reporting Items for Systematic Reviews and Meta-Analyses, a reporting standard updated in 2020. The flow diagram is only one piece of it, and its job is deliberately narrow. It will not tell you whether an effect is large, whether the studies were well designed, or whether the conclusion holds. It tells you how the pool of evidence was assembled. That is worth reading first because everything downstream inherits from that pool. A careful meta-analysis run on a selectively gathered set of studies still produces a selectively gathered answer, and the forest plot alone will never reveal it.
Structurally the diagram is a column of boxes joined by arrows, and the arrows carry counts. Records enter at the top. At each level some are removed and branch off to the side with a tally, and what remains flows down. The format is doing one useful thing: it asks the authors to let you trace a single number from the top of the page to the bottom without ever losing sight of where records went.
A worked example#
It is easier to see when you put numbers on it. Imagine a review of a blood-pressure drug that reports 8,200 records from three databases, plus 40 found by checking reference lists. After removing 1,900 duplicates, 6,340 records go to screening. Title and abstract screening excludes 6,100, leaving 240 reports sought for full-text review. Of those, 190 are excluded with reasons: wrong population (70), wrong comparator (55), no usable outcome data (40), and wrong study design (25). That leaves 50 studies in the review, 44 of them contributing to the meta-analysis.
Now audit it. Does 8,240 minus 1,900 equal 6,340? Yes. Does 240 minus 190 equal 50? Yes. Do the four exclusion reasons add to 190? They do. The subtraction closes at every level, the reasons are categorized rather than lumped, and the gap between 50 included and 44 pooled is small and explainable. This is what an auditable funnel looks like. You still do not know whether the drug works, but you know the authors are showing their work.
The four levels, and what to check at each#
Identification#
The top box reports how many records the search returned, usually split by source. Expect counts from bibliographic databases (MEDLINE, Embase, the Cochrane registers) and, in the 2020 layout, a separate stream for records found through other methods: citation searching, reference lists, trial registries, or contact with authors. Two questions matter here. Were multiple databases searched? A review resting on one index has a built-in blind spot, because no single database covers the whole literature. And is deduplication stated as a number? The same paper turns up in several databases, so removing duplicates before screening is expected and healthy, as long as you see the figure rather than an unexplained drop from a large count to a smaller one.
Screening#
Screening is the title and abstract pass. A reviewer, ideally two working separately, reads each summary and decides whether it is plausibly relevant. Most records die here, and that is normal: a broad search of a common condition can return tens of thousands of hits, and the great majority are obviously off topic on a five-second read. Bulk exclusion without individual reasons is acceptable at this level, because you are discarding the clearly irrelevant. The number to watch is the ratio of records reaching full-text review. If fifteen thousand records collapse to twelve, that funnel is either unusually precise or too narrow, and you deserve the search string to judge which.
Eligibility#
This is the level that separates a rigorous review from a loose one. Reports that survived screening are read in full and checked against the pre-specified inclusion and exclusion criteria, and every report removed here must carry a reason. Read those reasons closely, because they are the most informative counts on the diagram. Specific, categorized exclusions signal that the criteria were applied consistently. A single lumped figure, "excluded, n = 47" with no breakdown, tells you nothing and should lower your confidence, because it hides whether studies were dropped on principle or because their results were inconvenient. The reasons are also where you inspect the criteria themselves. If a review of a drug's benefit excludes trials that captured a relevant safety outcome, that is a choice worth understanding before you accept the summary estimate.
Inclusion#
The bottom box gives the number of studies in the qualitative synthesis and, where relevant, the number contributing to the quantitative synthesis or meta-analysis. Those two counts can differ for legitimate reasons, because a study can meet the criteria yet lack data in a poolable form. When they differ, the diagram should say by how much, and ideally why.
Patterns that should make you pause#
A few signatures are worth slowing down for. The clearest is arithmetic that does not close: the number entering a box does not equal the number leaving it plus those excluded. That is a bookkeeping failure at best and, at worst, a sign that records were handled off the page. Missing exclusion reasons at the eligibility level are another, because they strip away your ability to audit the most consequential decisions. Be cautious of a single narrow database with no attempt at other sources, which raises the odds that relevant work never entered the funnel at all. And be wary of an unusually sharp search that yields a tiny included set from a large base with no strategy you can inspect. None of these prove misconduct on their own. Each one simply moves the burden of trust off the diagram and onto your own skepticism, which is the reverse of what good reporting should do.
The habit worth building is to read the flow diagram before the forest plot. The forest plot tells you what the included studies found. The flow diagram tells you whether the right studies were included in the first place, and whether the reasons for leaving evidence out were stated plainly or left for you to guess. A search that shows its work is not automatically correct, but it is auditable, and auditable is the precondition for trusting anything that follows.
Sources and further reading
Questions and answers
Does a clean flow diagram mean the review is trustworthy?
No. It means the search and selection are auditable. The diagram says nothing about whether the included studies were well designed or whether the pooled estimate is reliable. For that you look at risk-of-bias assessment, heterogeneity, and the individual study methods.
What is the single most useful part to read?
The exclusion reasons at the full-text level. Categorized, specific reasons show the criteria were applied consistently. A single lumped number with no breakdown is the most common sign that the most important decisions cannot be checked.
Why do the "included" and "meta-analysis" counts sometimes differ?
Because a study can be eligible for the review yet report its data in a form that cannot be pooled statistically. A small, explained gap is normal. A large or unexplained one is worth a closer look.