The honest way to judge an ergogenic-aid claim is to ignore the confidence on the label and ask a single question: what does the evidence behind it actually look like? Apply that filter and the list of substances with real support for improving exercise performance turns out to be very short, with creatine and dietary nitrate among the few that survive. Even for those, the benefit is modest, depends on dose and timing, and tends to be smallest in the people who are already the fittest. This piece walks through how the grading works, from the strongest form of evidence down to the questions you can ask yourself.
Key points#
- Only a handful of supplements have good evidence for a direct performance benefit; a 2018 IOC consensus named caffeine, creatine, certain buffering agents, and nitrate.
- The strongest evidence is not one dramatic study but many randomized trials pooled and quality-checked, using tools such as AMSTAR-2.
- Statistical significance is not size. Real, replicated effects can still be too small to notice, and they often fade in trained athletes.
- Dose, timing, and the population tested all bound what a supplement can do, so a claim that promises the same result to everyone is overreaching.
- A performance effect says nothing about safety, and some products carry contamination risk for competitive athletes.
Start at the top of the evidence ladder#
The best evidence for a performance claim is rarely the study a marketer quotes. A single trial, however clean, can land on a chance result or a design quirk. Confidence grows only when many well-run randomized trials are pooled and weighed together, and when someone also asks how trustworthy each of those pooled reviews really is.
Sports scientists now use structured tools for that second step. A widely applied one is AMSTAR-2, a checklist that rates a systematic review as high, moderate, low, or critically low confidence based on how it handled bias, study selection, and its own analysis. Above that sits the umbrella review, which does not gather studies but gathers the meta-analyses themselves and grades them.
Two 2025 umbrella reviews on dietary nitrate show the machinery at work. A Sports Medicine umbrella review by Poon and colleagues pulled together 20 systematic reviews of nitrate and scored them with AMSTAR-2, finding that most rated low or critically low, a candid admission that even a genuine effect can rest on a shaky literature. A separate umbrella review in Nutrients by Tian and colleagues did the same for beetroot juice, the most common nitrate source, and rated most of its underlying reviews as moderate to high quality. When you read any claim, that layered structure, many randomized trials pooled and then quality-checked, is the shape you are hoping to find behind it. Its absence is itself information.
Why so few substances make the cut#
It is tempting to assume a short list means the reviewers are impossible to satisfy. The real reason is that most performance claims are built on thin foundations: one small study, an unblinded design, a lab measure that never translates to competition, or a result that evaporates when someone repeats the experiment.
The 2018 International Olympic Committee consensus statement in the British Journal of Sports Medicine, led by Ronald Maughan, surveyed this whole field and concluded that only a small set of supplements had good evidence for a direct performance benefit: caffeine, creatine, specific buffering agents, and nitrate. After decades of research and marketing, the substances that clear the bar can be counted on one hand. That is not the reviewers being fussy. It is grading evidence doing exactly its job, which is to ask whether a finding would still stand if you looked at all the data rather than the most flattering slice of it.
A significant effect is not a large one#
Once a substance has survived that scrutiny, the next trap is confusing statistical significance with practical size. The Poon review found that nitrate improved outcomes such as time to exhaustion and muscular endurance, but the pooled effect sizes were small to moderate, roughly a quarter to a half of a standard deviation depending on the measure. The Tian review described strength improvements as negligible in magnitude even where they reached significance.
For a reader, that gap is the whole point. A supplement can have a real, replicated effect that is still too small to feel in a recreational setting, or that shows up only in one narrow task. A claim that leads with "clinically proven" or a headline percentage, but never tells you the baseline, the exact task, or who was tested, is hiding the very numbers that would let you judge whether the effect matters to you.
The conditions the effect actually needs#
A supported substance only performs inside the conditions that were studied, and the nitrate reviews are precise about this. Poon and colleagues reported that benefits were more consistent above a threshold intake around 6 mmol per day and with several days of continued use rather than a single serving. Tian and colleagues described both acute pre-exercise dosing and multi-day protocols that reached defined nitrate levels.
Who was tested matters just as much as how much. Both nitrate reviews found larger effects in non-athletes and recreationally active people than in highly trained individuals, which fits a ceiling effect: the fitter a body already is, the less room a supplement has to improve it. Creatine shows a related pattern, with responses that differ from person to person. The IOC panel put the general lesson plainly, noting that responses vary widely between people because of factors including genetics, the gut microbiome, and habitual diet. A claim that promises everyone the same gain is contradicting the evidence it leans on.
Five questions that sort signal from noise#
You can run much of a reviewer's filter yourself without reading a single paper. A handful of questions do most of the work:
- Is the claim backed by pooled randomized trials, or by one study, testimonials, or a mechanism alone?
- Does it give an effect size and the exact task, or only a percentage with no baseline?
- Does it state the dose, the timing, and the population tested, and do those match you?
- Does it admit that responses vary, or promise one uniform result?
- Is the outcome real performance, or a laboratory stand-in that may not carry over?
Most marketed claims stumble on several of these at once. That does not prove a product does nothing. It signals that the benefit has not been established, which is a very different thing from established evidence of benefit.
Safety is a separate question#
Evidence that something helps performance says nothing about whether it is safe, and a claim guarantees neither. The IOC statement flagged a hazard aimed squarely at competitive athletes: some supplements are contaminated with substances banned under anti-doping codes, and an accidental dose can end a career. Reading a claim well means reading what it omits, including what is actually in the bottle. Anyone weighing a supplement should talk it through with a clinician who knows their history and goals.
Sources and further reading
Questions and answers
Are creatine and nitrate worth taking?
The evidence for a small performance benefit is stronger for these than for almost anything else on the market, but "stronger" still means modest, dose-dependent, and often minimal in well-trained people. Whether that trade is worth it depends on your goals, your training level, and a conversation with a clinician.
Why do so many supplements have glowing claims but weak evidence?
A confident claim is cheap to make and does not require pooled randomized trials behind it. Marketing can rest on a single small study, a lab measure, or testimonials, none of which meet the standard a graded review applies.
Does a bigger reported percentage mean a better supplement?
Not by itself. A percentage without the baseline, the task, and the population tested cannot be judged, and a statistically significant change can still be too small to notice in real training or competition.