Regulators do not approve a claim because a stack of studies is tall. When an assessor opens a clinical dossier, three questions run in parallel: was each study done well enough to be believed, does the evidence actually match the people and the decision in front of them, and does the whole body of work point the same way. A separate lever sits on top of all three. The height of the bar is set by what the product claims to do and how badly a wrong answer would hurt, so a wellness reminder and a life-sustaining implant are never asked to clear the same fence.
Key points#
- A grade is a judgment about whether the evidence supports a specific claimed use, not a count of studies.
- Study quality asks whether a result was earned; the design label alone is not the grade.
- Directness (relevance) asks whether the study answers the actual question being decided.
- Totality asks whether many studies, of different kinds, converge, including the ones that did not flatter the product.
- The required standard of proof rises with the intended use and the harm of an error.
Three questions behind every grade#
Structured appraisal systems such as GRADE break the reading of evidence into parts so that reviewers weigh the same features every time (Guyatt and colleagues, BMJ 2008). Three of those features do most of the work.
Was the study done well enough to believe it#
Quality is the question of whether a study earned the number it reports. An assessor looks at how participants were assigned to groups, whether the people measuring outcomes knew who received what, how many participants dropped out, and whether the final analysis counted everyone who started. A randomized, blinded trial that follows its pre-registered plan is built to resist the biases that push weaker designs toward flattering a treatment.
The common mistake is to read the design label as the grade. A randomized trial run carelessly, with unblinded outcome assessment and heavy dropout, can be less trustworthy than a carefully built observational study. Assessors read the conduct, not the badge. What counts is whether the methods actually controlled the specific ways that this kind of study tends to mislead.
Does it answer the right question#
A flawless study can still be the wrong evidence. If it enrolled a different population, tested a different version of the product, used a different dose, or measured an outcome that does not match the claim, then its precision is beside the point. Appraisers call this directness, and in the final judgment it often carries more weight than quality does.
The gap usually hides in the small print. A trial may have recruited younger and healthier volunteers than the people who will actually use the product. It may have measured a short-term laboratory marker when the claim is about something a patient feels over a year. A result that holds firmly in one setting cannot be assumed to transfer cleanly to another, and directness is where that caution turns into a regulatory question rather than an academic footnote.
Do the studies agree with each other#
Totality is the demand that the whole body of evidence point in a consistent direction. One striking trial is a claim awaiting confirmation, not a settled fact. Several studies, of different designs, run in different places, all landing near the same answer are far harder to dismiss as chance or the quirk of a single site. Assessors weigh consistency, the size of the effect, whether more of the product produces more of the response, and biological plausibility together, rather than reading any one figure in isolation.
Totality also means counting the evidence that never made the highlight reel. Reviewers look for trials that were registered and never reported, outcomes that were measured and then dropped, and subgroups in which the effect vanished. A dossier that shows only its best results has told the assessor less than it thinks, because experienced reviewers read the silences as carefully as the tables.
Why the same evidence can pass or fail#
Here is the part that surprises people: an identical body of evidence can be enough for one product and short of the mark for another. The reason is that the bar is set by the consequence of being wrong.
Picture two pieces of software that use the very same algorithm. In the first, it tidies a list of results for a clinician to read and confirm. In the second, it decides on its own whether that clinician is ever alerted to a dangerous value. The math is identical; the stakes are not. The higher the harm from an error, the more confirmation an assessor demands before accepting the claim.
This is why intended use is the hinge of the whole assessment, and why it is far more specific than the technology. The claimed use, written down in plain language, determines the evidence required. Widen the claim and the bar rises with it, because the product is now asking to be trusted in situations the narrower claim never covered. Regulatory frameworks differ across jurisdictions, but they share this logic: they sort products into tiers by the consequence of failure and ask for proof proportionate to it. A prediction model offered as a second opinion is held to a gentler standard than the same model offered to replace human review.
Benefit and risk, weighed on one scale#
A product clears the bar when its expected benefit, judged against the uncertainty that remains, outweighs its expected harm for the use being claimed. Regulators describe this explicitly as a structured benefit-risk judgment rather than a pass-fail checklist (FDA benefit-risk framework). Assessors will tolerate more residual uncertainty when the potential benefit is large and the existing options are poor, and less when the condition is mild and good treatments already exist.
That trade-off is why context changes the verdict. Preventive services aimed at healthy people are held to this same net-benefit reasoning: a recommendation depends not only on how strong the evidence is but on how the benefits and harms balance for the population in question (USPSTF grade definitions). Evidence that satisfies a regulator for a severe disease with no good treatment may fall well short for a minor complaint that other products already handle safely.
How to read, or build, a claim#
If you are assembling evidence for a claim, write the intended use first and let that one sentence discipline everything else, because every study will be judged against it. Match the evidence to the population and the decision you are claiming, not to whoever was easiest to enroll. Report the disappointing results next to the favorable ones, since their absence is the first thing a careful reviewer notices.
If you are reading someone else's claim, borrow the assessor's checklist. Was each study done well enough to believe. Does it match the people and the decision you care about. Does the whole body of evidence agree. Then ask what the claim is actually for, and how much harm a wrong answer would cause, because that is what decides how much proof is enough. A claim that cannot even name its intended use has not yet earned a grade. For anything touching your own health, take the specifics to a clinician who knows you.
Sources and further reading
Questions and answers
Is a randomized trial always the strongest evidence
Not automatically. A randomized design controls many biases by construction, but a trial run with unblinded outcomes, high dropout, or a population unlike the intended users can be weaker in practice than a well-conducted observational study. Reviewers grade the conduct and the fit, not the label.
Why do regulators ask for more than one good study
Because a single result, however impressive, can reflect chance or the peculiarities of one site or protocol. When studies of different designs and settings converge on a similar answer, that agreement is much harder to explain away, which is what totality of evidence is meant to capture.
Why does the required evidence change from product to product
Because the bar tracks the intended use and the harm of an error. A tool that supports a clinician who confirms every result is judged more leniently than one that acts on its own in a high-stakes decision, even when the underlying method is the same.