The gap between the ad and the label#
A cosmetic ad promises a smoother, younger, more rested face. The clinical trial behind it proved something much smaller and much more specific: that a defined number of people moved a defined amount on a defined scale, at one moment in time, in one small patch of skin. The whole skill of reading aesthetic evidence is learning to see that gap and to ask what the study genuinely tested. Botulinum toxin for the frown lines between the eyebrows, the glabellar lines, is a useful teaching case because its trials are well built and its approved claim is unusually easy to state.
Key points#
- An approved cosmetic indication is a narrow, sourced claim, not a general promise of a better appearance.
- Frown-line trials measure a four-point wrinkle score at maximum frown, not "younger" or "refreshed."
- Newer studies count a treatment as a success only when the patient and the evaluator both agree, which is a tougher bar than either rating alone.
- Blinding and a validated scale are what keep a subjective result honest.
- Every word of the label ("temporary," "moderate to severe," "glabellar," "adults") is a limit on what the evidence covers.
Start with the approved indication, and read it as a set of limits#
Before looking at any endpoint, read the sentence the regulator actually approved. For these products it is the temporary improvement in the appearance of moderate to severe glabellar lines in adults. That sentence looks like marketing, but every clause is a boundary on the evidence.
- Temporary. The benefit fades over months. The trials measured a peak and a duration, never a lasting change.
- Improvement in the appearance. The claim is visual and cosmetic. It is not a statement about skin biology or any health outcome.
- Moderate to severe. Only people who began at the middle or high end of the scale were studied. Mild lines were not the tested group.
- Glabellar lines. One region between the brows. An approval here is not evidence for the forehead, the corners of the eyes, or any other site that was not separately studied and cleared.
- Adults. The enrolled population, not everyone who might ask.
Marketing tends to smooth all five limits into one general pitch. Off-label use at other sites is common in practice, but common is not the same as studied and approved, and the two should stay separate when you weigh the evidence.
What the trial actually measured#
A pivotal aesthetic study does not score "rested" or "natural." It scores a change on a named instrument. For frown lines the standard tool is a four-point wrinkle severity scale, sometimes called the Facial Wrinkle Scale, usually paired with a photonumeric guide: 0 none, 1 mild, 2 moderate, 3 severe. Trials enroll only people who start at 2 or 3, because a scale cannot register improvement in a line that is already absent at rest.
Two design details do most of the interpretive work. First, the face is scored at maximum frown, while the person actively contracts the corrugator and procerus muscles. That is deliberate: the toxin acts on muscle activity, so the endpoint captures the dynamic lines it can affect and says little about static creases etched into the skin over years. Second, success is a specific threshold. A review of onabotulinumtoxinA and abobotulinumtoxinA in the clinical literature (PMC2861845) describes the pattern: randomized, placebo-controlled, double-blind studies in which a responder reaches a score of 0 or 1 at a set follow-up, typically around day 30, with the effect peaking in the first two to four weeks. The evidence has a shape, and it is a number, on a scale, at a time.
Why two raters beat one#
Recent approvals raised the standard. The FDA prescribing information for daxibotulinumtoxinA (2022) and its Drug Trials Snapshot define success at Week 4 as a score of none or mild plus at least a two-point improvement from baseline, and that result has to hold for both the investigator and the patient. Making an expert evaluator and the treated person each clear the same bar separately is a stricter test than either judgment on its own. It guards against a clinician who sees a change the patient does not feel, and against a hopeful patient who reports a change the evaluator cannot see. So when a summary says a treatment "worked," check whether success meant one rating or two ratings that had to agree before you believe it.
Why blinding and a validated scale carry the weight#
Cosmetic outcomes are unusually open to bias, because the result is subjective and the wish for it is strong. Two features keep the finding trustworthy.
A validated scale comes first. Validated means the instrument has been tested so that different trained raters give the same score to the same face, and the same rater stays consistent over time. The photonumeric reference photos exist so that "moderate" means the same thing across dozens of clinical sites. Without that discipline, a trial measures the rater's mood rather than the drug.
Blinding does the rest. In a double-blind trial, neither the injecting evaluator nor the patient knows who received active toxin and who received saline. This matters precisely because a wrinkle treatment creates expectation. An unblinded evaluator who knows a patient was treated tends to see success, and a patient who knows tends to report it. The placebo arm exists to subtract that expectation. When a network meta-analysis of botulinum toxin formulations (PubMed 36097079) pooled randomized trials to rank products against one another, its conclusions were only as sound as the blinded, scale-based endpoints feeding it: the share reaching 0 or 1, and the share improving by one or two points, all scored at maximum frown about a month out.
A checklist for any "clinically proven" cosmetic claim#
When a procedure claims clinical proof, six questions expose what the evidence really supports.
- What scale, and is it validated? A named, photonumeric, reliability-tested scale beats a vague "significant improvement."
- Who was enrolled? Baseline severity and body region define exactly what the result covers.
- Was it blinded and placebo-controlled? A subjective outcome without blinding invites the expectation effect.
- Who scored it, and did the ratings have to agree? An expert evaluator plus patient agreement is stronger than a single judgment.
- At what time point, and for how long? A peak effect and its duration are different facts, and one flattering timepoint can hide a short-lived result.
- Does the approved indication match the pitch made to me? If the promise is broader than the label, the extra breadth is not carrying trial-grade evidence.
The aim is not reflexive doubt. Glabellar-line botulinum toxin is among the better-evidenced procedures in aesthetic medicine, exactly because its trials were built around a validated scale, honest blinding, and a narrow, testable endpoint. That rigor is the reason to trust the specific claim, and the reason to stay cautious when a much broader promise borrows the same lab coat.
Sources and further reading
Questions and answers
Does an approved wrinkle treatment work on other lines too?
Not automatically. An approval for glabellar lines covers that one region and the severity range that was studied. Forehead lines and crow's feet need their own trials and their own clearance. Use at other sites may be reasonable clinically, but it is not backed by the same approval.
Why do trials score the face frowning rather than relaxed?
Because the toxin acts on muscle contraction. Scoring at maximum frown captures the dynamic lines the drug can actually change and avoids crediting it for static creases set into the skin over time.
What is the single most useful question to ask?
Ask what scale was used and whether it was validated. A named, reliability-tested scale, scored under blinding against placebo, tells you the result is about the treatment and not about hope.