Evidence explainer

Skin, musculoskeletal, and eye health

What an Approved Aesthetic Trial Actually Proves

An approved cosmetic claim rests on a narrow result. For frown-line botulinum toxin it is a four-point wrinkle score, agreed by patient and evaluator under blinding, on a fixed day, against placebo.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. The gap between the ad and the label
  2. Key points
  3. Start with the approved indication, and read it as a set of limits
  4. What the trial actually measured
  5. Why blinding and a validated scale carry the weight
  6. A checklist for any "clinically proven" cosmetic claim

The gap between the ad and the label#

A cosmetic ad promises a smoother, younger, more rested face. The clinical trial behind it proved something much smaller and much more specific: that a defined number of people moved a defined amount on a defined scale, at one moment in time, in one small patch of skin. The whole skill of reading aesthetic evidence is learning to see that gap and to ask what the study genuinely tested. Botulinum toxin for the frown lines between the eyebrows, the glabellar lines, is a useful teaching case because its trials are well built and its approved claim is unusually easy to state.

Key points#

Start with the approved indication, and read it as a set of limits#

Before looking at any endpoint, read the sentence the regulator actually approved. For these products it is the temporary improvement in the appearance of moderate to severe glabellar lines in adults. That sentence looks like marketing, but every clause is a boundary on the evidence.

Marketing tends to smooth all five limits into one general pitch. Off-label use at other sites is common in practice, but common is not the same as studied and approved, and the two should stay separate when you weigh the evidence.

What the trial actually measured#

A pivotal aesthetic study does not score "rested" or "natural." It scores a change on a named instrument. For frown lines the standard tool is a four-point wrinkle severity scale, sometimes called the Facial Wrinkle Scale, usually paired with a photonumeric guide: 0 none, 1 mild, 2 moderate, 3 severe. Trials enroll only people who start at 2 or 3, because a scale cannot register improvement in a line that is already absent at rest.

Two design details do most of the interpretive work. First, the face is scored at maximum frown, while the person actively contracts the corrugator and procerus muscles. That is deliberate: the toxin acts on muscle activity, so the endpoint captures the dynamic lines it can affect and says little about static creases etched into the skin over years. Second, success is a specific threshold. A review of onabotulinumtoxinA and abobotulinumtoxinA in the clinical literature (PMC2861845) describes the pattern: randomized, placebo-controlled, double-blind studies in which a responder reaches a score of 0 or 1 at a set follow-up, typically around day 30, with the effect peaking in the first two to four weeks. The evidence has a shape, and it is a number, on a scale, at a time.

Why two raters beat one#

Recent approvals raised the standard. The FDA prescribing information for daxibotulinumtoxinA (2022) and its Drug Trials Snapshot define success at Week 4 as a score of none or mild plus at least a two-point improvement from baseline, and that result has to hold for both the investigator and the patient. Making an expert evaluator and the treated person each clear the same bar separately is a stricter test than either judgment on its own. It guards against a clinician who sees a change the patient does not feel, and against a hopeful patient who reports a change the evaluator cannot see. So when a summary says a treatment "worked," check whether success meant one rating or two ratings that had to agree before you believe it.

Why blinding and a validated scale carry the weight#

Cosmetic outcomes are unusually open to bias, because the result is subjective and the wish for it is strong. Two features keep the finding trustworthy.

A validated scale comes first. Validated means the instrument has been tested so that different trained raters give the same score to the same face, and the same rater stays consistent over time. The photonumeric reference photos exist so that "moderate" means the same thing across dozens of clinical sites. Without that discipline, a trial measures the rater's mood rather than the drug.

Blinding does the rest. In a double-blind trial, neither the injecting evaluator nor the patient knows who received active toxin and who received saline. This matters precisely because a wrinkle treatment creates expectation. An unblinded evaluator who knows a patient was treated tends to see success, and a patient who knows tends to report it. The placebo arm exists to subtract that expectation. When a network meta-analysis of botulinum toxin formulations (PubMed 36097079) pooled randomized trials to rank products against one another, its conclusions were only as sound as the blinded, scale-based endpoints feeding it: the share reaching 0 or 1, and the share improving by one or two points, all scored at maximum frown about a month out.

A checklist for any "clinically proven" cosmetic claim#

When a procedure claims clinical proof, six questions expose what the evidence really supports.

The aim is not reflexive doubt. Glabellar-line botulinum toxin is among the better-evidenced procedures in aesthetic medicine, exactly because its trials were built around a validated scale, honest blinding, and a narrow, testable endpoint. That rigor is the reason to trust the specific claim, and the reason to stay cautious when a much broader promise borrows the same lab coat.

Sources and further reading

  1. Botulinum toxin type A for glabellar lines: science and clinical data (PMC)
  2. Network meta-analysis of BoNT/A for glabellar lines (PubMed)
  3. DAXXIFY Prescribing Information, FDA (2022)
  4. FDA Drug Trials Snapshot: DAXXIFY

Questions and answers

Does an approved wrinkle treatment work on other lines too?

Not automatically. An approval for glabellar lines covers that one region and the severity range that was studied. Forehead lines and crow's feet need their own trials and their own clearance. Use at other sites may be reasonable clinically, but it is not backed by the same approval.

Why do trials score the face frowning rather than relaxed?

Because the toxin acts on muscle contraction. Scoring at maximum frown captures the dynamic lines the drug can actually change and avoids crediting it for static creases set into the skin over time.

What is the single most useful question to ask?

Ask what scale was used and whether it was validated. A named, reliability-tested scale, scored under blinding against placebo, tells you the result is about the treatment and not about hope.