Evidence explainer

Medicines and drug development

Wearables in Clinical Trials: What Fit for Purpose Really Requires

A number from a wrist sensor becomes evidence only when the device and the measurement are shown to fit the exact question a trial asks. FDA's 2023 final guidance sets out what that proof looks like.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Key points
  2. Why a wearable is a stand-in, and why that matters
  3. The three proofs behind a fit-for-purpose number
  4. A short checklist for reading a wearable claim
  5. Consumer gadget versus regulated endpoint

When a study reports that a treatment improved "sleep quality" or "daily activity" as measured by a wrist device, the result is trustworthy only if that device and the measurement built on it have been shown to be fit for purpose for the precise question being asked. That short phrase carries the weight of the U.S. Food and Drug Administration's final guidance, Digital Health Technologies for Remote Data Acquisition in Clinical Investigations, released in December 2023 and posted to the Federal Register on December 22 of that year. Fit for purpose is not a slogan. It is a chain of three separate proofs, and a number that skips any link in that chain is closer to advertising than to evidence.

Key points#

Why a wearable is a stand-in, and why that matters#

In the FDA's language, a digital health technology is a system of hardware and software that measures or captures data for a health purpose. Inside a trial, that device is doing a job once handled by a trained observer or a clinic instrument. It is a stand-in. Swapping a person or a calibrated machine for a consumer-grade sensor is only fair if the sensor produces data of comparable meaning, not merely comparable convenience. The 2023 document finalizes a 2021 draft and was completed in part to satisfy a mandate Congress set in the Food and Drug Omnibus Reform Act, so these expectations now rest on statute rather than sitting as friendly advice.

The discipline the guidance imposes is simple to state and easy to blur: keep three claims apart that marketing tends to fuse into one.

The three proofs behind a fit-for-purpose number#

Does the hardware do what its spec says?#

The first proof is verification: evidence that the device and its software capture and process the raw signal correctly. A photoplethysmography sensor estimating heart rate, an accelerometer tallying steps, or a phone camera reading a test strip each has a technical specification, and verification shows that a given unit meets it under the conditions the trial will actually see. One detail is easy to miss. When a trial uses several device models or firmware versions, the sponsor has to show they agree with one another, because a step counted on one watch should mean the same as a step counted on another. Skin tone, motion, and how snugly the band sits are not fine print. They are the very conditions that decide whether a sensor performs to spec or falls short without anyone noticing.

Does the measurement mean anything clinically?#

Verification can pass while the endpoint stays hollow. Validation is the harder proof, and it comes in two layers. Analytical validation asks whether the algorithm's output faithfully reflects the physiological signal it is processing, judged against a reference standard. Clinical validation asks whether that output actually corresponds to how a patient feels, functions, or survives. A wrist device can track movement with real precision and still say nothing dependable about whether someone's disease has improved. The distance between measuring motion accurately and demonstrating that a patient got better is precisely where thin evidence tends to hide.

Can the intended patient really use it?#

The third proof is usability, and the guidance treats it as part of fit for purpose rather than a nicety. If participants cannot operate the device as designed, the data decay no matter how good the sensor is. That means checking whether the target population, including older adults or people living with the very condition under study, can wear, charge, and interact with the technology for the length of the trial. The guidance also notes that lacking a personal device should not by itself exclude someone from a study. That is both an equity principle and a data-quality one: a sample made up only of the technologically comfortable may not represent the patients a treatment is meant to help.

A short checklist for reading a wearable claim#

Because fit for purpose is always contextual, a handful of questions do most of the work when you meet a wearable result.

What exactly was measured, and against what reference? A named comparator, such as polysomnography for sleep or a supervised walk test for mobility, is reassuring. Silence about the reference standard is a warning sign.

Is the endpoint established or brand new? The guidance draws a firm line. Using a digital tool to replace manual capture of an accepted endpoint is one task. Using it to define a new endpoint is a heavier lift, and the FDA expects extra justification: how the novel measure relates to other endpoints, how reliable its data are, and how a known treatment effect would even show up through it. Novel digital endpoints can be genuinely useful, yet they carry a larger burden of proof that marketing language rarely acknowledges.

Do the people studied resemble the people the product targets? A device validated in young, healthy volunteers tells you little about how it behaves in the frail or the acutely ill.

Who was accountable when data went wrong? The guidance expects sponsors to plan for firmware updates, data loss, and device errors, and to restrict trial devices to participants and their caregivers. A trial that spells out these procedures is treating its data with the seriousness the standard requires.

Consumer gadget versus regulated endpoint#

The same hardware often runs both a consumer wellness feature and a regulated trial endpoint, and the two live under different rules. A consumer metric can be roughly directional and still be worth having for personal curiosity; nobody is deciding whether a drug works based on your morning readiness score. A trial endpoint that will inform whether a medical product is effective has to clear verification, validation, and usability before it counts. When a marketing page borrows the phrase "clinically validated" without showing that chain, it is leaning on a term the FDA reserves for a specific and demanding process. Knowing what the process requires is what lets a reader tell the two apart.

Sources and further reading

  1. FDA Final Guidance: Digital Health Technologies for Remote Data Acquisition in Clinical Investigations
  2. Federal Register Availability Notice (Dec 22, 2023)
  3. FDA: Digital Health Technologies (DHTs) for Drug Development

Questions and answers

Is a "clinically validated" wearable automatically trustworthy for any measurement?

No. Validation is tied to a specific measurement and context. A device shown to be valid for resting heart rate has not been shown to be valid for sleep stages, stress, or disease activity. Always ask which measurement, and against which reference standard, the claim refers to.

What is the difference between verification and validation?

Verification confirms the hardware and software capture the raw signal correctly against a technical spec. Validation confirms that the resulting measurement reflects the true physiological signal (analytical validation) and corresponds to something clinically meaningful (clinical validation). A device can pass verification and still fail validation.

Why does the FDA treat novel digital endpoints more cautiously?

Because a new endpoint has no established track record linking it to how patients feel, function, or survive. The guidance asks sponsors to justify how the novel measure relates to other endpoints, how reliable it is, and how a known effect would be detectable through it, which is a larger burden than digitizing an already-accepted measure.