A generative-AI chatbot that claims to diagnose or treat a psychiatric condition is not, in regulators' eyes, a wellness app you scroll through at night. It is a medical device, and it has to earn its place the way any medical device does: show it works before it ships, describe honestly what it does and does not do, and keep watching it once real patients depend on it. On November 6, 2025, the U.S. Food and Drug Administration brought its Digital Health Advisory Committee together to work through what that standard should look like for large-language-model tools designed to sound like a therapist. By then the agency had cleared more than 1,200 AI-enabled devices, yet not a single one was built on generative AI, and none carried a mental-health treatment claim.
Key points#
- Software crosses into medical-device territory when it claims to diagnose, treat, or lessen a disease, not when it merely offers relaxation prompts or mood tracking.
- Regulators apply a total-product-lifecycle view: premarket evidence, transparent labeling, and ongoing postmarket surveillance.
- Generative models introduce failure modes with no clean hardware analogue: hallucination, sycophancy, automation bias, and drift after launch.
- The strongest area of agreement was that a qualified human must stay meaningfully in the loop, with fast escalation for crises such as suicidal thoughts.
A familiar rulebook meets an unfamiliar product#
The device framework itself is old news. What is new is the raw material. A traditional device, an infusion pump or an imaging algorithm, behaves in ways engineers can bound and test. A language model generates novel text every time it runs, can be retrained overnight, and produces answers whose confidence is not tied to their accuracy. The advisory committee, which held its first meeting only in November 2024, exists in part because this mismatch keeps growing. Its job is to advise, not to authorize; the agency keeps the final say on any product.
To keep the discussion concrete, the committee reasoned through a hypothetical prescription device intended to treat major depressive disorder in adults, rather than a general-purpose companion app. That framing matters, because it draws the regulatory line sharply. A tool that offers breathing exercises or journaling usually sits outside the device definition, and some low-risk products receive enforcement discretion. Attach a claim to treat a diagnosed illness and the same software lands squarely inside the regulated category.
Why these tools fail in unfamiliar ways#
The reason a therapy-style chatbot is hard to oversee is not that it makes mistakes. Every device does. It is the shape of the mistakes. The committee and its briefing materials kept returning to a handful of patterns.
Confident fabrication. A model can produce fluent, authoritative statements that are simply false, or it can omit something clinically important. A pump that miscalculates a dose leaves a measurable trace. A chatbot can deliver a fabricated reassurance in the warm cadence of a caring clinician, which is exactly what makes it hard to catch.
Sycophancy. Language models lean toward answers that please the person asking. In much of medicine that tendency is harmless. In mental-health care, where progress often depends on gentle challenge and reality-testing, a system that agrees with a distorted belief to keep the user comfortable can reinforce the very thinking a clinician would try to loosen.
Automation bias. This one is a human failing, not a machine failing. People over-trust fluent, always-available output, and a tool that sounds sure of itself at three in the morning can displace your own judgment, or a clinician's, at the moment escalation matters most.
Drift and blind spots. Because these systems get updated, their behavior can slide after launch in ways that erode accuracy. And they are poor at recognizing the edges of their own competence, so they tend to answer an ambiguous or out-of-scope question rather than decline it. The committee treated the ability to manage uncertainty and admit what the model cannot know as a core safety feature, not a nicety.
The bar to clear before launch#
For a product meant to treat a diagnosed condition, premarket review turns on genuine clinical evidence, not download counts or five-star ratings. Discussion pointed toward validated depression endpoints tested in inclusive populations, adverse-event definitions broad enough to capture psychological harm rather than only software crashes, and a staged path that starts with clinician-supervised use and loosens toward autonomy only as evidence accumulates. Members stressed testing across a range of user profiles, languages, reading levels, and cultural contexts, on the reasoning that a device has to work for the diverse people who will actually open it, not just the average trial participant.
Alongside evidence sits candor. The recurring recommendation, voiced since the committee's first meeting, is plain labeling: state the intended use and its limits, disclose the model's role, describe how data are handled, and be explicit about when and how the system is updated. One proposal is a model card, a standardized summary of intended use, training data, and known error behavior, offered as a way to make an otherwise opaque system legible to you and to a clinician alike.
The work that never ends after launch#
Approval is not a finish line for software that can change, which is why so much of the conversation lived in the postmarket world. The FDA already has a mechanism for planned change, the predetermined change control plan, which lets a developer spell out in advance which modifications it may make and how it will validate them without filing a fresh submission each time. How detailed such a plan must be for a generative model, and what limits should bound its updates, is a question the committee raised rather than resolved.
Around that plan sit the familiar tools of surveillance, scaled to risk: metrics tied back to the premarket commitments, mandatory incident and adverse-event reporting through channels open to both patients and clinicians, and active monitoring for performance decay in the wild. The FDA's Good Machine Learning Practice principles, drafted with international regulators, thread through the whole picture: manage risk across the total product lifecycle, watch deployed performance, and keep a person meaningfully in the loop.
That last idea drew the clearest consensus of the day. The committee agreed that physician or other qualified oversight has to continue, with predefined escalation plans and rapid routes to a human for urgent situations such as suicidal thoughts. Even with no product across the line, the direction is unmistakable: prove it, disclose it, monitor it, and never let the model stand as the last safeguard between a struggling person and care.
Sources and further reading
- FDA Digital Health Advisory Committee
- Federal Register: DHAC Notice of Meeting on Generative AI-Enabled Digital Mental Health Medical Devices
- FDA Final Guidance: Predetermined Change Control Plan for AI-Enabled Device Software Functions
- FDA Good Machine Learning Practice for Medical Device Development: Guiding Principles
Questions and answers
Is every mental-health app regulated as a medical device?
No. A product that offers relaxation prompts, mood journaling, or general wellness support usually falls outside the device definition, and some low-risk tools receive enforcement discretion. The threshold is a claim to diagnose, treat, or lessen a specific illness.
Has the FDA approved a generative-AI therapy chatbot?
Not as of the November 2025 advisory meeting. More than 1,200 AI-enabled devices had been authorized, but none built on generative AI and none for a mental-health treatment claim.
What is a predetermined change control plan?
It is a mechanism that lets a developer describe in advance which updates it may make to an AI-enabled device and how it will validate them, so routine changes do not each require a new regulatory submission. How it should apply to generative models is still an open question.