Artificial intelligence has genuinely sped up the earliest steps of finding a drug, but the part that decides whether medicines reach patients, human trials, has barely moved. Picture the path from idea to pharmacy as a long pipeline. The front end, choosing a target and inventing candidate molecules, is where AI now does useful work. The back end, proving in people that a compound is both safe and effective, is where roughly nine out of ten programs have always died, and AI has not changed that arithmetic. Most of the noise you hear is really an argument about how much the front-end savings are worth once the back end is still standing in the way.
Key points#
- AI helps most with three early tasks: reading protein shapes, designing candidate molecules, and predicting how those molecules might behave.
- Those are all ways of ranking hypotheses faster. None of them replace the experiment that tests whether a drug actually works.
- Molecules that came from AI seem to clear early safety testing at a high rate, but so far they fail the efficacy stage at about the usual rate.
- As of the latest public data, no drug designed mainly by AI has finished the full path to approval.
- When you read a bold claim, ask what the model actually did, against which comparison, and measured by which outcome.
The three places AI earns its keep#
Start with what the evidence supports rather than what the marketing promises.
The first win is reading protein shape. A drug usually works by fitting into a pocket on a protein, so knowing that protein's three-dimensional shape gives chemists a head start. Software can now predict many of those shapes in hours, a task that once took years of laboratory work. That is a real acceleration, and it has limits worth naming. Predictions are most reliable for stiff, well-studied proteins and least reliable when the protein flexes, shifts between forms, or floats inside a fatty cell membrane the model represents poorly. A predicted shape is a good guess about geometry. It says nothing about whether latching onto that protein will actually move the disease, which is a separate and much harder problem.
The second win is inventing and refining molecules. Generative models can sketch enormous numbers of candidate structures and score each one against traits chemists care about, such as how tightly it might bind, whether it dissolves, or whether it carries chemical groups linked to harm. This turns a slow, hand-guided search into a fast, filtered one. The catch is familiar to you if you work with these systems: they shine when they have plenty of clean data to learn from and stumble when asked to reason about chemistry unlike anything in their training set.
The third win is predicting behavior, often bundled under the label ADMET, for absorption, distribution, metabolism, excretion, and toxicity. Forecasting these properties early lets a team abandon a weak candidate before pouring years of bench work into it. The value is genuine, but the forecasts are statistical. They tilt the odds toward a good molecule rather than certifying one.
The through line across all three is the same. AI is excellent at proposing, ranking, and prioritizing ideas. It is not a stand-in for the experiments that put those ideas to the test.
The wall AI cannot climb#
Now the numbers everyone quotes, read with the footnote and not just the headline.
One widely cited analysis of AI-originated molecules that reached human testing found early-phase safety success in the range of 80 to 90 percent, comfortably above the historic industry norm of roughly 40 to 65 percent. That figure is the crowd favorite. The first phase of testing mostly checks whether a drug is tolerated, so the kindest interpretation is that AI is genuinely good at drafting clean, well-behaved, drug-like molecules that do not misbehave at the first human doses.
The same analysis then reported a very different result one stage later. In the phase that tests whether the drug changes the disease, the AI cohort succeeded around 40 percent of the time, on a small sample, right in line with the historic average. This second stage has always been the graveyard of drug development, and the AI-designed candidates are dying there at the usual pace. The edge that looked so dramatic in early safety testing seems to evaporate exactly where the truly hard question starts.
Two cautions belong beside those figures. The pool of AI-designed drugs that have travelled deep into trials is still small, so the efficacy number could drift up or down as more programs report out. And as of the most recent public data, no drug discovered or designed mainly by AI has completed the full journey to regulatory approval. Some high-profile candidates have been shelved after longer-term data failed to confirm an early signal, which is an ordinary part of development and a plain reminder that early promise is not the same as proof.
The reason the back half of the pipeline resists compression is biological, not computational. A model can conjure a molecule that grips its target beautifully. Whether gripping that target meaningfully helps a varied population of real patients, without doing unacceptable harm, is something only a well-run trial can settle. That is the wall, and no current AI clears it.
The rules of the road did not loosen#
None of this softened the framework that governs testing in people. Good Clinical Practice, the standard known as ICH E6, was substantially revised, with the E6(R3) update reaching its final form in early 2025 and adoption spreading across major regions through the year. The revision deliberately modernizes trials around risk-based thinking, digital tools, and data-driven oversight, but it does not lower the evidence bar. The companion guidance on how trials are designed and how results are judged still applies. A study that uses AI to recruit patients, watch its sites, or crunch its data is still a trial, and it must still state in advance what would count as success and defend that analysis to a regulator.
How to judge a headline#
A few plain questions separate substance from spin. What, precisely, did the model do, and at which step of the pipeline? Is the reported advantage measured against a fair comparison or a flattering one? Is the endpoint a computer score, a preclinical readout, or an actual outcome in patients? And has the same honesty been applied to the failures as to the wins, since a platform that only publishes its survivors tells you almost nothing?
AI has earned a lasting place at the front of drug discovery. It makes the hunt for candidates faster and, in the right hands, sharper. What it has not done is repeal the math of clinical development, where safety and benefit in real people remain the deciding evidence. Treating AI as a powerful accelerant for the early stages is fair. Treating it as a bypass around the trials is not.
Sources and further reading
Questions and answers
Has AI invented an approved drug yet?
Not as of the latest public data. Several AI-originated molecules have entered human trials and a few have progressed, but none has completed the full path to regulatory approval, and some have been discontinued along the way.
Why do AI-designed drugs do so well in early testing but not later?
Early testing mainly checks safety and tolerability, and AI seems good at designing clean, drug-like molecules that behave predictably at first. Later testing asks whether the drug changes the disease in real patients, a biological question that no model can answer on its own, so the early advantage tends to fade.
Does AI make new medicines cheaper or faster overall?
It can shorten and de-risk the earliest steps, which is valuable. But the most expensive, failure-prone part of development is the human trials, and that stage has not sped up, so the total cost and timeline of bringing a drug to market remain largely unchanged for now.