All research

AI is speeding up drug discovery. What happens next?

Most AI effort in drug discovery still sits before the clinic, where the measurable cost leverage looks smaller than at Phase II, and the field still grades its models on metrics that say little about real development decisions.
EssaySina Neda

Counts of AI-associated drug programs in clinical trials range from about 70 to more than 170, depending on how strictly a tracker defines "AI-discovered". None has been approved yet. Both sides of the AI debate quote these numbers, and the counts alone can't settle anything. They fit a pipeline maturing on a normal schedule, and they also fit activity running ahead of evidence.

A Perspective in Nature Reviews Drug Discovery, published on August 7, is the most useful attempt I've read to get past the counting. The authors, who include Andreas Bender, Jack Scannell and David Shaywitz, write that evidence of AI's clinically relevant impact is "disappointingly limited." They aren't criticizing the models themselves. Their argument is that the field tends to work on whatever is computationally tractable and grades itself with measures that rarely get tested against a real development decision.

This essay walks through the paper's evidence, where I think it leads, and what would prove it wrong.

A good model can still lead to a bad decision

The conflation that drives everything else in the paper is between validating a model and improving a decision.

A model can be evaluated on its own against a held-out dataset, producing an AUC or an error score. It can also be evaluated by whether a project made better calls because the model existed. Most published work does the first. The second is the one with economic meaning, and the two come apart more often than the publication record suggests.

The paper illustrates this with two models that have nearly the same AUC but perform very differently in use, because one is used for selection and the other for deselection. Selection means picking a handful of compounds to advance from many. What matters there is precision at the top of the ranking, and a missed compound further down costs little. Deselection means filtering compounds out on safety grounds, where catching the true positives matters most: one false negative can be very costly, while a false positive only costs you an option. In the authors' worked example, models with similar aggregate scores differ by about twofold in one selection setting and two- to threefold in a deselection setting. A single aggregate metric averages over the distinction that decides which model you should actually deploy.

The same problem shows up in what counts as success. A ligand binds a target in an assay. A drug also has to reach and engage its target in people, with acceptable efficacy and safety. The paper notes that databases such as ChEMBL and PubChem hold more than a million bioactive ligands, against roughly a thousand marketed drugs.

The biggest cost lever is in Phase II

The paper's cost model points to Phase II as the stage where improvements matter most.

This matches older work. In a 2010 model, Paul and colleagues found Phase II attrition to be one of the largest levers on the cost of each new drug. The same model also puts preclinical discovery at roughly a third of that cost, and the 2026 authors say plainly that preclinical costs and timelines matter to individual companies and to the economics of a drug. My point is narrower: the current concentration of AI effort doesn't match where the measured leverage is largest.

The biomarker result deserves more attention than it gets. It is an effect of about the size the field keeps promising, and it comes from knowing which patients to enroll. That work sits with trial enrichment, real-world evidence, multi-omics and diagnostics teams, a segment that, as far as I can tell, has drawn less capital and attention than molecule design.

Much of the benchmark evidence doesn't transfer

Even within preclinical work, the paper shows that model results travel less well than benchmark tables suggest.

Across four commonly used ADME datasets (absorption, distribution, metabolism and excretion), only 0.1% of compounds and 0.6% of chemical scaffolds overlap. A model trained on one therefore has little demonstrated applicability over the others. The datasets barely share a chemical domain, so strong performance on one endpoint tells you little about whether the model stays reliable across the rest. The authors warn that these mismatched domains make it hard to optimize several properties at once. They also see a risk of over-relying on model generalization.

The best case for upstream AI

The strongest argument for the current approach is option value, and I don't think the paper engages with it enough.

If preclinical work gets cheaper and faster, a company can run more programs on the same budget. More shots on goal can raise expected output even if the success rate per program never changes. On that reading, faster preclinical work is a real gain that simply doesn't show up in cost per launch, and judging AI by Phase II attrition asks it to do something it wasn't designed to do.

The option-value argument weakens when the added programs are highly correlated, and that seems to be the case here. According to the paper, many early AI clinical programs build on established disease biology and chemistry, which makes the extra programs more alike than "more shots on goal" implies. The same point affects the headline clinical number. Jayatunga and colleagues found that 21 AI-discovered molecules completed Phase I with a success rate of 80% to 90%. Only 10 had completed Phase II, and 4 of those succeeded. That's about 40%, in line with historical industry rates. Phase I tests tolerability more than whether the biological hypothesis holds, and a high Phase I success rate on well-studied targets is not, by itself, evidence of an AI contribution.

There is real counter-evidence to weigh, though. Insilico's rentosertib, which the company says was both targeted and designed with generative AI, completed a 71-patient Phase IIa trial that met its safety endpoint and showed an exploratory lung-function signal, and it entered Phase III in July 2026. Takeda's zasocitinib, a TYK2 inhibitor that Nimbus Therapeutics designed with Schrödinger's physics-based computational platform, is under priority FDA review, with a decision expected in the first quarter of 2027.

What this means for software in drug development

Swea doesn't invest in drug programs or in companies that own their own pipelines. Swea's thesis covers software that runs regulated work in labs, manufacturing and quality. Still, the paper points at a few needs that sit close to that focus.

  1. Tools that record which development decisions a model influenced and what happened afterward, so a company can measure decision quality over time.
  2. Data infrastructure that connects lab and preclinical predictions to later outcomes, which is the evidence the paper says is scarce.
  3. Software for patient selection and trial enrichment, where the paper's biomarker result suggests the economic leverage is largest.

This is where I think value could collect, and none of it is proven. My read is that methods in this field spread quickly, so the durable advantages tend to be proprietary data with demonstrated predictive value and the clinical operations needed to show that a model changed a decision. That favors incumbents with clinical infrastructure, and partnerships in which pharma supplies the data and the trials, which is roughly where the industry has landed.

The real test arrives in the clinic

The question I'd apply to any AI claim in drug development is which decision the model changes, and how you would know.

Applied to this field, it becomes one observable question. Around 2031, the several dozen AI-originated programs now in Phase I and Phase II should have produced enough readouts to support a rate instead of a handful of anecdotes. The comparison that matters then is whether those programs fail at Phase II materially less often than matched conventional programs on comparable targets. Until that comparison exists, the central claim remains unproven.

Insist on the matching. I expect unmatched comparisons to appear first, and they will probably flatter AI programs because of which targets they chose, independent of anything the AI did.

What would change my mind

I'll treat this view as wrong if a target-matched comparison shows AI-originated programs failing at Phase II materially less often than conventional programs on comparable targets. Matching matters because AI-first pipelines concentrate on established biology and chemistry, and an unadjusted comparison would flatter them for reasons unrelated to AI. A Phase III success for rentosertib, whose target came from AI, would also make me revisit how much weight I put on the target-selection argument.

How sure I am about each claim

ClaimStatusMain source
Improving Phase II success has the largest effect on capitalized cost per launchKnownBender et al.
Biomarker-stratified projects cost slightly more than half as much per launchKnownBender et al.
Most projects at AI-first companies are preclinicalKnownBender et al.
AI-discovered molecules succeed 80% to 90% in Phase I and about 40% in Phase II, on 21 and 10 moleculesKnown, small sampleJayatunga et al.
Four common ADME datasets share 0.1% of compounds and 0.6% of scaffoldsKnownBender et al.
Counts of AI programs in clinical trials range from about 70 to more than 170Reported, definitions varyTrade trackers
The sector is allocating against its own best evidenceMy readThis essay
AI-originated programs will not show materially lower Phase II failure once matchedOpen questionThis essay

Sources

Materials used for the analysis, grouped by the role they play in the evidence.

Primary research

Company materials

News and trade coverage

Content on this website is for informational purposes only and is not investment, legal or tax advice, or an offer to sell or a solicitation of an offer to buy any security, including any interest in a fund.

© Swea Ventures