[{"data":1,"prerenderedAt":405},["ShallowReactive",2],{"navigation":3,"research-navigation":16,"research:\u002Fresearch\u002Fai-drug-discovery-bottlenecks":25},[4,8,12],{"title":5,"path":6,"stem":7},"Thesis","\u002Fthesis","1.thesis",{"title":9,"path":10,"stem":11},"Research","\u002Fresearch","2.research",{"title":13,"path":14,"stem":15},"About","\u002Fabout","3.about",[17],{"title":9,"path":10,"stem":18,"children":19,"page":24},"research",[20],{"title":21,"path":22,"stem":23},"AI is speeding up drug discovery. What happens next?","\u002Fresearch\u002Fai-drug-discovery-bottlenecks","research\u002Fai-drug-discovery-bottlenecks",false,{"id":26,"title":21,"author":27,"body":28,"date":349,"description":350,"extension":351,"kind":352,"meta":353,"navigation":354,"path":22,"seo":355,"sitemap":357,"sources":358,"stem":23,"updatedAt":349,"__hash__":404},"research\u002Fresearch\u002Fai-drug-discovery-bottlenecks.md","Sina Neda",{"type":29,"value":30,"toc":337},"minimark",[31,39,42,57,60,65,73,76,79,82,88,92,95,104,113,116,121,125,128,131,136,140,143,146,155,169,174,178,184,197,200,204,207,210,213,222,226,229,233],[32,33,35],"evidence",{"type":34},"thesis",[36,37,38],"p",{},"My view is that AI in drug discovery is disproportionately optimizing the preclinical stages. The measurable economic leverage there looks smaller than at Phase II, where most programs fail. Most projects at AI-first companies are still preclinical. Early clinical data show AI-discovered molecules doing very well in Phase I, but the small Phase II sample so far hasn't shown a corresponding improvement. Until a matched comparison shows AI-originated programs failing less often in Phase II, I'd treat the productivity case for AI in drug discovery as unproven.",[36,40,41],{},"Counts of AI-associated drug programs in clinical trials range from about 70 to more than 170, depending on how strictly a tracker defines \"AI-discovered\". None has been approved yet. Both sides of the AI debate quote these numbers, and the counts alone can't settle anything. They fit a pipeline maturing on a normal schedule, and they also fit activity running ahead of evidence.",[36,43,44,45,56],{},"A ",[46,47,51,52],"a",{"href":48,"rel":49},"https:\u002F\u002Fdoi.org\u002F10.1038\u002Fs41573-026-01496-2",[50],"nofollow","Perspective in ",[53,54,55],"em",{},"Nature Reviews Drug Discovery",", published on August 7, is the most useful attempt I've read to get past the counting. The authors, who include Andreas Bender, Jack Scannell and David Shaywitz, write that evidence of AI's clinically relevant impact is \"disappointingly limited.\" They aren't criticizing the models themselves. Their argument is that the field tends to work on whatever is computationally tractable and grades itself with measures that rarely get tested against a real development decision.",[36,58,59],{},"This essay walks through the paper's evidence, where I think it leads, and what would prove it wrong.",[61,62,64],"h2",{"id":63},"a-good-model-can-still-lead-to-a-bad-decision","A good model can still lead to a bad decision",[36,66,67,68,72],{},"The conflation that drives everything else in the paper is between ",[69,70,71],"strong",{},"validating a model and improving a decision",".",[36,74,75],{},"A model can be evaluated on its own against a held-out dataset, producing an AUC or an error score. It can also be evaluated by whether a project made better calls because the model existed. Most published work does the first. The second is the one with economic meaning, and the two come apart more often than the publication record suggests.",[36,77,78],{},"The paper illustrates this with two models that have nearly the same AUC but perform very differently in use, because one is used for selection and the other for deselection. Selection means picking a handful of compounds to advance from many. What matters there is precision at the top of the ranking, and a missed compound further down costs little. Deselection means filtering compounds out on safety grounds, where catching the true positives matters most: one false negative can be very costly, while a false positive only costs you an option. In the authors' worked example, models with similar aggregate scores differ by about twofold in one selection setting and two- to threefold in a deselection setting. A single aggregate metric averages over the distinction that decides which model you should actually deploy.",[36,80,81],{},"The same problem shows up in what counts as success. A ligand binds a target in an assay. A drug also has to reach and engage its target in people, with acceptable efficacy and safety. The paper notes that databases such as ChEMBL and PubChem hold more than a million bioactive ligands, against roughly a thousand marketed drugs.",[32,83,85],{"type":84},"my-read",[36,86,87],{},"The ligand numbers are the part I'd put in front of anyone evaluating a molecule-design company. Those million ligands aren't interchangeable answers, because they are spread across targets, binding sites and selectivity profiles. Still, a generative model that reliably produces novel potent binders has solved an abundant intermediate problem while the scarce end problem, producing a drug, stays open. When I look at a model, I want to know which decision it is meant to change and how its builders would know it changed.",[61,89,91],{"id":90},"the-biggest-cost-lever-is-in-phase-ii","The biggest cost lever is in Phase II",[36,93,94],{},"The paper's cost model points to Phase II as the stage where improvements matter most.",[32,96,98],{"type":97},"known",[36,99,100,101,103],{},"In the authors' model of capitalized cost per successful launch, improving the Phase II success rate has the greatest effect of any stage. Projects that use biomarkers to pick the right patients have a capitalized cost per launch only slightly more than half that of projects without them. The large majority of projects at AI-first drug discovery companies are still preclinical, with several dozen in Phase I and Phase II and very few in Phase III. (Bender et al., ",[53,102,55],{},", August 7, 2026.)",[36,105,106,107,112],{},"This matches older work. In a 2010 model, ",[46,108,111],{"href":109,"rel":110},"https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fnrd3078",[50],"Paul and colleagues"," found Phase II attrition to be one of the largest levers on the cost of each new drug. The same model also puts preclinical discovery at roughly a third of that cost, and the 2026 authors say plainly that preclinical costs and timelines matter to individual companies and to the economics of a drug. My point is narrower: the current concentration of AI effort doesn't match where the measured leverage is largest.",[36,114,115],{},"The biomarker result deserves more attention than it gets. It is an effect of about the size the field keeps promising, and it comes from knowing which patients to enroll. That work sits with trial enrichment, real-world evidence, multi-omics and diagnostics teams, a segment that, as far as I can tell, has drawn less capital and attention than molecule design.",[32,117,118],{"type":84},[36,119,120],{},"The paper establishes the economics. The capital-allocation conclusion is my inference. If the highest-value application of machine learning here looks more like clinical data science than generative chemistry, the sector has been allocating against its own best evidence. I can see why individual companies focus upstream anyway. Preclinical work lets you iterate weekly, publish results and show a board progress, so the choice makes sense for one company even if it adds up to a poor allocation across the industry.",[61,122,124],{"id":123},"much-of-the-benchmark-evidence-doesnt-transfer","Much of the benchmark evidence doesn't transfer",[36,126,127],{},"Even within preclinical work, the paper shows that model results travel less well than benchmark tables suggest.",[36,129,130],{},"Across four commonly used ADME datasets (absorption, distribution, metabolism and excretion), only 0.1% of compounds and 0.6% of chemical scaffolds overlap. A model trained on one therefore has little demonstrated applicability over the others. The datasets barely share a chemical domain, so strong performance on one endpoint tells you little about whether the model stays reliable across the rest. The authors warn that these mismatched domains make it hard to optimize several properties at once. They also see a risk of over-relying on model generalization.",[32,132,133],{"type":84},[36,134,135],{},"For anyone assessing a platform, I think this is the figure to ask about. It means strong results on one property prediction don't carry over to multi-property optimization without evidence. I'd also ask of any published benchmark whether its external test set shares scaffolds, assays or literature sources with the training data, because each of those would inflate the reported performance.",[61,137,139],{"id":138},"the-best-case-for-upstream-ai","The best case for upstream AI",[36,141,142],{},"The strongest argument for the current approach is option value, and I don't think the paper engages with it enough.",[36,144,145],{},"If preclinical work gets cheaper and faster, a company can run more programs on the same budget. More shots on goal can raise expected output even if the success rate per program never changes. On that reading, faster preclinical work is a real gain that simply doesn't show up in cost per launch, and judging AI by Phase II attrition asks it to do something it wasn't designed to do.",[36,147,148,149,154],{},"The option-value argument weakens when the added programs are highly correlated, and that seems to be the case here. According to the paper, many early AI clinical programs build on established disease biology and chemistry, which makes the extra programs more alike than \"more shots on goal\" implies. The same point affects the headline clinical number. ",[46,150,153],{"href":151,"rel":152},"https:\u002F\u002Fwww.sciencedirect.com\u002Fscience\u002Farticle\u002Fpii\u002FS135964462400134X",[50],"Jayatunga and colleagues"," found that 21 AI-discovered molecules completed Phase I with a success rate of 80% to 90%. Only 10 had completed Phase II, and 4 of those succeeded. That's about 40%, in line with historical industry rates. Phase I tests tolerability more than whether the biological hypothesis holds, and a high Phase I success rate on well-studied targets is not, by itself, evidence of an AI contribution.",[36,156,157,158,163,164,72],{},"There is real counter-evidence to weigh, though. Insilico's rentosertib, which the company says was both targeted and designed with generative AI, completed a ",[46,159,162],{"href":160,"rel":161},"https:\u002F\u002Fwww.nature.com\u002Farticles\u002Fs41591-025-03743-2",[50],"71-patient Phase IIa trial"," that met its safety endpoint and showed an exploratory lung-function signal, and it entered Phase III in July 2026. Takeda's zasocitinib, a TYK2 inhibitor that Nimbus Therapeutics designed with Schrödinger's physics-based computational platform, is under priority FDA review, with a ",[46,165,168],{"href":166,"rel":167},"https:\u002F\u002Fwww.takeda.com\u002Fnewsroom\u002Fnewsreleases\u002F2026\u002Ffda-priority-review-zasocitinib-psoriasis\u002F",[50],"decision expected in the first quarter of 2027",[32,170,171],{"type":84},[36,172,173],{},"Both are worth watching, and neither settles the question in my view. Rentosertib is the more interesting case because the target itself came from AI. Still, a small Phase IIa with safety as its primary endpoint isn't evidence of better Phase II odds yet. Zasocitinib would be an important milestone for computational design if approved, but it works on a target that already had an approved drug, Bristol Myers Squibb's deucravacitinib, since 2022. That makes it a good test of design quality and a weaker test of whether AI improves the odds of picking the right biology.",[61,175,177],{"id":176},"what-this-means-for-software-in-drug-development","What this means for software in drug development",[36,179,180,181,183],{},"Swea doesn't invest in drug programs or in companies that own their own pipelines. Swea's ",[46,182,34],{"href":6}," covers software that runs regulated work in labs, manufacturing and quality. Still, the paper points at a few needs that sit close to that focus.",[185,186,187,191,194],"ol",{},[188,189,190],"li",{},"Tools that record which development decisions a model influenced and what happened afterward, so a company can measure decision quality over time.",[188,192,193],{},"Data infrastructure that connects lab and preclinical predictions to later outcomes, which is the evidence the paper says is scarce.",[188,195,196],{},"Software for patient selection and trial enrichment, where the paper's biomarker result suggests the economic leverage is largest.",[36,198,199],{},"This is where I think value could collect, and none of it is proven. My read is that methods in this field spread quickly, so the durable advantages tend to be proprietary data with demonstrated predictive value and the clinical operations needed to show that a model changed a decision. That favors incumbents with clinical infrastructure, and partnerships in which pharma supplies the data and the trials, which is roughly where the industry has landed.",[61,201,203],{"id":202},"the-real-test-arrives-in-the-clinic","The real test arrives in the clinic",[36,205,206],{},"The question I'd apply to any AI claim in drug development is which decision the model changes, and how you would know.",[36,208,209],{},"Applied to this field, it becomes one observable question. Around 2031, the several dozen AI-originated programs now in Phase I and Phase II should have produced enough readouts to support a rate instead of a handful of anecdotes. The comparison that matters then is whether those programs fail at Phase II materially less often than matched conventional programs on comparable targets. Until that comparison exists, the central claim remains unproven.",[36,211,212],{},"Insist on the matching. I expect unmatched comparisons to appear first, and they will probably flatter AI programs because of which targets they chose, independent of anything the AI did.",[32,214,216,219],{"type":215},"open-question",[36,217,218],{},"The paper offers no timeline and declines to predict when results will arrive, so the 2031 horizon is my inference from how long clinical development takes.",[36,220,221],{},"Two assumptions could undermine the test itself. Attribution may be messy. If AI helped with lead optimization on a target chosen conventionally, it's unclear how the program should be counted, and companies have an incentive to count generously. And a properly target-matched control set may not exist in large enough numbers, given how closely AI-first pipelines cluster on well-established targets. If the comparison can't be made cleanly, the thesis stays untested, and the debate runs another cycle on the same evidence.",[61,223,225],{"id":224},"what-would-change-my-mind","What would change my mind",[36,227,228],{},"I'll treat this view as wrong if a target-matched comparison shows AI-originated programs failing at Phase II materially less often than conventional programs on comparable targets. Matching matters because AI-first pipelines concentrate on established biology and chemistry, and an unadjusted comparison would flatter them for reasons unrelated to AI. A Phase III success for rentosertib, whose target came from AI, would also make me revisit how much weight I put on the target-selection argument.",[61,230,232],{"id":231},"how-sure-i-am-about-each-claim","How sure I am about each claim",[234,235,236,252],"table",{},[237,238,239],"thead",{},[240,241,242,246,249],"tr",{},[243,244,245],"th",{},"Claim",[243,247,248],{},"Status",[243,250,251],{},"Main source",[253,254,255,267,276,285,296,305,316,327],"tbody",{},[240,256,257,261,264],{},[258,259,260],"td",{},"Improving Phase II success has the largest effect on capitalized cost per launch",[258,262,263],{},"Known",[258,265,266],{},"Bender et al.",[240,268,269,272,274],{},[258,270,271],{},"Biomarker-stratified projects cost slightly more than half as much per launch",[258,273,263],{},[258,275,266],{},[240,277,278,281,283],{},[258,279,280],{},"Most projects at AI-first companies are preclinical",[258,282,263],{},[258,284,266],{},[240,286,287,290,293],{},[258,288,289],{},"AI-discovered molecules succeed 80% to 90% in Phase I and about 40% in Phase II, on 21 and 10 molecules",[258,291,292],{},"Known, small sample",[258,294,295],{},"Jayatunga et al.",[240,297,298,301,303],{},[258,299,300],{},"Four common ADME datasets share 0.1% of compounds and 0.6% of scaffolds",[258,302,263],{},[258,304,266],{},[240,306,307,310,313],{},[258,308,309],{},"Counts of AI programs in clinical trials range from about 70 to more than 170",[258,311,312],{},"Reported, definitions vary",[258,314,315],{},"Trade trackers",[240,317,318,321,324],{},[258,319,320],{},"The sector is allocating against its own best evidence",[258,322,323],{},"My read",[258,325,326],{},"This essay",[240,328,329,332,335],{},[258,330,331],{},"AI-originated programs will not show materially lower Phase II failure once matched",[258,333,334],{},"Open question",[258,336,326],{},{"title":338,"searchDepth":339,"depth":339,"links":340},"",2,[341,342,343,344,345,346,347,348],{"id":63,"depth":339,"text":64},{"id":90,"depth":339,"text":91},{"id":123,"depth":339,"text":124},{"id":138,"depth":339,"text":139},{"id":176,"depth":339,"text":177},{"id":202,"depth":339,"text":203},{"id":224,"depth":339,"text":225},{"id":231,"depth":339,"text":232},"2026-10-02","Most AI effort in drug discovery still sits before the clinic, where the measurable cost leverage looks smaller than at Phase II, and the field still grades its models on metrics that say little about real development decisions.","md","Essay",{},true,{"title":21,"description":356},"Why AI drug discovery needs to show it improves clinical decisions and Phase II outcomes, not only preclinical speed and benchmark scores.",{"loc":22,"lastmod":349},[359,365,369,373,376,381,384,388,393,400],{"label":360,"url":48,"type":361,"role":362,"establishes":363,"published":364},"Bender et al., Artificial intelligence in drug discovery: what it is, where we stand and the path forward, Nature Reviews Drug Discovery (2026)","paper","primary","The Perspective's analysis of clinically relevant AI impact, decision-level model evaluation, development-stage economics, pipeline maturity and ADME dataset overlap.","2026-08-07",{"label":366,"url":151,"type":361,"role":362,"establishes":367,"accessed":368},"Jayatunga et al., How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons, Drug Discovery Today (2024)","Phase I success of 80% to 90% across 21 AI-discovered molecules and about 40% in Phase II across 10 molecules, in line with historical industry rates.","2026-09-30",{"label":370,"url":109,"type":361,"role":362,"establishes":371,"published":372},"Paul et al., How to improve R&D productivity: the pharmaceutical industry's grand challenge, Nature Reviews Drug Discovery (2010)","A model of cost per launch in which Phase II attrition is a major lever and preclinical discovery accounts for roughly a third of the cost of each new drug.","2010-02-19",{"label":374,"url":160,"type":361,"role":362,"establishes":375,"accessed":368},"Insilico Medicine and collaborators, A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial, Nature Medicine (2025)","A 71-patient, 12-week Phase IIa trial of rentosertib with safety as the primary endpoint and an exploratory lung-function signal at the highest dose.",{"label":377,"url":378,"type":379,"role":362,"establishes":380,"accessed":368},"Insilico Medicine, Insilico initiates Phase III clinical trial for rentosertib (2026)","https:\u002F\u002Finsilico.com\u002Fnews\u002Fxmjsn4l091-insilico-initiates-phase-iii-clinical-tr","company","The start of a Phase III trial of rentosertib in idiopathic pulmonary fibrosis in July 2026.",{"label":382,"url":166,"type":379,"role":362,"establishes":383,"accessed":368},"Takeda, U.S. FDA accepts New Drug Application under Priority Review for zasocitinib in moderate-to-severe plaque psoriasis (2026)","The FDA's acceptance of the zasocitinib application under priority review, with a decision expected in the first quarter of 2027.",{"label":385,"url":386,"type":379,"role":362,"establishes":387,"accessed":368},"Schrödinger, Case study: design of a highly selective, allosteric, picomolar TYK2 inhibitor","https:\u002F\u002Fwww.schrodinger.com\u002Fwp-content\u002Fuploads\u002F2024\u002F07\u002F23_530_TYK2-Case-Study_Mkt_R7-4_Digital.pdf","Nimbus Therapeutics' use of Schrödinger's computational, physics-based design platform in developing the TYK2 inhibitor now known as zasocitinib.",{"label":389,"url":390,"type":379,"role":362,"establishes":391,"published":392},"Bristol Myers Squibb, FDA approves Sotyktu (deucravacitinib) for moderate-to-severe plaque psoriasis (2022)","https:\u002F\u002Fnews.bms.com\u002Fnews\u002Fdetails\u002F2022\u002FU.S.-Food-and-Drug-Administration-Approves-Sotyktu-deucravacitinib-Oral-Treatment-for-Adults-with-Moderate-to-Severe-Plaque-Psoriasis\u002Fdefault.aspx","The first approval of a TYK2 inhibitor, in September 2022.","2022-09-09",{"label":394,"url":395,"type":396,"role":397,"establishes":398,"accessed":399},"Matterfact, Isomorphic Labs raises $2.1B; 173 AI drug programmes in trials, none approved","https:\u002F\u002Fwww.matterfact.com\u002Fnewsletter\u002F2026-06-04-ai-drug-discovery","news","secondary","The upper end of published counts of AI-associated programmes in clinical trials; it does not independently establish the underlying count.","2026-09-17",{"label":401,"url":402,"type":396,"role":397,"establishes":403,"accessed":368},"IntuitionLabs, AI-discovered drugs in clinical trials, 2026","https:\u002F\u002Fintuitionlabs.ai\u002Farticles\u002Fai-discovered-drugs-clinical-trials-2026","How counts of AI-designed drugs in human trials vary with the definition used, from about 70 to more than 170.","Ir-LlRdO1rD-uJL77VbtrVpJGmb_p4kOwDzxR-i6YkE",1791476336131]