This work presents a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail, and connects evaluation choices to real-world applications and case studies encountered in pharmaceutical research.
Abstract
Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model’s performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework on a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation) showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.
Multitask learning is a promising strategy in computational drug discovery, potentially improving predictive performance and generalization over traditional single-task models. MTL has shown particular value in absorption, distribution, metabolism, elimination, and toxicity (ADMET) and potency predictions, which are key for drug design. Yet, many existing Web servers rely on the same uncurated, decade-old data sets, creating an illusion of diversity. This work critically reviews open-source ADMET Web services, revealing extensive data redundancy and limited curation across the field. We introduce OneADMET, a meticulously curated data set of 738,161 compounds with 1,119,719 measurements spanning 44 ADMET end points and 1 489 biological activities. We report a unified ChemProp-based MTL model capable of handling hundreds of continuous tasks simultaneously, which has practical advantages for model deployment and maintenance. Additionally, we observed that these MTL models match or surpass single-task models in predictive accuracy. This study highlights the utility of large-scale MTL for pharmacokinetics profiling and contributes practical tools and data sets for the community.
P. Llompart, C. Minoletti, G. Marcou et al.· Journal of Medicinal Chemist...· 0 citations
Drug discovery is frequently limited by high attrition rates, and poor absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles are a major cause of late-stage failure. Therefore, precise ADMET property prediction is necessary to develop safe and effective drug candidates. Traditional experimental assays and rule-based computational procedures are limited by their poor predictive power, cost, and time, despite providing valuable insights. Innovative strategies to deal with these issues have been introduced by developments in artificial intelligence (AI), such as machine learning (ML), deep learning (DL), graph neural networks (GNNs), generative models, and multi-task learning (MTL). AI techniques can better generalize scaffolds, capture interdependencies between pharmacokinetic and toxicological endpoints, and model complex nonlinear relationships by leveraging large, diverse datasets. Explainable AI (XAI) enhances transparency by detecting biological and structural characteristics that are relevant to predictions, even if integrated pipelines combine predictive modeling with molecular creation and optimization. AI-driven ADMET prediction is becoming a vital tool in lowering attrition, speeding up candidate prioritization, and influencing the direction of rational drug development, despite persistent issues with data quality, regulatory acceptance, and synthetic viability.
Satyam Kumar Vishwash, Ram Babu Soni, Ratima Sood et al.· Current Computer - Aided Dru...· 0 citations
This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models and proposes leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches.
Lucille Tomin, Vida Bodaghi-Namileh, D. Schwartz et al.· Journal of Chemical Informat...· 0 citations
This work addresses one of the most pervasive obstacles to applying AI in real-world drug development by addressing conformal prediction framework tailored to label shift by weighting conformal scores using marginal label probability ratios and enhancing the trustworthiness of AI-driven predictions.
Hyeonsu Lee, Juyeong Kim, Erkhembayar Jadamba et al.· 0 citations
How machine learning, deep learning, natural language processing, and related computational methods are being applied across the drug discovery process is reviewed, with particular attention to AlphaFold-based protein structure prediction, AI-supported virtual screening, generative chemistry, retrosynthetic planning, digital pathology, and the use of real-world clinical data.
Yue Peng· International Journal of Bio...· 0 citations
Artificial intelligence is revolutionizing drug discovery by accelerating target identification, molecular design, virtual screening, and toxicity prediction, while tackling longstanding challenges like high costs and lengthy timelines in traditional pipelines. This review explores recent AI innovations—such as AlphaFold for protein structure prediction, generative models for de novo drug design, and graph neural networks for drug repurposing—alongside real-world case studies from companies like Exscientia, Insilico Medicine, and BenevolentAI, which have produced clinical candidates like DSP-1181 and rentosertib. Despite these advances, key hurdles persist, including data quality issues, model interpretability, synthetic feasibility for complex molecules, and integration with experimental workflows, underscoring the need for explainable AI, better datasets, and ethical frameworks to bridge research gaps. Looking ahead, hybrid AI-experimental approaches and collaborations between pharma giants and AI startups promise to deliver safer, more personalized therapies faster.
P. Jadhav, R. Pingale, Kanchan Gajanan Gawai et al.· International journal for ad...· 0 citations