Aug 2026· Journal of Chemical Information and Modeling· Vol 66 16, pp.
9761-9783
· 0 citations· 99 references
Medicine
TL;DR
This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models and proposes leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches.
Abstract
Machine learning applications in preclinical drug development have been focused on automated covariate selection in pharmacometric modeling and high-throughput screening processes early in drug discovery. While inherent drug property prediction has made significant improvements in the past decade, fusing early target-based drug discovery methods to preclinical stage pharmacokinetic (PK) property predictions has been limited. This scoping review investigates the current state of PK property prediction of small molecules in drug discovery using machine learning methods and a combination of machine learning and mechanistic models. We identified major obstacles hindering the development of superior prediction models for small molecule behavior in biological systems. These encompass data accessibility, quantity, and quality, architectural constraints such as poor interpretability and model inherent assumptions, and the lack of robust evaluation and uncertainty assessment methods. To mitigate data-related constraints, we advocate for the use of collaborative federated learning frameworks. Furthermore, we propose leveraging the pattern recognition capabilities of deep learning models in conjunction with the biological interpretability provided by mechanistic approaches to strike an optimal balance between accuracy and biological explainability guided by the intended application of the prediction model. Addressing these limitations will advance reliable modeling pipelines and enable effective extrapolation to novel chemical space, additional species, and emerging drug development scenarios.
Drug discovery is frequently limited by high attrition rates, and poor absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles are a major cause of late-stage failure. Therefore, precise ADMET property prediction is necessary to develop safe and effective drug candidates. Traditional experimental assays and rule-based computational procedures are limited by their poor predictive power, cost, and time, despite providing valuable insights. Innovative strategies to deal with these issues have been introduced by developments in artificial intelligence (AI), such as machine learning (ML), deep learning (DL), graph neural networks (GNNs), generative models, and multi-task learning (MTL). AI techniques can better generalize scaffolds, capture interdependencies between pharmacokinetic and toxicological endpoints, and model complex nonlinear relationships by leveraging large, diverse datasets. Explainable AI (XAI) enhances transparency by detecting biological and structural characteristics that are relevant to predictions, even if integrated pipelines combine predictive modeling with molecular creation and optimization. AI-driven ADMET prediction is becoming a vital tool in lowering attrition, speeding up candidate prioritization, and influencing the direction of rational drug development, despite persistent issues with data quality, regulatory acceptance, and synthetic viability.
Satyam Kumar Vishwash, Ram Babu Soni, Ratima Sood et al.· Current Computer - Aided Dru...· 0 citations
In the early stages of drug discovery, predicting drug-target affinity is a crucial task. Due to the vast scale of genomic and chemical spaces, traditional biological methods are time-consuming, labor-intensive, and resource-demanding. As a result, machine learning-based computational methods have emerged to narrow down the pool of drug candidates. However, machine learning approaches still face several challenges in practical applications, particularly the scarcity of labeled samples and poor model generalization capability. To address these issues, this paper proposes a novel drug-target affinity prediction model, termed MetaBayes-DTA, based on an uncertainty-aware meta-learning framework. The model integrates the few-shot rapid adaptation capability of meta-learning with an uncertainty quantification mechanism to enhance prediction accuracy and reliability. MetaBayes-DTA is evaluated on two benchmark datasets, DAVIS and KIBA. Experimental results demonstrate that the proposed model outperforms existing methods.
Naihan Shi, Yanpeng Zhao, Wanying Li et al.· 2026 IEEE 27th China Confere...· 0 citations
This work presents a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail, and connects evaluation choices to real-world applications and case studies encountered in pharmaceutical research.
Srijit Seal, Akshat Shirish Zalte, David Alencar Araripe et al.· bioRxiv· 0 citations
Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.
Ju Hyung Lee, S. Choi, Utku Ozbulak et al.· Journal of Cheminformatics· 0 citations
This work addresses one of the most pervasive obstacles to applying AI in real-world drug development by addressing conformal prediction framework tailored to label shift by weighting conformal scores using marginal label probability ratios and enhancing the trustworthiness of AI-driven predictions.
Hyeonsu Lee, Juyeong Kim, Erkhembayar Jadamba et al.· 0 citations
How machine learning, deep learning, natural language processing, and related computational methods are being applied across the drug discovery process is reviewed, with particular attention to AlphaFold-based protein structure prediction, AI-supported virtual screening, generative chemistry, retrosynthetic planning, digital pathology, and the use of real-world clinical data.
Yue Peng· International Journal of Bio...· 0 citations