Skip to content
Open access

Deep Learning Pipeline for Accelerating Virtual Screening in Drug Discovery.

2026 · Methods in molecular biology · Vol 3021, pp. 337-368 · 0 citations
Medicine

TL;DR

This chapter explores the systematic deep learning pipeline for virtual screening, covering data acquisition, molecular representation learning, model architectures, training strategies, and uncertainty estimation, and discusses advanced model optimization techniques, including curriculum learning, transfer learning, and adversarial training, which enhance predictive accuracy and robustness.

Abstract

Deep learning has revolutionized virtual screening in drug discovery, offering unprecedented improvements in hit identification, lead optimization, and molecular property prediction. Traditional virtual screening methods, including structure-based docking and ligand-based screening, often suffer from computational inefficiencies, poor generalization, and reliance on predefined molecular descriptors. Deep learning addresses these limitations by leveraging graph neural networks (GNNs), transformer-based models, generative AI, and reinforcement learning to discover novel drug candidates more efficiently. This chapter explores the systematic deep learning pipeline for virtual screening, covering data acquisition, molecular representation learning, model architectures, training strategies, and uncertainty estimation. We discuss advanced model optimization techniques, including curriculum learning, transfer learning, and adversarial training, which enhance predictive accuracy and robustness. Despite significant progress, challenges remain, particularly in data quality, model interpretability, generalization to novel chemical spaces, and computational cost. Emerging trends, such as self-supervised learning, quantum computing for molecular simulations, and AI-integrated automated laboratories, are paving the way for the next generation of AI-driven drug discovery. By integrating machine learning with experimental validation, AI-powered virtual screening is set to accelerate early-stage drug discovery, reducing costs and improving the efficiency of therapeutic development.

Read PDF

Similar papers

Review Open access Aug 2026

Geometric Deep Learning‐Based Drug Design Models for Small‐Molecule Drug Discovery

Deep neural network (DNN)‐based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure‐based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non‐Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein–ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small‐molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)‐equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry‐aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit‐to‐lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high‐quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi‐resolution learning, and self‐supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI‐enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next‐generation therapeutics.

A. Srivastav, Unnati Modi, Rahul Kumar et al. · 0 citations
Review Aug 2026

Advancing cancer drug discovery through the integration of machine learning and high-throughput screening.

This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.

K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al. · 0 citations
Review Open access 2025

AI-Assisted Drug Discovery: Emerging Technologies and Challenges

This study reviews AI-driven drug discovery methods, presents a structured AI pipeline from data collection to candidate selection, and evaluates performance using metrics such as prediction accuracy, screening efficiency, lead optimization success, and toxicity reduction.

Joseph Robin · 0 citations
Review Aug 2026

Recent Advances in Deep Learning-Based Drug-Target Binding Affinity Prediction

It is indicated that although many methods report strong performance on standard benchmarks, their effectiveness is often influenced by dataset bias and limited evaluation settings, and most methods exhibit reduced performance in cold-start scenarios, highlighting challenges in generalization.

Jafin Khan, Md Hossain Shuvo · 0 citations
Open access Jul 2026

Smiles-based bioactivity prediction through molecular encoder selection and data augmentation.

Quantitative prediction of inhibitor potency can accelerate early-stage drug discovery. Recently, data-driven approaches have gained widespread interest in drug discovery, as evidenced by a growing number of benchmarking challenges and open competitions. In this context, we developed a machine learning-based methodology that can find the most effective way of predicting IC50 values against ASK1 from SMILES, for "Jump AI(.py) 2025: 3rd AI Drug Discovery Competition", hosted by the Korea Pharmaceutical and Bio-Pharma Manufacturers Association (KPBMA) on the Dacon platform. Applying our methodology achieved the highest overall predictive performance among all participating teams. Beyond this competition setting, we present a compact SMILES-based modeling workflow comprising (i) a pre-trained encoder, (ii) regression models, (iii) data augmentation, and (iv) hyperparameter tuning. We systematically compared molecular representations from sequence- and graph-based models, including ChemBERTa-2 and MolCLR. Across encoder-regressor combinations, ChemBERTa-77 M-MLM embeddings paired with support vector regression (SVR) yielded the strongest predictive performance. Embedding-level mix-up augmentation and SVR hyperparameter tuning further improved predictive performance. Our findings highlight that careful SMILES preprocessing and encoder selection have a critical influence on IC50 values and provide a reproducible benchmark for single-target bioactivity prediction, thus contributing to a more efficient drug discovery process. Scientific Contribution In this study, we propose a machine learning methodology for predicting the IC50 values of ASK1 inhibitors from SMILES representations, with a systematic comparison of molecular encoders and regression models. Our results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation. This strategy would be particularly useful for settings with imbalanced labels or limited data, and could be applied more broadly to IC50 prediction for other kinase inhibitors.

Ju Hyung Lee, S. Choi, Utku Ozbulak et al. · 0 citations