Skip to content

Modeling Structure-Activity Relationships with Machine Learning to Identify DPP4 Inhibitors as potential Therapeutics for Type 2 Diabetes

Jul 2026 · Journal of Computational Biophysics and Chemistry · pp. 1-33 · 0 citations

TL;DR

Findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery.

Abstract

The development of potent dipeptidyl peptidase-4 (DPP4) inhibitors remains a promising therapeutic strategy for the management of type 2 diabetes mellitus (T2DM). In the present study, an integrated computational workflow incorporating machine learning-based quantitative structure-activity relationship (QSAR) modeling, ligand-based virtual screening, molecular docking, molecular dynamics (MD) simulations, and binding free energy calculations was employed to identify novel DPP4 inhibitors. A curated dataset of experimentally validated DPP4 inhibitors was obtained from the ChEMBL database and subjected to systematic preprocessing and molecular descriptor generation. Several machine learning regression algorithms were initially evaluated to identify the most suitable predictive models. The best-performing tree-based algorithms were subsequently optimized and combined using Ridge Stacking and Weighted Average ensemble strategies. Among the developed models, the optimized Ridge Stacking ensemble demonstrated the highest predictive performance, achieving an R 2 of 0.746, an RMSE of 0.819, and a Pearson correlation coefficient of 0.864, indicating strong predictive accuracy and good generalization capability. The robustness of the model was further confirmed through 10-fold cross-validation, bootstrap validation, residual analysis, and applicability domain assessment. The validated ensemble model was then used to screen 95 compounds identified through ligand-based virtual screening. Among these candidates, CP20 exhibited the highest predicted pIC 50 value and was selected for further evaluation together with the reference inhibitor omarigliptin. Molecular docking, structural interaction fingerprinting, molecular dynamics simulations, and MM/GBSA and MM/PBSA binding free energy analyses demonstrated that CP20 formed stable interactions with key catalytic residues of DPP4 and maintained favorable conformational stability throughout the simulation. Collectively, these findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery. Experimental validation is warranted to confirm its biological activity and therapeutic potential.

View source

Similar papers

Open access Jul 2026

Machine learning–driven identification of PIM2 kinase inhibitors through QSAR modeling and molecular dynamics simulations

The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.

A. Fahira, M. Shahab, Zaheer Ud Din et al. · 0 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations
Open access Jul 2026

Machine Learning–Driven Discovery of Dietary Polyphenol DPP-4 Inhibitors via Molecular Docking

T2DM is a chronic metabolic disorder of rising global prevalence, in which DPP-4 serves as a key therapeutic target through its role in incretin hormone degradation. This study aimed to identify polyphenolic compounds derived from dietary sources as potential DPP-4 inhibitors through an in silico pipeline integrating machine learning (ML) and molecular docking. The bioactivity dataset for DPP-4 was retrieved from ChEMBL (CHEMBL284), processed into a binary classification dataset, and represented using 2048-bit Morgan fingerprints combined with five RDKit descriptors. Six ML models were integrated into a soft-voting ensemble, achieving a Matthews correlation coefficient (MCC) of 0.8334 and an AUC-ROC of 0.9739. Screening of 162 compounds from the Phenol-Explorer database yielded 11 potential active inhibitors. Molecular docking identified hesperetin (−8.548 kcal/mol) as the leading candidate, followed by pelargonidin, daidzein, and naringenin, with key binding residues including Ser209, Glu205/206, Tyr631, Arg125, and Tyr662.

Nur Laily Harfita, Ahmad Faisal Nasution, Zuliana Amalia et al. · 0 citations
#graph neural networks Open access Aug 2026

Discovery of a potent TDP1 inhibitor through machine learning-driven predictive modeling combined with structure-based virtual screening and experimental validation

An integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation identifies AO65 as a promising lead for further TDP1-focused investigation.

Huang Zeng, Manyi Zhang, Bo Qiu et al. · 0 citations
Open access Aug 2026

In Silico Identification of Novel Leads as Potential DPP-IV Inhibitors as Antidiabetic Agents Using Virtual Screening

Purpose: Dipeptidyl peptidase-IV (DPP-IV) is a validated therapeutic target for type 2 diabetes mellitus due to its role in incretin hormone degradation. This study aimed to identify novel small-molecule DPP-IV inhibitors from the ECBD database using an integrated virtual screening and molecular dynamics (MD) approach, acknowledging that experimental validation is necessary to confirm biological activity. Methods: Structure-based virtual screening of 5,500 ECBD compounds was performed using molecular docking to identify high-affinity ligands, followed by interaction analysis with key catalytic residues. Pharmacokinetic suitability was evaluated through in-silico ADMETox profiling. The top-ranked hits were further subjected to 100 ns MD simulations to assess complex stability and conformational effects on the DPP-IV active site. Binding free energies were calculated using the MM/GBSA method, and docking reliability was validated through redocking experiments. Results and Discussion: Five compounds EOS34295, EOS4915, EOS9480, EOS2567, and EOS62725 exhibited stronger noncovalent binding affinities than vildagliptin and formed stable interactions with essential residues, including GLU205, GLU206, ASN710, and PHE357. ADMETox predictions indicated favorable oral drug-likeness. MD simulations revealed stable protein–ligand complexes, with lower RMSD values for the hit ligands (1.20–1.51 Å) compared with vildagliptin (1.90 Å). RMSF analysis showed consistent flexibility without destabilizing fluctuations. Additional hydrogen-bond interactions with PHE357 emerged during MD simulations, indicating enhanced binding persistence. MM/GBSA analysis confirmed stronger binding energies for EOS34295 (−56.58 kcal/mol), EOS62725 (−37.90 kcal/mol), and EOS2567 (−31.27 kcal/mol) relative to vildagliptin (−23.88 kcal/mol). Conclusion: This study identifies EOS34295, EOS2567, and EOS62725 emerged as promising DPP-IV inhibitory hits with superior binding stability, supporting their potential as antidiabetic leads for future experimental validation.

M. Turkar, Rahul Ahirwar, R. Sahu et al. · 0 citations