Skip to content

An integrated deep learning and molecular dynamics pipeline for the discovery of novel CDK8 inhibitors against acute myeloid leukaemia.

Aug 2026 · SAR and QSAR in environmental research (Print) · pp. 1-35 · 0 citations · 48 references
Medicine

TL;DR

An integrated computational pipeline combining neural network-based potency prediction with molecular dynamics simulations for CDK8 inhibitor discovery was developed and novel molecular structures beyond the training distribution were generated.

Abstract

Cyclin-dependent kinase 8 (CDK8) has emerged as a promising therapeutic target for acute myeloid leukaemia (AML). We developed an integrated computational pipeline combining neural network-based potency prediction with molecular dynamics (MD) simulations for CDK8 inhibitor discovery. A curated dataset of 1,200 unique CDK8 inhibitors was assembled from ChEMBL. The optimal neural network architecture (two hidden layers, 512→128 units) with dropout regularization achieved a test set r2 of 0.47 and RMSE of 0.87. Y-randomization testing confirmed genuine structure-activity relationships. Using an extrapolation-focused training strategy combined with a SELFIES-based genetic recombination algorithm, we generated novel molecular structures beyond the training distribution. SHAP analysis revealed fingerprint bits as critical determinants of potency. Applicability‑domain analysis confirmed that the novel hit falls within validated chemical space. The identified candidate exhibited a predicted pIC50 of 11.97, substantially exceeding the most potent training compound (pIC50 = 10.09) and compound 12 (pIC50 = 7.47). Molecular docking revealed Moldock scores of -143.9 kcal/mol for the novel hit versus -113.9 kcal/mol for compound 12. MD simulations demonstrated stable binding with the novel hit forming a highly stable hydrogen bond with Asp98. MM-GBSA calculations showed superior binding free energy for the novel hit (-99.10 vs. -47.77 kcal/mol).

View source

Similar papers

#graph neural networks Open access Aug 2026

Discovery of a potent TDP1 inhibitor through machine learning-driven predictive modeling combined with structure-based virtual screening and experimental validation

An integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation identifies AO65 as a promising lead for further TDP1-focused investigation.

Huang Zeng, Manyi Zhang, Bo Qiu et al. · 0 citations
Aug 2026

Deep learning-assisted virtual screening of a large chemical library for selective GSK3β inhibitors.

GSK3BMTPred, a multitask deep neural network model, was developed for simultaneous prediction of inhibitor classification and inhibitory potency and identified compounds showing stable interactions with key Adenosine Triphosphate (ATP) residues and favorable predicted absorption, distribution, metabolism, excretion, and toxicity properties.

Tanmaykumar Varma, Pradnya Kamble, Prabha Garg · 0 citations
Open access Jul 2026

Explainable Machine Learning-Guided Virtual Screening and Molecular Docking for Identification of Novel FYN Kinase Inhibitors

FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.

Ahmet Turan Demir · 4 citations
Open access Aug 2026

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Jurica Novak · 0 citations