Aug 2026· Pharmaceuticals· Vol 19, pp. 1209· 0 citations· 62 references
Medicine
TL;DR
This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.
Abstract
Background: Anaplastic Lymphoma Kinase (ALK) is an oncogenic receptor tyrosine kinase implicated in several cancers. Despite the clinical success of ALK inhibitors, acquired resistance continues to drive the search for novel chemotypes. We developed a multiclass machine learning framework to classify ALK inhibitory activity using a curated ChEMBL dataset. Methods: Models were built using 2D molecular descriptors together with MACCS and ECFP4 fingerprints. Three widely used algorithms, Support Vector Machine (SVM), Random Forest (RF), and XGBoost, were applied for model development. Results: RF and XGBoost models demonstrated the best performance, achieving accuracies of ~0.75–0.79 with consistently high ROC–AUC values, particularly for fingerprint-based features. Bemis–Murcko scaffold analysis identified enriched chemotypes and underexplored scaffolds for further prioritization. The validated models were subsequently used to screen the Maybridge library, and compounds predicted to possess potential ALK inhibitory activity were prioritized for further computational evaluation. Applicability-domain filtering confirmed that the selected compounds occupied the predicted ALK inhibitor chemical space across multiple activity classes. The shortlisted compounds were subsequently evaluated by molecular docking to characterize their binding modes and interactions. Three candidate hits (SCR00078, SCR00073, and AW01085) were selected for further evaluation using 500 ns molecular dynamics simulations alongside the reference inhibitor Brigatinib. Simulation analyses revealed stable protein–ligand complexes and reduced conformational fluctuations relative to apo ALK, while MM/PBSA calculations identified SCR00078 and AW01085 as the most favorable binders. Conclusions: This integrated ML-to-simulation workflow prioritizes structurally novel candidate hits with predicted ALK inhibitory activity and provides an effective strategy for scaffold discovery and hit prioritization.
The natural product compounds CNP0456830 and CNP0467494 exhibited the lowest binding free energies for both EGFR and PIK3CA, identifying them as the most promising dual-target inhibitors.
Si-miao Lu, Yi Zhu, Yong-tao Han et al.· Diseases of the esophagus· 0 citations
Anaplastic Lymphoma Kinase (ALK), a tyrosine receptor kinase is of immense importance in non small cell lung cancer (NSCLC). Therefore, there is a need to design novel derivatives with the intention to overcome the limitations of resistance and debilitating side effects associated with current FDA-approved ALK inhibitors. This work therefore employed artificial intelligence and computational techniques to screen a library of ALK tyrosine kinase inhibitors, retrieved from ChEMBL database, with their corresponding IC50 in nM. The compounds’ descriptors were obtained using the Padel descriptors software and screened to reduce dimensionality and remove redundancy and multicollinearity. The compounds’ descriptors and their corresponding IC50 in nM were imported to Google Colab workspace with the necessary Python packages for machine learning (ML) models building. The best model was used to predict the bioactivity of new derivatives of 5-FDA approved drugs and TPX-1301, taking into account their drug-likeness properties for initials screening. The binding affinities and modes of lead compounds at the binding domain of ALK tyrosine kinase receptors were predicted using molecular docking while the binding free energies were obtained using MM-GB/SA calculations. Among all the trained models, the artificial neural network showed the most promising results, with a coefficient of determination (R2) value of 0.84, a root mean square error (RMSE) value of 0.27, a mean squared error (MSE) value of 0.22 and a mean absolute error (MAE) of 0.26. on the training data and an R2 value of 0.62, RMSE of 0.73, MSE of 0.53 and MAE of 0.55 for the test data, an indication of its reliability in making prediction. Cross-docking of cognate lorlatinib against ALK (4CLI) model yielded a binding mode closely aligned with the native conformer of lorlatinib, exhibiting an RMSD of 0.13 Å. Additionally, Induced Fit Docking (IFD) and Prime MM-GB/SA calculations indicated that briga_15 possesses the highest IFD Score of -675.27 kcal/mol, MM-GB/SA binding energy of -61.60 kcal/mol and a predicted pIC50 value of 8.57. The binding of briga_15 is facilitated by critical hydrogen bond networks with His1124 and Met1199 of ALK. Within the crizotinib series, crizo_35 demonstrated an exceptional IFD score of -671.16 kcal/mol, a predicted experimental pIC50 value of 7.75, and a binding energy of -77.71 kcal/mol, with water-mediated hydrogen bond networks involving Asp1203 and Lys1150 of ALK. These findings suggest that briga_15 and crizo_35 are promising leads warranting further optimization, synthesis, and biological evaluation.
O. Oyeneyin, N. Ipinloju, N. Gumede· Discover Chemistry· 0 citations
FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.
Ahmet Turan Demir· Intelligent Systems Research...· 4 citations
The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.
A. Fahira, M. Shahab, Zaheer Ud Din et al.· Journal of Genetic Engineeri...· 0 citations
This work demonstrates that integrating docking-derived pharmacophores with conformational ensemble-based machine learning provides an effective approach for discovering novel inhibitors against underexplored kinase targets.
G. Shakhatreh, M. Taha, S. Daoud· RSC Advances· 0 citations
An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.
Jurica Novak· International Journal of Mol...· 0 citations