Skip to content
Open access

From Descriptor Learning to Binding Stability: An Explainable Machine Learning Pipeline for EGFR Double-Mutant Inhibitor Discovery

Aug 2026 · International Journal of Molecular Sciences · Vol 27 · 0 citations · 72 references
Medicine

TL;DR

An integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Abstract

Drug resistance arising during cancer development and progression remains a major challenge in the treatment of epidermal growth factor receptor (EGFR)-driven tumors, particularly those harboring the clinically relevant T790M/L858R double mutation. In this study, we developed an integrated computational workflow combining explainable machine learning, virtual screening, molecular dynamics simulations, and binding free-energy calculations to identify novel inhibitors of this drug-resistant EGFR variant. An XGBoost regression model was trained using scaffold-aware cross-validation, Bayesian hyperparameter optimization, and sequential feature selection, resulting in a compact model based on 16 molecular descriptors. The model demonstrated robust predictive performance on external validation data, while SHAP analysis identified descriptors related to the local electronic environment, fragment distribution, and molecular topology as the primary contributors to activity prediction. The optimized model was subsequently applied to screen compounds from the Enamine REAL database. Top-ranked candidates were evaluated using explicit-solvent molecular dynamics simulations and MM/GBSA binding free-energy calculations. Several compounds formed stable protein–ligand complexes and maintained key interactions with residues known to be important for EGFR inhibition, including Lys745, Met790, and Leu718. These results demonstrate that the proposed workflow can efficiently prioritize computational candidates of drug-resistant EGFR mutants and may support the development of new therapeutic strategies for overcoming resistance in EGFR-driven cancers.

Read PDF

Similar papers

Open access Jul 2026

Machine learning–driven identification of PIM2 kinase inhibitors through QSAR modeling and molecular dynamics simulations

The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.

A. Fahira, M. Shahab, Zaheer Ud Din et al. · 0 citations
Open access Jul 2026

Machine Learning-Driven Discovery of Novel HER2 Inhibitors Through Integrated Virtual Screening and Molecular Dynamics Simulations

Background: HER2 is a key oncogenic gene in breast cancer, involved in tumor progression, metastasis, and therapeutic resistance. This study aimed to find new HER2 inhibitors using a hybrid of machine learning (ML) and structure-based virtual screening (VS), combined with molecular dynamics (MD) simulations on various scaffolds. Methods: Four supervised molecular fingerprint classification models were trained on a dataset of 10,000 validated compounds from ChEMBL. Random Forest was the top model for screening a large compound library. Selected compounds underwent molecular docking in the HER2 ATP binding site, ADMET, drug likeness, toxicity analysis, and 200 ns MD simulations. Methods like PCA, FEL, hydrogen-bond analysis, DCCM, RDF, salt-bridge analysis, and MM/PBSA were used to assess binding stability. Results: Virtual screening identified three compounds, CHMEBL193865 (Lead-1), CHMEBL46740 (Lead-2), and CHMEBL151318 (Lead-3)—with better binding affinity and interaction profiles than the reference inhibitor. MD simulations showed stable protein–ligand complexes with RMSD values of 2.32–2.76 Å. Among these, Lead-2 was the most structurally stable, and Lead-1 had the most favorable binding free energy. All three compounds showed good drug likeness, ADMET properties, and low predicted toxicity. Conclusions: These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.

Alhumaidi B. Alabbas, Safar M. Alqahtani · 0 citations
Open access Aug 2026

Integrating molecular dynamics and machine learning to identify potential apo-state conformational and solvent-exposure signatures associated with resistant KRAS mutants

A computational framework integrating molecular dynamics (MD)-derived structural, energetic, thermodynamic, and contact-based descriptors with machine learning may inform the design of inhibitors targeting secondary KRAS resistance mutations, pending validation in additional structurally independent mutant systems.

Katarzyna Mizgalska, Konstancja Urbaniak, Denis Imbody et al. · 0 citations
Jul 2026

Modeling Structure-Activity Relationships with Machine Learning to Identify DPP4 Inhibitors as potential Therapeutics for Type 2 Diabetes

Findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery.

Iqra Anwar, T. Chohan, Drakhshaan et al. · 0 citations
Open access Jul 2026

Explainable Machine Learning-Guided Virtual Screening and Molecular Docking for Identification of Novel FYN Kinase Inhibitors

FYN kinase is a non-receptor protein tyrosine kinase involved in various cancers and neurodegenerative diseases; however, no selective FYN inhibitor has been approved yet. Here we introduce the explainable Machine Learning (ML) coupled with virtual screening and Molecular Docking (MD) pipeline for fast prediction of new FYN kinase inhibitors. In this study, we constructed the training set of 906 molecules active against FYN kinase from the ChEMBL database. Molecules were encoded with Extended-Connectivity Fingerprints (ECFP4). The classification models Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) were developed, and the latter showed the better performance in test (AUC=0.8118) and 5-fold cross-validation (AUC=0.8297). Based on the SHapley Additive exPlanations (SHAP) values obtained via TreeExplainer, nitrogen-containing heterocycles and hydrogen bond acceptors have been identified as the most important molecular substructures. Using the optimal XGBoost classifier, screening of 2,000 approved drugs has been performed, resulting in 470 hit molecules (23.5% hit rate). Five best molecules were further submitted to the MD procedure using AutoDock Vina to dock to FYN kinase domain (PDB RCSB: 2DQ7), showing binding energies in the interval of -9.57 to -6.32 kcal/mol. Dasatinib Anhydrous (CHEMBL1421) was the second strongest binder (-8.49 kcal/mol), effectively interacting with the ATP binding site. Although CHEMBL1171837 was the strongest binder (-9.57 kcal/mol), it was caught in the ADMET profiling. According to ADMET profiling, the top one inhibitor (CHEMBL1421) satisfies Lipinski’s rule of five and Veber rules. Analysis of hydrogen bond and hydrophobic interactions revealed hydrogen bonding with ASP148, LYS39, and ASN86 and hydrophobic interactions with ALA147, ILE80, and GLY88. Validation by self-docking procedure (self-docking or STS) showed low Root Mean Square Deviation (RMSD)<2.0 Å with a binding affinity of -11.53 kcal/mol. This work highlights how explainable ML can be used in combination with structure-based docking to expedite the drug discovery process against FYN kinase and can be applied to other kinase targets.

Ahmet Turan Demir · 4 citations
Open access Jul 2026

Identification of novel EGFR inhibitors for glioblastoma through pharmacophore-guided virtual screening and molecular dynamics

The pharmacophore-based screening and docking analysis identified eleven promising EGFR-binding compounds, of which eight demonstrated optimal ADMET characteristics and stable interactions within the active site during molecular dynamics simulations, suggesting their potential efficacy as EGFR inhibitors.

M. Moulay, M. Mahmoud, Reem M. Farsi et al. · 0 citations