Discovery of a potent TDP1 inhibitor through machine learning-driven predictive modeling combined with structure-based virtual screening and experimental validation
An integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation identifies AO65 as a promising lead for further TDP1-focused investigation.
Abstract
Tyrosyl-DNA phosphodiesterase I (TDP1) repairs topoisomerase I (TOP1)–mediated DNA damage and is a promising anticancer target, particularly in combination with TOP1 inhibitors. However, the discovery of potent and drug-like TDP1 inhibitors remains challenging due to the limited structural diversity of known active compounds. Here, we developed an integrated computational framework combining machine learning (ML), deep learning (DL), and structure-based docking with experimental validation. A curated dataset of 2040 compounds (857 active, 1183 inactive) was assembled and analyzed by scaffold composition. A total of 40 binary classification models were constructed using six ML algorithms and a deep neural network (DNN), each paired with five molecular fingerprint representations, along with five graph neural network architectures (GCN, GAT, MPNN, AttentiveFP, and FPGNN). The SVM::RDKitDes model performed best (AUC = 0.89, F1 = 0.78, BA = 0.80), with robustness confirmed by Y-scrambling and randomized-split analyses, and SHAP analysis identified 20 key descriptors of TDP1 inhibition. The model was deployed as a web application (http://drugpred.top:5050) and standalone desktop applications (.exe) are available at https://github.com/zenghuang8006/TDP1-inhibitor-prediction. The validated model was applied to screen 201 231 compounds, followed by drug-likeness filtering and hierarchical docking, yielding 16 candidates. Biological evaluation identified compound AO65 as a potent TDP1 inhibitor (IC50 = 0.80 ± 0.02 µM), and quantum chemical calculations and docking elucidated its electronic properties and binding within the catalytic domain. This work demonstrates the value of integrating ML-driven prediction with structure-based approaches and identifies AO65 as a promising lead for further TDP1-focused investigation.
The proposed workflow efficiently reduced a large chemical space to a focused set of TNKS1 inhibitor candidates while substantially reducing the experimental screening burden, highlighting the value of integrating consensus ML, SBVS, and experimental validation to accelerate early-stage hit discovery for TNKS1 and other therapeutic targets.
M. Bilotta, Adriana Gargano, R. Rocca et al.· Pharmaceuticals· 0 citations
Background: HER2 is a key oncogenic gene in breast cancer, involved in tumor progression, metastasis, and therapeutic resistance. This study aimed to find new HER2 inhibitors using a hybrid of machine learning (ML) and structure-based virtual screening (VS), combined with molecular dynamics (MD) simulations on various scaffolds. Methods: Four supervised molecular fingerprint classification models were trained on a dataset of 10,000 validated compounds from ChEMBL. Random Forest was the top model for screening a large compound library. Selected compounds underwent molecular docking in the HER2 ATP binding site, ADMET, drug likeness, toxicity analysis, and 200 ns MD simulations. Methods like PCA, FEL, hydrogen-bond analysis, DCCM, RDF, salt-bridge analysis, and MM/PBSA were used to assess binding stability. Results: Virtual screening identified three compounds, CHMEBL193865 (Lead-1), CHMEBL46740 (Lead-2), and CHMEBL151318 (Lead-3)—with better binding affinity and interaction profiles than the reference inhibitor. MD simulations showed stable protein–ligand complexes with RMSD values of 2.32–2.76 Å. Among these, Lead-2 was the most structurally stable, and Lead-1 had the most favorable binding free energy. All three compounds showed good drug likeness, ADMET properties, and low predicted toxicity. Conclusions: These findings support further in vitro and in vivo testing for developing new therapeutics against HER2-overexpressing breast cancer, highlighting two scaffolds with promising lead optimization potential.
Alhumaidi B. Alabbas, Safar M. Alqahtani· Pharmaceuticals· 0 citations
These findings introduce H_1 as a computationally prioritized, putative MLK4-binding lead and provide a hypothesis-generating framework for MLK4-targeted scaffold prioritization, while recognizing that experimental activity and kinome selectivity profiling remain necessary before H_1 can be described as a confirmed MLK4 inhibitor or MLK4-selective compound.
Afnan A. Alzaghari, S. Daoud, Husam Nassar et al.· Journal of Pharmaceutical In...· 0 citations
Findings identify CP20 as a promising lead scaffold for the development of novel DPP4 inhibitors and demonstrate the effectiveness of an ensemble machine learning-guided computational framework for accelerating antidiabetic drug discovery.
Iqra Anwar, T. Chohan, Drakhshaan et al.· Journal of Computational Bio...· 0 citations
The proto-oncogene serine/threonine kinase PIM2 is a critical regulator of cell proliferation, survival, and tumor progression and represents an attractive therapeutic target for several cancers. In this study, an integrated machine learning–guided computational pipeline was developed to identify potential PIM2 inhibitors by combining quantitative structure–activity relationship (QSAR) modeling, virtual screening, molecular docking, molecular dynamics (MD) simulations, and pharmacokinetic prediction. Bioactivity data for PIM2 inhibitors were retrieved from the ChEMBL database, yielding 5953 compounds. After data cleaning, structural standardization, and removal of duplicates and invalid entries, a curated dataset of 1584 compounds was obtained for QSAR modeling. To address dataset imbalance, the Synthetic Minority Oversampling Technique (SMOTE) was applied before model development. Twelve molecular fingerprint descriptors were generated and used to construct 180 QSAR models using five machine learning algorithms, including Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Regression (SVR), k-Nearest Neighbors (KNN), and Multilayer Perceptron (MLP). Among these models, the Random Forest–fingerprint model demonstrated the best predictive performance, achieving a mean R2 of 0.971 with low prediction errors (RMSE = 0.271; MAE = 0.125) across training, testing, and cross-validation datasets. The optimized model was subsequently applied to virtual screening of multiple chemical libraries, including FDA-approved drugs, natural product databases, and commercial compound collections. Several promising candidates were identified, including TCMBANKIN000009 (emetine), Amb28533044 (4,6′-Anhydrooxysporidinone), NPC170963 (Lysophosphatidylcholine (15:0)), NPC262768 (Endosulfan), and NPC469603 (8-hydroxyircinialactam A). Molecular docking showed that these compounds bind within the ATP-binding pocket of PIM2 kinase, forming interactions with key residues such as Lys62, Asp125, Asp128, and Glu168. Subsequent molecular dynamics simulations confirmed the stability of selected complexes, demonstrating reduced residue fluctuations, stable protein compactness, and persistent intermolecular interactions during the simulation. Furthermore, ADMET prediction suggested favorable pharmacokinetic and toxicity profiles for several compounds. Collectively, these findings highlight the potential of the identified molecules as promising PIM2 inhibitor candidates, providing valuable leads for future experimental validation and anticancer drug development.
A. Fahira, M. Shahab, Zaheer Ud Din et al.· Journal of Genetic Engineeri...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 15, 2026
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.