Automated molecular property prediction (AutoMPP), an automated machine learning‐based pipeline that automates model selection and systematically evaluates fingerprint combinations across 75 molecular property prediction tasks, is established as a robust and adaptable framework for molecular property prediction.
Abstract
Accurate molecular property prediction is fundamental to drug discovery and is critically governed by molecular representations. While most existing approaches primarily focus on small molecules, extending reliable prediction to structurally complex macrocyclic compounds remains challenging due to their conformational flexibility and nonlocal interactions. To bridge this gap, we developed automated molecular property prediction (AutoMPP), an automated machine learning‐based pipeline that automates model selection and systematically evaluates fingerprint combinations across 75 molecular property prediction tasks. The results demonstrate that multifingerprint fusion significantly improves predictive robustness, with a four‐fingerprint combination achieving superior performance across diverse molecular tasks. Using this optimized representation strategy, AutoMPP outperforms leading models Uni‐Mol and fingerprints and graph neural networks (FP‐GNN), securing top performance on 68% (51/75) of tasks. Notably, AutoMPP generalizes effectively to macrocycles, such as cyclic peptide, attaining a Pearson correlation coefficient of 0.794, outperforming Uni‐Mol (0.692) and FP‐GNN (0.681). Furthermore, by integrating SHapley Additive exPlanations, the framework offers chemist‐intelligible insights into the structural determinants. Together, these results establish AutoMPP as a robust and adaptable framework for molecular property prediction, capable of identifying task‐specific optimal fingerprint combinations and learning architectures for both small molecules and complex macrocycles.
This review provides a systematic overview of recent advances in SSL-based molecular property prediction and analyzes how multimodal molecular representation learning by integrating sequence, graph, three-dimensional structure, and textual information can improve the quality and expressiveness of molecular representations.
Shuning Yang, Lei Deng· Journal of Chemical Informat...· 0 citations
Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.
Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al.· 1 citation
ABSTRACT Accurately predicting drug–target affinity (DTA) is crucial for accelerating virtual screening and guiding lead optimization in drug discovery. However, current computational approaches face a critical trade‐off: interaction‐free models lack fine‐grained binding details, while interaction‐based models overlook higher‐order contextual and functional patterns. This limitation hinders both prediction performance and real‐world generalization. To overcome this, we propose MF‐Net, a unified hierarchical multiscale fusion framework that integrates sequence‐, atomic‐, and fragment‐level representations to model drug–target interactions across complementary scales. MF‐Net achieves state‐of‐the‐art performance on the PDBBind v2016 benchmark and demonstrates strong early enrichment across multiple virtual screening datasets. Additionally, ADP‐Glo assays confirm that the MF‐Net‐guided virtual screening pipeline identifies seven novel nanomolar inhibitors targeting hematopoietic progenitor kinase 1 (HPK1). Among them, one compound achieves sub‐nanomolar activity (IC50 = 0.41 nM), outperforming the positive control inhibitor Sunitinib. These results demonstrate that MF‐Net not only excels on standard benchmarks but also delivers tangible lead discovery outcomes, underscoring its practical value for structure‐based drug design.
Shuo Liu, Xiang Zhang, Haixia Feng et al.· Advancement of science· 0 citations
It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations
Drug-induced liver injury (DILI) is a major cause of drug development failure and post-marketing withdrawal. Accurate computational prediction of hepatotoxicity is hindered by complex biological mechanisms and scarce labeled toxicity data. Although pretrained molecular language models like ChemBERTa perform well in molecular property prediction, their generalization ability for small-sample DILI prediction remains underexplored. Here, we systematically compared traditional molecular fingerprint-based machine learning methods and ChemBERTa-based models for DILI classification on the DILIst dataset. Canonical SMILES from PubChem were used to generate Morgan fingerprints and ChemBERTa embeddings. We evaluated Random Forest, XGBoost, full fine-tuning, frozen encoder transfer learning, and embedding-based classifiers under both random and scaffold data splits. Results showed that Morgan fingerprints combined with Random Forest achieved the best performance, with ROC-AUC of 0.783 and PR-AUC of 0.849 under random split. Scaffold split markedly degraded the performance of all models, indicating poor generalization to unseen chemical scaffolds. ChemBERTa embedding-based classifiers outperformed end-to-end fine-tuning, suggesting that pretrained representations are better used as fixed feature extractors under limited labeled DILI data. Further SHAP analysis detected key toxicity-related molecular fragments, and t-SNE showed insufficient latent-space separation between DILI-positive and negative compounds. Our results confirm that traditional fingerprint-based machine learning remains highly competitive for small-sample hepatotoxicity prediction, and this work provides a reliable computational framework for early drug safety assessment.
Wanying Li, Naihan Shi, Song He et al.· 2026 IEEE 27th China Confere...· 0 citations
A novel Dual-Attention Multimodal framework for Graphs and Sequence-based representations, so-called DAM-GS, which provides a promising solution for molecular property prediction with broad applications in drug discovery and computational molecular science.
Bay Van Nguyen, Vinh Truong, Ha Duong Thi Hong et al.· Journal of Chemical Informat...· 0 citations