Atomic charge is crucial in drug design for analyzing reactive sites and interactions between ligands and targets. While quantum mechanical methods offer high accuracy, they are generally computationally costly. Conversely, empirical approaches, while computationally efficient, frequently suffer from lack of precision and generalizability. Recent a number of machine learning-based models have been developed for atomic charge predictions, but they struggle with accurately representing molecular structures and capturing the chemical environments affecting atomic charges, thus limiting their generalization and accuracy. To overcome these limitations, we propose LumiCharge, a novel atomic charge prediction framework that incorporates high-order spherical harmonics convolutions and explicitly models multibody interactions. In constructing this model, we employ a strategy that integrates both high- and low-order information, enhancing its geometric spatial perception capability, which is currently underexplored in the field. Benchmark evaluations demonstrate that LumiCharge outperforms state-of-the-art (SOTA) models by 30%-60% across diverse data sets. Additionally, in cross-scale experiments, LumiCharge demonstrates exceptional extrapolation capability and robustness across molecules of varying sizes, effectively overcoming the limitations imposed by molecular sizes. On an external halogen-containing test set, LumiCharge achieves an RMSE of 0.055e, meeting practical application requirements. Finally, a case study of virtual screening for the androgen receptor (AR) target further validates its outstanding accuracy compared to the OPLS3e force field and other deep learning (DL)-based baseline models, highlighting its exceptional generalization capacity and practical utility in real-world scenarios.
Qun Su, Hui Zhang, Qiaolin Gou et al.· Journal of Physical Chemistr...· 2 citations
Accurate prediction of the peptide-protein interaction (PepPI) is crucial for developing peptide-based therapeutics and vaccines. However, this computational task has traditionally faced significant challenges, such as the scarcity of structure data along with the corresponding label of the binding affinity for bound complexes. To address these challenges, we introduce PepBAN, a deep learning framework for modeling PepPI predictions. PepBAN incorporates two technical advancements: (1) adopting the protein language model ESM-2 to characterize proteins and ESM-2 or a graph-based foundation model for peptides without structure data and (2) leveraging the conditional domain adversarial learning to enhance generalization across a broad range of protein targets, especially when there are limited binding data. At the core of PepBAN is a bilinear attention network (BAN) that effectively learns the pattern of pairwise local interactions, enables the identification of key residues participating in the peptide-protein interactions, and offers an intuitive approach to interpret the underlying mechanisms of PepPIs via analyzing attention weights. Our numerical experiments demonstrated that PepBAN outperformed the previous state-of-the-art models across several well-established benchmark studies. Furthermore, we evaluated PepBAN's applicability in predicting cyclic peptide-protein interactions, a task that poses significant challenges due to the presence of noncanonical amino acids. These nonstandard residues require specialized handling, which most existing sequence-based PepPI prediction models did not adequately address, and we adopt an atom-resolved molecular graph approach to process cyclic peptides. Despite this complexity, PepBAN demonstrated a clear advantage by achieving a superior prediction performance and offering a distinct edge in tackling the emerging chemical space of cyclic peptides, which has great potential for novel therapeutic development. In summary, PepBAN serves as a valuable tool for advancing peptide-based drug and therapeutic development.
Shuaiyan Li, Xiaorui Wang, Yuchen Zhu et al.· Journal of Chemical Informat...· 2 citations
The androgen receptor (AR) represents a pivotal therapeutic target for prostate cancer. However, existing orthosteric ligand-binding pocket (LBP) antagonists [e.g., enzalutamide (ENZ)] encounter significant obstacles due to resistance-conferring mutations in the LBP. Allosteric antagonists targeting the BF3 site exhibit great potential in overcoming such resistance but have low inhibitory efficacy. In our study, we employed an integrated computational modeling strategy, including Gaussian-accelerated molecular dynamics (GaMD), MM/GBSA free-energy calculations, and elastic network model (ENM)-based signaling communication pathway analyses. This approach is used to probe the cooperativity of allosteric BF3 antagonists [e.g., VPC-13808 (VPC)] with diverse orthosteric LBP ligands [e.g., ENZ and testosterone (TES)] in suppressing AR activity. Herein, four types of AR systems were examined: AR bound to LBP agonist (AR·TES), LBP antagonists (e.g., AR·ENZ), and combinations of LBP agonist/antagonist with BF3 antagonist (e.g., AR·TES·VPC and AR·ENZ·VPC). Results indicate that BF3 antagonists can synergize with the LBP antagonist to amplify conformational flexibility in H12 and induce anticorrelated dynamics of H12 with H3 and H4. This induces the downward movement of H12 and its displacement away from H3/H4, triggering the wide opening of the AF2 binding cleft and substantially reducing the coactivator recruitment. Furthermore, the BF3 antagonist can interact with specific residues (e.g., F673, F826, L830, and Y834) and cooperate with the LBP agonist or antagonist to allosterically perturb the AF2 conformation. Multiple short- and/or long-range BF3→AF2 and LBP→AF2 signaling transition pathways are involved, such as F673→Y834→L722→L812→L744→V746→L873→ENZ→L880/V889/V891. These mechanistic insights establish the foundation for developing novel AR BF3 antagonist and LBP-BF3 combination therapies, suggesting a promising avenue for enhancing the efficacy and overcoming the resistance in castration-resistant prostate cancer treatment.
Xiaotian Kong, Yushan Zou, Peng Cao et al.· Journal of Chemical Informat...· 1 citation
Atomic charge is a fundamental quantum chemical property essential for advancing drug design and discovery. Although quantum mechanics (QM) methods offer the highest level of accuracy, their computational demands scale quadratically with the number of atoms, limiting their practicality for large-scale applications. In light of this, empirical and semiempirical methods have been introduced to improve computational efficiency, albeit often at the expense of accuracy. The advent of artificial intelligence has witnessed a growing application of machine learning (ML) techniques to accelerate atomic charge predictions. However, existing ML models often suffer from low accuracy and limited generalization capabilities. To address these challenges, we introduce an advanced equivariant graph attention neural network specifically engineered to model long-range atomic electrostatic interactions with high precision. This model introduces a sophisticated global graph attention mechanism, enabling it to capture charge contributions across multiple scales. By utilizing a combination of structural symmetry-preserving transformations and multiscale attention, our approach not only preserves the inherent symmetries of molecular structures but also substantially improves the model's accuracy, generalization, and robustness in complex scenarios. Our empirical analyses demonstrate that, compared to leading baseline models, the proposed model improves charge prediction accuracy by over 40% on average across various charge-calculation schemes. Remarkably, the model achieves superior performance on the external RESP (restrained electrostatic potential) test data sets, with a 54.6% improvement over the baseline. Additionally, we evaluated our charge model under the setting of virtual screening, where it outperforms both the OPLS3 charges and baseline deep learning models across all evaluation metrics, highlighting its extensive potential for scientific discovery.
Qiaolin Gou, Qun Su, Jike Wang et al.· Journal of Chemical Informat...· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
Accurate modeling of protein-peptide interactions is essential for understanding fundamental biological processes and designing peptide-based drugs. However, predicting the complex structures of these interactions remains challenging, primarily due to the high conformational flexibility of peptides. To support a fair and systematic evaluation of recent deep learning (DL) approaches, we introduce PepPCBench, a benchmarking framework tailored to assess protein folding neural networks (PFNNs) in protein-peptide complex prediction. As part of this framework, we curated PepPCSet, a data set of 261 experimentally resolved complexes with peptides ranging from 5 to 30 residues. We benchmark five full-atom PFNNs, including AlphaFold3 (AF3), AlphaFold-Multimer (AFM), Chai-1, HelixFold3 (HF3), and RoseTTAFold-All-Atom (RFAA), using comprehensive evaluation metrics. Our benchmarking reveals meaningful performance differences among these methods and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy. While AF3 shows strong performance in structure prediction, further analysis indicates that confidence metrics correlate poorly with experimental binding affinities, underscoring the need for improved scoring strategies and generalizability. By providing a reproducible and extensible framework, PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction.
Silong Zhai, Huifeng Zhao, Jike Wang et al.· Journal of Chemical Informat...· 13 citations· ⚡1
MHC-II neoantigens play a critical role in immunotherapy, either as direct effectors or through their influence on CD8+ T cells. However, only a small fraction of tumor DNA mutations qualify as functional neoantigens, and current prediction tools often lack accuracy, leading to the low immunogenicity of predicted neoantigens in vivo. Here, we present EpiMII, a Graph Neural Network model for MHC-II epitope design, which learns from the structural features of epitopes to predict their sequences. To train EpiMII, we constructed a reliable, large dataset containing 142,934 MHC-II epitope structures. This approach achieves a 4.2x improvement over ProteinMPNN, with a sequence recovery rate of 78.0% for known MHC-II epitopes in the Protein Data Bank. As a case study, we designed a neoantigen from hepatocellular carcinoma. All five designed epitopes significantly activated CD4+ T cells in vitro and induced secretion of IFN-γ and TNF-α. Notably, one epitope treatment significantly reduced tumor volume in mice in vivo. EpiMII offers a novel and efficient approach for identifying MHC-II epitopes/neoantigens, potentially contributing to vaccine development.
Jiayi Yuan, Xiaowei Xu, Ze-Yu Sun et al.· bioRxiv· 1 citation
Recently, various self‐supervised learning (SSL) methods based on 3D graph neural networks (GNNs) have been developed to comprehensively represent the structural information of molecules in 3D space; this is essential for discovering new drugs. However, existing methods fail to comprehensively characterize the 3D structures of molecules and neglect the electronic structural information that significantly influences key properties such as molecular reactivity, strong electrostatic interactions, and chemical adsorption. Therefore, here, a novel molecular representation learning method is constructed, Q‐GEM, incorporating quantum and geometric structural information enhancement, based on the quantum chemical property database QuanDB and SSL methods. Q‐GEM comprises a GNN embedded with the molecular electronic and complete 3D geometrical structural information as well as several well‐designed multiscale SSL tasks, achieving superior absolute molecular conformation prediction and conformational discrimination. The Q‐GEM achieved state‐of‐the‐art performance in 12 out of 13 prediction tasks on the MoleculeNet dataset, with an average performance improvement of 3.3% and 2.0% for classification and regression prediction tasks, respectively. Moreover, an average performance improvement of 5.2% is achieved in three localized quantum chemical properties, fully demonstrating the excellent performance of Q‐GEM in distinguishing molecular electronic structures. The Q‐GEM represents a novel, powerful breakthrough for accurate molecular property prediction.
Zhijiang Yang, Liangliang Wang, Tengxin Huang et al.· Advancement of science· 5 citations
To fulfill functions for differentially regulating the downstream signaling pathways, functional ligands (i.e., agonists or antagonists) targeting nuclear receptors (NRs) are designed to stabilize different conformations (active or inactive) of the proteins. However, in practical applications, it is usually difficult to determine the molecular category of an NR ligand because these molecules all bind in the same location of an NR protein, namely, the ligand-binding pocket (LBP). Considering that ligands with different properties (agonists or antagonists) prefer to bind with differential conformations of NRs, it is possible to identify the molecular type of a given ligand through the differential binding environment (active or inactive conformations) of the protein-ligand interaction. Therefore, in this study, we established a unified model (NRIGN) based on the deep graphic architecture to discriminate agonists and antagonists targeting 26 successful or in-clinical-trial NR targets. Our result shows that NRIGN achieves an excellent prediction accuracy (ACC >0.95) and is robust enough to be applied in various real-world scenarios, such as predicting the molecular type of ligands in crystallized NR structures, ligands with multiple NR activities, and ligands with their types altered by target mutations. The proposed model is expected to promote rational design of drugs targeting NR proteins.
Kaimo Yang, Dejun Jiang, Qirui Deng et al.· Journal of Chemical Informat...· 2 citations
Photodynamic therapy (PDT) is a clinically approved therapeutic modality that has demonstrated significant potential for cancer treatment, and triplet photosensitizers (PSs) play a key role in its efficacy. Despite deep learning having emerged as a next-generation tool for material discovery, existing methods mainly target a limited subset of triplet PSs, such as thermally activated delayed fluorescence (TADF) materials, neglecting the critical intersystem crossing (ISC) between the high-lying singlet and triplet states (ΔESnTn). To overcome this limitation, we compiled a comprehensive dataset (∼1.90 × 109) of triplet PSs encompassing various ISC mechanisms. Then, we proposed a novel strategy that incorporates two models: a fragment-based model (Frag-MD) and a character-based model (MD), both integrating a conditional transformer, recurrent neural networks, and reinforcement learning. In silico experiments revealed that the Frag-MD model outperforms the MD model in generating larger conjugated motifs with higher average ring numbers and atom counts; while the MD model generates twice as many unique motifs and excels in novelty and diversity, as evaluated by conditional and MOSES metrics. Therefore, our approach is highly effective for modifying conjugated motifs and designing novel triplet PSs. Notably, the recently reported high-efficiency triplet PSs have been re-identified through ablation experiments using our proposed models, which target ΔESnTn and significantly outperform traditional baselines, achieving a prediction accuracy of 73% versus 4%. Our approach holds the potential to establish a new paradigm for discovering novel PSs applicable in PDT.
Kepeng Chen, Xiaoting Zhang, Jike Wang et al.· Chemical Science· 3 citations
Drug repositioning (DR) identifies new therapeutic uses for approved drugs, reducing development burdens and offering safer treatment options for patients. While high-throughput technologies generate complex, large-scale multiomics data, existing DR tools struggle to comprehensively analyze the resulting biological networks. To address this challenge, we present DRHIN, an integrated, interactive web server for DR over heterogeneous information networks (HINs) using advanced deep learning techniques. DRHIN integrates transcriptomics, proteomics, and microbiome data, incorporating eight biological entities and 19 association types to build diverse HINs and elucidate the underlying molecular mechanisms. It includes 19 state-of-the-art graph representation algorithms, enabling flexible training, comparison, and evaluation of heterogeneous network data. The platform provides a code-free portal supporting three key predictive tasks: discovering drug-disease associations, repurposing existing drugs for new indications, and identifying potential therapies for specific diseases, making analyses accessible and reproducible. Leveraging high-performance computing, DRHIN efficiently processes million-scale networks, ensuring practical applicability in real-world scenarios. The web server is freely accessible at http://drhin.tianshanzw.cn.
Bowei Zhao, Dongxu Li, Yue Yang et al.· Journal of Chemical Informat...· 6 citations· ⚡1
Structure-based machine learning algorithms have been utilized to predict the properties of protein-protein interaction (PPI) complexes, such as binding affinity, which is critical for understanding biological mechanisms and disease treatments. While most existing algorithms represent PPI complex graph structures at the atom-scale or residue-scale, these representations can be computationally expensive or may not sufficiently integrate finer chemical-plausible interaction details for improving predictions. Here, we introduce MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently. This framework maps proteins onto a concise CG-scale complex graph, where nodes represent CG beads and edges encode chemically plausible interactions. The GNN-based encoder is tailored to extract high-quality representations from this graph, efficiently capturing the overall properties of the protein complex structure. Extensive experiments on three different downstream PPI property prediction tasks demonstrate that MCGLPPI achieves competitive performance compared with the counterparts at the atom- and residue-scale, but with only a third of the computational resource consumption. Furthermore, the CG-scale pre-training on protein domain-domain interaction structures enhances its predictive capabilities for PPI tasks. MCGLPPI offers an effective and efficient solution for PPI overall property predictions, serving as a promising tool for the large-scale analysis of biomolecular interactions.
Yang Yue, Shu Li, Yihua Cheng et al.· bioRxiv· 14 citations
The integration of organic synthesis with enzymatic catalysis offers a promising route toward efficient and sustainable construction of complex molecules. While organic synthesis enables diverse transformations, enzymatic catalysis enhances stereoselectivity under mild conditions, improving cost-effectiveness and environmental impact. However, current enzymatic synthesis planning algorithms face challenges in formulating robust hybrid organic–enzymatic strategies. Key issues include the difficulty in devising hybrid planning approaches and the reliance on template-based enzyme recommendations, which limits their adaptability across diverse scenarios. Here we show ChemEnzyRetroPlanner, an open-source hybrid synthesis planning platform that combines organic and enzymatic strategies with AI-driven decision-making. The platform features advanced computational modules, including hybrid retrosynthesis planning, reaction condition prediction, plausibility evaluation, enzymatic reaction identification, enzyme recommendation, and in silico validation of enzyme active sites. A central innovation is the RetroRollout* search algorithm, which outperforms existing tools in planning synthesis routes for organic compounds and natural products across multiple datasets. ChemEnzyRetroPlanner provides an intuitive graphical interface and programmatic APIs for scalability, while leveraging the chain-of-thought strategy and the Llama3.1 model to autonomously activate hybrid synthesis strategies for diverse scenarios. The results indicate that this fully automated, open-source system holds potential value for improving the efficiency and sustainability of molecular synthesis. The integration of organic and enzymatic synthesis enhances molecule construction efficiency. Here, the authors present ChemEnzyRetroPlanner, an AI-driven platform for automated hybrid synthesis planning, improving synthesis route efficiency and sustainability.
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.