Skip to content

Category

federated learning

367 papers

EpiMII: Integrating Structure and Graph Neural Networks for MHC-II Epitope and Neoantigen Design

MHC-II neoantigens play a critical role in immunotherapy, either as direct effectors or through their influence on CD8+ T cells. However, only a small fraction of tumor DNA mutations qualify as functional neoantigens, and current prediction tools often lack accuracy, leading to the low immunogenicity of predicted neoantigens in vivo. Here, we present EpiMII, a Graph Neural Network model for MHC-II epitope design, which learns from the structural features of epitopes to predict their sequences. To train EpiMII, we constructed a reliable, large dataset containing 142,934 MHC-II epitope structures. This approach achieves a 4.2x improvement over ProteinMPNN, with a sequence recovery rate of 78.0% for known MHC-II epitopes in the Protein Data Bank. As a case study, we designed a neoantigen from hepatocellular carcinoma. All five designed epitopes significantly activated CD4+ T cells in vitro and induced secretion of IFN-γ and TNF-α. Notably, one epitope treatment significantly reduced tumor volume in mice in vivo. EpiMII offers a novel and efficient approach for identifying MHC-II epitopes/neoantigens, potentially contributing to vaccine development.

Jiayi Yuan, Xiaowei Xu, Ze-Yu Sun et al. · 1 citation

Q‐GEM: Quantum Chemistry Knowledge Fusion Geometry‐Enhanced Molecular Representation for Property Prediction

Recently, various self‐supervised learning (SSL) methods based on 3D graph neural networks (GNNs) have been developed to comprehensively represent the structural information of molecules in 3D space; this is essential for discovering new drugs. However, existing methods fail to comprehensively characterize the 3D structures of molecules and neglect the electronic structural information that significantly influences key properties such as molecular reactivity, strong electrostatic interactions, and chemical adsorption. Therefore, here, a novel molecular representation learning method is constructed, Q‐GEM, incorporating quantum and geometric structural information enhancement, based on the quantum chemical property database QuanDB and SSL methods. Q‐GEM comprises a GNN embedded with the molecular electronic and complete 3D geometrical structural information as well as several well‐designed multiscale SSL tasks, achieving superior absolute molecular conformation prediction and conformational discrimination. The Q‐GEM achieved state‐of‐the‐art performance in 12 out of 13 prediction tasks on the MoleculeNet dataset, with an average performance improvement of 3.3% and 2.0% for classification and regression prediction tasks, respectively. Moreover, an average performance improvement of 5.2% is achieved in three localized quantum chemical properties, fully demonstrating the excellent performance of Q‐GEM in distinguishing molecular electronic structures. The Q‐GEM represents a novel, powerful breakthrough for accurate molecular property prediction.

Zhijiang Yang, Liangliang Wang, Tengxin Huang et al. · 5 citations

Effective generation of heavy-atom-free triplet photosensitizers containing multiple intersystem crossing mechanisms based on deep learning

Photodynamic therapy (PDT) is a clinically approved therapeutic modality that has demonstrated significant potential for cancer treatment, and triplet photosensitizers (PSs) play a key role in its efficacy. Despite deep learning having emerged as a next-generation tool for material discovery, existing methods mainly target a limited subset of triplet PSs, such as thermally activated delayed fluorescence (TADF) materials, neglecting the critical intersystem crossing (ISC) between the high-lying singlet and triplet states (ΔESnTn). To overcome this limitation, we compiled a comprehensive dataset (∼1.90 × 109) of triplet PSs encompassing various ISC mechanisms. Then, we proposed a novel strategy that incorporates two models: a fragment-based model (Frag-MD) and a character-based model (MD), both integrating a conditional transformer, recurrent neural networks, and reinforcement learning. In silico experiments revealed that the Frag-MD model outperforms the MD model in generating larger conjugated motifs with higher average ring numbers and atom counts; while the MD model generates twice as many unique motifs and excels in novelty and diversity, as evaluated by conditional and MOSES metrics. Therefore, our approach is highly effective for modifying conjugated motifs and designing novel triplet PSs. Notably, the recently reported high-efficiency triplet PSs have been re-identified through ablation experiments using our proposed models, which target ΔESnTn and significantly outperform traditional baselines, achieving a prediction accuracy of 73% versus 4%. Our approach holds the potential to establish a new paradigm for discovering novel PSs applicable in PDT.

Kepeng Chen, Xiaoting Zhang, Jike Wang et al. · 3 citations

Revisiting Protein-Protein Docking: A Systematic Evaluation Framework

Protein-protein interactions play pivotal roles in a wide range of biological processes. Determining the atomic-level structures of protein-protein complexes is indispensable for elucidating macromolecular interaction mechanisms and advancing structure-based drug design. Protein-protein docking, as one of the leading computational approaches for predicting complex structures, has seen considerable progress but requires rigorous evaluation in practical applications. In this study, we proposed a comprehensive benchmarking framework to evaluate 11 docking methods spanning traditional (HDOCK, PatchDock, PIPER, ZDOCK) and deep learning (DL)-based (EquiDock, ElliDock, EBMDock, GeoDock, DiffDock-PP, AlphaFold-Multimer, AlphaFold3) approaches. Our framework incorporates the classical DockingBenchmark 5.5 data set for evaluating flexible docking, introduces a newly curated data set (AACBench) for antibody-antigen complex docking, and establishes the PPCBench data set to examine the out-of-distribution (OOD) generalization capabilities of DL-based methods. In docking against apo structures, AlphaFold3 achieves a superior top-5 success rate of 77.98%, whereas the traditional approach HDOCK reaches merely 12.84%, despite its highest top-5 success rate of 85.24% when docking against holo structures. For antibody-antigen docking, AlphaFold3 remains the most accurate method (top-5 success rate: 31.78%) and substantially outperforms AlphaFold-Multimer in modeling the CDR-H3 loop. In OOD generalization tests, all DL-based models exhibit markedly reduced performance on the PPCBench data set. Overall, our work establishes a unified benchmarking framework that enables systematic evaluation of docking methods across diverse tasks and provides critical insights into the strengths and limitations of current docking strategies, thereby informing future developments in protein-protein docking research.

Linlong Jiang, Ke Zhang, Kai Zhu et al. · 3 citations

MetalloDock: Decoding Metalloprotein-Ligand Interactions via Physics-Aware Deep Learning for Metalloprotein Drug Discovery.

Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDock, the first DL-based docking framework specially designed for metalloprotein targets. By innovatively integrating an autoregressive spatial decoding engine with a physics-constrained geometric generation paradigm, MetalloDock can precisely reconstruct metal coordination geometries and accurately capture metal-ligand interactions, which enhance both the accuracy of metalloprotein-ligand docking and binding affinity prediction. Extensive evaluations on our custom-built benchmark data set demonstrate that MetalloDock outperforms existing methods, including AlphaFold3, in docking success rate and virtual screening performance for metalloprotein targets. In real-world applications, MetalloDock successfully identified multiple novel hit compounds in a virtual screening campaign targeting the prostate-specific membrane antigen. Additionally, it enabled rational drug design for acidic polymerase endonuclease, leading to the discovery of potent inhibitors. These results highlight the broad applicability of MetalloDock in accelerating metalloprotein-targeted drug discovery and provide a standardized framework for future evaluation of metalloprotein-specific docking algorithms.

Hui Zhang, Xujun Zhang, Qun Su et al. · 5 citations

Computational and AI-Driven Ecosystem for Structure-Based Covalent Drug Discovery.

ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.

Shi Li, Hongyan Du, Xujun Zhang et al. · 4 citations

DRHIN: An Integrated and Interactive Web Server for Drug Repositioning

Drug repositioning (DR) identifies new therapeutic uses for approved drugs, reducing development burdens and offering safer treatment options for patients. While high-throughput technologies generate complex, large-scale multiomics data, existing DR tools struggle to comprehensively analyze the resulting biological networks. To address this challenge, we present DRHIN, an integrated, interactive web server for DR over heterogeneous information networks (HINs) using advanced deep learning techniques. DRHIN integrates transcriptomics, proteomics, and microbiome data, incorporating eight biological entities and 19 association types to build diverse HINs and elucidate the underlying molecular mechanisms. It includes 19 state-of-the-art graph representation algorithms, enabling flexible training, comparison, and evaluation of heterogeneous network data. The platform provides a code-free portal supporting three key predictive tasks: discovering drug-disease associations, repurposing existing drugs for new indications, and identifying potential therapies for specific diseases, making analyses accessible and reproducible. Leveraging high-performance computing, DRHIN efficiently processes million-scale networks, ensuring practical applicability in real-world scenarios. The web server is freely accessible at http://drhin.tianshanzw.cn.

Bowei Zhao, Dongxu Li, Yue Yang et al. · 6 citations · ⚡1
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

Structure-based machine learning algorithms have been utilized to predict the properties of protein-protein interaction (PPI) complexes, such as binding affinity, which is critical for understanding biological mechanisms and disease treatments. While most existing algorithms represent PPI complex graph structures at the atom-scale or residue-scale, these representations can be computationally expensive or may not sufficiently integrate finer chemical-plausible interaction details for improving predictions. Here, we introduce MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently. This framework maps proteins onto a concise CG-scale complex graph, where nodes represent CG beads and edges encode chemically plausible interactions. The GNN-based encoder is tailored to extract high-quality representations from this graph, efficiently capturing the overall properties of the protein complex structure. Extensive experiments on three different downstream PPI property prediction tasks demonstrate that MCGLPPI achieves competitive performance compared with the counterparts at the atom- and residue-scale, but with only a third of the computational resource consumption. Furthermore, the CG-scale pre-training on protein domain-domain interaction structures enhances its predictive capabilities for PPI tasks. MCGLPPI offers an effective and efficient solution for PPI overall property predictions, serving as a promising tool for the large-scale analysis of biomolecular interactions.

Yang Yue, Shu Li, Yihua Cheng et al. · 14 citations
#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

Designing effective mRNA sequences for therapeutics remains a formidable challenge. Inspired by successes in protein design, language models (LMs) are now being applied to RNA, but progress is often impeded by the lack of comprehensive training data. Existing models are frequently limited to UTR or CDS regions, restricting their application for complete mRNA sequences. We introduce mRNABERT, a robust, all-in-one mRNA designer pre-trained on the largest available mRNA dataset. To enhance performance, we propose a dual tokenization scheme with a cross-modality contrastive learning framework to integrate semantic information from protein sequences. On a comprehensive benchmark, mRNABERT demonstrates state-of-the-art performance, outperforming previous models in the majority of tasks for 5’ UTR and CDS design, RNA-binding protein (RBP) site prediction, and full-length mRNA property prediction. It also surpasses large protein models in several related tasks. In conclusion, mRNABERT’s superior performance across these diverse tasks signifies a substantial leap forward in mRNA research and therapeutic development. Designing complete mRNA sequences for new vaccines and therapies is a complex challenge. Here, the authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks.

Ying Xiong, Aowen Wang, Yu Kang et al. · 22 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

Molecular representation learning (MRL) has shown promise in accelerating drug development by predicting chemical properties. However, imperfectly annotation among datasets pose challenges in model design and explainability. In this work, we formulate molecules and corresponding properties as a hypergraph, extracting three key relationships: among properties, molecule-to-property, and among molecules, and developed a unified and explainable multi-task MRL framework, OmniMol. It integrates a task-related meta-information encoder and a task-routed mixture of experts (t-MoE) backbone to capture correlations among properties and produce task-adaptive outputs. To capture underlying physical principles among molecules, we implement an innovative SE(3)-encoder for physical symmetry, applying equilibrium conformation supervision, recursive geometry updates, and scale-invariant message passing to facilitate learning-based conformational relaxation. OmniMol achieves state-of-the-art performance in properties prediction, reaches top performance in chirality-aware tasks, demonstrates explainability for all three relations, and shows effective performance in practical applications. Our code is available in our https://github.com/bowenwang77/OmniMol public repository. AI models for drug discovery often struggle with real-world, incomplete data. Here, the authors present OmniMol, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 9 citations
#machine learning Open access May 2025

Token-Mol 1.0: tokenized drug design with large language models

The integration of large language models (LLMs) into drug design is gaining momentum; however, existing approaches often struggle to effectively incorporate three-dimensional molecular structures. Here, we present Token-Mol, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens. Built on a transformer decoder and trained with causal masking, Token-Mol introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream applications. The model surpasses existing methods, improving molecular conformation generation by over 10% and 20% across two datasets, while outperforming token-only models by 30% in property prediction. In pocket-based molecular generation, it enhances drug-likeness and synthetic accessibility by approximately 11% and 14%, respectively. Notably, Token-Mol operates 35 times faster than expert diffusion models. In real-world validation, it improves success rates and, when combined with reinforcement learning, further optimizes affinity and drug-likeness, advancing AI-driven drug discovery. In this work the authors present Token-Mol, a token-only 3D drug design model, which deploys the Gaussian cross-entropy (GCE) loss function for regression tasks. It exhibits superior performance in molecular conformation generation, property prediction, and pocket-based generation, thus opening up new avenues for drug design.

Jike Wang, Rui Qin, Mingyang Wang et al. · 30 citations · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.