ConspectusProstate cancer (PCa) is the most prevalent malignancy among men worldwide, with its pathogenesis and progression heavily reliant on the sustained activation of the androgen receptor (AR) signaling pathway. The AR, a transcription factor of nuclear receptor superfamily, serves as the most privileged therapeutic target in PCa, as evidenced by the clinical efficacy of first- and second-generation AR antagonists. Current clinically available AR antagonists exclusively target the ligand binding pocket (LBP), suppressing tumor proliferation through competitive inhibition of androgen binding and subsequent blockade of AR signaling transduction. However, their therapeutic utility is invariably limited by acquired resistance mechanisms, including point mutations that alter LBP specificity, AR gene amplification leading to receptor overexpression, and the emergence of constitutively active splice variants that bypass ligand-dependent activation. Thus, the development of novel AR antagonists featuring innovative mechanisms and structural scaffolds is imperative to overcome resistance to antiandrogen therapy. However, the AR exhibits significant structural flexibility, and the lack of antagonist-bound crystal structures has hindered structure-based rational drug design. In this Article, we summarize our advances in elucidating the molecular mechanisms underlying AR conformational regulation and highlight our progress in the structure-based design and development of novel AR antagonists. First, our molecular dynamic (MD) studies collectively elucidate the molecular mechanisms by which the AR ligand binding domain (LBD) regulates its functional states through dynamic conformational changes mediated by distinct allosteric pathways when bound to agonists or antagonists, providing atomic-level insights and structural basis for drug development. Then, we successfully identified structurally diverse lead compounds targeting the LBP through various integrated approaches combining MD simulations, structure-based virtual screening (SBVS), and systematic biological evaluation. These compounds exhibited potent activity against clinically relevant AR mutations F877L, W742C, T878A, and H875Y, demonstrating their potential to overcome mutations-driven resistance. Further, we explored non-LBP mediated strategies for AR antagonism, including: (1) targeting the allosteric binding sites on LBD; (2) identification of novel druggable binding sites; and (3) targeting alternative domains beyond the LBD. As a paradigm-shifting example, we proposed inhibition of AR LBD dimerization as a novel mechanism of action for LBP-targeting AR antagonists. Building upon this insight, we characterized a promising pocket at the dimer interface, designated the Dimerization Interface Pocket (DIP), and developed first-in-class antagonists specifically targeting this site, which exhibit exceptional therapeutic potential. Collectively, these multipronged strategies not only highlight the power of computation-driven approaches in drug discovery but also yield a diverse pipeline of resistance-targeting candidates, directly addressing the unmet clinical need in advanced PCa.
Xin Chai, Tingjun Hou, Dan Li· Accounts of Chemical Researc...· 0 citations
Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDock, the first DL-based docking framework specially designed for metalloprotein targets. By innovatively integrating an autoregressive spatial decoding engine with a physics-constrained geometric generation paradigm, MetalloDock can precisely reconstruct metal coordination geometries and accurately capture metal-ligand interactions, which enhance both the accuracy of metalloprotein-ligand docking and binding affinity prediction. Extensive evaluations on our custom-built benchmark data set demonstrate that MetalloDock outperforms existing methods, including AlphaFold3, in docking success rate and virtual screening performance for metalloprotein targets. In real-world applications, MetalloDock successfully identified multiple novel hit compounds in a virtual screening campaign targeting the prostate-specific membrane antigen. Additionally, it enabled rational drug design for acidic polymerase endonuclease, leading to the discovery of potent inhibitors. These results highlight the broad applicability of MetalloDock in accelerating metalloprotein-targeted drug discovery and provide a standardized framework for future evaluation of metalloprotein-specific docking algorithms.
Hui Zhang, Xujun Zhang, Qun Su et al.· Journal of the American Chem...· 5 citations
It is evidenced that many elaborately designed molecules that can interact well with the binding pocket of their target fail to exhibit activity in wet-lab experiments. This may associate with the interacting process of drug-target recognition. To efficiently characterize the drug-target interacting process, various enhanced sampling technologies have been proposed; yet, very few studies have systemically investigated whether the settings of these simulations are favorable to characterize the purposed tasks. Here, by comparing two popular enhanced sampling technologies, namely, the well-temped metadynamics and random acceleration molecular dynamics (RAMD), we systemically investigate the strategies to efficiently characterize the dissociating process of protein-ligand interactions. Two target families are employed for the analysis, including the kinase family (represented by TRK1) that represents the interaction-pathway obvious systems and the nuclear receptor family (represented by THRβ) that represents the interaction-pathway unobvious systems. Our results suggest that (1) in terms of maintaining stability of the protein structure, MetaD at various simulation conditions and RAMD with a large random force are good choice; (2) drug residence time derived from both MetaD and RAMD based on various parameters shows reasonable correlation to the experimental binding strength of the ligands, but RAMD usually runs with much less simulation time; and (3) both enhanced sampling methods result in reasonably consistent pathway preference for the two target families. Taken together, it will be much time-saving to utilize RAMD with high random force for interaction pathway exploration for both the pathway obvious and unobvious systems if the protein keeps stable in the simulation; otherwise, MetaD with a high bias factor is proposed to balance the computational accuracy and efficiency for the exploration.
Zhiliang Jiang, Mingyun Shen, Zhe Wang et al.· Journal of Chemical Physics· 1 citation
Molecular glues, including protein degraders and protein-protein interaction (PPI) stabilizers, have emerged as a new paradigm of drug design for regulating interactions between biomacromolecules; yet it is still a challenge for rational design of molecular glues. KRAS, as a prevalent oncogenic driver, is notoriously difficult to target by traditional small molecular drugs due to its challenging binding surface and frequent mutations. Although the small molecular drug RMC7977 has been designed as a PPI stabilizer for stabilizing the inherently weak RAS-CYPA interaction, the precise molecular mechanism underlying its stabilization effect and selectivity difference requires a deeper understanding. To this end, we leverage an integrated computational strategy combining molecular dynamics (MD) simulation, end-point binding free-energy calculation, and enhanced sampling technologies to elucidate the dynamic characteristics of RAS-ligand-CYPA interactions. Our result exhibits a high correlation between the predicted binding affinities and the experimental observations, demonstrating that RMC7977, acting as a strong PPI stabilizer, significantly enhances the stability of the KRAS-CYPA interaction, where, by delicately remodeling the protein-protein interface, the drug optimizes various interactions. Moreover, the results also uncover the dynamic process of stabilizer-mediated KRAS-CYPA stabilization and the mechanistic origin of the binding selectivity. This study provides essential molecular-level insights into RMC7977's function and offers a valuable computational framework for evaluating the stabilization effect of ligands targeting the KRAS-CYPA and other challenging PPI systems.
Kexin Xu, Mingyun Shen, Zhe Wang et al.· Journal of Chemical Informat...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.
Shi Li, Hongyan Du, Xujun Zhang et al.· Accounts of Chemical Researc...· 4 citations
Structure-based machine learning algorithms have been utilized to predict the properties of protein-protein interaction (PPI) complexes, such as binding affinity, which is critical for understanding biological mechanisms and disease treatments. While most existing algorithms represent PPI complex graph structures at the atom-scale or residue-scale, these representations can be computationally expensive or may not sufficiently integrate finer chemical-plausible interaction details for improving predictions. Here, we introduce MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently. This framework maps proteins onto a concise CG-scale complex graph, where nodes represent CG beads and edges encode chemically plausible interactions. The GNN-based encoder is tailored to extract high-quality representations from this graph, efficiently capturing the overall properties of the protein complex structure. Extensive experiments on three different downstream PPI property prediction tasks demonstrate that MCGLPPI achieves competitive performance compared with the counterparts at the atom- and residue-scale, but with only a third of the computational resource consumption. Furthermore, the CG-scale pre-training on protein domain-domain interaction structures enhances its predictive capabilities for PPI tasks. MCGLPPI offers an effective and efficient solution for PPI overall property predictions, serving as a promising tool for the large-scale analysis of biomolecular interactions.
Yang Yue, Shu Li, Yihua Cheng et al.· bioRxiv· 14 citations
Designing effective mRNA sequences for therapeutics remains a formidable challenge. Inspired by successes in protein design, language models (LMs) are now being applied to RNA, but progress is often impeded by the lack of comprehensive training data. Existing models are frequently limited to UTR or CDS regions, restricting their application for complete mRNA sequences. We introduce mRNABERT, a robust, all-in-one mRNA designer pre-trained on the largest available mRNA dataset. To enhance performance, we propose a dual tokenization scheme with a cross-modality contrastive learning framework to integrate semantic information from protein sequences. On a comprehensive benchmark, mRNABERT demonstrates state-of-the-art performance, outperforming previous models in the majority of tasks for 5’ UTR and CDS design, RNA-binding protein (RBP) site prediction, and full-length mRNA property prediction. It also surpasses large protein models in several related tasks. In conclusion, mRNABERT’s superior performance across these diverse tasks signifies a substantial leap forward in mRNA research and therapeutic development. Designing complete mRNA sequences for new vaccines and therapies is a complex challenge. Here, the authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks.
Ying Xiong, Aowen Wang, Yu Kang et al.· Nature Communications· 22 citations· ⚡1
Proteolysis-targeting chimeras (PROTACs) present a potentially effective strategy against various diseases via selective proteolysis. How to increase the efficacy of PROTACs remains challenging. Here, we explore the necessity of the linker, which has been deemed as an integral part of heterobifunctional PROTACs. Adopting single amino acid-based degradation signals, we find that the linker is not a required feature of the PROTACs. Notably, the linker-free PROTAC, Pro-BA, exhibits superior efficacy over its linker-bearing counterparts in degrading EML4-ALK and inhibiting lung cancer cell growth, as Pro-BA induces a stronger interaction between the target and the E3 ubiquitin ligase. Pro-BA is a water-soluble, orally administered degrader that significantly inhibits the tumor growth in a xenograft mouse model. The broad applicability of this linker-free PROTAC strategy is further validated through the development of BCR-ABL degrader. Our study introduces a design paradigm for PROTACs, potentially facilitating the advancement of more efficient therapeutic degraders. Linkers are traditionally seen as important for PROTAC activity. Here, the authors demonstrate that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.
Antibodies are crucial for medical applications, yet traditional methods for designing sequences are inefficient. This study introduces AntiBMPNN, an advanced deep‐learning framework that leverages an antibody‐specific 3D dataset, a fine‐tuned message‐passing neural network (MPNN), a frequency‐based scoring function, and AlphaFold 3 to achieve highly accurate antibody sequence design. AntiBMPNN surpasses ProteinMPNN with a perplexity of 1.5 and over 80% sequence recovery. Its scoring function, combined with AlphaFold 3, effectively prioritizes sequences based on structural recovery, positional stability, and biochemical or complex properties. Experimental validation highlights a 75% success rate in single‐point antibody design. AntiBMPNN consistently outperforms AbMPNN, AntiFold, and ProteinMPNN in designing complementarity determining regions (CDR) 1‐3, yielding stronger binding affinities. For CDR1 of huJ3 (anti‐HIV nanobody), it achieves a half maximal effective concentration (EC₅₀) of 9.2 nM (nanomolar), better than ProteinMPNN (135.2 nM) and AntiFold (59.3 nM), and comparable to AbMPNN (6.6 nM). For CDR2 of the D6 nanobody (targeting CD16), AntiBMPNN reaches 0.3 nM, outperforming AbMPNN (2.3 nM), AntiFold (0.7 nM), and ProteinMPNN (0.7 nM). In CDR3 of huJ3, it achieves 1.7 nM, surpassing AbMPNN (51.2 nM), with no detectable activity from AntiFold or ProteinMPNN. These findings confirm that AntiBMPNN‐designed sequences for J3 and D6 outperform the originals, highlighting its potential to improve therapeutic antibody design.
Ze-Yu Sun, Jiayi Yuan, Divya Jaiswal et al.· Advancement of science· 9 citations
The discovery of CAP-Gly domain-containing linker protein 1(CLIP1)-Leukocyte tyrosine kinase (LTK) as an oncogenic fusion reveals a unique dependency not only on LTK kinase activity but also on CLIP1-mediated multimerization, a noncatalytic function that drives oncogenic signaling. While this fusion is currently targeted with anaplastic lymphoma kinase inhibitors, their exclusive focus on kinase inhibition leaves the scaffolding function intact, necessitating a complete protein clearance strategy. Here, we report the AI-guided development of a first-in-class proteolysis-targeting chimera (PROTAC) designed to selectively degrade the CLIP1-LTK fusion protein. By integrating deep learning models for ternary complex prediction with structure-based molecular optimization, we designed DCL05, an orally bioavailable degrader of CLIP1-LTK fusion protein, achieving picomolar degradation potency (DC50 = 40 pM) and robust antitumor activity. DCL05 consistently outperformed existing kinase inhibitors across a broad spectrum of LTK resistance-associated mutations, both in vitro and in vivo. Collectively, our study explores resistance-associated contexts of LTK and establishes a structure-guided PROTAC development pipeline, providing a promising therapeutic strategy for overcoming acquired resistance in kinase-driven cancers.
Shicheng Chen, Haiting Duan, S. Zhong et al.· Proceedings of the National...· 0 citations
Enzymatic reactions play an emerging role in a broad spectrum of scientific and industrial applications. The inherent complexity of enzymes, such as their substrate specificity, conformational flexibility, and the vast diversity of reactions involved, poses substantial challenges for the advanced computational prediction of enzymatic reactions with desirable accuracy. Moreover, existing approaches are mostly tailored for a specific sub-task, such as substrate prediction or binding site annotation, which limits their applicability. In this study, we introduce ERAM, a task-agnostic multimodal learning framework capable of addressing a broad range of downstream applications with both accuracy and efficiency. ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data. In enzyme retrieval tasks, ERAM achieves an improvement of 28.31% in mean average precision compared with the state-of-the-art (SOTA) method, CREEP. In substrate prediction tasks, ERAM outperforms the SOTA method ESP, achieving average improvements of 35.53% and 22.97% in Matthews correlation coefficient across two datasets. Additionally, ERAM exhibits commendable interpretability by assigning higher attention weights to binding sites, resulting in lower false-positive rates (42.36%) and higher overlap scores (70.59%) in the unsupervised binding site prediction task compared to RXNAA Mapper. By learning embeddings of substrates, enzymes, and products within a unified knowledge graph latent space, ERAM demonstrates its potential as a versatile and effective tool for enzyme catalysis research.
Multi-target drugs hold great promise for treating complex diseases, yet existing methodologies predominantly rely on ligand-based approaches, which lack sufficient biological context and are often confined to specific target pairs, resulting in limited generalizability. Here, we introduce LaMGen, a general-purpose multi-target drug design framework powered by large language models (LLMs). Built on MTD2025, a dataset comprising over 600,000 quantum-accurate molecular conformations and 700,000 multi-target associations, LaMGen directly yields energy-favorable conformations with quantum-level accuracy. The framework integrates ESM-C protein embeddings, rotation-aware ligand tokens, and a TriCoupleAttention module to capture multi-level target–ligand interactions. Across independent benchmarks, LaMGen outperforms diffusion-based model across multiple properties, generating molecules in an average of 0.44 s, while preserving high conformational plausibility. Retrospective analyses demonstrate that LaMGen not only can reproduce molecules identical to known actives, but also consistently produces structurally novel candidates with conserved core scaffolds and superior binding affinities. Designing effective multi-target therapeutics remains a major challenge, as existing ligand- or protein-centric methods struggle to generate biologically contextualized, spatially valid 3D molecules, particularly for triple-target systems. This study introduces LaMGen, an LLM-powered framework that leverages large-scale protein-ligand data and rotation-aware molecular encoding to rapidly produce chemically plausible multi-target candidates, achieving strong zero-shot generalization, superior molecular quality, and robust performance across dual- and triple-target design tasks.
Qun Su, Qiaolin Gou, Hui Zhang et al.· Nature Communications· 1 citation
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.