Targeted covalent drugs have demonstrated remarkable potential in disease treatment over the past decades. However, existing methods for covalent drug design are often limited to serine and cysteine, ignoring other potentially ligandable binding sites. Statistical analyses indicate that over 95% of binding pockets contain covalent-binding residues, suggesting that all ligands that targeting these pockets possess the potential to be modified into covalent ligands. To achieve this goal, we introduced CovalentLab, an interactive computational platform that integrates ligand-based and warhead-based strategies into a unified workflow for the rational design of covalent ligands. Leveraging a covalent binding site prediction model constructed on ESM-2 with LoRA fine-tuning, CovalentLab enables the prediction and ranking of nine classes of covalent-binding residues in proteins according to their reactivity and facilitates systematic warhead attachment to ligands using 210 electrophilic groups or user-defined warheads. Using this platform, a comprehensive library of more than 100,000 covalent molecules across 95 targets was generated. Notably, CovalentLab has been successfully applied to various essential real-world targets, identifying wet-laboratory-validated bioactive compounds ranging from TRK orthosteric inhibitors to GAC allosteric inhibitors. By bridging gaps in covalent drug discovery, CovalentLab offers a versatile, publicly accessible resource to expand the druggable targets and accelerate the development of targeted covalent therapies.
Xi Xue, Xiangying Liu, Xue Liu et al.· JACS Au· 4 citations· ⚡1
Recently, various self‐supervised learning (SSL) methods based on 3D graph neural networks (GNNs) have been developed to comprehensively represent the structural information of molecules in 3D space; this is essential for discovering new drugs. However, existing methods fail to comprehensively characterize the 3D structures of molecules and neglect the electronic structural information that significantly influences key properties such as molecular reactivity, strong electrostatic interactions, and chemical adsorption. Therefore, here, a novel molecular representation learning method is constructed, Q‐GEM, incorporating quantum and geometric structural information enhancement, based on the quantum chemical property database QuanDB and SSL methods. Q‐GEM comprises a GNN embedded with the molecular electronic and complete 3D geometrical structural information as well as several well‐designed multiscale SSL tasks, achieving superior absolute molecular conformation prediction and conformational discrimination. The Q‐GEM achieved state‐of‐the‐art performance in 12 out of 13 prediction tasks on the MoleculeNet dataset, with an average performance improvement of 3.3% and 2.0% for classification and regression prediction tasks, respectively. Moreover, an average performance improvement of 5.2% is achieved in three localized quantum chemical properties, fully demonstrating the excellent performance of Q‐GEM in distinguishing molecular electronic structures. The Q‐GEM represents a novel, powerful breakthrough for accurate molecular property prediction.
Zhijiang Yang, Liangliang Wang, Tengxin Huang et al.· Advancement of science· 5 citations
To fulfill functions for differentially regulating the downstream signaling pathways, functional ligands (i.e., agonists or antagonists) targeting nuclear receptors (NRs) are designed to stabilize different conformations (active or inactive) of the proteins. However, in practical applications, it is usually difficult to determine the molecular category of an NR ligand because these molecules all bind in the same location of an NR protein, namely, the ligand-binding pocket (LBP). Considering that ligands with different properties (agonists or antagonists) prefer to bind with differential conformations of NRs, it is possible to identify the molecular type of a given ligand through the differential binding environment (active or inactive conformations) of the protein-ligand interaction. Therefore, in this study, we established a unified model (NRIGN) based on the deep graphic architecture to discriminate agonists and antagonists targeting 26 successful or in-clinical-trial NR targets. Our result shows that NRIGN achieves an excellent prediction accuracy (ACC >0.95) and is robust enough to be applied in various real-world scenarios, such as predicting the molecular type of ligands in crystallized NR structures, ligands with multiple NR activities, and ligands with their types altered by target mutations. The proposed model is expected to promote rational design of drugs targeting NR proteins.
Kaimo Yang, Dejun Jiang, Qirui Deng et al.· Journal of Chemical Informat...· 2 citations
Photodynamic therapy (PDT) is a clinically approved therapeutic modality that has demonstrated significant potential for cancer treatment, and triplet photosensitizers (PSs) play a key role in its efficacy. Despite deep learning having emerged as a next-generation tool for material discovery, existing methods mainly target a limited subset of triplet PSs, such as thermally activated delayed fluorescence (TADF) materials, neglecting the critical intersystem crossing (ISC) between the high-lying singlet and triplet states (ΔESnTn). To overcome this limitation, we compiled a comprehensive dataset (∼1.90 × 109) of triplet PSs encompassing various ISC mechanisms. Then, we proposed a novel strategy that incorporates two models: a fragment-based model (Frag-MD) and a character-based model (MD), both integrating a conditional transformer, recurrent neural networks, and reinforcement learning. In silico experiments revealed that the Frag-MD model outperforms the MD model in generating larger conjugated motifs with higher average ring numbers and atom counts; while the MD model generates twice as many unique motifs and excels in novelty and diversity, as evaluated by conditional and MOSES metrics. Therefore, our approach is highly effective for modifying conjugated motifs and designing novel triplet PSs. Notably, the recently reported high-efficiency triplet PSs have been re-identified through ablation experiments using our proposed models, which target ΔESnTn and significantly outperform traditional baselines, achieving a prediction accuracy of 73% versus 4%. Our approach holds the potential to establish a new paradigm for discovering novel PSs applicable in PDT.
Kepeng Chen, Xiaoting Zhang, Jike Wang et al.· Chemical Science· 3 citations
Abstract Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein–cyclic peptide (P–cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson–Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P–cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P–cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P–cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P–cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P–cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of −0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = −0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P–cp interactions.
Protein-protein interactions play pivotal roles in a wide range of biological processes. Determining the atomic-level structures of protein-protein complexes is indispensable for elucidating macromolecular interaction mechanisms and advancing structure-based drug design. Protein-protein docking, as one of the leading computational approaches for predicting complex structures, has seen considerable progress but requires rigorous evaluation in practical applications. In this study, we proposed a comprehensive benchmarking framework to evaluate 11 docking methods spanning traditional (HDOCK, PatchDock, PIPER, ZDOCK) and deep learning (DL)-based (EquiDock, ElliDock, EBMDock, GeoDock, DiffDock-PP, AlphaFold-Multimer, AlphaFold3) approaches. Our framework incorporates the classical DockingBenchmark 5.5 data set for evaluating flexible docking, introduces a newly curated data set (AACBench) for antibody-antigen complex docking, and establishes the PPCBench data set to examine the out-of-distribution (OOD) generalization capabilities of DL-based methods. In docking against apo structures, AlphaFold3 achieves a superior top-5 success rate of 77.98%, whereas the traditional approach HDOCK reaches merely 12.84%, despite its highest top-5 success rate of 85.24% when docking against holo structures. For antibody-antigen docking, AlphaFold3 remains the most accurate method (top-5 success rate: 31.78%) and substantially outperforms AlphaFold-Multimer in modeling the CDR-H3 loop. In OOD generalization tests, all DL-based models exhibit markedly reduced performance on the PPCBench data set. Overall, our work establishes a unified benchmarking framework that enables systematic evaluation of docking methods across diverse tasks and provides critical insights into the strengths and limitations of current docking strategies, thereby informing future developments in protein-protein docking research.
Linlong Jiang, Ke Zhang, Kai Zhu et al.· Journal of Chemical Informat...· 3 citations
Abstract Macrocycles have gained significant attention in drug design owing to their distinctive structural and physicochemical features. Despite the abundance of available experimental data, there remains a need for a centralized resource to support macrocycle-based drug discovery. Here, we present Macrocycle-DB, the most extensive online database dedicated to macrocycles, featuring 45 525 compounds, including 76 approved drugs and 105 clinical candidates that target 2533 proteins. The database offers comprehensive structural information, experimental bioactivity data, and physicochemical properties for each macrocycle, along with co-crystal structures to visualize protein–ligand interactions. Additionally, Macrocycle-DB provides specialized descriptors, scaffold and linker details for synthetic macrocycles, and high-quality downloadable datasets to facilitate computational drug design. Macrocycle-DB is freely accessible at https://macro-db.dpbio.tech/ and https://macro-db.cn/.
ConspectusProstate cancer (PCa) is the most prevalent malignancy among men worldwide, with its pathogenesis and progression heavily reliant on the sustained activation of the androgen receptor (AR) signaling pathway. The AR, a transcription factor of nuclear receptor superfamily, serves as the most privileged therapeutic target in PCa, as evidenced by the clinical efficacy of first- and second-generation AR antagonists. Current clinically available AR antagonists exclusively target the ligand binding pocket (LBP), suppressing tumor proliferation through competitive inhibition of androgen binding and subsequent blockade of AR signaling transduction. However, their therapeutic utility is invariably limited by acquired resistance mechanisms, including point mutations that alter LBP specificity, AR gene amplification leading to receptor overexpression, and the emergence of constitutively active splice variants that bypass ligand-dependent activation. Thus, the development of novel AR antagonists featuring innovative mechanisms and structural scaffolds is imperative to overcome resistance to antiandrogen therapy. However, the AR exhibits significant structural flexibility, and the lack of antagonist-bound crystal structures has hindered structure-based rational drug design. In this Article, we summarize our advances in elucidating the molecular mechanisms underlying AR conformational regulation and highlight our progress in the structure-based design and development of novel AR antagonists. First, our molecular dynamic (MD) studies collectively elucidate the molecular mechanisms by which the AR ligand binding domain (LBD) regulates its functional states through dynamic conformational changes mediated by distinct allosteric pathways when bound to agonists or antagonists, providing atomic-level insights and structural basis for drug development. Then, we successfully identified structurally diverse lead compounds targeting the LBP through various integrated approaches combining MD simulations, structure-based virtual screening (SBVS), and systematic biological evaluation. These compounds exhibited potent activity against clinically relevant AR mutations F877L, W742C, T878A, and H875Y, demonstrating their potential to overcome mutations-driven resistance. Further, we explored non-LBP mediated strategies for AR antagonism, including: (1) targeting the allosteric binding sites on LBD; (2) identification of novel druggable binding sites; and (3) targeting alternative domains beyond the LBD. As a paradigm-shifting example, we proposed inhibition of AR LBD dimerization as a novel mechanism of action for LBP-targeting AR antagonists. Building upon this insight, we characterized a promising pocket at the dimer interface, designated the Dimerization Interface Pocket (DIP), and developed first-in-class antagonists specifically targeting this site, which exhibit exceptional therapeutic potential. Collectively, these multipronged strategies not only highlight the power of computation-driven approaches in drug discovery but also yield a diverse pipeline of resistance-targeting candidates, directly addressing the unmet clinical need in advanced PCa.
Xin Chai, Tingjun Hou, Dan Li· Accounts of Chemical Researc...· 0 citations
Voltage-gated sodium channels (VGSCs/Navs) are essential targets for the treatment of numerous neurological, muscular, and cardiac disorders. Despite the increasing clinical interest in subtype-selective modulators, current public databases provide fragmented and inconsistent information on VGSC-related compounds and targets, particularly lacking coverage on peptides. To address this limitation, we developed NavDB, a specialized and open-access database focusing on VGSC modulators and targets. NavDB integrates 8023 curated data records covering 5168 compounds, including small molecules, toxins, drugs, and peptides, along with comprehensive annotations on biological activity, druggability, and structural feature. NavDB also features advanced functions such as text-based and structure-based search, peptide similarity matching, and AI-powered property prediction. Moreover, the database offers high-quality 3D visualizations of targets and peptides, with disulfide bond and signal peptide annotations. All data are freely downloadable to support both experimental and computational drug discovery. NavDB is publicly available at: http://cadd.zju.edu.cn/navdb/.
Gaoang Wang, Jiahui Yu, Haiyi Chen et al.· Journal of Chemical Informat...· 0 citations
Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDock, the first DL-based docking framework specially designed for metalloprotein targets. By innovatively integrating an autoregressive spatial decoding engine with a physics-constrained geometric generation paradigm, MetalloDock can precisely reconstruct metal coordination geometries and accurately capture metal-ligand interactions, which enhance both the accuracy of metalloprotein-ligand docking and binding affinity prediction. Extensive evaluations on our custom-built benchmark data set demonstrate that MetalloDock outperforms existing methods, including AlphaFold3, in docking success rate and virtual screening performance for metalloprotein targets. In real-world applications, MetalloDock successfully identified multiple novel hit compounds in a virtual screening campaign targeting the prostate-specific membrane antigen. Additionally, it enabled rational drug design for acidic polymerase endonuclease, leading to the discovery of potent inhibitors. These results highlight the broad applicability of MetalloDock in accelerating metalloprotein-targeted drug discovery and provide a standardized framework for future evaluation of metalloprotein-specific docking algorithms.
Hui Zhang, Xujun Zhang, Qun Su et al.· Journal of the American Chem...· 5 citations