Skip to content

Category

protein folding

562 papers

Precision Design of Fluorogenic Probes via Orthogonal Tuning of Binding and Photophysics for Isoform-Selective ALDH2 Imaging.

Fluorogenic probes that report enzyme activity are essential for studying biological functions. However, designing them for targets with low catalytic turnover and narrow substrate specificity remains a significant challenge. Here, we present a precision design framework that separates the requirements for sensitivity and selectivity by integrating molecular docking, quantum chemical modeling of fluorogenic mechanisms, and targeted fine-tuning of the probe structures. As a proof of concept, we developed A5, a fluorogenic substrate for aldehyde dehydrogenase 2 (ALDH2) that exhibits high isoform selectivity and a >240-fold signal enhancement over the standard NADH assay. A5 enables quantitative imaging of ALDH2 activity across multiple biological scales─in blood samples, live cells, and intact mouse brains─and supports the identification of small-molecule activators with therapeutic potential in an Alzheimer's disease model. This work establishes a modular strategy for creating activity-based probes tailored to challenging enzymatic targets, with broad applications in precision imaging, drug discovery, and mechanistic biochemistry.

Rongrong Tao, Yu Chen, Taorui Yang et al. · 4 citations
#protein folding Jun 2025

Unraveling the Efficacy of AR Antagonists Bearing N-(4-(Benzyloxy)phenyl)piperidine-1-sulfonamide Scaffold in Prostate Cancer Therapy by Targeting LBP Mutations.

Point mutations in the androgen receptor (AR) are significant drivers of resistance in prostate cancer (PCa), posing a great challenge to the development of effective treatment strategies. Building on our previous discovery of the suboptimal AR antagonist T1-12, we developed LT16, which contains an N-(4-(benzyloxy)phenyl)piperidine-1-sulfonamide scaffold through structural optimization and comprehensive screening against T878A-mutated AR. LT16 outperformed existing antiandrogens by fully antagonizing clinical AR mutations and effectively suppressing castration- and enzalutamide-resistant LNCaP cells proliferation in vitro. Mechanically, LT16 was found to disrupt AR nuclear translocation, hinder AR homodimerization, and suppress transcription of AR-regulated genes by competitive binding to the ligand binding pocket. Further in vivo experiments demonstrated that LT16 significantly reduced both regular- and enzalutamide-resistant LNCaP tumor volume and serum prostate-specific antigen levels in mice. These findings position LT16 as a promising and innovative therapeutic for advanced PCa, particularly in cases where resistance to current therapies is a concern.

Xin Chai, Xinyue Wang, Lvtao Cai et al. · 2 citations

PepBAN: A Deep Learning Framework with Bilinear Attention and Adversarial Learning for Peptide-Protein Interaction Prediction

Accurate prediction of the peptide-protein interaction (PepPI) is crucial for developing peptide-based therapeutics and vaccines. However, this computational task has traditionally faced significant challenges, such as the scarcity of structure data along with the corresponding label of the binding affinity for bound complexes. To address these challenges, we introduce PepBAN, a deep learning framework for modeling PepPI predictions. PepBAN incorporates two technical advancements: (1) adopting the protein language model ESM-2 to characterize proteins and ESM-2 or a graph-based foundation model for peptides without structure data and (2) leveraging the conditional domain adversarial learning to enhance generalization across a broad range of protein targets, especially when there are limited binding data. At the core of PepBAN is a bilinear attention network (BAN) that effectively learns the pattern of pairwise local interactions, enables the identification of key residues participating in the peptide-protein interactions, and offers an intuitive approach to interpret the underlying mechanisms of PepPIs via analyzing attention weights. Our numerical experiments demonstrated that PepBAN outperformed the previous state-of-the-art models across several well-established benchmark studies. Furthermore, we evaluated PepBAN's applicability in predicting cyclic peptide-protein interactions, a task that poses significant challenges due to the presence of noncanonical amino acids. These nonstandard residues require specialized handling, which most existing sequence-based PepPI prediction models did not adequately address, and we adopt an atom-resolved molecular graph approach to process cyclic peptides. Despite this complexity, PepBAN demonstrated a clear advantage by achieving a superior prediction performance and offering a distinct edge in tackling the emerging chemical space of cyclic peptides, which has great potential for novel therapeutic development. In summary, PepBAN serves as a valuable tool for advancing peptide-based drug and therapeutic development.

Shuaiyan Li, Xiaorui Wang, Yuchen Zhu et al. · 2 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

Accurate modeling of protein-peptide interactions is essential for understanding fundamental biological processes and designing peptide-based drugs. However, predicting the complex structures of these interactions remains challenging, primarily due to the high conformational flexibility of peptides. To support a fair and systematic evaluation of recent deep learning (DL) approaches, we introduce PepPCBench, a benchmarking framework tailored to assess protein folding neural networks (PFNNs) in protein-peptide complex prediction. As part of this framework, we curated PepPCSet, a data set of 261 experimentally resolved complexes with peptides ranging from 5 to 30 residues. We benchmark five full-atom PFNNs, including AlphaFold3 (AF3), AlphaFold-Multimer (AFM), Chai-1, HelixFold3 (HF3), and RoseTTAFold-All-Atom (RFAA), using comprehensive evaluation metrics. Our benchmarking reveals meaningful performance differences among these methods and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy. While AF3 shows strong performance in structure prediction, further analysis indicates that confidence metrics correlate poorly with experimental binding affinities, underscoring the need for improved scoring strategies and generalizability. By providing a reproducible and extensible framework, PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction.

Silong Zhai, Huifeng Zhao, Jike Wang et al. · 13 citations · ⚡1

CarsiDock-Cov: A deep learning-guided approach for automated covalent docking and screening

The interest in covalent drugs has resurged in recent decades, spurring the development of numerous specialized computational docking tools to facilitate covalent ligand design and screening. Herein, we present CarsiDock-Cov, a new paradigm distinguishing itself as the first deep learning (DL)-guided approach for covalent docking. CarsiDock-Cov retains the core components of its non-covalent predecessor, leveraging a DL model pretrained on millions of docking complexes to predict protein–ligand distance matrices, along with a dedicated-designed geometric optimization procedure to convert these distances into refined binding poses. Additionally, it incorporates several key enhancements specifically tailored to optimize the protocol for covalent docking applications. Our approach has been extensively validated on multiple public datasets regarding the docking and screening of covalent ligands, and the results indicate that our approach not only achieves comparably improved applicability compared to its non-covalent predecessor, but also exhibits competitive performance against various state-of-the-art covalent docking tools. Collectively, our approach represents a significant advance in covalent docking methodology, offering an automated and efficient solution that shows considerable promise for accelerating covalent drug discovery and design.

Chao Shen, Hongyan Du, Xujun Zhang et al. · 12 citations

EpiMII: Integrating Structure and Graph Neural Networks for MHC-II Epitope and Neoantigen Design

MHC-II neoantigens play a critical role in immunotherapy, either as direct effectors or through their influence on CD8+ T cells. However, only a small fraction of tumor DNA mutations qualify as functional neoantigens, and current prediction tools often lack accuracy, leading to the low immunogenicity of predicted neoantigens in vivo. Here, we present EpiMII, a Graph Neural Network model for MHC-II epitope design, which learns from the structural features of epitopes to predict their sequences. To train EpiMII, we constructed a reliable, large dataset containing 142,934 MHC-II epitope structures. This approach achieves a 4.2x improvement over ProteinMPNN, with a sequence recovery rate of 78.0% for known MHC-II epitopes in the Protein Data Bank. As a case study, we designed a neoantigen from hepatocellular carcinoma. All five designed epitopes significantly activated CD4+ T cells in vitro and induced secretion of IFN-γ and TNF-α. Notably, one epitope treatment significantly reduced tumor volume in mice in vivo. EpiMII offers a novel and efficient approach for identifying MHC-II epitopes/neoantigens, potentially contributing to vaccine development.

Jiayi Yuan, Xiaowei Xu, Ze-Yu Sun et al. · 1 citation

One-Shot Rational Design of Covalent Drugs with CovalentLab

Targeted covalent drugs have demonstrated remarkable potential in disease treatment over the past decades. However, existing methods for covalent drug design are often limited to serine and cysteine, ignoring other potentially ligandable binding sites. Statistical analyses indicate that over 95% of binding pockets contain covalent-binding residues, suggesting that all ligands that targeting these pockets possess the potential to be modified into covalent ligands. To achieve this goal, we introduced CovalentLab, an interactive computational platform that integrates ligand-based and warhead-based strategies into a unified workflow for the rational design of covalent ligands. Leveraging a covalent binding site prediction model constructed on ESM-2 with LoRA fine-tuning, CovalentLab enables the prediction and ranking of nine classes of covalent-binding residues in proteins according to their reactivity and facilitates systematic warhead attachment to ligands using 210 electrophilic groups or user-defined warheads. Using this platform, a comprehensive library of more than 100,000 covalent molecules across 95 targets was generated. Notably, CovalentLab has been successfully applied to various essential real-world targets, identifying wet-laboratory-validated bioactive compounds ranging from TRK orthosteric inhibitors to GAC allosteric inhibitors. By bridging gaps in covalent drug discovery, CovalentLab offers a versatile, publicly accessible resource to expand the druggable targets and accelerate the development of targeted covalent therapies.

Xi Xue, Xiangying Liu, Xue Liu et al. · 4 citations · ⚡1
#computer vision May 2025

A Unified Deep Graph Model for Identifying the Molecular Categories of Ligands Targeting Nuclear Receptors

To fulfill functions for differentially regulating the downstream signaling pathways, functional ligands (i.e., agonists or antagonists) targeting nuclear receptors (NRs) are designed to stabilize different conformations (active or inactive) of the proteins. However, in practical applications, it is usually difficult to determine the molecular category of an NR ligand because these molecules all bind in the same location of an NR protein, namely, the ligand-binding pocket (LBP). Considering that ligands with different properties (agonists or antagonists) prefer to bind with differential conformations of NRs, it is possible to identify the molecular type of a given ligand through the differential binding environment (active or inactive conformations) of the protein-ligand interaction. Therefore, in this study, we established a unified model (NRIGN) based on the deep graphic architecture to discriminate agonists and antagonists targeting 26 successful or in-clinical-trial NR targets. Our result shows that NRIGN achieves an excellent prediction accuracy (ACC >0.95) and is robust enough to be applied in various real-world scenarios, such as predicting the molecular type of ligands in crystallized NR structures, ligands with multiple NR activities, and ligands with their types altered by target mutations. The proposed model is expected to promote rational design of drugs targeting NR proteins.

Kaimo Yang, Dejun Jiang, Qirui Deng et al. · 2 citations
#computer vision Nov 2025

Improving the predictive performance of binding affinities and poses for protein–cyclic peptide complexes through fine-tuned MM/PBSA(GBSA)-based methods

Abstract Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein–cyclic peptide (P–cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson–Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P–cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P–cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P–cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P–cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P–cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of −0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = −0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P–cp interactions.

Huifeng Zhao, Jianxiang Huang, Gaoqi Weng et al. · 8 citations

Revisiting Protein-Protein Docking: A Systematic Evaluation Framework

Protein-protein interactions play pivotal roles in a wide range of biological processes. Determining the atomic-level structures of protein-protein complexes is indispensable for elucidating macromolecular interaction mechanisms and advancing structure-based drug design. Protein-protein docking, as one of the leading computational approaches for predicting complex structures, has seen considerable progress but requires rigorous evaluation in practical applications. In this study, we proposed a comprehensive benchmarking framework to evaluate 11 docking methods spanning traditional (HDOCK, PatchDock, PIPER, ZDOCK) and deep learning (DL)-based (EquiDock, ElliDock, EBMDock, GeoDock, DiffDock-PP, AlphaFold-Multimer, AlphaFold3) approaches. Our framework incorporates the classical DockingBenchmark 5.5 data set for evaluating flexible docking, introduces a newly curated data set (AACBench) for antibody-antigen complex docking, and establishes the PPCBench data set to examine the out-of-distribution (OOD) generalization capabilities of DL-based methods. In docking against apo structures, AlphaFold3 achieves a superior top-5 success rate of 77.98%, whereas the traditional approach HDOCK reaches merely 12.84%, despite its highest top-5 success rate of 85.24% when docking against holo structures. For antibody-antigen docking, AlphaFold3 remains the most accurate method (top-5 success rate: 31.78%) and substantially outperforms AlphaFold-Multimer in modeling the CDR-H3 loop. In OOD generalization tests, all DL-based models exhibit markedly reduced performance on the PPCBench data set. Overall, our work establishes a unified benchmarking framework that enables systematic evaluation of docking methods across diverse tasks and provides critical insights into the strengths and limitations of current docking strategies, thereby informing future developments in protein-protein docking research.

Linlong Jiang, Ke Zhang, Kai Zhu et al. · 3 citations

Macrocycle-DB: a comprehensive database for macrocycle-based drug discovery

Abstract Macrocycles have gained significant attention in drug design owing to their distinctive structural and physicochemical features. Despite the abundance of available experimental data, there remains a need for a centralized resource to support macrocycle-based drug discovery. Here, we present Macrocycle-DB, the most extensive online database dedicated to macrocycles, featuring 45 525 compounds, including 76 approved drugs and 105 clinical candidates that target 2533 proteins. The database offers comprehensive structural information, experimental bioactivity data, and physicochemical properties for each macrocycle, along with co-crystal structures to visualize protein–ligand interactions. Additionally, Macrocycle-DB provides specialized descriptors, scaffold and linker details for synthetic macrocycles, and high-quality downloadable datasets to facilitate computational drug design. Macrocycle-DB is freely accessible at https://macro-db.dpbio.tech/ and https://macro-db.cn/.

Minchuan Jiang, Tianyu Liu, Muzammal Hussain et al. · 6 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.