2025· Advances in Neural Information Processing Systems 38· 0 citations· 61 references
TL;DR
This work proposes the first Bayesian flow formulation for protein backbone orientations by recasting orientation modeling as an equivalent hyperspherical generation problem with antipodal symmetry and delivers consistently exceptional performance in both peptide and antibody design tasks.
Abstract
Designing functional proteins is a critical yet challenging problem due to the intricate interplay between backbone structures, sequences, and side-chains. Current approaches often decompose protein design into separate tasks, which can lead to accumulated errors, while recent efforts increasingly focus on all-atom protein design. However, we observe that existing all-atom generation approaches suffering from an information shortcut issue, where models inadvertently infer sequences from side-chain information, compromising their ability to accurately learn sequence distributions. To address this, we introduce a novel rationalized information flow strategy to eliminate the information shortcut. Furthermore, motivated by the advantages of Bayesian flows over differential equation–based methods, we propose the first Bayesian flow formulation for protein backbone orientations by recasting orientation modeling as an equivalent hyperspherical generation problem with antipodal symmetry. To validate, our method delivers consistently exceptional performance in both peptide and antibody design tasks. Our code, checkpoint, and designed PDBs can be found in https://github.com/GenSI-THUAIR/ProBayes .
This work introduces ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics.
D. P. Kulathunga, Divyanshu Shukla, D. Potoyan· bioRxiv· 0 citations
This review focuses on coordinate- and residue-frame-based diffusion approaches for generating protein structures, paying particular attention to geometric equivariance, conditioning strategies, all-atom modelling and interaction-aware design.
Wen-Ran Li, Xavier F. Cadet, David Medina-Ortiz et al.· International Journal of Mol...· 0 citations
This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.
Guodong Min, Huan Peng· Methods in molecular biology· 0 citations
Protein switches are proteins that can respond to biochemical stimuli by rearranging their structural elements, essential for cells to transduce signals. The de novo design of such proteins requires amino acid sequences whose energy landscapes support multiple stimulus-dependent conformations, yet most current de novo protein design pipelines are optimized for single stable structures. Existing multi-state inverse-folding methods can design sequences compatible with multiple backbones, but they assume that suitable backbone ensembles are already available, often requiring expert knowledge. We introduce Diff-Switch, a framework for sampling switch-like backbone ensembles from pretrained protein diffusion models. Given a reference backbone structure and domain decomposition, our method preserves local domain geometry while encouraging diversity in global domain arrangements along user-specified collective variables, inspired by metadynamics. We implement this objective through a controlled diffusion sampler with reward-tilting for local similarity between the ensemble members and history-dependent bias in collective-variable space to avoid repeated sampling of the same global arrangement. The resulting ensembles provide candidate conformational states for downstream multi-state inverse folding. Across our evaluation set of 20 diverse proteins, using conformations from these generated ensembles improves the success rate of finding switch-compatible sequences over baseline sampling. We further apply the method to a real-world protein switch design task and characterize the resulting designs.
Alireza Omidi, Jiajun He, Jennifer M. Bui et al.· bioRxiv· 0 citations
TTS-Design is proposed, a test-time compute scaling framework that enhances protein sequence design without retraining models or relying on larger training datasets, and can consistently improve sequence recovery and structural reliability across different backbone models, without retraining or increasing model size.
Zizhe Jin, Yi Zheng, Huan Yee Koh et al.· 0 citations
It is demonstrated that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools, and shown that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions.
G. Rajagopal, Søren C. Spina, Joe Bailey et al.· bioRxiv· 0 citations