Skip to content
Open access

Aligning protein-generative models to experimental fitness with ProteinDPO.

Aug 2026 · Nature Methods · 0 citations · 47 references
Medicine

TL;DR

This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.

Abstract

Biological generative models can predict biological functions without task-specific training data but often under-perform specialized models. This is due to a fundamental 'alignment gap', where the rules learned during unsupervised training are not related to the function of interest. Here we demonstrate how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences. Our aligned model, ProteinDPO, achieves stability prediction competitive to task-specific models and consistently outperforms unsupervised and fine-tuned versions of the model. Notably, ProteinDPO generalizes beyond its training data to enable stabilization and improved binding affinity prediction of large multichain protein complexes. When applied to stabilization of the hemagglutinin trimer, a primary component of influenza vaccines, ~80% of designs achieve increased or similar stability compared with the native hemagglutinin and up to 32 °C improvements from recently emerged mammalian strains. Our results demonstrate how to augment generative models with biophysical information and, more broadly, provide a general framework for the alignment of biological foundation models.

Read PDF

Similar papers

Preprint Jul 2026

Exploring the Alignment of Generation and Understanding in Protein Structure Modeling

Understanding and generation are often treated as two separate paradigms in training deep neural networks, despite the fact that both are trained with closely related objectives such as denoising and masked prediction. While prior studies have shown that generative models often learn suboptimal representations for understanding tasks in vision, it is less understood whether a similar gap exists in the protein domain. In this work, we systematically investigate this question by benchmarking state-of-the-art protein generative models on widely-used protein understanding tasks, and observe that these models exhibit consistently poor performance compared to existing protein encoders. Furthermore, inspired by the Representation Alignment (REPA) framework, we propose to explicitly align generative protein diffusion models with pretrained protein understanding models during training. Experiments on the MotifBench demonstrate that representation alignment significantly improves functional protein generation, boosting the MotifBench score of Protpardelle-1c from 39.2 to 47.1, corresponding to a 20% relative improvement. Our results suggest that representation alignment provides a general and effective mechanism for bridging understanding and generation in protein structure modeling.

Junde Xu, Yuansheng Huang, Zijun Gao et al. · 0 citations
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 2 citations
Open access Aug 2026

MAXWELL: Calibrating the probabilistic outputs of protein language models to the mutation-induced stability change landscape

MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability, is introduced.

Mingchen Li, Xiaoran Cheng, Fan Jiang et al. · 0 citations
Open access Jul 2026

ProteinDock: A physics-informed layer to improve protein-protein docking reliability

It is demonstrated that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools, and shown that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions.

G. Rajagopal, Søren C. Spina, Joe Bailey et al. · 0 citations