Jul 2026· Protein Science· Vol 35· 0 citations· 51 references
Medicine
Abstract
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
It is concluded that molecular dynamics has an important place in improving the physicality of existing protein structure prediction paradigms, leading to the development of the Subspace Relaxation Operator (SRO).
Colin Baker, Pranav Mahableshwarkar, Ritambhara Singh et al.· 0 citations
Protein sequences are constrained not only by the need to fold into stable structures, but also by specific functional requirements imposed by natural selection. Yet predictions of how amino-acid changes affect proteins typically collapse these constraints into a single scalar score. Quantitatively separating these effects at scale remains an open challenge, with direct relevance spanning protein design to understanding the molecular mechanisms of disease. Inverse-folding (IF) models have emerged as fast, unsupervised predictors of folding energy changes (ΔΔG), but because they learn statistical correspondences between structure and sequence, they can conflate conservation driven by function with conservation driven by stability. Here, we show that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites. Applying the best-performing blend together with an evolutionary language model, we decompose each variant’s evolutionary cost into folding energy and dark energy, the latter capturing functional constraints beyond folding stability. With this decomposition, and without the need for supervision, we find that disease gain-of-function variants show a distinct functional signature from loss-of-function variants. In particular, we identify oncogenic drivers as largely preserving stability while exhibiting high dark energy, as opposed to tumor suppressors which are predominantly destabilized, paving the way to a mechanistic understanding of driver mutations in cancer. Together, these results provide a scalable framework for accurate ΔΔG prediction and mechanistic disentanglement of variant effects.
Ezequiel A. Galpern, Xavier Soler Sanchis, Charles W. J. Pugh et al.· bioRxiv· 0 citations
Recent advances in de novo protein design have enabled the generation of diverse novel proteins. However, a fundamental challenge remains: even when an amino acid sequence is designed with the target structure as the most stable conformation, there is currently no reliable computational method for assessing whether the target structure is sufficiently stabilized relative to alternative conformations. While experimental realization of the intended fold requires the target structure to be thermodynamically favored by a large free‐energy gap, the absence of a quantitative measure of folding stability makes it difficult to distinguish reliable from unreliable designs. Here, we propose the Residue‐attributed Interpretable Neural network for predicting Absolute folding free energy by Merging structure and sequence Information (RINAMI), a machine learning model that predicts the absolute folding free energy (ΔG) of proteins from their three‐dimensional structures and amino acid sequences. RINAMI integrates structure‐ and sequence‐based representations derived from ProteinMPNN and Evolutionary Scale Modeling 2 (ESM2) using a multi‐head cross‐attention mechanism that contextualizes sequence‐derived signals within the structural environment. Benchmarking RINAMI on both natural and designed proteins from the Mega‐scale and Maxwell datasets shows that it outperforms the tested existing approaches, achieving higher correlations with experimental measurements and improved or comparable prediction errors. An ablation study supports the contribution of sequence–structure integration for predictive accuracy. In addition, RINAMI exhibits strong interpretability by capturing key physicochemical effects, including the destabilizing effect of buried hydrophilic residues, the stabilizing effect of buried hydrophobic residues, and the characteristics of cysteine. Together, these results establish RINAMI as an accurate and interpretable framework for ΔG prediction and provide a practical computational tool for evaluating and prioritizing protein designs prior to experimental testing.
Naoki Tomita, G. Chikenji· Protein Science· 0 citations
An improved force field is developed, derived from its parent, Amber ff24EXP-GA, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins.
This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.
Wengan He, Yongsheng Luo, Lihong Jiang et al.· 0 citations
It is argued that incorporating frustration into computational and experimental strategies will be essential to move beyond purely stability-driven approaches toward the rational engineering of functional proteins.
Franco L. Simonetti, Eli J. Draizen, Rocío Espada et al.· Biochimica et Biophysica Act...· 0 citations