This work investigates the effectiveness of two different computational methods as triaging tools for prioritizing target positions and identifying specific mutations that are likely to improve protein thermodynamic stability and proposes a hybrid mutation prioritization and selection strategy that achieves better accuracy than either method alone.
Abstract
Directed evolution for protein engineering, as currently practiced in the biotechnology and pharmaceutical industries, is both tedious and expensive. Computationally driven protein design has the potential to expedite the engineering process and generate high-quality variants at a lower cost than traditional approaches. We investigated the effectiveness of two different computational methods as triaging tools for prioritizing target positions and identifying specific mutations that are likely to improve protein thermodynamic stability. Our benchmarking study used a comprehensive dataset consisting of 174,945 mutations across 180 distinct proteins and evaluated the ESM (Evolutionary Scale Modeling) protein language model alongside a physics-based method, MM/GBSA (Molecular Mechanics Generalized Born Surface Area). We found prediction biases in each method but also determined that these biases can be mitigated by applying the two methods in a complementary manner. We propose a hybrid mutation prioritization and selection strategy that achieves better accuracy than either method alone. Through re-ranking, the combined prioritization strategy attained a higher overall average ROC (receiver operating characteristic) AUC (area under curve) of 0.743 across the dataset compared to either MM/GBSA alone (0.685) or ESM Log Odds alone (0.597). The integrated framework can be adapted and applied to newer AI and physics-based models as the field advances.
MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.
J. Koo, Young-Ho Park, Sun-Uk Kim· Signal Transduction and Targ...· 0 citations
Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.
Daniel Siegismund, Mario Wieser, E. Natali et al.· bioRxiv· 0 citations
MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability, is introduced.
Mingchen Li, Xiaoran Cheng, Fan Jiang et al.· bioRxiv· 0 citations
UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.
Hong Tan, Shenggeng Lin, Yi Xiong· Chemical Science· 0 citations
EvoMOBO is established as a modular framework for multi-objective protein engineering using experimental or mechanism-derived labels using simulation-derived mechanistic descriptors, with experiments reserved for final validation.
Kai Wen, Sirui Wang, Yixin Sun et al.· bioRxiv· 0 citations
It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations