Skip to content
Open access

Prediction of Distal Mutation Effects in Enzymes via Integration of Molecular Dynamics Descriptors and Zero-Shot Model.

Aug 2026 · Journal of Chemical Theory and Computation · Vol 22 16, pp. 8607-8616 · 0 citations · 66 references
Medicine

TL;DR

This work presents an integrated framework that combines molecular dynamics-derived descriptors with the zero-shot prediction model GEMS to identify beneficial distal mutations, offering an efficient and generalizable strategy for enzyme engineering.

Abstract

Identifying beneficial distal mutations remains a key challenge in enzyme engineering, as such residues can regulate activity through long-range dynamical coupling and allosteric communication. Although protein language models effectively capture sequence and evolutionary information, they lack explicit representation of conformational dynamics in the presence of the substrate, limiting their ability to detect distal regulatory sites. Here, we present an integrated framework that combines molecular dynamics (MD)-derived descriptors with the zero-shot prediction model GEMS to identify beneficial distal mutations. Compared to the pure zero-shot model, MD-derived descriptors efficiently capture key distal residues involved in dynamical coupling and allosteric communication with the active site. Consequently, these constraints enable the zero-shot model to predict distal mutations more precisely. By integrating sequence, evolutionary, and dynamic information, our approach expands the diversity of candidate sites while maintaining a manageable screening scale, offering an efficient and generalizable strategy for enzyme engineering.

Read PDF

Similar papers

Open access Aug 2026

A unified predictor of protein stability changes across all mutation types via implicit structure learning

UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.

Hong Tan, Shenggeng Lin, Yi Xiong · 0 citations
Open access Aug 2026

The physicochemical basis of protein evolution: property-informed evolutionary models (PRIME)

Abstract Standard probabilistic models of coding sequence evolution effectively identify where and when selection acts but remain agnostic to the mechanistic realization of these forces. We introduce PRIME (PRoperty Informed Models of Evolution), a framework of codon-level maximum likelihood methods—including global (G-PRIME), episodic (E-PRIME), and site-specific (S-PRIME) implementations—that explicitly model amino acid exchangeability as a function of physicochemical properties. By parameterizing attributes such as molecular volume, hydropathy, and secondary structure propensities, PRIME aims to resolve the biophysical basis of selective constraint across both the sequence and the phylogeny. At the site level, S-PRIME leverages an explicit biophysical taxonomy to categorize residues as conserved, neutral, or changing for specific properties, resolving selective signals that are missed by traditional rate-based metrics. Our analysis of a benchmark of 24 diverse datasets and a genome-wide screen of 18,944 mammalian genes demonstrates that consideration of biophysical realism can yield substantial improvements in model fit, acting synergistically with rate variation to explain complex evolutionary patterns. We find that physicochemical constraints at individual sites can be reliably detected in datasets with sufficient information redundancy (substitutions per unique amino acid; AUC=0.91), with sensitivity exceeding 90% in data-rich alignments. E-PRIME reveals a distinct hierarchy in biophysical constraints: while core packing and beta-sheet scaffolds are rigidly conserved, alpha-helix propensity and surface electrostatics serve as the primary substrates for adaptive tuning. Furthermore, PRIME importance weights align with aspects of the primary semantic axes of deep learning representations (ESM-2) and capture key features of experimental fitness landscapes. By transforming abstract evolutionary rates into interpretable biophysical rules, PRIME provides a useful framework for characterizing the mechanistic drivers of protein diversity.

Hannah Kim, Konrad Scheffler, Anton Nekrutenko et al. · 0 citations
Open access Aug 2026

Accelerating protein engineering: an integrated framework combines protein language models and epistatic landscape modeling

MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.

J. Koo, Young-Ho Park, Sun-Uk Kim · 0 citations
Review Jul 2026

Protein function evolution through the lens of conformational dynamics: A single-molecule perspective.

Together, these insights position conformational dynamics at the center of understanding and engineering the evolutionary logic of protein function, opening the door to study how proteins are tuned to operate under the nonequilibrium conditions of living cells.

Sixto M. Herrera, Elías Manríquez-Benítez, Exequiel Medina · 0 citations
Open access Aug 2026

Elucidating enzyme–substrate specificity through co-folding foundation model

Boltz2ESI is introduced, an end-to-end framework that predicts enzyme–substrate interactions by leveraging structural knowledge learned by a biomolecular foundation model and consistently outperforms state-of-the-art sequence-based and rigid-docking approaches.

Xiwei Cheng, Seonghwan Seo, C. Huh et al. · 0 citations
Aug 2026

Evaluating Mechanical-Embedding ML/MM for Predicting Mutation Effects in Chorismate Mutase Catalysis.

Predicting how mutations alter enzyme catalysis remains a central challenge in enzymology and enzyme engineering. Although quantum mechanics/molecular mechanics (QM/MM) simulations can in principle compute the activation free energy associated with enzymatic reactions, their high computational cost limits systematic studies across many variants. Here, we benchmark a mechanical-embedding machine learning potential/molecular mechanics (ML/MM) protocol for predicting mutation effects on chorismate mutase catalysis, a model system extensively studied both experimentally and computationally. In this framework, the QM-region potential energy surface is represented by an actively learned machine learning potential, while QM/MM electrostatic interactions are treated classically using partial charges predicted from instantaneous geometries for the QM region. Combined with umbrella sampling, the ML/MM approach enables efficient estimation of activation free energies and direct comparison with experimental kinetics. The method shows reasonable correlations with experiment across both nonpolar and polar active-site mutations and is quantitatively accurate for nonpolar mutations despite their narrow energetic range (<1 kcal mol-1). However, it substantially underestimates the activation free energy for polar mutations. The results highlight both the promise and limitations of mechanical-embedding ML/MM approaches for predicting mutation effects on enzyme catalysis.

Zi-Chen Sun, Yi-Fan Li, Wenqiang Cui et al. · 0 citations