Aug 2026· Journal of Chemical Theory and Computation· Vol 22, pp. 8131-8144· 0 citations· 99 references
TL;DR
A physics-based framework for the relationship between the electrostatic energy and the electrostatic field generated by the MM environment, MIREANN exhibits powerful transferability in reproducing the free energy barriers predicted by QM/MM-MD with errors less than 0.5 kcal mol–1 for various enzyme variants without the need to retrain the MLP.
Abstract
To accelerate computationally expensive quantum mechanical (QM) calculations, integrating machine learning potentials (MLPs) with molecular mechanics (MM) has become a promising approach, which can be employed to perform multiscale ML/MM-MD simulations and estimate reaction free energy for enzymatic catalysis. However, the ML/MM frameworks face the challenge of intuitive chemical interpretability in QM/MM electrostatic coupling, which limits their transferability in unlearned MM environments. Therefore, their practical application in enzyme engineering is hindered by the bottleneck of MLP retraining for different enzyme environments. In this work, guided by linear response theory, the MLPs trained in the gas phase were naturally adapted to QM/MM systems by integrating an MM atom electrostatic field-induced recursively embedded atom neural network (MIREANN) into multiscale coupling. The performance and broad applicability of our model are exhibited by nanosecond-scale ML/MM-MD on diverse enzymatic systems (e.g., cyclooxygenase, hydrolases), which achieves QM/MM accuracy at a speed close to molecular mechanics. Since MIREANN provides a physics-based framework for the relationship between the electrostatic energy and the electrostatic field generated by the MM environment, it exhibits powerful transferability in reproducing the free energy barriers predicted by QM/MM-MD with errors less than 0.5 kcal mol–1 for various enzyme variants without the need to retrain the MLP. Our model is expected to be applied to a wider range of chemical and biological systems, such as enzyme or drug design, to promote the in-depth development of related research.
Predicting how mutations alter enzyme catalysis remains a central challenge in enzymology and enzyme engineering. Although quantum mechanics/molecular mechanics (QM/MM) simulations can in principle compute the activation free energy associated with enzymatic reactions, their high computational cost limits systematic studies across many variants. Here, we benchmark a mechanical-embedding machine learning potential/molecular mechanics (ML/MM) protocol for predicting mutation effects on chorismate mutase catalysis, a model system extensively studied both experimentally and computationally. In this framework, the QM-region potential energy surface is represented by an actively learned machine learning potential, while QM/MM electrostatic interactions are treated classically using partial charges predicted from instantaneous geometries for the QM region. Combined with umbrella sampling, the ML/MM approach enables efficient estimation of activation free energies and direct comparison with experimental kinetics. The method shows reasonable correlations with experiment across both nonpolar and polar active-site mutations and is quantitatively accurate for nonpolar mutations despite their narrow energetic range (<1 kcal mol-1). However, it substantially underestimates the activation free energy for polar mutations. The results highlight both the promise and limitations of mechanical-embedding ML/MM approaches for predicting mutation effects on enzyme catalysis.
Zi-Chen Sun, Yi-Fan Li, Wenqiang Cui et al.· Journal of Chemical Theory a...· 0 citations
The application of machine learning potentials (MLPs) to accurately simulate enzymatic reactions remains challenging. This is primarily due to the high dimensionality and structural heterogeneity of enzyme systems, as well as the need to incorporate off-equilibrium and transition state conformations into the training data. Integrating MLPs into a quantum mechanical/molecular mechanical (QM/MM) framework through either direct learning or Δ-learning, together with reaction-specific training strategies, can help overcome these limitations. In this work, we present an integrated workflow for MLP/ΔMLP-assisted QM/MM simulations that achieves ab initio (ai) or density functional theory (DFT)-level accuracy in predicting enzyme reaction thermodynamics. Developed within CHARMM and tightly integrated with the mlp_qmmm Python package, the workflow automates training-data generation, data sanitization, MLP/ΔMLP training, and deployment of trained models in QM/MM molecular dynamics (MD) simulations. In addition, a low-overhead interface implemented in CHARMM enables efficient model inference on both CPU and GPU during simulations. The workflow further incorporates an iterative model refinement strategy that systematically improves predictive performance through successive rounds of sampling, high-level labeling, and retraining. The capabilities of the approach are demonstrated using the hydride-transfer reaction catalyzed by four variants of dihydrofolate reductase (DHFR). Compared with conventional ai/DFT-QM/MM simulations, the ΔMLP-assisted approach achieves more than 500-fold acceleration while maintaining subkcal/mol accuracy in predicted reaction free energies and free energy barriers. Iterative refinement further improves the underlying energy and force predictions, leading to more accurate thermodynamic and structural properties obtained from QM/MM simulations with only a modest additional computational cost. Overall, this work establishes a scalable and extensible workflow for the systematic development and iterative refinement of reaction-specific MLP/ΔMLP models, enabling highly accurate and computationally efficient simulations of enzyme-catalyzed reactions.
Abdul Raafik Arattu Thodika, S. Panda, Xiaoliang Pan et al.· Journal of Chemical Theory a...· 0 citations
The integration of QM/MM methods with molecular dynamics (MD) simulations has become a powerful tool for elucidating enzymatic reaction mechanisms at atomistic resolution, providing valuable insights for biocatalyst design and drug development. However, the high computational cost of QM methods, amplified by the extensive configurational sampling inherent in MD, remains a key bottleneck in studying complex enzymatic processes. Hybrid machine learning/molecular mechanics (ML/MM) methods, which replace the quantum calculations with machine-learned interatomic potentials, offer a promising alternative toward near-QM/MM accuracy at substantially reduced computational cost. This Perspective surveys the central methodological challenges in developing ML/MM frameworks, including the generation of high-quality reference data and the treatment of multiscale coupling. Emerging applications in biosystems are highlighted, and future directions are outlined toward accurate, efficient, and broadly transferable ML/MM models to support next-generation workflows for biomolecular simulations.
Xinhu Sha, Chenyu Wu, Daiqian Xie et al.· Journal of Physical Chemistr...· 0 citations
Neural network potentials (NNPs) can provide insight into biological processes at atomic resolution. Training these NNPs requires large and diverse datasets of molecules, conformations, and configurations. However, so far little attention has been paid to the description of solvation, despite its importance for biomolecular systems. This work lays the foundation for NNPs where solvation is an integral part of the model. Following a quantum-mechanics/molecular-mechanics (QM/MM) formalism with an electrostatic embedding scheme, systems are decomposed into a QM zone with the solute(s), which is electrostatically coupled to the point charges from surrounding solvent molecules (MM zone). Using an accelerated sampling approach, we generate the biomolecular multiscale simulation (BMS25) dataset with over 50,000 topologies and more than 1.5 million unique conformations of peptides and miniproteins as well as small molecules and transition states from chemical reactions. The dataset includes energies, gradients, and multipoles of solute molecules as well as gradients on solvent molecules at the
ω
B97M-D4/ma-def2-TZVPP level of theory, enabling the development of multiscale NNPs for simulating large biomolecular systems.
Moritz Thürlemann, Felix Pultar, Igor Gordiy et al.· Scientific Data· 0 citations
We offer a practical and conceptual introduction to some of the current approaches to modelling enzymatic reaction mechanisms, ranging from quantum mechanics (QM), molecular mechanics (MM), and hybrid QM/MM approaches to enhanced sampling methods, knowledge-based approaches, and machine learning (ML) advances. We discuss how static and dynamic QM/MM approaches, as well as multi-PES strategies, have contributed to understanding the role of conformational diversity, electrostatic preorganization and solvent participation in the determination of catalytic barriers and reaction paths. We focus on how advanced sampling techniques and data-driven collective variables have enabled the exploration of rare events and reaction coordinates, as well as how knowledge- and rule-based approaches have facilitated the interpretation and hypothesis generation for different families of enzymes. Recent developments in ML potentials, ML collective variables, and committor-based sampling are presented as innovative methods that have been able to address some of the current challenges in accuracy, sampling efficiency, and the identification of low-dimensional representations of reaction coordinates. A case study of α-amylase demonstrates how the combination of these strategies leads to a comprehensive understanding of enzyme reactivity, from the chemical to the conformational level. Collectively, these developments contribute to a predictive understanding of enzymatic catalysis, which will have extensive implications in enzyme engineering, sustainable chemistry, and drug discovery. Advances in high performance computing, automated simulation pipelines and data formats will likely make multiscale simulation more accessible and reproducible. Simultaneously, the combined application of mechanistic knowledge, ML, and experimental validation will hopefully advance the discovery and optimization of biocatalysts with well-defined properties, tailored to meet pressing societal needs, such as plastic biodegradation, carbon sequestration, sustainable synthesis and personalised medicine.
Rui P. P. Neves, João T. S. Coimbra, Pedro Paiva et al.· Chemical Science· 0 citations
Electrostatic embedding improved every accuracy and correlation metric for TYK2 but performed comparably to the classical and mechanical-embedding baselines for CDK2, thrombin, p38 and JNK1, and standard single-molecule energy and charge benchmarks were not good predictors of this target-dependent outcome.