Transferable Collective Variable to accelerate Protein-Ligand (Un)Binding Transitions via Explainable Machine Learning and Intriguing Role of Ligand Solvation
This study presents a method to derive optimized CV from transition state region (TS) via an interpretable machine learning (ML) model, Elastic Net, which greatly accelerate ligand binding-unbinding transitions and achieves rapid free energy surface (FES) convergence across diverse systems.
Abstract
The process of drug unbinding is of immense importance in the field of biophysics and therapeutics. The behavior of these systems is greatly influenced by their thermodynamic and kinetic properties. Therefore, it is crucial to accurately estimate the ligand binding free energies and rate of ligand dissociation, yet these processes are often governed by rare event transitions that lie beyond the reach of standard brute-force molecular dynamics simulations. While enhanced sampling simulations offer a solution, their efficacy is strictly contingent upon the selection of appropriate collective variables (CVs) which is non-trivial for complex systems like protein-ligand complexes. In this study, we present a method to derive optimized CV from transition state region (TS) via an interpretable machine learning (ML) model, Elastic Net. By employing some physically intuitive order parameters, the derived optimized CV from the TS-region greatly accelerate ligand binding-unbinding transitions and achieves rapid free energy surface (FES) convergence across diverse systems including buried and solvent exposed active sites such as Trpsin-benzamidine complex, host-guest systems and sodium epoxidase etc. Intriguingly significant contribution of the ligand hydration is found in the optimized CV which depicts crucial role of solvent in driving ligand binding-unbinding transitions. The estimated binding free energies for different protein-ligand complexes match quite well with experiments, while maintaining a low computational cost. The derived optimized CV is also used to calculate the ligand residence times across different systems and calculated residence times are within the experimental range for all systems, again with very little computational costs. Moreover, we show that the optimized CV constructed from TS region via an interpretable ML model is transferable across diverse systems, offering a robust and scalable framework for drug discovery and investigation of complex biomolecular recognition.
The kinetics of protein-ligand binding systems are increasingly recognized as a key determinant of drug efficacy, yet remain far harder to compute than binding affinities. Existing kinetics methods either bias the dynamics along a collective variable (CV), demanding careful system-specific CV design, or use path sampling, which keeps the dynamics unbiased but can struggle to converge rates out of deep free-energy wells and often relies on hand-engineered descriptors. By combining the `best of both worlds', we propose a method to compute accurate kinetics for general ligand-unbinding problems at modest computational expense and minimal fine tuning, building on the AI for Molecular Mechanism Discovery (AIMMD) path sampling framework. To avoid the need for feature engineering, we opt for modelling the committor with a single descriptor-free, equivariant graph neural network shared across all systems. We also partially flatten deep bound-state wells with a static, basin-restricted bias potential. This improves convergence by lifting the path sampling state boundary out of regions, where the committor is hard to learn, while leaving the reactive region strictly unbiased. Across host-guest and protein-ligand systems spanning roughly 17 orders of magnitude in residence time, the method robustly recovers rates in line with reference and experimental values. Simultaneously, and without further sampling, it also reconstructs the underlying unbinding mechanisms. We additionally find that accurate rates do not require globally accurate committor models, allowing for efficient kinetics estimation even in a low-data training regime. Requiring little system-specific setup, our approach offers an efficient and broadly generalizable route to binding kinetics, and its shared committor architecture lays crucial groundwork for probing structure-kinetics relationships across ligand series in drug discovery.
This work not only establishes a pioneering paradigm for interpretable ML-driven force field refinement but also provides the first feature engineering solution incorporating chemical, physical, and structural information specifically designed for the machine learning of energetic molecular crystals.
Qi He, Pengju Wang, Xudong He et al.· Molecules· 0 citations
PML is introduced, a novel computational framework that describes a binding interface as a family of multiscale manifolds, and results indicate that much of what determines binding strength is encoded in the shape of the interface itself, and that a single geometric description serves both classes without hand-tailored features.
Xingjian Xu, Zhe Su, Guo-Wei Wei et al.· arXiv.org· 0 citations
It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.
R. Khalil, Elena Frasnetti, Han Kurt et al.· Journal of Physical Chemistr...· 0 citations
This work assesses ML potentials for exploring RNA conformations using the adenine–adenine dinucleoside monophosphate (ApA) dimer, a fundamental RNA building block, and parametrized ML potentials based on the equivariant MACE architecture and informed by both ab initio and semiempirical property data.
Leonardo Medrano Sandonas, Macarena Tolmos Nehme, L. F. Cofas-Vargas et al.· Journal of Chemical Theory a...· 0 citations
Characterizing the solvation of open-shell transition metal ions remains a challenge for classical force fields due to ligand field electronic effects. Here, we present AMOEBA + NN, a machine-learning-augmented polarizable potential, to investigate Cu2+ solvation in NH3 and H2O and mixed environments. The model, trained on quantum-mechanical (QM) association energies, accurately reproduces the Cu2+ energy landscape across diverse coordination geometries. Molecular dynamics simulations demonstrate that AMOEBA + NN successfully captures the electronically driven Jahn–Teller (JT) distortion. Evidence from radial distribution functions (RDFs) and geometry optimizations reveals characteristic axial elongation, which is absent in conventional classical descriptions. Kinetic analysis shows a highly dynamic Cu2+–NH3 environment with a rapid ligand exchange (residence time of 66.52 ps), contrasting with the H2O system where binding is two to three orders of magnitude more stable. Furthermore, the calculated hydration free energy of −477.59 ± 0.50 kcal mol−1 shows excellent agreement with experimental data (within 1.31 kcal mol−1). This work provides a unified, computationally efficient framework for describing the structural, dynamic, and thermodynamic observables of transition metal coordination.
Zhecheng He, Yanxing Wang, Shubham Chatterjee et al.· Chemical Science· 0 citations