Jul 2026· Journal of Chemical Theory and Computation· Vol 22, pp. 7558 - 7569· 0 citations· 57 references
Medicine
TL;DR
This work employs a dual-coordinate approach where both solutes are explicitly present but do not interact with one another, and utilizes a harmonic ″anchor″ restraint to a central atom on each molecule to enforce spatial overlap without modifying the internal intramolecular dynamics of either solute.
Abstract
The integration of machine-learned interatomic potentials (MLIPs) into free energy simulations (FES) offers the promise of near-quantum mechanical accuracy at a reduced computational cost. However, employing MLIPs for alchemical relative free energy calculations remains challenging since MLIPs are typically not trained to handle the unphysical intermediate states required for such transformations. In previous work, we demonstrated how atoms and molecules can be gradually decoupled in systems fully described by MLIPs by manipulating the neighbor list and introducing an artificial offset to interatomic distances. Here, we extend this approach to general alchemical transformations between two states and apply the methodology to computing relative solvation free energies (RSFEs) between arbitrary solute pairs in systems fully described by an unmodified MLIP. We employ a dual-coordinate approach where both solutes are explicitly present but do not interact with one another. By introducing λ-dependent distance offsets to the neighbor list, we perform a simultaneous transformation: smoothly decoupling the first solute from the solvent while coupling the second solute to the same environment. To enforce spatial overlap without modifying the internal intramolecular dynamics of either solute, we utilize a harmonic ″anchor″ restraint to a central atom on each molecule. We demonstrate the robustness of this method using the MACE-OFF23(S) potential across a diverse set of small molecule pairs, including transformations between species with no chemical similarity (e.g., toluene to tetrahydrofuran). The method’s accuracy is validated through thermodynamic cycle closure by combining direct RSFE calculations with absolute solvation free energy (ASFE) calculations. The methodology achieves excellent internal consistency, with cycle closure errors less than or equal to ±0.11 kcal/mol, well within the estimated statistical uncertainty. By requiring only a single energy evaluation per simulation step and avoiding architecture-specific modifications and retraining, this anchor-based dual-coordinate approach provides a highly adaptable and efficient route for relative alchemical FES with modern MLIPs.
The convergence of molecular dynamics simulations and machine-learned interatomic potentials (MLIPs) promises density functional theory (DFT) level accuracy at near-classical force-field computational costs. However, while the average fidelity to reference energies and forces approaches perfection, several failure modes limit MLIP reliability in production simulations. These include spurious bond formation, inconsistent reproduction of long-range interactions, and inconsistent spin-state references. Here, the origins of these behaviors are studied by benchmarking the UMA, ORB, MACE, and AIMNet2 models against reference DFT bond dissociation curves for an illustrative range of species. These benchmarks reveal that models without explicit atomic charge resolution predict spurious stable bonds between like-charged halide anions, effectively transmuting two Cl- ions into neutral Cl2. Models with atomic partial charge equilibration correctly predict repulsion in these systems. Conversely, several secondary limitations are exposed in these benchmarks, including inconsistent agreement with unrestricted DFT (uDFT) versus restricted DFT (rDFT) energies and inconsistent core-region treatment. This comparative analysis suggests that, while artifacts related to core repulsion and asymptotic electrostatics are readily repairable through improved physical priors and better data curation, the issue of spurious bond formation is intrinsic to the inability of global charge specification to disambiguate similar local geometries at different charge and spin states.
Ericka Roy Miller, V. Sathyaseelan, Dylan M Gilley et al.· Journal of Chemical Theory a...· 0 citations
Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.
Accurate solvation free energies from molecular dynamics simulations require efficient sampling of coupled slow variables, including solvent coordinates, solute conformational modes, and the alchemical coordinate $\lambda$. Here, we develop a $\lambda$-dynamics framework that combines mass scaling, on-the-fly probability enhanced sampling (OPES), and driven adiabatic free energy dynamics (d-AFED) to address these sampling challenges within a unified protocol. For rigid organic solutes, Hamiltonian replica exchange with mass scaling is first used to quantify the effect of octanol solvent relaxation. Reducing all octanol atomic masses by a factor of ten accelerates convergence by more than fivefold while preserving equilibrium solvation free energies. These calculations then provide reference benchmarks for $\lambda$-OPES, a dual-bias $\lambda$-dynamics strategy that combines the"standard"and"explore"variants of OPES to promote transitions along the alchemical coordinate. This approach reaches convergence on timescales comparable to replica exchange, but without predefined $\lambda$ windows or multiple parallel simulations. For flexible $N$-acetyl amino-acid amide solutes, $\lambda$-OPES is coupled with d-AFED on selected backbone and side-chain dihedrals to enable simultaneous alchemical and conformational enhanced sampling. This combined strategy improves agreement with experimental octanol-water partition coefficients and reduces the mean absolute error from 0.75 log units with $\lambda$-OPES alone to 0.30 log units with $\lambda$-OPES-d-AFED. Overall, this work establishes an integrated enhanced sampling protocol for solvation free energy calculations across rigid organic solutes and flexible peptide-like solutes, and provides a foundation for the application of alchemical free energy methods to larger and more conformationally complex systems.
Gabriela B. Correa, C. Abreu, Nishanth N Nair et al.· 0 citations
Multiscale approaches combining machine-learned interatomic potentials with classical molecular mechanics force fields are emerging as a scalable alternative to QM/MM approaches, in which a quantum mechanical calculation is embedded in a molecular mechanics environment. They enable quantum-level accuracy for localized chemistry at substantially reduced cost. However, reliable coupling across region boundaries and application in heterogeneous environments remains a central challenge. Here, we extend the buffer region embedding strategy (BuRNN) for hybrid machine-learned interaction potentials/molecular mechanics (MLIP/MM) simulations in complex environments. The buffer region elevates the interactions of the inner region with its immediate surroundings to the MLIP level, whereas the interactions within the buffer region and with the outer region are still described at the MM level. We validate the approach across four test systems of increasing complexity: methanol/water mixtures benchmarked against experimental total X-ray structure factors, a functionalized fullerene designed to describe a covalent boundary between the buffer and outer region, aqueous heme b with an axial cysteine ligand to probe coordinative bond dissociation, and the resting-state of the cytochrome P450 1A2 enzyme. Across these systems, BuRNN reproduces experimental observables or the underlying QM reference data and yields stable dynamics, including challenging metal–ligand interactions. We compare explicit and implicit link-atom treatments at the buffer–outer boundary and find that explicit capping provides more robust uncertainty estimates, which is critical for active-learning model generation. These results establish BuRNN as a practical approach to perform MLIP/MM simulations in biomolecular systems.
M. Caspary, Radek Crha, Edgar Galicia-Andrés et al.· Journal of Chemical Informat...· 0 citations
High-entropy alloys (HEAs) exhibit exceptional structural and functional properties arising from their complex local chemical environments, and their vast compositional space offers considerable flexibility to further tune and optimize these properties. Atomistic simulations based on density functional theory (DFT) have played a central role in elucidating the thermodynamic, mechanical, magnetic, and defect-related properties of HEAs. However, DFT simulations are severely limited by the intrinsic chemical and configurational complexity of these alloys, particularly because reliable predictions require extensive statistical sampling over chemically diverse configurations and access to extended spatial and temporal scales. In this review, we summarize recent advances in atomistic simulations of HEAs, with particular emphasis on machine learning interatomic potentials (MLIPs), which extend beyond conventional DFT approaches. We discuss how MLIPs enable statistically robust simulations with near-DFT accuracy while dramatically reducing computational cost, thereby allowing explicit treatment of chemical short-range order, vibrational contributions to Gibbs energies, point defects, diffusion, dislocation behavior, grain boundaries, and hydrogen absorption in chemically complex alloys. Particular attention is devoted to the role of local chemical environments, many-body interactions, and configurational sampling in determining HEA properties. We further review recent developments in universal/foundation MLIPs trained on chemically diverse datasets and discuss their potential for rapid and transferable atomistic simulations of HEAs across broad compositional and configurational spaces. We discuss current limitations and open challenges, including transferability to highly distorted defect configurations, treatment of magnetic and charge degrees of freedom, incorporation of finite-temperature excitations, and construction of representative training datasets for chemically and structurally complex systems. This review aims to provide a comprehensive perspective on the ongoing transition from conventional DFT-based simulations toward scalable, statistically rigorous, and predictive atomistic modeling frameworks for HEAs and related compositionally complex materials.
Yuji Ikeda, Xiang Xu, Pranav Kumar et al.· Journal of Materials Science· 0 citations
Electrostatic embedding improved every accuracy and correlation metric for TYK2 but performed comparably to the classical and mechanical-embedding baselines for CDK2, thrombin, p38 and JNK1, and standard single-molecule energy and charge benchmarks were not good predictors of this target-dependent outcome.