Skip to content
Open access

CGNEP‐MB‐pol: A Single‐Site Coarse‐Grained Machine Learning Potential for Water

Aug 2026 · Materials Genome Engineering Advances · Vol 4 · 0 citations · 49 references

TL;DR

A CG machine learning potential for water, named CGNEP‐MB‐pol, is introduced, which integrates a one‐molecule to one‐bead mapping with the neuroevolution potential (NEP) framework and MB‐pol reference data, thereby aiming to alleviate, within the liquid‐water and ice‐Ih states examined here, the state dependence and limited transferability that often constrain conventional CG models.

Abstract

Coarse‐grained (CG) molecular dynamics (MD) can greatly extend the accessible time and length scales for water, provided that the reduced model captures key structural, dynamical, and thermodynamic properties. Here, we introduce a CG machine learning potential (MLP) for water, named CGNEP‐MB‐pol, which integrates a one‐molecule to one‐bead mapping with the neuroevolution potential (NEP) framework and MB‐pol reference data, thereby aiming to alleviate, within the liquid‐water and ice‐Ih states examined here, the state dependence and limited transferability that often constrain conventional CG models. Without altering the NEP descriptor or network architecture, the model is trained on CG coordinates and labels derived from atomistic configurations through a two‐step dataset refinement strategy, combining multi‐state training based on forces, energies, and virials. CGNEP‐MB‐pol remains compatible with the GPUMD + NEP workflow while delivering a substantial speedup over its all‐atom counterpart. It reproduces mapped reference forces, radial distribution functions of liquid water and ice, and the overall temperature dependence of self‐diffusion and viscosity. Two‐phase simulations capture ice growth, melting, and solid–liquid coexistence, yielding a melting temperature of approximately 274 K. This study establishes a practical route for constructing CG‐MLPs from high‐fidelity atomistic reference models.

Read PDF

Similar papers

Open access Aug 2026

Multiscale and Multi‐Timestep Switching of Multiple Machine Learning Force Fields for Artificial Intelligence‐Driven Materials Simulations

Molecular dynamics (MD) is essential for investigating atomic‐scale processes in materials and molecular systems, but the cost of high‐accuracy machine learning force field simulations still limits accessible system sizes and timescales. Here, we propose a practical model‐switching strategy for Deep Potential (DP)‐based MD simulations that alternates between independently trained DP models with different cutoff radii: a standard 6 Å model for higher accuracy and a reduced‐cutoff 4 Å model for faster inference. The method was implemented in LAMMPS/DeePMD and evaluated using solid‐phase anatase TiO 2 and liquid‐phase polyethylene glycol (PEG). For anatase TiO 2 , the 1:3 4–6 Å switching scheme preserved radial distribution function (RDF) correlations of 0.996 or higher relative to the 6 Å baseline while achieving a 1.24‐fold speedup. For PEG, the switching scheme maintained RDF correlations of 0.996 or higher with a 1.18‐fold speedup. Additional optimization using network‐size reduction and mixed‐precision inference achieved a 2.53‐fold speedup with RDF correlations of 0.975–0.988. Constant particle‐number, pressure, and temperature (NPT) simulations remained stable, whereas constant particle‐number, volume, and energy (NVE) simulations revealed system‐dependent energy‐drift behavior, particularly for aggressively optimized models. These results demonstrate that DP model switching provides a simple and practical route for accelerating structural MD simulations while highlighting the need for validation when strict energy conservation is required.

Ryuya Kanda, Megumu Yamazaki, Yuta Yoshimoto et al. · 0 citations
Open access Aug 2026

Machine learning modeling reveals unconventional nucleation mechanism in phase‐change material GeTe

GeTe, a prototypical phase‐change material, has attracted pronounced attention for next‐generation storage and neuromorphic computing, yet its nucleation mechanism remains an ongoing pursuit. The challenge stems from the limited scales of conventional simulations and the intrinsic structural disorder of nucleation. To bridge the scale gap, we developed a machine learning potential (ML potential) trained on an extensive density functional theory dataset, enabling large‐scale molecular dynamics simulations that capture the complete crystallization pathway with quantum‐level fidelity. By employing a unified short‐ and medium‐range structural framework, we simplified the complexity and disorder inherent to nucleation, allowing us to unravel the intricate atomic environment, pinpoint essential structural motifs, and track their dynamic evolution. Through this approach, our simulations reveal an unconventional nucleation process: the initial formation of Ge‐rich clusters with defective octahedral coordination, followed by their Te‐mediated assembly into the final rock‐salt structure. This sequential mechanism presents a different scenario from classical nucleation theory's expectation of alternating Ge/Te incorporation, wherein the Te sublattice preferentially forms a face‐centered‐cubic structure ahead of Ge ordering. The observed two‐stage nucleation process adds a valuable perspective on the structural ordering kinetics in PCMs, helping to explain their fast crystallization characteristics at the atomic level.

Siqi Tang, Qundao Xu, Shaojie Yuan et al. · 0 citations
Preprint Jul 2026

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

This work introduces implicit machine learning force fields, which replace explicit stacks of neural network layers with self-consistent fixed-point equations, and demonstrates this across three major classes of graph neural networks: invariant, equivariant Cartesian tensor, and SO(3)-equivariant spherical-tensor architectures.

J. Maess, Leon Werner, J. Frank et al. · 0 citations
Open access Jul 2026

Machine Learning Potential for Ga–In Alloy Melting

Radar‐metric analysis shows that PBE + D3‐based models provide the most balanced overall performance for density, liquid structure, and diffusion, whereas LDA‐based models give the best predictions for Ga and EGaIn.

Chen Hua, Jing Liu · 1 citation
Open access Jul 2026

NNP/CG-MM: Embedding of All-Atom Neural Network Potentials into a Coarse-Grained Molecular Mechanics Environment

Neural network potentials (NNPs), or neural network-based force fields, are gaining widespread attention for their ability to model complex chemical, materials, and biophysical systems. In NNPs, the total energy of the system can be decomposed into atom-centered components, where the energies and forces are described using a deep neural network. However, despite the flexibility and accuracy of NNPs, they are typically more computationally intensive than classical molecular dynamics. There have been recent advances in embedding NNPs into a molecular mechanics (MM)-based environment to maximize efficiency while retaining the accuracy of NNPs. In this work, we propose a method called NNP/CG-MM, in which an all-atom NN force field is systematically embedded into a coarse-grained molecular mechanics environment (CG-MM). Coarse-graining (CG) involves constructing a simplified representation of a larger fine-grained (FG) system with the goal of significantly accelerating computations while maintaining the accuracy of the FG system when projected onto the CG variable distributions. The NNP-CG coupling terms are constructed using the multiscale CG force-matching (MS-CG) method. The scheme is tested on liquids and in capturing features of the hydrophobic effect, where three-body correlations in the CG solvent can play an important role.

Kuntal Ghosh, G. A. Voth · 0 citations
Preprint Jul 2026

AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms

We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD). This is in contrast with existing generative models that remove degrees of freedom by operating on coarse-grained, vacuum, or implicit solvent systems. Operating at this resolution allows for post-processing through force field energy evaluations and MD simulations, and enables the prediction of relevant properties in a gray-box manner (as ensemble averages of potential energy evaluations over generated samples). We demonstrate the utility of this paradigm on absolute hydration free energy (AHFE), producing estimates 4-10x faster and with comparable accuracy to standard GPU-based MD. By generating uncorrelated samples from alchemical Boltzmann distributions, we create more accurate, interpretable, and refinable ensemble predictions with calibrated uncertainty estimates, unlike regression methods which are entirely black-box predictors. Our approach also yields predictable benefits from increasing train- and test-time compute, realized by scaling model size and generating more samples, respectively. We believe that this approach demonstrates the utility of high-resolution ensemble generation for free energy estimation, with future potential to replace MD in tasks such as the prediction of lipophilicity, membrane permeability, or absolute binding free energy (ABFE) -- whose grounding and interpretability may be critical for the development of new drugs and materials.

Emmanuel Bengio, Sanjeev Raja, Y. Pang et al. · 1 citation