Skip to content
Open access

Probing Machine Learning Interatomic Potentials on Ion Transport Properties

Jun 2026 · Advanced Intelligent Discovery · 0 citations · 49 references

TL;DR

A benchmark study systematically benchmark six state‐of‐the‐art MLPs, namely CHGNet, EquiformerV2 in two training variants, MatterSim, SevenNet, and MACE, on representative Li‐ and Na‐based superionic conductors to highlight the importance of accurately capturing atomic forces in reasonably producing ion transport properties in all‐solid‐state battery materials.

Abstract

Machine learning interatomic potentials (MLPs) are promising for accelerating the simulation of ion transport in all‐solid‐state battery materials, but their accuracy across diverse material compositions and symmetries remains unquantified. Here, we systematically benchmark six state‐of‐the‐art MLPs, namely CHGNet, EquiformerV2 in two training variants, MatterSim, SevenNet, and MACE, on representative Li‐ and Na‐based superionic conductors. By comparing predicted atomic forces, diffusion coefficients (DCs) from MLP‐driven molecular dynamics simulations, and second‐ and third‐order interatomic force constants (IFCs) against density functional theory (DFT) and ab initio molecular dynamics (AIMD), we assess the predictive reliability of each model across these properties. EquiformerV2 models, especially the variant trained only on the OMAT dataset, exhibit the lowest force prediction errors and yield DCs in the closest agreement with AIMD. Continually, MLPs with more accurate force predictions produce more reliable diffusion metrics. However, while second‐order IFCs are reasonably captured, all models struggle to reproduce third‐order (anharmonic) IFCs, highlighting significant challenges in modeling anharmonic lattice dynamics with current MLPs. This benchmark study highlights the importance of accurately capturing atomic forces in reasonably producing ion transport properties in all‐solid‐state battery materials and provides practical guidance for universal MLP selection for ion transport simulations and future improvement and development.

Read PDF

Similar papers

Open access Jul 2026

Full-stack quantification of variability in predicting ion transport properties using machine-learned interatomic potentials

Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is, therefore, crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods and improving the MD sampling statistics.

Tawfiqur Rakib, Lucas K. Wagner, Elif Ertekin · 0 citations
Open access Jul 2026

Foundational Machine‐Learning Interatomic Potential for Simulating Chemically Complex Ni‐Based Superalloys

For decades, atomistic simulation of chemically complex Ni‐based superalloys has remained beyond practical reach. Here, we apply the GRACE foundational machine‐learning interatomic potential to predict chemical ordering and stacking‐fault energetics in the and phases of CMSX‐4, a commercial multicomponent Ni‐based superalloy. After benchmarking against structural and thermodynamic reference data, we use hybrid Monte‐Carlo/molecular dynamics sampling to study the impact of local chemical order on planar‐fault energies. GRACE reproduces elemental equilibrium lattice parameters within of DFT references, while underestimating melting temperatures of ordered Ni–Al phases by up to . The simulations reveal local chemical ordering in the phase and the expected sublattice occupancies in the phase. In the phase, the short‐range order raises the shear barriers by approximately while leaving the intrinsic stacking fault energy of unchanged. In the phase, alloying raises the complex and superlattice intrinsic stacking fault energies by approximately relative to stoichiometric Al. These results show that pretrained foundational potentials enable atomistic simulations of chemically complex multicomponent superalloys at scales inaccessible to direct first‐principles calculations.

Aditya Vishwakarma, Sarath Menon, Fritz Körmann et al. · 0 citations
Open access Aug 2026

Multiscale and Multi‐Timestep Switching of Multiple Machine Learning Force Fields for Artificial Intelligence‐Driven Materials Simulations

Molecular dynamics (MD) is essential for investigating atomic‐scale processes in materials and molecular systems, but the cost of high‐accuracy machine learning force field simulations still limits accessible system sizes and timescales. Here, we propose a practical model‐switching strategy for Deep Potential (DP)‐based MD simulations that alternates between independently trained DP models with different cutoff radii: a standard 6 Å model for higher accuracy and a reduced‐cutoff 4 Å model for faster inference. The method was implemented in LAMMPS/DeePMD and evaluated using solid‐phase anatase TiO 2 and liquid‐phase polyethylene glycol (PEG). For anatase TiO 2 , the 1:3 4–6 Å switching scheme preserved radial distribution function (RDF) correlations of 0.996 or higher relative to the 6 Å baseline while achieving a 1.24‐fold speedup. For PEG, the switching scheme maintained RDF correlations of 0.996 or higher with a 1.18‐fold speedup. Additional optimization using network‐size reduction and mixed‐precision inference achieved a 2.53‐fold speedup with RDF correlations of 0.975–0.988. Constant particle‐number, pressure, and temperature (NPT) simulations remained stable, whereas constant particle‐number, volume, and energy (NVE) simulations revealed system‐dependent energy‐drift behavior, particularly for aggressively optimized models. These results demonstrate that DP model switching provides a simple and practical route for accelerating structural MD simulations while highlighting the need for validation when strict energy conservation is required.

Ryuya Kanda, Megumu Yamazaki, Yuta Yoshimoto et al. · 0 citations
Preprint Jul 2026

Benchmarking Universal Machine Learning Force Fields for Molecular Dynamics of Lunar Regolith Minerals

Universal machine-learning interatomic potentials provide a promising route for accelerating molecular dynamics simulations of materials, but their transferability to lunar regolith-relevant silicates, oxides, and hydrogen-bearing surface species remains elucidated. Here, we benchmark six foundation models, MACE-MH, MatterSim, SevenNet-0, UPET, UMA, and NequIP-OAM-L, using NVT molecular dynamics simulations of four representative lunar minerals: forsterite, fayalite, ilmenite, and anorthite. Structural fidelity is evaluated using temperature stability, bond-distance statistics, bond-angle distributions, and partial radial distribution functions, with comparison to crystallographic reference data. The models reproduce Si--O, Mg--O, Al--O, and Ca--O local environments reasonably well, while Fe--O and Ti--O coordination environments show broader distributions and larger short-timescale fluctuations, highlighting the need for further validation and fine tuning with additional ground truth data for Fe- and Ti-bearing lunar phases. Hydroxylated surface tests show consistent O--H bond-distance distributions across models and minerals, suggesting that these foundation models may provide useful starting points for screening surface hydroxyl stability and volatile-related processes. Performance benchmarks on a single NVIDIA RTX 4090 show that SevenNet-0, MatterSim, and UPET provide the highest throughput among the six tested models, MACE-MH remains practical at intermediate cost, and UMA and NequIP-OAM-L extend the comparison to newer foundation potentials at higher runtime cost and memory demand. These results provide an initial benchmark for applying universal foundation models to lunar mineral simulations and identify key directions for future ab initio validation, model fine-tuning, and applications to lunar volatile evolution, space weathering, ISRU, and polar sample return studies.

Ziyu Huang, K. Nomura · 0 citations
Preprint Aug 2026

Data-Efficient Construction of Material-Specific Machine-Learning Interatomic Potentials from Ab Initio Molecular Dynamics Trajectories

Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.

Jonas Hänseroth, Christian Dreßler · 0 citations
Preprint Jul 2026

Dyna-Mat: End-to-end benchmarking of foundation machine learning interatomic potentials in finite-temperature ensembles

Foundation machine learning interatomic potentials (MLIPs) are increasingly being used as drop-in replacements for first-principles calculations, enabling simulations of materials at length and time scales that were previously inaccessible. However, due to lack of ground truth data, their accuracy on structural and dynamical observables in finite thermodynamic ensembles is yet to be established. Here, we introduce Dyna-Mat-v1.0, a benchmark dataset of condensed-phase first-principles molecular dynamics trajectories designed to test foundation MLIPs at realistic finite-temperature conditions. Using this dataset, we evaluate 15 foundation MLIPs across four model tiers by comparing both single-point energy and force errors on first-principles configurations and observables generated from MLIP-driven trajectories. We find that"on average"models with lower single-point force errors also yield lower errors for structural and dynamical observables. However, there are individual systems for which low force errors lead to qualitative failures in the predicted structure. Pressure remains poorly described across most models, pointing to limitations in the density functional theory stress labels available in current large-scale training datasets. Finally, we construct an accuracy-cost Pareto frontier to identify the best trade-offs for molecular dynamics with foundation MLIPs, finding that the latest generation of cross-trained models is close to Pareto-optimal according to the accuracy metrics considered here. Overall, Dyna-Mat-v1.0 shows that end-to-end finite-temperature validation is essential for quantifying the predictive behaviour of foundation MLIPs, and provides a simple, scalable route for assessing them beyond static and harmonic benchmarks relevant to materials design.

Mikołaj J Gawkowski, Nongnuch Artrith, Silvia Bonfanti et al. · 1 citation