This work introduces MLIP Studio, an open and free platform that brings more than 60 universal MLIPs into a unified interactive interface for molecules and materials, and demonstrates that MLIP-based pre-optimization can reduce subsequent DFT optimization effort by ~33$\times$.
Abstract
Universal machine learning interatomic potentials (MLIPs) are foundation AI models transforming atomistic simulations, but their practical use remains hindered by fragmented software ecosystems, dependency conflicts, and the lack of accessible benchmarking tools. These models approach first-principles density functional theory (DFT) accuracy at a fraction of the computational cost. We introduce MLIP Studio (available at https://mlipstudio.iisc.ac.in), an open and free platform that brings more than 60 universal MLIPs into a unified interactive interface for molecules and materials. The platform enables end-to-end MLIP-driven workflows, including property prediction, geometry optimization, vibrational and equation-of-state analysis, spin-state determination, custom model deployment, and high-throughput benchmarking against reference data. Automated parity plots and sortable error tables facilitate rapid identification of element-wise outliers and problematic data points. We demonstrate that MLIP-based pre-optimization can reduce subsequent DFT optimization effort by ~33$\times$. Additionally, the application enables benchmarking of computational performance. Through a comprehensive case study involving the 2D magnetic material CrCl$_3$ on a sapphire substrate, we show how cross-model comparisons of various properties and potential-energy landscapes can guide task-specific MLIP selection. Overall, MLIP Studio lowers the barrier to the reliable use of foundation models in end-to-end research workflows, benchmarking, and education in computational chemistry and materials science.
The Active Learning Framework (ALF), an open-source Python package designed to streamline the design and deployment of MLIP training datasets on High Performance Computing resources, is introduced, illustrating ALF’s effectiveness in compiling datasets that capture essential chemical and structural regimes.
V. Grizzi, P. Lohr, Nikita Fedik et al.· Journal of Chemical Theory a...· 0 citations
AI2Pot is presented, a scalable and unified MLIP framework that seamlessly integrates model training, evaluation, and large-scale MD simulations with PyTorch-compatible ecosystem, and offers an user-friendly end-to-end framework for the developing, training, and deploying MLIPs for large scale MD.
Hanyu Liu, Linggang Zhu, Xuanguang Zhang et al.· 0 citations
Dyna-Mat-v1.0 shows that end-to-end finite-temperature validation is essential for quantifying the predictive behaviour of foundation MLIPs, and provides a simple, scalable route for assessing them beyond static and harmonic benchmarks relevant to materials design.
Mikołaj J Gawkowski, Nongnuch Artrith, Silvia Bonfanti et al.· 1 citation
Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.
This work demonstrates how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation.
Timo Reents, Marnik Bercx, Giovanni Pizzi· 0 citations
The limits of equivariant MLIPs are examined, and a family of foundation potentials in the NequIP and Allegro equivariant MLIP architectures are presented which achieve leading inference speeds and strong scalability as well as excellent accuracies across a range of community benchmarks.
Seán R. Kavanagh, Chuin Wei Tan, Menghang Wang et al.· 0 citations