This work benchmarks 23 mainstream open-source MLIPs on a low-cost NVIDIA DGX Spark, using a fixed 192-atom system under a unified ASE-based pipeline, and evaluates three dimensions: predictive accuracy, MD simulation throughput, and atomic scalability.
Abstract
Most MLIP benchmarks reward static accuracy while ignoring inference efficiency and hardware scalability -- driving model bloat with unclear real-world value. We benchmark 23 mainstream open-source MLIPs on a low-cost NVIDIA DGX Spark (128 GB native memory, capped at 80 GB to mimic ordinary lab hardware), using a fixed 192-atom system under a unified ASE-based pipeline. We evaluate three dimensions: predictive accuracy, MD simulation throughput, and atomic scalability. Our results expose a sharp accuracy-efficiency trade-off: large SOTA models deliver only 3-5 meV/atom more accuracy than lightweight ones, but lose orders of magnitude in throughput -- in the worst case, becoming only marginally faster than DFT itself. Lightweight MLIPs, by contrast, sit on the Pareto frontier and run on modest hardware. The lesson is that single-dimensional benchmarks mislead the field, and that future MLIP development should value efficiency and scalability alongside accuracy.
DPA4C is introduced, an equivariant potential whose architecture and compressed CUDA operators are co-designed under deployment constraints to pursue accuracy and efficiency together and brings quantum-trained universal accuracy into a regime of speed and system size previously associated with empirical potentials.
Tian Li, Jianming Xue, Linfeng Zhang et al.· 0 citations
This work introduces MLIP Studio, an open and free platform that brings more than 60 universal MLIPs into a unified interactive interface for molecules and materials, and demonstrates that MLIP-based pre-optimization can reduce subsequent DFT optimization effort by ~33$\times$.
Manas Sharma, Sudeep N. Punnathanam, A. Rajan· 1 citation
The Active Learning Framework (ALF), an open-source Python package designed to streamline the design and deployment of MLIP training datasets on High Performance Computing resources, is introduced, illustrating ALF’s effectiveness in compiling datasets that capture essential chemical and structural regimes.
V. Grizzi, P. Lohr, Nikita Fedik et al.· Journal of Chemical Theory a...· 0 citations
The limits of equivariant MLIPs are examined, and a family of foundation potentials in the NequIP and Allegro equivariant MLIP architectures are presented which achieve leading inference speeds and strong scalability as well as excellent accuracies across a range of community benchmarks.
Seán R. Kavanagh, Chuin Wei Tan, Menghang Wang et al.· 0 citations
Foundation machine learning interatomic potentials (MLIPs) deliver near-ab-initio accuracy at a fraction of the computational cost, yet their promise for Metal-organic Frameworks (MOFs) remains largely unrealized as large unit cells make first-principles training data expensive to generate, fine-tuned models are scarce, and experimentally grounded benchmarks are scarcer still. We introduce uMOF, a three-part contribution addressing this gap. First, we release the largest and most accurate density functional theory dataset for MOFs to date, computed at the r$^2$SCAN-D4 level of theory across 85524 configurations spanning 19950 unique frameworks and 79 elements, covering empty and gas-loaded structures, geometry optimizations, equations of state, and finite-temperature molecular dynamics. Second, we release a literature-mined benchmark of 3986 verified property values (3146 experimental) extracted from 626 papers by a seven-stage, checkpointed multi-pass large language model pipeline, linked to more than 650 crystallographic information files. Third, we release two universal MLIPs for MOFs, uMOF-MH and uMOF-POLAR, fine-tuned from two architecturally distinct MACE foundation models on the uMOF dataset. On near-equilibrium, ``Tier-1''properties (bulk modulus, phonon-derived heat capacity) the uMOF models perform comparably to existing foundation and fine-tuned baselines. On harder, dynamics-sensitive properties like gas adsorption enthalpies via Widom insertion and adsorption isotherms, the uMOF models outperform every baseline we test, including MOF-specialized gas-capture models trained on datasets up to three orders of magnitude larger, cutting error by more than 80% to within experimental uncertainty. We trace this advantage to the physical diversity of the training data and to level of theory where a small (1.7%) fraction of MD simulations is decisive for MLIP stability.
T. J. Inizan, Prathami Divakar Kamath, A. Elena et al.· 0 citations
AI2Pot is presented, a scalable and unified MLIP framework that seamlessly integrates model training, evaluation, and large-scale MD simulations with PyTorch-compatible ecosystem, and offers an user-friendly end-to-end framework for the developing, training, and deploying MLIPs for large scale MD.
Hanyu Liu, Linggang Zhu, Xuanguang Zhang et al.· 0 citations