Skip to content
Open access

Critical benchmarking of machine-learned interatomic potentials for intermolecular and noncovalent interactions

Sep 2026 · Machine Learning: Science and Technology · Vol 7 · 0 citations · 127 references
Physics

Abstract

Accurate benchmarking of intermolecular interaction energies is central to evaluating quantum chemical methods and guiding the development of reliable machine-learned interatomic potentials (MLIPs) for chemical and biological applications. In this work, we benchmark five MLIPs, namely AIMNet2(2023), AIMNet2(2025), MACE-OFF23(M), MACE-OMol, and UMA-S-OMol, across twenty-one datasets spanning hydrogen-bonded, dispersion- and π-dominated, sigma-hole, ionic and charge transfer, and repulsive nonequilibrium interactions, with reference values at or near CCSD(T)/complete basis set (CBS) accuracy. AIMNet2(2025) is a continually pretrained variant of AIMNet2(2023) that retains the original architecture but incorporates an additional 3.8 million structures specifically curated to improve the description of noncovalent interactions (NCIs). Across these chemically diverse test sets, AIMNet2(2025) delivers consistent and systematic improvements over its predecessor, with the most pronounced gains observed in the hydrogen-bonded, sigma-hole, and repulsive regimes, while remaining broadly competitive with the substantially larger MACE-OMol and UMA-S-OMol models. Nevertheless, the supramolecular S12L and L7 benchmarks show only marginal improvement, with all evaluated MLIPs exhibiting large errors driven by a small number of pathological complexes. Two factors beyond intrinsic model quality significantly influence the reported performance. First, partial overlap between training and benchmark data, quantified here through systematic overlap detection, inflates apparent accuracy for all models, most strongly for those trained on the OMol25 data. Second, differences in the density functional theory reference level used for MLIP training establish distinct and irreducible error floors relative to the CCSD(T)/CBS targets, meaning that superior benchmark performance may in part reflect closer proximity of the training functional to the reference method rather than stronger modeling capability. Sigma-hole interactions emerge as the interaction category with the lowest training-benchmark overlap across all models and therefore provide the most discriminating test of genuine generalization. Together, these findings demonstrate that meaningful MLIP evaluation must carefully account for data provenance, reference theory consistency, and the distinction between interpolation and true out-of-distribution generalization, particularly as standard NCI benchmark sets become increasingly absorbed into large-scale training datasets.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.