Skip to content
Preprint

Reconstructing local environments from concise atomistic representations

Jul 2026 · 1 citation · 74 references
Physics

TL;DR

This work investigates the inverse problem of recovering atomic structures from local invariant descriptors, and shows that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features.

Abstract

Symmetry-based representations of local atomic structure, such as the power spectrum or bispectrum, are routinely used to characterize the structural diversity of datasets and as input features for atomistic machine learning. Although these descriptors systematically incorporate increasingly complex geometric correlations, it remains unclear if a given feature can be mapped back to a discrete point cloud, whether such a reconstruction is unique, and how changes in the descriptor are reflected in the underlying atomic geometry. The choice and discretization of the radial and angular bases, as well as the high dimensionality of the resulting feature vectors -- which may contain hundreds or thousands of components -- make this interpretation even more challenging. In this work, we investigate the inverse problem of recovering atomic structures from local invariant descriptors. We show that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features. Even representations that are formally incomplete or locally ill-conditioned can be inverted to accurate geometric reconstructions of atomic environments across molecular and material datasets. Our reconstruction framework provides a general algorithmic means of identifying approximate degeneracies of invariant descriptors and recovering distinct atomic environments that cannot be distinguished by a given representation. Finally, by reconstructing atomic configurations from descriptors, we examine how perturbations in invariant descriptors of different correlation orders translate into structural distortions.

View source

Similar papers

Preprint Jul 2026

Rem3Di: Learning smooth, chiral 3D molecular descriptors from atomistic foundation models

Rem3Di is introduced, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening and provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.

Steffen Wedig, Felix Burton, Rokas Elijošius et al. · 0 citations
Open access Feb 2026

Machine learning of electronic structure and atomistic properties from the external potential.

This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.

Jigyasa Nigam, T. Smidt, G. Dusson · 2 citations
Preprint Jul 2026

Anisotropic representations for E(3)-equivariant machine learning coarse-grained potentials

A novel anisotropic machine learning CG potential is introduced that extends the point particle representation of atomic nuclei to massive ellipsoidal beads with orientation-dependent features, enabling the learning of energies, forces, and torques directly from atomistic data.

V. Shankar, Emil Annevelink · 0 citations
Preprint Jul 2026

Using large language models to probe the limits of atom-centered structural descriptors

Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies''reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.

M. Domina, Michele Ceriotti · 1 citation
Preprint Aug 2026

FUCrIMODo: structure recovery from atomistic descriptors via multi-stage genetic algorithms

Data-driven approaches to materials discovery rely on numerical representations of atomic structures as input for machine learning models. Inverting these descriptors - recovering atomic structures from their representations - is essential for most generative material design pipelines, yet it remains challenging, particularly for periodic systems. Existing inversion methods are either tailored to specific invertible descriptors or require candidate structures with similar atomic arrangements and compositions, limiting the exploration of novel regions in chemical and configurational space. Here, we propose a generalizable, similarity-driven sampling approach, powered by a novel stage-wise optimization strategy, to recover atom types, atomic positions, and unit cell shapes directly from a descriptor. Our approach requires only descriptor features and parameters as input without any prior structural knowledge. The capability of our method is demonstrated by the averaged Smooth Overlap of Atomic Positions (SOAP) descriptor.

Louis Boehm, M. Kuban, Claudia Draxl · 0 citations
Book Open access Aug 2026

Capturing Motif Topological Diversity via Geometry-Adaptive Riemannian Molecular Representation Learning

Molecular properties are often governed by a small number of local substructures, or motifs, whose topologies can vary drastically across molecules. Existing molecular representation learning approaches typically embed all motifs into a single Euclidean or fixed-curvature space, which fails to capture the motif-level topological heterogeneity and leads to geometric mismatch, impairing property prediction. To address this challenge, we propose a geometry-adaptive Riemannian framework for molecular representation learning, which explicitly models motifs as the basic units and learns their embeddings across multiple constant-curvature spaces. Each motif is adaptively aligned with the geometric space that best fits its intrinsic topology, enabling simultaneous modeling of cyclic, hierarchical, and tree-like structures. Motif embeddings are then aggregated into molecule-level representations, emphasizing functional substructures while suppressing irrelevant background. Extensive experiments on benchmark molecular property prediction datasets demonstrate that our approach outperforms state-of-the-art baselines, shows strong generalization under distribution shifts, and provides interpretable motif-level insights, offering a general and scalable framework for scientific molecular modeling. Our code is available at https://github.com/qimuya/mo-mi-r.

Fei Liu, Wen-Kai Lu, Feilong Wang et al. · 0 citations