This work investigates the inverse problem of recovering atomic structures from local invariant descriptors, and shows that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features.
Abstract
Symmetry-based representations of local atomic structure, such as the power spectrum or bispectrum, are routinely used to characterize the structural diversity of datasets and as input features for atomistic machine learning. Although these descriptors systematically incorporate increasingly complex geometric correlations, it remains unclear if a given feature can be mapped back to a discrete point cloud, whether such a reconstruction is unique, and how changes in the descriptor are reflected in the underlying atomic geometry. The choice and discretization of the radial and angular bases, as well as the high dimensionality of the resulting feature vectors -- which may contain hundreds or thousands of components -- make this interpretation even more challenging. In this work, we investigate the inverse problem of recovering atomic structures from local invariant descriptors. We show that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features. Even representations that are formally incomplete or locally ill-conditioned can be inverted to accurate geometric reconstructions of atomic environments across molecular and material datasets. Our reconstruction framework provides a general algorithmic means of identifying approximate degeneracies of invariant descriptors and recovering distinct atomic environments that cannot be distinguished by a given representation. Finally, by reconstructing atomic configurations from descriptors, we examine how perturbations in invariant descriptors of different correlation orders translate into structural distortions.
Rem3Di is introduced, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening and provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.
Steffen Wedig, Felix Burton, Rokas Elijošius et al.· 0 citations
This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.
Jigyasa Nigam, T. Smidt, G. Dusson· Journal of Chemical Physics· 2 citations
A novel anisotropic machine learning CG potential is introduced that extends the point particle representation of atomic nuclei to massive ellipsoidal beads with orientation-dependent features, enabling the learning of energies, forces, and torques directly from atomistic data.
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies''reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.
Data-driven approaches to materials discovery rely on numerical representations of atomic structures as input for machine learning models. Inverting these descriptors - recovering atomic structures from their representations - is essential for most generative material design pipelines, yet it remains challenging, particularly for periodic systems. Existing inversion methods are either tailored to specific invertible descriptors or require candidate structures with similar atomic arrangements and compositions, limiting the exploration of novel regions in chemical and configurational space. Here, we propose a generalizable, similarity-driven sampling approach, powered by a novel stage-wise optimization strategy, to recover atom types, atomic positions, and unit cell shapes directly from a descriptor. Our approach requires only descriptor features and parameters as input without any prior structural knowledge. The capability of our method is demonstrated by the averaged Smooth Overlap of Atomic Positions (SOAP) descriptor.
Molecular properties are often governed by a small number of local substructures, or motifs, whose topologies can vary drastically across molecules. Existing molecular representation learning approaches typically embed all motifs into a single Euclidean or fixed-curvature space, which fails to capture the motif-level topological heterogeneity and leads to geometric mismatch, impairing property prediction. To address this challenge, we propose a geometry-adaptive Riemannian framework for molecular representation learning, which explicitly models motifs as the basic units and learns their embeddings across multiple constant-curvature spaces. Each motif is adaptively aligned with the geometric space that best fits its intrinsic topology, enabling simultaneous modeling of cyclic, hierarchical, and tree-like structures. Motif embeddings are then aggregated into molecule-level representations, emphasizing functional substructures while suppressing irrelevant background. Extensive experiments on benchmark molecular property prediction datasets demonstrate that our approach outperforms state-of-the-art baselines, shows strong generalization under distribution shifts, and provides interpretable motif-level insights, offering a general and scalable framework for scientific molecular modeling. Our code is available at https://github.com/qimuya/mo-mi-r.
Fei Liu, Wen-Kai Lu, Feilong Wang et al.· Proceedings of the 32nd ACM...· 0 citations