Oct 2026· Acta Crystallographica. Section B: Structural Science, Crystal Engineering and Materials· Vol 82· 0 citations
Medicine
TL;DR
This work assesses 464 ICSD structures from 1913 to 1929 using two independent tools: Eir v1.4.2 (bond valence sums, global instability index) and MaplePy (Madelung part of lattice energy).
Abstract
The earliest entries in the Inorganic Crystal Structure Database (ICSD) are both primary historical documents and active computational inputs. They record how the Braggs, Goldschmidt, Vegard, and others established the structural foundations of inorganic chemistry; they are also served, without distinction, to automated pipelines, machine-learning training sets, and high-throughput workflows. We assess 464 ICSD structures from 1913 to 1929 using two independent tools: Eir v1.4.2 (bond valence sums, global instability index) and MaplePy (Madelung part of lattice energy). The pre-1925 median global instability index is 0.077 valence unit (v.u.), below that of modern rock salt reference data (mean 0.192 v.u.); the Braggs' 1913 NaCl gives 0.005 v.u., Bragg's 1922 corundum 0.001 v.u. Early crystallographers were doing simple structures and getting them right. The problem is not quality but metadata. A small number of entries carry genuine topological errors (BeO in NaCl type, SnTe in zinc blende) whose corrective information already exists in database Comments fields but is not machine-readable, not reflected in quality classifications, and relatively invisible to downstream users. We propose six automated fitness-for-purpose flags that surface these cases without removing or modifying any historical record.
A correctly solved crystal structure should agree with the experimental data, and its geometry should correspond to a local minimum on the potential energy surface (PES). The idea of verifying crystal structure solutions by comparing them with their geometry-optimized versions was introduced 15 years ago. Recent developments in machine learning interatomic potentials (MLIPs) have made it possible to replace computationally expensive density functional theory (DFT) calculations with AI/neural-network-based alternatives. MLIPs can reach DFT-comparable precision with a substantial gain in speed. We selected one promising MLIP, Universal Models for Atoms, trained on the Open Molecular Crystals 2025 dataset, and processed a prefiltered subset of 216 919 structures from the Cambridge Structural Database. Due to the limitations of the MLIP available when this study commenced, ionic compounds, salts and metal-containing structures were excluded. The current methodology cannot process disordered structures, and available computational resources limit the maximum unit-cell volume that can be treated to 4000 Å3. All structures in the dataset were geometry optimized using the MLIP, and similarity descriptors were calculated to quantify the differences between the original and optimized structures. Automatic analysis was followed by the manual identification of issues indicated by the descriptors' values. We detected anomalies in experimental structures that had already passed all prior validation, as well as limitations in the reliability of the MLIP PES calculations. For 1867 crystal structures, bond-pattern change was observed, while 3331 structures showed a root-mean-square Cartesian displacement greater than 0.25 Å. Future improvements to the methodology and extension to systems not covered by this study are discussed.
This work demonstrates how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation.
Timo Reents, Marnik Bercx, Giovanni Pizzi· 0 citations
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.
The prediction of the structural stability of octet $AB$-type binary compounds is a classical materials informatics problem. The challenge is to capture the relative stability of 4-fold coordinated atoms in zincblende ($\beta$-ZnS) structure and 6-fold coordinated atoms in rocksalt (NaCl) structure, modulated by charge transfer and atomic-size differences. Previous structure maps and machine-learning approaches used atomic features such as valence-electron count, ionization potential and atomic radii, using either physical intuition or symbolic regression. Here, we demonstrate that explicitly incorporating the domain knowledge of the interatomic bonds can significantly and systematically improve the prediction of $\beta$-ZnS/NaCl stability. We encode this bonding information through a coarse-grained representation of the local electronic structure obtained by a recursive solution of a tight-binding bond model. The underlying pairwise Hamiltonians are taken from downfolded eigenstates of density-functional theory calculations for diatomic molecules and thereby include domain knowledge of the bond between specific $A-B$ pairs. The benefit of this description is demonstrated with an ensemble of independently trained Kernel Ridge or symbolic regression models combined with sequential feature selection. The obtained models are compared to a previous symbolic-regression model using the same set of \emph{ab initio} calculations for octet binaries as training data. We find a significant improvement in the prediction of the formation energy difference of $AB$ compounds as compared to previous works and demonstrate that an increasing amount of bond-informed recursion features improves the predictive accuracy.
Rohan D. Kumar, Mariano Forti, A. Naik et al.· 0 citations
This work introduces a verified and interpretable pathway for high‐throughput screening of orthorhombic perovskites and provides fundamental insight into the descriptor‐property connections governing different perovskite polymorphs.
Q. Fatima, A. A. Haidry, Usaid Ahmad et al.· Advanced Theory and Simulati...· 0 citations