Aug 2026· Advanced Theory and Simulations· Vol 9· 0 citations· 46 references
TL;DR
This work introduces a verified and interpretable pathway for high‐throughput screening of orthorhombic perovskites and provides fundamental insight into the descriptor‐property connections governing different perovskite polymorphs.
Abstract
Accurate prediction of the properties of all relevant polymorphs is necessary for rational design of stable, high‐performance perovskite materials. While the discovery of cubic perovskites has been hastened by machine learning (ML), the orthorhombic (Pnma) phase needed for device performance at room temperature has been little studied due to its complicated structural distortions. This work addresses this gap by developing a specialized ML framework for orthorhombic ABX3 halide perovskites. We assembled a curated dataset of 3000 Pnma structures and developed a set of 67 descriptors specifically describing octahedral tilting, anisotropic distortion, and electrical effects. A thorough benchmark of twelve algorithms revealed XGBoost to be the best model, yielding credible predictions for formation energy (test R2 = 0.939) and band gap (test R2 = 0.619). Interestingly, the recursive feature removal showed that a low‐dimensional electronic subspace controls formation energy, while the prediction of the band gap requires a complicated interplay of structural and electronic features. In contrast, prediction of thermodynamic stability (energy above hull) remained a challenge, revealing the limitations of existing compositional descriptors. This work introduces a verified and interpretable pathway for high‐throughput screening of orthorhombic perovskites and provides fundamental insight into the descriptor‐property connections governing different perovskite polymorphs.
Predicting electrical conductivity in perovskite and double-perovskite materials remains challenging as this property depends on various electronic, chemical, and structural factors. In this work, we evaluate classical machine-learning models trained on different descriptor sets including DFT band gap, non-orbital compositional descriptors, orbital-related descriptors, SOAP structural fingerprints, and a reduced mixed descriptor set to predict DFT-derived transport conductivity. The band gap provides a strong baseline but is insufficient to fully predict the target. The best overall performance is obtained using non-orbital compositional descriptors with Random Forest regression, while orbital-related descriptors achieve nearly comparable accuracy, confirming the importance of valence-electron characteristics. A compact mixed descriptor set preserves nearly the full predictive power of the larger descriptor spaces, showing that accurate prediction can be achieved using a small number of physically motivated variables. In contrast, SISSO showed much lower accuracy, suggesting that sparse symbolic expressions are insufficient to capture the nonlinear relationships underlying our target values. These results demonstrate that physically informed classical machine-learning models can provide an effective surrogate framework for reproducing DFT/BoltzTraP-derived conductivity trends in perovskite and double-perovskite materials.
Fatemeh Mohammad Dezashibi, F. Roshani· Scientific Reports· 0 citations
Molecular passivation plays a crucial role in improving the efficiency and stability of perovskite optoelectronic devices. However, quantitative evaluation of passivation materials remains challenging because interfacial binding strength and lattice distortion must be considered simultaneously. Here, we develop a dual descriptor machine learning framework based on a density functional theory derived dataset to evaluate binding energy (BE) and lattice distortion value (LDV) from molecular structural descriptors. Among the evaluated algorithms, random forest achieves the best performance for both targets, with a root mean squared error of 0.524 and a correlation coefficient of 0.976 for BE prediction, and a root mean squared error of 0.048 and a correlation coefficient of 0.848 for LDV prediction. Feature analysis reveals that BE is primarily governed by descriptors related to ammonium group electronic effects and molecular polarity, whereas LDV is influenced by a broader set of electronic and steric features, reflecting the more complex origin of lattice distortion. Validation using representative modifiers further shows that the strongest binding does not necessarily correspond to the most desirable passivation behavior. This work establishes a physically informed dual descriptor strategy for evaluating perovskite passivation materials and suggests that promising modifiers should combine sufficient interfacial binding, moderate molecular polarity, and limited lattice perturbation.
Yao Lu, Jie Dong, Juan Meng et al.· RSC Advances· 0 citations
Perovskite oxides have emerged as an important class of material with promising energy applications owing to their compositional and structural flexibility, which enables stabilization of both low- and high-symmetry phases and gives rise to diverse physical properties. Under ambient conditions, most perovskites adopt low-symmetry structures characterized by octahedral tilting and B-site displacements. Despite their importance, computational studies have largely focused on the ideal cubic phase as modeling these distortions remains challenging. The difficulty stems from the absence of a quantitative framework capable of capturing composition-dependent distortions that can occur through multiple non-equivalent atomic displacement modes, often requiring computationally expensive large supercells to explore the structural landscape. Consequently, the influence of distortions on the stability and properties of low-symmetry perovskites remains insufficiently understood. In this work, we develop an efficient computational framework for the rapid construction and exploration of composition-dependent structural models across both low- and high-symmetry phases. Using $\textit{symmetry constrained templates}$ and $\textit{unconstrained supercell templates}$, we systematically investigate 15 representative compositions to uncover relationships between composition, supercell size and shape, and distortion patterns. Based on these insights, we propose a robust and computationally inexpensive protocol for rapid structural exploration and assess the influence of different distortion modes on key physical properties.
Panupol Untarabut, Sylvian Cadars, F. Pascale et al.· 1 citation
The Northeast Materials Database is leveraged to develop machine learning models that predict magnetic materials with targeted Curie temperatures from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators.
F. Uçar, Nida Katı· Scientific Reports· 0 citations
Molybdenum carbide (MoC) nanoparticles (NPs) have attracted extensive interest due to their superior catalytic performance, yet studying the properties of realistic, experimental-scale models by means of first principles-based methods remains computationally unfeasible. Here, an efficient on-the-fly machine learning force field (MLFF) workflow was employed to overcome this difficulty. By sampling bulks, slabs, and clusters at the cost of thousands of DFT single-point calculations only, the present approach reliably predicts the properties of large-scale NPs. Our results revealed that subnanometric cubic δ-MoC clusters are energetically stable, whereas metastable hexagonal α-MoC clusters exhibit greater structural flexibility. Furthermore, a phase transition crossover diameter at ∼4.3 nm was identified, beyond which bulk-like α-MoC NPs replace δ-MoC ones as the most stable morphology. This rationalizes prior experimental observations that cubic δ-MoC phases are prevalent at small sizes while hexagonal α-MoC phases dominate the large particle size regime. The present study not only provides an efficient workflow to build reliable MLFFs for realistic transition metal carbide NPs but also provides critical insights into the size-dependent morphology to guide the rational design of MoC-based nanocatalysts and related materials.
Wei Cao, F. Viñes, Francesc Illas· Nanoscale· 0 citations
An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.
D. Pundhir, Ashok Kumar· Applied Physics A· 0 citations