Skip to content
Open access

Interpretable machine learning for Curie temperature prediction of magnetic materials: compositional descriptors, shap analysis, and a perovskite case study

Aug 2026 · Scientific Reports · 0 citations

TL;DR

The Northeast Materials Database is leveraged to develop machine learning models that predict magnetic materials with targeted Curie temperatures from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators.

Abstract

The rational design of magnetic materials with targeted Curie temperatures ( $$T_C$$ ) remains a central challenge in materials science. In this study, we leverage the Northeast Materials Database (NEMAD), comprising 33,668 experimentally reported magnetic compounds, to develop machine learning models that predict $$T_C$$ from composition-derived descriptors rooted in molecular-level elemental properties, supplemented by a small set of coarse crystal-system and structure-family indicators. After rigorous data curation—including deduplication of 13,186 unique compounds and extraction of 49 physics-informed features encompassing electronegativity, atomic radii, ionization energy, and valence electron statistics—we benchmark four gradient-boosted and ensemble algorithms. XGBoost achieves the best predictive performance with a five-fold cross-validated $$R^2$$ of 0.79 and a mean absolute error of 70.2 K. SHapley Additive exPlanations (SHAP) analysis reveals that the fractional content of 3 d magnetic elements (Mn, Fe, Co, Ni, Cr) dominates the prediction landscape, contributing an average of 85 K to the SHAP value, followed by the average first ionization energy and minimum atomic number of the constituents. A dedicated deep-dive into 2,724 perovskite-structured compounds demonstrates that the global model generalizes well to this technologically important family ( $$R^2$$ = 0.73). Finally, we screen 20 hypothetical perovskite compositions and identify candidates—notably Sr $$_2$$ FeMoO $$_6$$ and La $$_{0.7}$$ Ba $$_{0.3}$$ MnO $$_3$$ —predicted to exhibit room-temperature ferromagnetism. These findings provide an interpretable, data-driven bridge between molecular-scale electronic descriptors and macroscopic magnetic ordering, offering a practical tool for accelerated discovery of high- $$T_C$$ magnetic materials.

Read PDF

Similar papers

Jul 2026

Predicting dielectric constants of crystalline materials using explainable machine learning and composition-aware feature engineering

An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.

D. Pundhir, Ashok Kumar · 0 citations
Aug 2026

The Critical Role of Tilting Descriptors in Data‐Driven Property Prediction of Orthorhombic Perovskites

This work introduces a verified and interpretable pathway for high‐throughput screening of orthorhombic perovskites and provides fundamental insight into the descriptor‐property connections governing different perovskite polymorphs.

Q. Fatima, A. A. Haidry, Usaid Ahmad et al. · 0 citations
Open access Jul 2026

Accelerating Bulk Modulus Design of High-Entropy Alloys Through Explainable Machine Learning and SHAP-Driven Insights

This work presents an interpretable machine learning (ML) system that uses composition- and physics-based descriptors to predict the bulk moduli of high-entropy alloys (HEAs). Extra Trees, Random Forest, Gradient Boosting, AdaBoost, and LightGBM are five ensemble ML algorithms that were systematically shaped and refined by hyperparameter fine-tuning. With a test R2 of about 0.852 and an RMSE and MAE of about 5.49 GPa and 1.5 GPa, respectively, Extra Tree outperformed the other optimized models, indicating good generalization capacity for untested HEA compositions. The computational efficiency results showed that LightGBM had the fastest prediction speed (~4.24 ms), whereas Extra Trees had the shortest training time (~17.3 s). The majority of the optimized models had statistically equal prediction performance (p > 0.05), according to statistical validation using paired t-test analysis, even though residual error distributions for the Extra Tree model established consistent and unbiased predictions. To enhance the interpretability of the model, SHAP-based explainable analysis was performed, which included SHAP importance, dependence, and waterfall plots. The SHAP results revealed that the primary determinants impacting bulk modulus behavior in HEAs were Zr content, mean electronegativity, Al content, bond strength, and melting-temperature-related parameters. The proposed framework enables the rapid identification and design of next-generation HEAs by permitting precise and computationally efficient bulk modulus prediction, as well as physically significant insights into descriptor–property connections.

Sandeep Jain, Naresh Kumar Wagri, Sunil Dohare et al. · 0 citations
Aug 2026

Machine learning prediction of organic compound melting points informed by condensed-phase and electronic descriptors.

Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.

Frank T. Mtetwa, N. Giles, W. Wilding et al. · 0 citations
Jul 2026

Data-Driven Exploration of the Polyethylene Catalyst Chemical Space via Machine Learning.

A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.

Xuefeng Li, Haoke Qiu, Hanwen Pei et al. · 0 citations
Preprint Aug 2026

Machine Learning Guided Discovery of Corundum High Entropy Oxides

Early thinking in the field of high entropy oxides (HEOs) emphasized their likely abundance, with combinatorial arguments hinting at a myriad of new materials. The experimental reality has proven more challenging: the stability of HEOs cannot be straightforwardly predicted based on ionic radii, lattice geometry, and charge-balancing considerations alone. In this work, we employ machine learning interatomic potentials (MLIPs) to predict the synthesizability of HEOs of the form $A_2$O$_3$ derived from a selection of trivalent cations. From nearly 500 possible compositions, we identify 16 promising candidates for experimental validation with solid-state and combustion synthesis. We discover three new HEOs in the corundum structure, including (Al,Cr,Fe,Rh,Sc)$_2$O$_3$, and one novel cation-ordered phase, (Al,Fe,Ga,Sc)$_2$O$_3$. By far the most common synthesis outcome was a mixture of competing phases, sometimes involving redox reactions. Our results also reveal profound synthesis method dependence for the final product, where qualitatively equivalent outcomes between the two synthesis methods were only observed for 3 of the 16 tested compositions. We conclude that the occurrence rate of HEOs is far rarer than initially believed and that machine learning approaches can effectively guide us to the"needle in the haystack".

Abraham A. Mancilla, O. A. Dicks, S. Aamlid et al. · 0 citations