Machine-learning-guided prediction and mechanistic insights into support-acidity-regulated N2 selectivity over cu-based catalysts for acetonitrile oxidation.
Aug 2026· Journal of Colloid and Interface Science· Vol 725, pp.
141357
· 0 citations· 47 references
Medicine
Abstract
Selective conversion of nitrogen-containing species into harmless molecular nitrogen (N2) remains a key challenge in the catalytic oxidation of nitrogen-containing volatile organic compounds (NVOCs). Machine learning (ML) provides an effective approach for predicting catalytic performance and identifying key descriptors from complex literature-derived datasets. Herein, a literature-derived catalyst database was constructed to predict N2 selectivity during NVOC oxidation and clarify the factors governing nitrogen transformation. Thirteen descriptors related to catalyst composition, structural properties, support acidity, and reaction conditions were used to train eight ML models. Among them, the ExtraTrees model exhibited the best predictive performance, with a coefficient of determination of 0.958 and a root mean square error of 7.638 on the test set. Shapley additive explanations and partial dependence plots revealed that oxygen concentration, reactant concentration, reaction temperature, gas hourly space velocity, and support acidity were the dominant factors affecting N2 selectivity, with support acidity identified as the key catalyst-related descriptor. Guided by this descriptor-level insight, Cu/M and CuFe/M catalysts (M = SiO2, ZSM-5, and Al2O3) were prepared and evaluated for acetonitrile oxidation. The catalytic and spectroscopic results confirmed the predicted role of support acidity, showing that different supports regulate CH3CN adsorption, CN activation, and the evolution of hydrolysis and oxidation related nitrogen-containing intermediates, thereby affecting nitrogen-product distributions and N2 selectivity. This work integrates interpretable machine-learning prediction with targeted external validation and mechanistic analysis, providing mechanistic insight into support-acidity-regulated nitrogen transformation and guidance for designing NVOC oxidation catalysts with high N2 selectivity.
The electrochemical nitrogen reduction reaction (NRR) offers a promising route for sustainable ammonia synthesis under ambient conditions. Yet, its practical development remains constrained by low catalytic activity, poor selectivity against the competing hydrogen evolution reaction (HER), and the limited availability of rigorously validated experimental data. In recent years, machine learning (ML) has emerged as a powerful tool for accelerating NRR catalyst discovery by mining heterogeneous literature datasets, revealing structure–property relationships, and identifying unexplored catalyst electrolyte operating windows. This review presents a comprehensive account of ML applications in NRR, covering curated experimental databases, feature engineering based on atomic, structural, and DFT‐derived descriptors. Particular emphasis is placed on ML‐guided insights into single‐atom, dual‐atom, alloy, oxide, nitride, and defect‐engineered catalysts. Descriptor‐driven design rules are used to identify promising catalyst motifs through DFT screening and gradient‐boosted regression. The review critically examines the limitations of current ML‐assisted NRR research, including data scarcity, and the mechanistic overlap between NRR and HER. In this context, uncertainty quantification and interpretability tools such as SHAP analysis and partial dependence plots are highlighted as valuable strategies for improving model reliability. Finally, emerging closed‐loop DFT‐ML‐experiment workflows, active learning, and ML models are discussed as essential pathways toward practical, scalable ammonia electrosynthesis.
Predictive modeling of catalytic biomass gasification suffers from simplified categorical representations of catalysts. This study developed a dual system machine learning framework coupling catalyst properties with gasification outputs. The Bayesian-optimized gradient boosting was utilized to train a catalytic model on 149 experimental datasets, achieving a carbon conversion efficiency prediction accuracy of R2 = 0.937. Principal component analysis was applied to compress seven catalyst descriptors (including specifies surface area and active site size) into a single performance score (PC1), explaining 47.3% of the variance. This score serves as the continuous input for the multi-output gasification model to predict CH4, H2, CO, and CO2 yields (R2 = 0.940, 0.976, 0.916, and 0.920, respectively). Independent validation using pine sawdust and calcium oxide confirmed that the model predicts syngas composition with less than five percentage points of absolute deviation. This framework provides a tool for catalyst screening and process parameter optimization in biomass clean energy conversion systems.
Yadong Ge, Hongru Li, Zaixin Li et al.· Bioresource Technology· 0 citations
Biodiesel production over metal-doped biochar and activated carbon (AC) catalysts involves complex nonlinear interactions among feedstock characteristics, catalyst descriptors, and operating conditions, making accurate yield prediction a challenging task. While machine learning (ML) has shown potential in process modeling, existing studies lack catalyst-aware frameworks that integrate material and process descriptors within a unified representation for biodiesel yield prediction. Furthermore, current approaches are often limited by small datasets and insufficient model interpretability, restricting their ability to support reliable catalyst screening and process optimization. To address these challenges, this study develops an explainable ML framework for biodiesel yield prediction using a literature-derived dataset of metal-doped biochar and AC catalyst systems. The framework integrates catalyst, feedstock, and operating-condition descriptors, augments sparse experimental data through curve digitization, evaluates six ML models, and applies SHAP and CatBoost-based explainability analysis. The neural network model achieved the highest predictive accuracy on unseen data, with RMSE of 3.27%, MAE of 1.64%, and R2 of 0.95, whereas linear regression showed the weakest performance, highlighting the nonlinear behavior of the catalytic system. Validation using an independent experimental dataset further confirmed model generalization. Explainability analysis identified alcohol-to-oil ratio, reaction time, catalyst amount, reaction temperature, and feedstock acid value as the key factors governing biodiesel yield. The proposed ML framework provides an accurate and interpretable approach for catalyst screening and data-driven optimization of sustainable biodiesel production processes.
Menna Ebrahim, Mostafa Mahmoud, Fatma H. Ashour et al.· RSC Advances· 0 citations
CO2 methanation is a key route for CO2 valorization and renewable hydrogen storage. Machine learning (ML) provides an analytical framework for understanding catalytic systems. However, conventional ML models usually use one-hot encoding, which splits intrinsically related catalytic components into independent sparse binary variables and limits the understanding of complex relationships among catalysts, reaction operating conditions, and CO2 conversion. In this study, an interpretable Categorical Boosting model using native categorical encoding (CatBoost-NCE) was applied to a dataset containing 4026 experimental entries. By preserving native categorical descriptors, the model enabled direct quantitative analysis of active metals, supports, promoters, and preparation methods, while maintaining reliable predictive performance across three independent data splits (mean test R2 = 0.918 ± 0.003). PDP and SHAP analyses showed that operating conditions were the dominant factors, while catalyst composition and preparation descriptors also made substantial contributions. Among categorical catalyst descriptors, the relative influence followed the order support > active metal ≈ preparation method > promoter. Beyond global feature interpretation, active metal-conditioned temperature analysis revealed distinct temperature response patterns among different catalyst systems. Metal-conditioned support substitution analysis further showed that support effects were closely coupled with active metal. For Ni-based catalysts, CeO2 exhibited a positive and directionally stable substitution effect. The model guided X-Ni–CeO2 catalysts showed good agreement with the predicted conversion trend, with an external validation R2 of 0.876. Characterization and in situ DRIFTS measurements further supported the interactions between active metal and support, and elucidated the evolution of reaction intermediates. ReaxFF molecular dynamics simulations performed on bare Ni clusters provided atomistic insights into plausible reaction events and pathways over intrinsic Ni active sites during CO2 methanation. This study converts accumulated catalyst data into actionable design rules and provides a reliable tool for the data-driven discovery of heterogeneous catalysts.
Oxidase-like (OXD-like) nanozymes are vital biomimetic catalysts that regulate oxygen activation and reactive oxygen species generation, whereas oxygen reduction reaction (ORR) catalysts are key cathodic catalysts in energy conversion devices. Although both systems share the O2-to-H2O2/H2O pathways, machine-learning (ML) studies of them have diverged, with ORR limited by idealized data ecosystems and OXD-like catalysis hindered by heterogeneous datasets. Accordingly, with ORR and OXD-like catalysts as representative examples, this review centers on data-driven ML-assisted materials design and highlights the differentiated data ecosystems within chemical science. We provide a detailed discussion of database construction, feature selection, and ML workflows. Furthermore, we show that data origin critically shapes ML studies, with theory-derived ORR databases often being overly idealized and literature-derived OXD databases highly noisy, thereby affecting structural representations, descriptor dimensionality, high-throughput feasibility, and research paradigms. Overall, this review provides methodology- and perspective-oriented guidance for chemistry and materials science from the perspective of ORR and OXD-like catalysts. We aim to outline a data-centered ML-assisted materials design logic that could inform broader materials science research. This logic integrates automated high-throughput experimentation, iterative active learning, and high-throughput characterization, and provides a progressive framework linking experimental protocol optimization, high-performance material discovery, and the extraction of intrinsic principles.
Juan Zhang, Zhengen Gao, Liang Zhang et al.· Small· 0 citations
Data-driven machine learning (ML) methods are now indispensable for analyzing and predicting material properties, offering critical insights for the design and synthesis of new materials. This study leverages ML approaches to identify and analyze the key factors influencing the thermodynamic properties of metal hydrides. Utilizing a comprehensive dataset of 721 metal hydrides from the experimental literature, we developed ML models that demonstrated robust predictive capabilities, achieving a coefficient of determination (R2) of 0.912 for equilibrium hydrogen plateau pressure and 0.881 for dehydrogenation enthalpy. Feature importance analysis revealed three critical descriptors, specific volume per atom for a given composition, magnetic transition metal fraction, and mean shear modulus, that significantly impact the thermodynamic properties of hydrogen storage materials. This research provides valuable insights into the fundamental factors governing hydrogen storage in metal hydrides and presents a promising approach for the design and discovery of novel hydrogen storage materials.