Skip to content

Data-Driven Exploration of the Polyethylene Catalyst Chemical Space via Machine Learning.

Jul 2026 · Journal of Physical Chemistry Letters · 0 citations · 45 references
Medicine

TL;DR

A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.

Abstract

The discovery of highly active polyethylene (PE) catalysts demands a systematic understanding of structure-condition-activity relationships in a vast chemical space. In this Letter, we present a data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation. From a curated data set of 507 catalysts (bis(phenoxyimine) and bis(imino)pyridine ligands, seven metals), a gradient boosting regression (GBR) model achieves a test R2 of 0.91, outperforming convolutional and graph neural networks. SHAP analysis identifies topological (Chi2v), electronic (EState_VSA), and hydrophobic (SlogP_VSA) descriptors as governing activity and reveals a classical volcano-type temperature dependence, fundamentally governed by the Sabatier principle. A virtual library of 665 685 structures, constructed via combinatorial fragment assembly, extends the known chemical space substantially. High-throughput screening, coupled with SCscore filtering, yields 1090 synthetically accessible candidates with predicted activities exceeding 2 × 107 g mol-1 h-1. Substructure analysis uncovers metal-dependent design rules, in which early transition metals favor electron-deficient aromatics while late metals profit from moderately sized alkyls. This work establishes a practical route from experimental data to actionable catalyst designs.

View source

Similar papers

Aug 2026

Multiple-Kernel Ridge Regression for Learning the Structure-Electronic Property Relationships of Pyranoazacoronene COFs

Covalent organic frameworks (COFs) are highly ordered, porous organic materials whose reticular construction from tailored nodes and linkers enables atomic-level control over structure and function. The design space of COFs is vast with virtually unlimited combinations of nodes, linkers, and functional groups. Interpretable machine learning (ML) offers a pathway to navigate this complexity by identifying the structural features that govern materials performance, yet interpretability often comes at the cost of predictive accuracy. In this work, we introduce a novel multiple-kernel learning framework that achieves both accuracy and mechanistic insight. A multiple-kernel ridge regression (MKRR) model was trained on band gaps predicted from GFN1-xTB level theory for a data set of 232 theoretical pyranoazacoronene (PAC) COFs produced from eight different conjugated linkers and 29 functional groups. Modifying these building units alone produced a range of band gaps between 0.4–2 eV. Manual analysis of the theoretical band gaps versus the linker indicates that breaking the conjugation pathway by altering the bond angle or by introducing a σ-bond increases the band gap while increasing the length of the linker decreases the band gap. All functional groups appear to reduce the band gap with three specific electron withdrawing groups reducing the band gap near 0.4 eV. For the ML, the building units were represented with three independent kernels that encoded the local environments of each node, linker, and functional group calculated from the Smooth Overlap of Atomic Positions (SOAP). After decomposing each kernel’s contribution to the model’s global predictions, we found that the MKRR model successfully captures the underlying structure–property relationships that influence the band gap. These results demonstrate that MKRR is an effective and interpretable framework for understanding and designing functional COFs.

Alathea E. Davies, O. Adesina, Isabella M. Valdez et al. · 1 citation
Open access Aug 2026

Machine Learning-Driven Refinement of Reactive Force Fields via Hierarchical “Center-Environment” Features for Energetic Molecular Crystals

This work not only establishes a pioneering paradigm for interpretable ML-driven force field refinement but also provides the first feature engineering solution incorporating chemical, physical, and structural information specifically designed for the machine learning of energetic molecular crystals.

Qi He, Pengju Wang, Xudong He et al. · 0 citations
Open access Jul 2026

P2MAT: A machine learning (ML) driven software for Property Prediction of MATerial.

This study presents a data-driven machine learning approach to predict the melting points of organic compounds, leveraging both 2D and 3D molecular descriptors and indicates that ML models can significantly improve melting-point predictions, providing a robust tool for the scientific community.

Md Kamruzzaman, Alexander Landera, N. Menon et al. · 1 citation
Jul 2026

Machine-Learning-Guided Genetic Inverse Design of Single-Atom Electrocatalysts for CO2 Reduction

A genome-inspired materials intelligence framework (GIMI) for inverse design in high-dimensional compositional spaces is proposed, enabling targeted exploration of complex compositional space and accelerating the discovery of high-performance catalysts.

Chen Zhu, H. Yang, Haifeng Wang et al. · 0 citations
Aug 2026

Rapid Generative Discovery of High‐Energy Molecules in Low Data Regimes Using Minimal Computational Resources

This work presents a novel approach toward high‐energy molecules by combining long short‐term memory (LSTM) networks for molecular generation and attentive graph neural networks (GNN) for property predictions by combining fixed SHA‐256 embeddings with partially trainable representations.

Siddharth Verma, A. Alankar · 0 citations