Boosting the Prediction Accuracy of Glass Transition Temperature in Polyimides: A Hybrid Machine Learning Approach Integrating Morgan Fingerprints and Molecular Descriptors
Aug 2026· Molecular Informatics· Vol 45· 0 citations· 42 references
Medicine
Abstract
The glass transition temperature (Tg) of polyimides is a critical parameter determining their processability and application performance. Traditional experimental methods for measuring Tg are time‐consuming and costly, while existing machine learning prediction models predominantly rely on manually defined molecular descriptors, which often fail to fully capture detailed molecular structural information, limiting their prediction accuracy and generalization capability. To address this, this study proposes a hybrid feature engineering strategy combining Morgan fingerprints and molecular descriptors to comprehensively represent the chemical structure of polyimides. Based on a dataset of 1257 polyimide samples from a public database, we systematically compared six feature selection methods and employed multiple mainstream machine learning algorithms for modeling. The results show that the CATB model performed best, achieving a coefficient of determination (R2) of 0.882 and a mean absolute error (MAE) of 17.34 °C on an independent test set, with fivefold cross‐validation further confirming the model's robustness. SHAP interpretability analysis revealed the significant influence of key features such as the number of rotatable bonds, ether bonds, and ether‐linked oxyethylene units on Tg, providing clear guidance for molecular design. External validation demonstrated the model's strong generalization ability. This study not only achieves high‐precision and robust Tg prediction but also highlights the importance of hybrid feature strategies in polymer property modeling, offering a data‐driven foundation for the rational design of polyimides.
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
Frank T. Mtetwa, N. Giles, W. Wilding et al.· Journal of Chemical Physics· 0 citations
The glass transition temperature (Tg) is a critical descriptor governing the morphological stability, emitter orientation, and interfacial integrity of amorphous thin films in organic electronics. However, experimental Tg measurements suffer from high resource costs and interlaboratory variability, while machine learning models are bottlenecked by scarce, noisy data sets. Here, we establish a physics-based atomistic molecular dynamics (MD) protocol to predict the Tg of 160 diverse organic electronic materials. To study computational throughput and predictive accuracy, we systematically benchmarked nine configurations spanning system sizes (5,000, 10,000, and 15,000 atoms) and cooling step relaxation times (5, 10, and 15 ns). Extracted via an automated, bias-free hyperbolic fitting scheme, our preferred standalone workflow (15,000 atoms, 15 ns) yields a correlation of R2 = 0.89 and a mean absolute error (MAE) of 10.9 K relative to experiment. Structural descriptor analysis confirms that accuracy remains uniform regardless of molecular weight or heteroatom density, establishing this transferable workflow as a digital sieve to accelerate the discovery of next-generation organic electronics.
Hadi Abroshan, Paul Winget, H. Kwak et al.· Journal of Physical Chemistr...· 0 citations
This study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.
Hakim Faraji, Julio Brito Santana, R. Rodríguez-Ramos et al.· ACS Omega· 0 citations
Machine learning (ML) models trained on bulk-crystal descriptors are increasingly used to prescreen catalysts by predicting adsorption energies, yet reported performances often rely on random K-fold cross-validation that permits the same bulk formula to appear in both training and test sets. We construct a reproducible benchmark that fuses 936 CatApp DFT adsorption energies with bulk descriptors from the Materials Project for H*, O*, and OH* on metal and alloy surfaces. We compare random K-fold cross-validation with GroupKFold grouped by parsed formula, the latter mimicking the realistic task of predicting adsorption on entirely new catalyst compositions. Under formula-grouped evaluation, random CV materially overestimates apparent generalization performance, with the largest and most robust effects for H* and OH* (protocol-inflation gaps up to approximately 0.8). The H* and OH* results are based on only 20 and 31 unique formulas, so their GroupKFold Spearman point estimates should be read as directional evidence rather than quantitative estimates. O* shows a smaller and statistically fragile protocol-inflation signal and, even where composition-plus-bulk features improve Random Forest and Ridge, the usable signal is best described as a very coarse pre-filter within a limited domain. Bulk descriptors are adsorbate-dependent: they improve O* prediction for Random Forest and Ridge, but degrade H* and OH*—a qualitative, directional observation given the small formula counts—whose binding is poorly captured by bulk crystal descriptors, consistent with the established view that it is governed by surface-localized electronic structure. These results outline a realistic performance boundary for bulk-to-surface ML in this benchmark: O* can be very coarsely prioritized from bulk descriptors within a limited domain, whereas H* and OH* are unlikely to be quantitatively predicted from bulk descriptors alone and would benefit from surface-aware models. We therefore recommend that bulk-to-surface adsorption-energy benchmarks report formula-grouped cross-validation alongside random cross-validation as a more robust and transparent practice.
The paper presents an approach to constructing predictive models for the physical and mechanical properties of elastomeric composites using machine learning methods. The relevance of the study is driven by the need to accelerate the development of new materials and reduce the labor intensity of full-scale experiments. An automated machine learning algorithm is proposed, encompassing stages of input data unification, feature space formation, and comparative analysis of regression models (Random Forest, Gradient Boosted Decision Trees, Gradient Tree-Boosting Tweedie, Poisson Regression, Light Gradient Boosting Machine и Stochastic Dual Coordinate Ascent). During experimental validation on datasets containing formulation data with varying content of sulfur, natural rubber (NR), and zinc oxide, predictions were made for theoretical density, Karrer plasticity, brittleness temperature, and curing temperature. It was established that ensemble methods demonstrate the highest predictive capability; however, model accuracy significantly depends on sample representativeness. Intervals of the studied parameters (particularly the 140–150 °C range for curing temperature) characterized by increased prediction uncertainty were identified, requiring additional algorithm calibration. The obtained results confirm the effectiveness of the proposed approach for formulation optimization and identification of hidden dependencies in the “composition-property” system.
M. Maslova, V. Kablov, A. Rybanov· Computational nanotechnology· 0 citations