Aug 2026· ChemPhotoChem· Vol 10· 0 citations· 19 references
Abstract
Aggregation‐induced emission (AIE) has revolutionized the design of photoluminescent materials by enabling strong solid‐state emission from molecularly nonemissive compounds. However, rational prediction of AIE properties remains challenging because photophysical behavior depends not only on molecular structure but also on aggregate‐state packing and measurement conditions. This study develops a quantitative and interpretable machine learning (ML) framework for predicting experimentally reported emission energies of AIE‐active molecules using continuous physicochemical descriptors derived from molecular structures. A dataset of 590 AIE luminogens—including conjugated organics, donor–acceptor (D–A) systems, silicon‐containing luminogens, and transition‐metal complexes—was analyzed using Gaussian process regression (GPR) combined with SHapley Additive exPlanations (SHAP). The optimized descriptor‐based model achieved moderate predictive performance (test
R
2
= 0.58) and provided chemically interpretable structure–property trends. Feature attribution indicated that nitrogen‐ and sulfur‐containing motifs, electrotopological‐state descriptors, Burden–CAS–University of Texas (BCUT) descriptors, and stereodefined vinylene units are statistically associated with lower emission energies within the present dataset. Morgan fingerprint baseline models showed higher random‐split accuracy, whereas leave‐one‐cluster‐out validation revealed cluster‐dependent degradation for structurally separated regions. This work therefore provides an interpretable initial screening strategy for AIE luminogens while clarifying the need for future models incorporating measurement conditions, solid‐state structural descriptors, and electronic‐structure‐informed features.
The rational design of aggregation-induced emission (AIE) luminogens presents a significant challenge in molecular photophysics, requiring approaches that connect mechanistic understanding with practical molecular screening. This work uses the electronic-energy difference between the S1/S0 minimum-energy conical intersection (MECI) and the vertically accessed Franck–Condon S1 state, ΔE(MECI – FCS1), as an efficient descriptor of conical-intersection accessibility. We compile quantum-chemical labels for 228 structure-matched polycyclic aromatic molecules and develop a dual-model strategy that combines an interpretable fingerprint-based model with a Uni-Mol model for rapid property prediction. In an external panel of literature luminogens, the predicted values are significantly lower for AIE molecules than those for aggregation-caused-quenching (ACQ) molecules, supporting an empirical operating threshold near 0.6 eV. This threshold provides an efficient first-pass guide for prioritizing candidates with accessible CI channels, substantially narrowing the pool for subsequent validation. The resulting workflow connects mechanistic insight, interpretable design rules, and high-throughput screening, providing an efficient platform for accelerating the discovery of novel AIEgens.
Wenjie Zhang, Ping-An Yin, Qi Ou et al.· Journal of Chemical Theory a...· 0 citations
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
Frank T. Mtetwa, N. Giles, W. Wilding et al.· Journal of Chemical Physics· 0 citations
This work presents a machine learning-symbolic regression (ML-SR) strategy to develop a physically interpretable formula for predicting low pressure CO2 adsorption capacity in hypothetical metal-organic frameworks (hMOFs). Four ML models were trained on a small dataset of 1,000 samples, and five key descriptors-largest cavity diameter, pore limiting diameter, void fraction, gravimetric surface area, and number of hydrogen atoms-were identified through SHAP and feature importance analyses. Symbolic regression was then employed to derive a concise adsorption formula, Q=aA, where a represents an adsorption baseline (mmol/g) and A is a dimensionless adsorption number incorporating four structural descriptors. We interpret A as the ratio between an adsorption binding force and a diffusion driving force, revealing how pore topology and surface chemistry jointly influence adsorption. Validation against a comprehensive dataset of 137,652 hMOFs demonstrates that this formula achieves over 70% prediction accuracy for 62,448 structures, confirming strong applicability within defined structural and operational ranges. Unlike conventional black box ML models, the proposed physics-guided expression enables efficient prediction and provides clearer insight into adsorption mechanisms.
Yimin Shao, Sheng-Ling Ma, Shenghong Ju et al.· 0 citations
Two-dimensional transition metal dichalcogenides (TMDs) are promising gas-sensing materials, but adsorption behavior across doped host-dopant-gas spaces remains difficult to predict and interpret. Here, we develop a descriptor-informed machine-learning framework for adsorption-energy prediction and regime-level screening on doped TMDs. A final data set of 354 first-principles adsorption entries was constructed for six hazardous gases, four TMD hosts, and substitutional metal dopants using 12 adsorption-configuration-independent descriptors. Among nine regression models, the boosted-tree models GBR and XGB provided the most reliable predictions, with held-out test R2 values above 0.95. SHAP analysis highlights gas-phase zero-point energy, dopant valence electron count, and molecular dipole moment as influential model-level descriptors. By integrating regression, descriptor-level interpretation, adsorption-regime classification, and holdout validation, this work provides a thermodynamics-guided reference for rapidly screening doped TMD candidates for gas-sensing applications.
A design principle is proposed for advanced metal nitride HEDMs: prioritizing high nitrogen-to-metal ratios, light metal elements, and structures wherein nitrogen atoms are spatially separated by the metal matrix, which minimizes N-N bonds and favors dominant M-N bonding.
Yaozhong Liu, Huifang Du, Caimu Wang et al.· Chinese Physics B· 0 citations
Rationally selection of precursor additives during perovskite crystallization is essential for obtaining high‐quality films and achieving high‐performance perovskite solar cells (PSCs). However, the discovery and optimization of effective additives still rely heavily on time‐consuming and costly trial‐and‐error experiments. In this work, we proposed a machine learning (ML) assisted screening strategy for precursor additives by integrating process parameters, material physicochemical properties, and molecular descriptors into a unified feature system. Among the five ML algorithms evaluated, the random forest model achieved the best predictive performance with relative errors below 5% on the external dataset. SHapley Additive exPlanations (SHAP) analysis further quantified the contribution of key features to device efficiency, offering guidance for additive structural design and property optimization. Following the established prescreening rules, two additives, [1,2,4]triazolo[1,5‐
a
]pyridine‐6‐carboxylic acid (6‐CATPy) and 4‐hydroxybenzenesulfonamide (4‐HBSA), were selected for experimental closed‐loop verification. Both additives effectively improved the preferred crystallization orientation, surface morphology, and defect passivation of perovskite films, resulting in enhanced film quality and power conversion efficiencies (PCEs) of 24.07% and 25.44%, respectively. These findings validate the proposed data‐driven screening strategy and demonstrate its potential for accelerating the rational discovery and optimization of precursor additives for high‐performance PSCs.
Zhimin Feng, Kuo Wang, Di Huang et al.· Rare Metals· 0 citations