Skip to content
Open access

A Leakage-Aware Benchmark Study of Machine Learning Models for Deep Eutectic Solvent Property Prediction

Aug 2026 · ACS Omega · Vol 11, pp. 48552 - 48562 · 0 citations · 24 references
Medicine

TL;DR

This study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.

Abstract

Predicting the physicochemical properties of deep eutectic solvents (DESs) remains challenging due to the large combinatorial design space and the complex, composition- and temperature-dependent interactions governing their behavior. While machine learning (ML) has been widely applied to DES property prediction, reported performance is often sensitive to data set structure, feature representation, and validation design, raising questions about the reliability and transferability of existing models. In this work, we present a systematic evaluation of descriptor-based ML models for DES property prediction using a curated and high-confidence data set spanning 2003–2026. A unified feature representation combining molecular descriptors, molar composition, and temperature is employed, and model performance is assessed under a hierarchy of validation protocols designed to control for data leakage and progressively increase extrapolation difficulty. The results show that predictive performance is strongly dependent on the validation design. Density and refractive index exhibit relatively stable behavior, while surface tension shows moderate predictability. In contrast, electrical conductivity and viscosity display a strong dependence on temperature and limited contribution from descriptor-based features. Under extrapolative validation, performance for these properties deteriorates substantially, indicating limited transferability. Additional analyses based on feature-space distance and similarity-based baselines indicate that prediction errors are strongly associated with training-domain proximity and that local similarity in the current feature space is insufficient for reliable prediction outside observed data regions. These findings suggest that the primary limitation arises from the representational capacity of static descriptor-based features rather than model choice alone. Overall, this study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.

Read PDF

Similar papers

Open access Jul 2026

Machine Learning-Based Prediction of CO2 Solubility in Deep Eutectic Solvents

Deep eutectic solvents, as an emerging class of green solvents, have demonstrated great potential in gas absorption and separation owing to their favorable physicochemical properties. However, accurate prediction of CO2 solubility in deep eutectic solvents across a wide range of temperatures and pressures remains a major challenge, limiting their optimization in carbon capture applications. In this work, two input representations, Simplified Molecular Input Line Entry System (SMILES)-based structural coding and physicochemical descriptors, were comparatively evaluated for CO2 solubility prediction in deep eutectic solvents. The dataset includes predominantly choline chloride-based deep eutectiv solvents, together with selected betaine-based and ammonium salt-based systems, spanning both hydrophilic and limited hydrophobic subclasses. The dataset contains 2,648 experimental measurements corresponding to 93 independent hydrogen bond acceptor-hydrogen bond donor (HBA-HBD) systems under different temperatures, pressures, and compositions. Four machine learning algorithms—extreme gradient boosting, random forest, deep neural network, and convolutional neural network—were evaluated using two input representations. All measurements associated with the same HBA–HBD pair were retained within the same data subset. Among the evaluated model–input combinations, the structural-coding-based random forest model achieved the highest test-set performance, with an R2 of 0.971 under the adopted random split. This study provides an exploratory comparison of machine learning strategies for CO2 solubility prediction within the deep eutectic solvent chemical space represented by the collected dataset.

Yiwen Wang, Shijia Peng, C. Nwaoha et al. · 0 citations
Aug 2026

Machine learning guided discovery of deep eutectic solvents for NH 3 capture: Experimental validation and mechanism

Deep eutectic solvents (DESs) have considerable potential for NH 3 capture, but traditional solvent screening methods are unable to identify appropriate DES efficiently. One thousand nine hundred fifty‐nine experimental solubility data points for 72 DESs were used to construct and compare multiple machine learning models based on σ‐profile descriptors to predict NH 3 solubility in DESs. CatBoost achieved the best performance ( R 2  = 0.993, RMSE = 0.079). Nested cross‐validation and independent test sets confirmed that the model has good physical consistency and cross‐system generalization. SHAP analysis further quantified the contributions of key features. The final model was then employed to predict the NH 3 solubilities of 1140 DESs, and the highest‐ranked systems were selected for subsequent characterization and absorption experiments. The results of the gas absorption performance experiment are in excellent agreement with the predictions of the model. Finally, quantum chemical calculations were used to clarify the microscopic mechanisms underlying DES formation and NH 3 interaction.

Lu Gao, Ruixin Li, Lili Wang et al. · 0 citations
Open access Aug 2026

When high accuracy misleads in literature-derived machine learning for deep eutectic solvent recommendation

Deep eutectic solvents (DESs) are widely used in analytical sample preparation, yet selecting suitable systems for specific analytes remains challenging due to the large combinatorial design space and reliance on empirical screening. Machine learning (ML) has been proposed as a data-driven alternative, but its reliability under literature-derived data constraints is unclear. In this study, a literature-derived experiment-level dataset for DES-based pesticide extraction was reconstructed and evaluated under a leakage-controlled, DOI-grouped validation framework. After screening, cleaning, and descriptor eligibility filtering, the final modeling table comprised 757 records from 94 studies, organized into 626 (DOI, analyte) groups. Following a comprehensive leakage audit, the modeling pipeline was rebuilt using a split-first, training-only strategy with leakage-clean descriptors. The hybrid model, combining global classification and pairwise preference learning, improved the record-level classification performance (ROC-AUC, MCC) relative to the baseline, but did not provide a stable group-level ranking benefit under the present literature-derived data structure. In restricted comparable multi-candidate groups, ranking performance decreased and the hybrid model frequently underperformed, with no statistically significant advantage observed. Sanity baseline analysis further showed that the baseline model captured non-trivial structure beyond random and simple heuristic ranking in comparable groups, whereas the hybrid formulation did not provide any stable additional benefit. These results are explained by sparse comparative structure, study-bounded target definition, and limited descriptor representation. Overall, this work demonstrates that in literature-derived DES extraction settings, improvements in predictive accuracy do not necessarily translate into reliable recommendation capability, highlighting the need for data-centric evaluation, comparable candidate reporting, and structured experimental metadata for ML-driven chemical recommendations.

Hakim Faraji, J. Méndez-Pérez, R. Rodríguez-Ramos et al. · 1 citation
Jul 2026

Predicting dielectric constants of crystalline materials using explainable machine learning and composition-aware feature engineering

An explainable machine-learning framework was developed for dielectric constant prediction using 52,168 crystalline materials extracted from the Joint Automated Repository for Various Integrated Simulations (JARVIS-DFT) database, demonstrating the complementary roles of electronic structure and elemental chemistry.

D. Pundhir, Ashok Kumar · 0 citations
Open access Aug 2026

A dual-descriptor machine learning framework for evaluating perovskite passivation materials

Molecular passivation plays a crucial role in improving the efficiency and stability of perovskite optoelectronic devices. However, quantitative evaluation of passivation materials remains challenging because interfacial binding strength and lattice distortion must be considered simultaneously. Here, we develop a dual descriptor machine learning framework based on a density functional theory derived dataset to evaluate binding energy (BE) and lattice distortion value (LDV) from molecular structural descriptors. Among the evaluated algorithms, random forest achieves the best performance for both targets, with a root mean squared error of 0.524 and a correlation coefficient of 0.976 for BE prediction, and a root mean squared error of 0.048 and a correlation coefficient of 0.848 for LDV prediction. Feature analysis reveals that BE is primarily governed by descriptors related to ammonium group electronic effects and molecular polarity, whereas LDV is influenced by a broader set of electronic and steric features, reflecting the more complex origin of lattice distortion. Validation using representative modifiers further shows that the strongest binding does not necessarily correspond to the most desirable passivation behavior. This work establishes a physically informed dual descriptor strategy for evaluating perovskite passivation materials and suggests that promising modifiers should combine sufficient interfacial binding, moderate molecular polarity, and limited lattice perturbation.

Yao Lu, Jie Dong, Juan Meng et al. · 0 citations
Open access Aug 2026

Boosting the Prediction Accuracy of Glass Transition Temperature in Polyimides: A Hybrid Machine Learning Approach Integrating Morgan Fingerprints and Molecular Descriptors

The glass transition temperature (Tg) of polyimides is a critical parameter determining their processability and application performance. Traditional experimental methods for measuring Tg are time‐consuming and costly, while existing machine learning prediction models predominantly rely on manually defined molecular descriptors, which often fail to fully capture detailed molecular structural information, limiting their prediction accuracy and generalization capability. To address this, this study proposes a hybrid feature engineering strategy combining Morgan fingerprints and molecular descriptors to comprehensively represent the chemical structure of polyimides. Based on a dataset of 1257 polyimide samples from a public database, we systematically compared six feature selection methods and employed multiple mainstream machine learning algorithms for modeling. The results show that the CATB model performed best, achieving a coefficient of determination (R2) of 0.882 and a mean absolute error (MAE) of 17.34 °C on an independent test set, with fivefold cross‐validation further confirming the model's robustness. SHAP interpretability analysis revealed the significant influence of key features such as the number of rotatable bonds, ether bonds, and ether‐linked oxyethylene units on Tg, providing clear guidance for molecular design. External validation demonstrated the model's strong generalization ability. This study not only achieves high‐precision and robust Tg prediction but also highlights the importance of hybrid feature strategies in polymer property modeling, offering a data‐driven foundation for the rational design of polyimides.

Peishuai Xing, Xiaodong Guo, Yang Wang et al. · 0 citations