Skip to content
Open access

When high accuracy misleads in literature-derived machine learning for deep eutectic solvent recommendation

Aug 2026 · RSC Advances · 1 citation · 25 references
Medicine

Abstract

Deep eutectic solvents (DESs) are widely used in analytical sample preparation, yet selecting suitable systems for specific analytes remains challenging due to the large combinatorial design space and reliance on empirical screening. Machine learning (ML) has been proposed as a data-driven alternative, but its reliability under literature-derived data constraints is unclear. In this study, a literature-derived experiment-level dataset for DES-based pesticide extraction was reconstructed and evaluated under a leakage-controlled, DOI-grouped validation framework. After screening, cleaning, and descriptor eligibility filtering, the final modeling table comprised 757 records from 94 studies, organized into 626 (DOI, analyte) groups. Following a comprehensive leakage audit, the modeling pipeline was rebuilt using a split-first, training-only strategy with leakage-clean descriptors. The hybrid model, combining global classification and pairwise preference learning, improved the record-level classification performance (ROC-AUC, MCC) relative to the baseline, but did not provide a stable group-level ranking benefit under the present literature-derived data structure. In restricted comparable multi-candidate groups, ranking performance decreased and the hybrid model frequently underperformed, with no statistically significant advantage observed. Sanity baseline analysis further showed that the baseline model captured non-trivial structure beyond random and simple heuristic ranking in comparable groups, whereas the hybrid formulation did not provide any stable additional benefit. These results are explained by sparse comparative structure, study-bounded target definition, and limited descriptor representation. Overall, this work demonstrates that in literature-derived DES extraction settings, improvements in predictive accuracy do not necessarily translate into reliable recommendation capability, highlighting the need for data-centric evaluation, comparable candidate reporting, and structured experimental metadata for ML-driven chemical recommendations.

Read PDF