Skip to content
Open access

Prediction of solubility of hydrogen in chemicals using QSPR-based machine learning approach: a comparative study

Jul 2026 · Journal of Computer-Aided Molecular Design · Vol 40 · 0 citations · 97 references
Medicine

TL;DR

For the first time, a comprehensive and predictive MLR-based QSPR model has been developed for this target and the predictive capability of the MLR-QSPR model was acceptable for training set.

Abstract

The ‘Quantitative Structure–Property Relationship’ (QSPR) method has been used for the prediction of solubility of hydrogen (x) in different chemicals. The dataset consists of 3761 datapoints including 100 unique chemicals at the wide ranges of T and P. An MLR-model, the simplest form of machine learning algorithm, was constructed using the selected descriptors to predict x in various chemicals. For the first time, a comprehensive and predictive MLR-based QSPR model has been developed for this target. The dataset was divided into a training set including 2570 datapoints and to a test set including 1191 datapoints. The advantages of this approach are thoroughly discussed and compared with other available models which were developed with other ML algorithms. Unlike previous models, internal validation was performed on the MLR-QSPR model. According to the results of statistical parameters (R2 = 0.96 and Q2LOO-CV = 0.96), the predictive capability of the MLR-QSPR model was acceptable for training set.

Read PDF

Similar papers

Open access Aug 2026

Prediction of the Molecular Lipophilicity of an Alkylphenol Family Using Quantum Chemistry and QSPR Methods

This study examined the relationship between the water/octanol partition coefficient of eighteen alkylphenols and molecular descriptors derived from quantum-chemical calculations using a quantitative structure-property relationship (QSPR) approach. The experimental database was divided into a training set of fourteen compounds and a test set of four compounds. Three descriptors, the electronic energy (ET), the energy of the Lowest Unoccupied Molecular Orbital (ELUMO) and the energy gap (ΔEgap) were used to develop a multiple linear regression model. The model showed a strong association between the experimental lipophilicity values and the selected descriptors, with R = 0.9850, R² = 0.9703, a standard deviation of 0.0853, and F = 108.9532. Internal validation using leave-one-out cross-validation and property randomisation, together with external validation using the test set and the Tropsha criteria, was applied to assess model stability and predictive performance. The applicability domain was evaluated using a Williams plot. Within the studied dataset and the defined applicability domain, the model showed close agreement between experimental and predicted water/octanol partition coefficients. The reported validation results indicate that the selected quantum-chemical descriptors can be used to model lipophilicity within this alkylphenol series. Predictions for additional alkylphenols should, however, remain restricted to compounds that fall within the model’s defined applicability domain.

Fatogoma Diarrassouba, K. Bamba, N. Ziao · 0 citations
Review Aug 2026

A glimpse at molecular descriptor selection algorithms in QSAR studies: including advantages and problems

Abstract One of the most frequently used methods in computational drug design is quantitative structure-activity relationship (QSAR). The main purpose of QSAR modeling is to estimate the relationship between chemical structures and biological activity in a group of molecules. In this method, molecules that have the greatest impact and the least side effects can be identified and extracted among huge numbers of molecular compounds. The molecular descriptors play a crucial role in QSAR design, and contain the physical, chemical, and geometric information. This information is called a feature that acts as an input to the QSAR model. Today, with the design of various applications to calculate molecular descriptors, the information obtained for each chemical structure is rising day by day which could lead to serious problems such as redundancy and over fitting. To solve this problem, researchers have used various techniques such as feature selection to improve the results of the model. The important point is that if the features are not properly selected, the QSAR model will fail. Up to now, different algorithms have been proposed to select the descriptors, which there are two main categories, supervised and unsupervised. The main purpose of this paper is to review feature selection methods in QSAR studies.

Fahimeh Motamedi, S. Zareian, S. Sardari et al. · 0 citations
Open access Jul 2026

LogPpred: An AI-Based Predictive Model for Accurate Estimation of Molecular LogP

Lipophilicity, commonly described by the n-octanol/water partition coefficient (LogP), is a key physicochemical property influencing the pharmacokinetic behavior of small molecules. Reliable LogP estimation during the early stages of drug discovery is essential to support molecular design and prioritize compounds with favorable ADMET properties. In this work, we report the development of LogPpred, an AI-based predictor of molecular lipophilicity. Starting from a curated dataset of 13,536 molecules with experimentally determined LogP values, multiple machine learning algorithms and molecular representations were systematically evaluated. The best-performing model, based on Gaussian Process Regression and RDKit molecular descriptors, achieved a mean absolute error (MAE) of 0.34 on the independent internal test set. External validation on a fully independent OECD-derived dataset yielded an MAE of 0.84, demonstrating good generalization capability across diverse chemical space. Furthermore, experimental LogP determination of an additional set of independently selected compounds confirmed the predictive reliability of the model, yielding an MAE of 0.66. Comparative analyses showed that LogPpred outperformed several widely used LogP prediction tools. Applicability domain analysis further supported the reliability of the model, with MAE values improving to 0.28 and 0.62 for in-domain compounds in the internal and external validation sets, respectively. Overall, LogPpred represents a robust, accurate, and transparent tool for the early assessment of molecular lipophilicity in medicinal chemistry and drug discovery.

Lisa Piazza, Lara Sortino, Alessio Costa et al. · 0 citations
Open access Aug 2026

Revisiting the LSER Approach in the Era of Machine Learning: Insights from IAM Chromatography

The present study demonstrates the integration of the linear solvation energy relationship (LSER) concept with machine learning (ML) methodologies to improve the predictive and interpretative capabilities of chromatographic retention modeling. Immobilized artificial membrane (IAM) chromatography was employed as a model biochromatographic system, and an in-house library of 993 structurally diverse compounds, with experimentally determined chromatographic hydrophobicity index of IAM (CHIIAM), was used to train LSER-ML models. LSER descriptors were calculated using Absolv and extended with ionization-state descriptors to evaluate the applicability of the LSER framework for realistic in silico virtual screening scenarios. Several regression algorithms were tested, including linear, neighborhood-based, kernel-based, and ensemble tree-based models. Among them, the support vector regression with the radial basis function kernel (SVR-RBF) demonstrated the most balanced performance across all validation metrics of R2train = 0.884, R2test = 0.853, and Q2cv = 0.811, achieving predictive errors (RMSEtrain = 5.546, RMSEtest = 4.609, and RMSEcv = 6.989) close to the analytical uncertainty. Model interpretability was achieved using SHapley Additive exPlanations (SHAP), which confirmed the mechanistic relevance of the Abraham descriptors and the dominant contribution of hydrophobic volume and hydrogen-bonding properties to IAM retention. Applicability domain was verified with a Williams plot (±3 standardized residuals and leverage threshold h*). The results indicate that the proposed LSER-ML approach provides an interpretable, robust, and generalizable tool for modeling membrane-mimetic chromatographic systems and can be effectively applied in virtual screening and property-based molecular design.

W. Nisterenko, K. Greber, Magdalena Kierkowicz et al. · 0 citations
Open access Aug 2026

Machine learning-based quantitative structure–activity relationship model for antibiotic prediction and discovery

Antimicrobial resistance (AMR) is a critical global health problem that has become increasingly alarming in recent years. The discovery of new antibiotics is one approach for alleviating AMR, however, screening for novel drugs is time consuming and expensive. To accelerate antibiotic discovery, the integration of machine learning algorithms with Quantitative Structure–Activity Relationship (QSAR) calculations could provide a rapid solution. Thus, this study combines a QSAR model and machine learning algorithms to predict antibacterial activities of potential novel drugs based on chemical information. Information on compounds that are reportedly active and inactive against bacteria was downloaded from the PubChem database and manually curated to create positive and negative datasets. The decision tree (DT), support vector machine (SVM), and naïve Bayesian (NB) algorithms were employed to predict the antibacterial activities of chemical compounds from their Simplified Molecular Input Line Entry System (SMILES) information. The models were then evaluated quantitatively and tuned. DT and SVM exhibited comparable predictive performance and outperformed the NB model, achieving accuracy, precision, sensitivity, and AUC-ROC values exceeding 0.90. DT was chosen for further analysis because of its simplicity and effectiveness. This revealed that descriptors relating to the electrotopology and β-lactam structures of compounds were the top contributors to model predictability. The model was then further tested against different classes of antibiotics and achieved high accuracy in all classes. The model is freely available as a web application at: https://antibacterial-predictor-model-ocogzqyibervrqb7trvfev.streamlit.app/.

Jiratchaya Nakbang, Chonthicha Arbsuwan, S. Prom-on et al. · 0 citations