Skip to content
Open access

Machine learning-guided screening and validation of antioxidant small molecules from Clausena lansium

Jul 2026 · Food Chemistry: X · Vol 38, pp. 104234 · 0 citations · 68 references
Medicine

Abstract

To discover natural antioxidants from Clausena lansium (Lour.) Skeels for food applications, we established a 474-compound database and applied a multiscale workflow integrating ensemble machine learning, molecular docking, structural clustering, molecular dynamics simulations, and quantum chemical calculations. The ensemble models achieved ROC-AUC values of 0.85–0.99 across eight antioxidant assays. Multi-criteria screening yielded 47 candidates, and 30 structurally diverse candidates underwent molecular dynamics simulations. Quercetin 3-arabinoside was prioritized, exhibiting stable non-covalent binding within the Keap1 Kelch domain during a 500 ns simulation. Quantum chemical calculations showed a HOMO–LUMO gap of 4.0589 eV, 25.5% lower than that of vitamin C. Given standard availability, its structural isomer Avicularin was assessed and exhibited dose-dependent antioxidant activity, with DPPH and ABTS IC50 values of 26.84 and 6.94 μM, respectively, and a FRAP value of 1.327 ± 0.218 mmol FeSO4 equivalents L−1 at 25 μM. This study supports antioxidant discovery from edible fruits.

Read PDF

Similar papers

Open access Jul 2026

Consensus ensemble learning with augmented molecular representation for accelerated virtual screening of kelulut honey phytochemicals targeting histamine H1 receptor

Traditional antihistamine discovery has relied on experimental screening followed by stepwise structural optimization. However, this approach is effective yet low and resource-intensive. Thus molecular docking and virtual screening were introduced to allow rapid evaluation of large compound libraries. This study aims to develop an H1R-specific ensemble model for predicting binding affinity scores between compounds extracted from Malaysian Kelulut Honey and H1R. In this work, four ensemble models were constructed in the KNIME Analytics Platform to be trained for binding affinity prediction. The best model was deployed to forecast the binding affinity scores of ligands present in Kelulut honey towards histamine H1 receptor (H1R). Among four ensemble models, Extreme Gradient Boosting (XGBoost) demonstrated superior robustness, yielding an R-squared (R2) value of 0.818, 0.830, and 0.799 in the training, test, and external validation set, respectively. From the predictive model, it was found that scaled molecular refractivity (SMR) from 2D descriptors and unit weight.wlambda2 and unit weight.wlambda3, from 3D structural descriptors contribute the most to the binding affinity prediction. The predicted binding affinity was then followed by redocking in PyRx for validation of ligands identified from LC-QTOF-MS. Based on consensus scoring, XGBoost promoted the prediction score with the highest increment of 20.9% for CID73642 with the incorporation of multidimensional descriptors. The advantages of consensus ML-driven prediction over traditional VS provide insights for H1R-ligand interactions via feature importance, underscoring its potential to accelerate and optimize the effectiveness of drug development across various diseases.

Wei-Feng Tan, Syajaratul Hassan, R. Edros et al. · 0 citations
Open access Jul 2026

Machine Learning–Driven Discovery of Dietary Polyphenol DPP-4 Inhibitors via Molecular Docking

T2DM is a chronic metabolic disorder of rising global prevalence, in which DPP-4 serves as a key therapeutic target through its role in incretin hormone degradation. This study aimed to identify polyphenolic compounds derived from dietary sources as potential DPP-4 inhibitors through an in silico pipeline integrating machine learning (ML) and molecular docking. The bioactivity dataset for DPP-4 was retrieved from ChEMBL (CHEMBL284), processed into a binary classification dataset, and represented using 2048-bit Morgan fingerprints combined with five RDKit descriptors. Six ML models were integrated into a soft-voting ensemble, achieving a Matthews correlation coefficient (MCC) of 0.8334 and an AUC-ROC of 0.9739. Screening of 162 compounds from the Phenol-Explorer database yielded 11 potential active inhibitors. Molecular docking identified hesperetin (−8.548 kcal/mol) as the leading candidate, followed by pelargonidin, daidzein, and naringenin, with key binding residues including Ser209, Glu205/206, Tyr631, Arg125, and Tyr662.

Nur Laily Harfita, Ahmad Faisal Nasution, Zuliana Amalia et al. · 0 citations
Aug 2026

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Min-Li Gan, Raoqing Liu, Hongsheng Liu et al. · 0 citations
Open access Jul 2026

Machine Learning Integrated Designing and Screening of 8-Hydroxyquinoline Based Metallo-β-Lactamase Inhibitors

The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.

Anthony M. Baudino, Kari L. Stone · 0 citations
Open access Aug 2026

Chemoinformatic prioritization of bioactive compounds associated with the bread matrix based on ChEMBL pharmacological annotations and machine learning

Bread and cereal-based products represent not only a source of macronutrients but also a complex food matrix containing low-molecular-weight phenolic, aromatic and fermentation-derived compounds. A number of these molecules have been associated with antioxidant, anti-inflammatory, antifungal and receptor-mediated effects in experimental studies and pharmacological databases. However, these records mainly describe test systems for individual compounds and do not demonstrate the corresponding effects directly in bread products. The objective was to develop a reproducible chemoinformatics framework for internal prioritization of compounds associated with bread and cereal matrices based on structural features and ChEMBL pharmacological annotations. The study included 43 curated compounds. After structural standardization and InChIKey deduplication, 2,058 molecular features were calculated, including physicochemical descriptors and topological fingerprints. Functional labels were generated from ChEMBL records for four activity classes: antioxidant, anti-inflammatory, antifungal, and AhR-modulating. Logistic regression, random forest, CatBoost, and XGBoost were compared using nested repeated multilabel cross-validation; no independent external dataset was used in this study. Random forest showed the best balance of performance and robustness, with a macro average precision of 0.8041 (95 % CI: 0.7766–0.8320) and a macro ROC-AUC of 0.8056. The anti-inflammatory class showed the most stable performance, whereas the antioxidant class was more heterogeneous. The highest-ranking compounds were protocatechuic, caffeic, gallic and ferulic acids, quercetin, and (+)-catechin. Evaluation with molecular scaffold-based splitting reduced macro average precision to 0.6245, indicating limited transferability to structurally distant compounds. The proposed approach should be regarded as a tool for preliminary candidate selection for further external and biological validation, rather than as experimental confirmation of their functional effects in the bread matrix.

M. Kuznetsov, I. Nikitin · 0 citations