Experimental analysis showing that the ensemble model predictions align with in vitro results provides initial support for the potential of ML to help reduce experimental complexity, though further validation across additional cell lines and coumarins is warranted.
Abstract
Coumarins are a group of naturally occurring compounds that have garnered significant interest for their potential anticancer properties, characterized by targeted effects and low toxicity. Despite this promise, the therapeutic application of certain coumarins is limited by insufficient data and inconsistencies in experimental results. To tackle these challenges, this study aims to explore the use of regression-based machine learning (ML) models to infer dose-exposure time windows corresponding to approximately 50% cell viability for natural coumarins esculetin, umbelliprenin, auraptene and galbanic acid across multiple human cancer cell lines. Extensive in vitro data from 64 peer-reviewed studies were extracted based on inclusion criteria, focusing on dose, time, and cell viability for four coumarins. Eight highly utilized ML regression algorithms, including Gradient Boosting Regressor (GBR), Histogram-based Gradient Boosting Regressor (HGBR), Random Forest Regressor (RFR), Extreme Gradient Boosting Regressor (XGBR), Support Vector Regressor (SVR), and regularized linear models Lasso, Ridge, and Elastic Net were employed to analyze the dataset and elucidate dose-response relationships. To validate ML predictions, in vitro experiments on a human colon carcinoma cell line, specifically not included in the original dataset, were conducted. Auraptene and galbanic acid were extracted from Ferula szowitsiana using thin layer chromatography, and LoVo cells were treated for 24, 48, and 72 h and evaluated for viability and apoptosis. The data analysis demonstrated that esculetin, umbelliprenin, and auraptene consistently exhibited robust dose-response dynamics, with notable time-dependent enhancements in their anticancer activity. Galbanic acid, however, showed cytotoxicity primarily at higher concentrations and with prolonged exposure. Interestingly, the predictions made by the gradient boosting ensemble models regarding the effects of auraptene and galbanic acid were validated through in vitro experiments on LoVo cells. Predictions from the remaining non-ensemble models, although in some cases consistent with the experimental outcomes, were not considered reliable for validation due to their comparatively low R2 scores. This research establishes grounds and proof of principle for understanding the efficacy of coumarins utilizing ML, highlighting their potential in contemporary oncology. Experimental analysis showing that the ensemble model predictions align with in vitro results provides initial support for the potential of ML to help reduce experimental complexity, though further validation across additional cell lines and coumarins is warranted.
Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.
D. Kumari, Aiman, Sakshi Singh et al.· Annual International Compute...· 0 citations
Colorectal-cancer (CRC) is the third leading cause of mortality due to cancer, thus there is a need for innovative-therapeutic-agents to enhance the efficacy of current treatments and improve outcomes. Here we performed different machine learning (MI) approaches (e.g., random forest, support vector machines, convolutional neural networks, CNN,…), with capable of handling the complex relationship between target markers, and CRC were utilized to select an approapriate with higher efficancy agent and then investigated the therapeutic impact of PGP in CRC. RNAseq and the integrative systems biology technique followed by MI algoritisms were applied to identify differentially expressed genes (DEGs) followed by validation in a large cohort of patients and then PGP was selected for in vitro and in vivo studies. Antiproliferative-activity of (Punica granatum var. pleniflora (PGP)) was tested in 2 and 3D cell-culture models. The effect of PGP on migratory-behaviors and apoptosis was determined using a wound-healing-assay and AnnexinV/PI staining, respectively. The expression were assessed using q-RT-PCR. Molecular-Pathology and histopathological-assessment was used followed by evaluation of oxidative-stress-markers. Metabolomics for assessment of chemical and active components of the PGP extract were determined by LC-MS/MS. The result illustrated a total of 856 upregulated/downregulated-genes in patients. Among the high top-score genes, fibrotic/inflammatory pathways were detected and further validated in 65 patients. PGP inhibited cell-growth and migration in cells by modulating CyclinD1, Survivin, and E-cadherin. Furthermore, PGP increased apoptosis. Moreover, PGP significantly decreased tumor-size in an animal-xenograft CRC via perturbation of fibrosis-markers, Col1A/ACTA2. PGP reduced inflammation, and oxidative-stress via modulation of SOD/Cat/total thiol. Phytochemical profiling showed a total of 28 and 43 compounds including Corilagin, Ellagic acid, Gallic Acid and Quercetin-hexoside, which have anti-cancer properties. The results demonstrated the therapeutic potential of PGP in tumor-growth reduction, indicating its potential value as a new approach in the treatment of colorectal-cancer.
Aida Yavari Kondori, Mehrdad Moetamani Ahmadi, Seyede Elnaz Nazari, Fereshteh Asgharzadeh, Elisa Giovannetti, Majid Khazaei, Amir Avan. Drug Development Using Machine Learning Approches: The Therapeutic Impact of Golnar on Inhibition of Tumor Growth in Colorectal Cancer [abstract]. In: Proceedings of AACR Drug Discovery and Development (AACR D3) Conference; 2026 Jul 21-24; Boston, MA. Philadelphia (PA): AACR; Clin Cancer Res 2026;32(14_Suppl):Abstract nr A037.
Aida Yavari Kondori, Mehrdad Ahmadi, S. Nazari et al.· Clinical Cancer Research· 0 citations
Despite recent therapeutic advances, treatment options for advanced, therapy-resistant, and metastatic prostate cancer (PCa) remain limited. Here, we developed and prospectively validated an integrated chemoinformatics and machine learning (ML) workflow with ligand-based similarity filtering. Validation on independent external data sets showed that this applicability-domain-guided integration strategy can reduce false positives and improve virtual screening performance. Screening of DrugBank identified five repurposing candidates with confirmed antiproliferative activity in both 2D and 3D PCa models. Among them, the antifungal agent chlormidazole emerged as the most promising candidate, displaying tumor-selective and predominantly cytostatic activity associated with p57 upregulation, reduced Rb phosphorylation, and G1 arrest. Chlormidazole also enhanced the antiproliferative activity of docetaxel in both models, achieving comparable efficacy at substantially lower docetaxel concentrations. These findings identify chlormidazole as a promising repurposing candidate for PCa and demonstrate the value of integrating chemoinformatics with ML for drug repurposing and virtual screening.
Leonardo Bernal, Luca Pinzi, T. Martinelli et al.· Journal of Medicinal Chemist...· 0 citations
This study developed a comprehensive strategy integrating computational predictions with in vitro experimental validations that reverses CDDP resistance in PROC by downregulating JUN, which dismantles intracellular pro-survival networks.
Chen Wang, Junfeng Guo, Taiyang Ye et al.· Frontiers in Pharmacology· 0 citations
This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.
K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al.· Future Medicinal Chemistry· 0 citations
Natural regulatory T cells (nTregs) are a key immunosuppressive component of the tumor microenvironment (TME), but their assessment typically requires specialized molecular assays. This study aimed to develop and validate a deep learning-based pathomics model to predict nTregs infiltration directly from routine hematoxylin and eosin (H&E)-stained whole slide images (WSI) of breast cancer and evaluate its prognostic significance.
Data from 1097 breast cancer patients in The Cancer Genome Atlas (TCGA) were analyzed. A cohort of 928 patients with complete RNA-seq data was used to establish the prognostic value of nTregs (estimated by ImmuneCellAI). A subset of 791 patients with matched high-quality H&E-stained WSI was randomly split into training (n = 633) and validation (n = 158) cohorts. A total of 1488 quantitative pathomic features were extracted. After feature selection via minimum Redundancy Maximum Relevance (mRMR) and Recursive Feature Elimination (RFE), a Gradient Boosting Machine (GBM) classifier was trained to predict high versus low nTregs status, generating a continuous Pathomics Score (PS). The PS was validated against FOXP3 immunohistochemistry (IHC) and evaluated for its association with overall survival (OS). Multi-omics analyses explored the underlying biology of PS-defined groups.
High nTregs infiltration was an independent predictor of poor OS (HR = 1.58, 95% CI 1.10–2.28,
p
= 0.013). The GBM model achieved an area under the receiver operating characteristic curve (AUC-ROC) of 0.81 (95% CI 0.78–0.85) in the training cohort and 0.72 (95% CI 0.63–0.80) in the validation cohort. The PS showed a strong correlation with FOXP3
+
cell density (
p
< 0.001) and was independently associated with worse OS (HR = 1.68, 95% CI 1.15–2.46,
p
= 0.008). Patients with high PS exhibited a distinct transcriptomic signature enriched for immune activation pathways (e.g., estrogen responses) and upregulated immune checkpoint genes (e.g., CD276, TNFSF4, TNFSF9), alongside an immunosuppressive microenvironment characterized by increased nTregs and M2-like macrophage estimates.
We developed and validated a pathomics model that non-invasively predicts nTregs infiltration and patient prognosis from standard H&E images. The PS serves as a novel, accessible digital biomarker that captures the complexity of an inflamed yet immunosuppressive TME and has the potential to augment clinical decision-making, particularly in resource-limited settings.
Yuanbing Xu, Yanming Pan, Manlu Cui et al.· Breast Cancer Research· 0 citations