Skip to content
Open access

PredictRx: AI based decision support tool for molecular screening for breast cancer drug recommendation

Jul 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 54 references
Medicine

TL;DR

PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication discovery shows how AI-driven predictive modeling can speed up molecular screening and early-stage breast cancer medication discovery.

Abstract

Introduction Breast cancer remains one of the leading causes of cancer-related mortality rate worldwide, and the identification of effective drug combinations is an essential requirement in pharmaceutical research. The integration of Artificial Intelligence (AI) in processing large volumes of chemical and biological data combines molecular representation, predictive modeling and structured support within a single accessible tool, which accelerates early-stage candidate identification for breast cancer research while promoting reproducibility, transparency and user centered design. Aim The current research focuses on developing and designing “PredictRx” which is an artificial intelligence based driven decision support tool which tends to benefit healthcare practioners to analyze the combination of drug which can be utilized for breast cancer patients. Methodology PredictRx was developed using molecular descriptors, physicochemical properties, and drug interaction datasets collected from publicly available biomedical databases. The tool integrates in total six supervised and unsupervised learning techniques to examine the structural similarities between compounds and predict the potential drug interactions for breast cancer. Various machine learning techniques, including Random Forest, Support Vector Machine, Logistic Regression, K-Means Clustering, DBSCAN, and Agglomerative Clustering, to analyse structural similarities and predict potential drug interactions and synergy patterns. Model performance was evaluated using Classification matrix, Silhouette Score, Calinski-Harabasz Index, and Davies-Bouldin Index. The tool was deployed as a browser-accessible web application for real-time interaction and visualization. Result The results suggests that Random Forest has the highest predictive performance accuracy of 1, and Agglomerative clustering delivered strongest scores (Silhouette Score: 0.6946; Davies-Bouldin Index: 0.2457). The current tool was deployed as a browser accessible web tool with possibility of real time interaction and result visualization. PredictRx is a distinctive easy to use, and interpretable screening tool focused on drug compatibility and synergy analysis. EDA further identified molecular weight, lipophilicity, and structural similarity as important contributors to drug compatibility prediction. Conclusion PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication discovery. The technology facilitates the effective identification of appropriate drug combinations and offers a scalable foundation for upcoming AI-assisted pharmaceutical research by combining clustering, classification, molecular representation, and visualization into a single interpretable platform.

Read PDF

Similar papers

Conference Jul 2026

Explainable Multi-Omic Machine Learning Framework for Predicting Drug Response in Breast Cancer

Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.

D. Kumari, Aiman, Sakshi Singh et al. · 0 citations
Review Open access Jul 2026

Advanced Artificial Intelligence and data science in bioinformatics-driven drug discovery for cancer: Pathways toward shorter and less toxic treatment

Recent literature on the application of artificial intelligence (AI) and data science within bioinformatics-driven cancer drug discovery is synthesized, examining how these tools are reshaping target identification, molecular design, biomarker discovery, and treatment personalization.

Yejide Eniola Dabiri · 0 citations
Open access Jul 2026

Machine Learning-Based Multi-Cancer Diagnostic System for Early Detection and Accurate Classification Across Diverse Cancer Types

Cancer is one of the major causes of death worldwide, mostly owing to late discovery and hence restricted treatment choices. Existing screening approaches are primarily invasive and often associated with complicated, long and expensive procedures. In biomedicine and bioinformatics, several research groups have examined the use of machine learning methods to solve the important challenge of categorizing cancer patients into high- and low-risk categories. These methodologies have thus been used to mimic the onset and treatment of cancer. The ability of ML algorithms to detect important characteristics in complex datasets further highlights their importance. Many of these approaches like as Decision Trees, Logistic Regression (LR), Support Vector Machines and K-Nearest Neighbours have been widely employed in cancer research to generate prediction models that aid decision makers to make better and more trustworthy decisions. ML methods are indeed able to improve our understanding of cancer formation, but need adequate validation to be regarded for application in ordinary clinical practice. Hence, an ML approach was utilized to simulate the progression of cancer. The prediction models shown here are based on several ML approaches and a broad variety of input features and Data Samples. The proposed framework incorporates data preprocessing, feature selection, and advanced classification algorithms to enhance diagnostic accuracy and facilitate timely clinical decision-making. The study emphasizes the potential of artificial intelligence in advancing precision oncology and improving healthcare outcomes.

Sakshi Singh, Saurav Kumar, Yusuf Perwej et al. · 0 citations
Review Aug 2026

Advancing cancer drug discovery through the integration of machine learning and high-throughput screening.

This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.

K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al. · 0 citations
Conference Aug 2026

Predictive Artificial Intelligence Approaches for Accurate Drug Classification in Pharmaceutical Industry

Drug classification plays a critical role in medicine as it aids in selecting the best medicines for an individual’s needs based on their individual characteristics and history. Computational methods are increasingly used in drug discovery to build structure-activity models for large chemical databases. This study introduces a scalable ML approach for drug classification based on scaffolds using SMILES from the ChEMBL database. The approach uses RDKit for physicochemical descriptor extraction, Bemis-Murcko scaffolds for target construction and employs feature selection, encoding, RobustScaler normalization and SMOTE for balancing classes. Models include Random Forest, XGBoost, and a Stacking Ensemble, with accuracy, precision, recall, F1-score, and ROC-AUC as evaluation metrics. The experimental findings show that the Stacking Ensemble outperforms Random Forest (88.06%) and XGBoost (85.91%), achieving an accuracy of 89.15% and a ROC-AUC of 98.29%, suggesting better generalization. The results demonstrate that ensemble and tree-based learning methods outperform traditional models and LSTM in scaffold classification. The novel approach provides a rapid, scalable, and precise framework that increases the efficiency of virtual screening and offers a reliable approach to AI-based decision-making in drug discovery.

Nazia F Shaik · 0 citations
Open access Jul 2026

A Comparative Analysis of Machine Learning Models for Cancer Types Classification Using RNA-Seq Gene Expression Data

Cancer is a significant global health concern, and scientists must set the right tumor identifier for accurate diagnosis and personalized treatment plans. While RNA-Seq gene expression data provides critical molecular information, it has two major challenges that can affect ML systems and hinder their performance: its large dimensionality and a propensity for class imbalance. This study offers a comprehensive and data-driven comparison of five machine learning classifiers that demonstrate superior performances in multi-class cancer diagnosis using the RNA-Seq-based TCGA dataset with Random Forest, Support Vector Machine (SVM), and Gradient Boosting k-Nearest Neighbors (k-NN) and Multilayer Perceptron (MLP). Mutual information was used to select important features that reduce the dimensions of the data, and then sensitivity analysis showed the successful performance of the method. To overcome these class imbalance problems, the Synthetic Minority Over-sampling Technique (SMOTE) method was used. The model performance was evaluated using a comprehensive testing framework that integrated 5-fold cross-validation with several evaluation metrics such as balanced accuracy, precision, recall, and F1-score, and confusion matrices and ROC curves. The research group performed ablation studies to better understand the individual effects of correct feature selection and SMOTE on their process. These outcomes show that the MLP classifier reached the highest balanced accuracy (0.9917) after resolving methodological avarice. The assessment demonstrates that SMOTE has a significant impact on improving classification results of minority classes due to its positive effects on recall and F1 score metrics. This study provides insights into selected gene features associated with known oncogenic pathways and, therefore, advanced biological interpretations per se of the obtained results. The work develops a framework that is reproducible and allows multi-metric evaluation of ML-based cancer classification, but exposes two key issues leading to data leakage and validation failures. Our research shows how the promise of AI-driven precision medicine can help build approaches to automated and accurate cancer typing.

Peter Makieu, Sahr Foday, Alfred Santigie Turay · 0 citations