PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication discovery shows how AI-driven predictive modeling can speed up molecular screening and early-stage breast cancer medication discovery.
Abstract
Introduction Breast cancer remains one of the leading causes of cancer-related mortality rate worldwide, and the identification of effective drug combinations is an essential requirement in pharmaceutical research. The integration of Artificial Intelligence (AI) in processing large volumes of chemical and biological data combines molecular representation, predictive modeling and structured support within a single accessible tool, which accelerates early-stage candidate identification for breast cancer research while promoting reproducibility, transparency and user centered design. Aim The current research focuses on developing and designing “PredictRx” which is an artificial intelligence based driven decision support tool which tends to benefit healthcare practioners to analyze the combination of drug which can be utilized for breast cancer patients. Methodology PredictRx was developed using molecular descriptors, physicochemical properties, and drug interaction datasets collected from publicly available biomedical databases. The tool integrates in total six supervised and unsupervised learning techniques to examine the structural similarities between compounds and predict the potential drug interactions for breast cancer. Various machine learning techniques, including Random Forest, Support Vector Machine, Logistic Regression, K-Means Clustering, DBSCAN, and Agglomerative Clustering, to analyse structural similarities and predict potential drug interactions and synergy patterns. Model performance was evaluated using Classification matrix, Silhouette Score, Calinski-Harabasz Index, and Davies-Bouldin Index. The tool was deployed as a browser-accessible web application for real-time interaction and visualization. Result The results suggests that Random Forest has the highest predictive performance accuracy of 1, and Agglomerative clustering delivered strongest scores (Silhouette Score: 0.6946; Davies-Bouldin Index: 0.2457). The current tool was deployed as a browser accessible web tool with possibility of real time interaction and result visualization. PredictRx is a distinctive easy to use, and interpretable screening tool focused on drug compatibility and synergy analysis. EDA further identified molecular weight, lipophilicity, and structural similarity as important contributors to drug compatibility prediction. Conclusion PredictRx shows how AI-driven predictive modeling which can speed up molecular screening and early-stage breast cancer medication discovery. The technology facilitates the effective identification of appropriate drug combinations and offers a scalable foundation for upcoming AI-assisted pharmaceutical research by combining clustering, classification, molecular representation, and visualization into a single interpretable platform.
Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.
D. Kumari, Aiman, Sakshi Singh et al.· Annual International Compute...· 0 citations
Recent literature on the application of artificial intelligence (AI) and data science within bioinformatics-driven cancer drug discovery is synthesized, examining how these tools are reshaping target identification, molecular design, biomarker discovery, and treatment personalization.
Yejide Eniola Dabiri· Magna Scientia Advanced Rese...· 0 citations
Cancer is one of the major causes of death worldwide, mostly owing to late discovery and hence restricted treatment choices. Existing screening approaches are primarily invasive and often associated with complicated, long and expensive procedures. In biomedicine and bioinformatics, several research groups have examined the use of machine learning methods to solve the important challenge of categorizing cancer patients into high- and low-risk categories. These methodologies have thus been used to mimic the onset and treatment of cancer. The ability of ML algorithms to detect important characteristics in complex datasets further highlights their importance. Many of these approaches like as Decision Trees, Logistic Regression (LR), Support Vector Machines and K-Nearest Neighbours have been widely employed in cancer research to generate prediction models that aid decision makers to make better and more trustworthy decisions. ML methods are indeed able to improve our understanding of cancer formation, but need adequate validation to be regarded for application in ordinary clinical practice. Hence, an ML approach was utilized to simulate the progression of cancer. The prediction models shown here are based on several ML approaches and a broad variety of input features and Data Samples. The proposed framework incorporates data preprocessing, feature selection, and advanced classification algorithms to enhance diagnostic accuracy and facilitate timely clinical decision-making. The study emphasizes the potential of artificial intelligence in advancing precision oncology and improving healthcare outcomes.
Sakshi Singh, Saurav Kumar, Yusuf Perwej et al.· International Journal of Lat...· 0 citations
This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.
K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al.· Future Medicinal Chemistry· 0 citations
Drug classification plays a critical role in medicine as it aids in selecting the best medicines for an individual’s needs based on their individual characteristics and history. Computational methods are increasingly used in drug discovery to build structure-activity models for large chemical databases. This study introduces a scalable ML approach for drug classification based on scaffolds using SMILES from the ChEMBL database. The approach uses RDKit for physicochemical descriptor extraction, Bemis-Murcko scaffolds for target construction and employs feature selection, encoding, RobustScaler normalization and SMOTE for balancing classes. Models include Random Forest, XGBoost, and a Stacking Ensemble, with accuracy, precision, recall, F1-score, and ROC-AUC as evaluation metrics. The experimental findings show that the Stacking Ensemble outperforms Random Forest (88.06%) and XGBoost (85.91%), achieving an accuracy of 89.15% and a ROC-AUC of 98.29%, suggesting better generalization. The results demonstrate that ensemble and tree-based learning methods outperform traditional models and LSTM in scaffold classification. The novel approach provides a rapid, scalable, and precise framework that increases the efficiency of virtual screening and offers a reliable approach to AI-based decision-making in drug discovery.
Nazia F Shaik· International Conference on...· 0 citations
Cancer is a significant global health concern, and scientists must set the right tumor identifier for accurate diagnosis and personalized treatment plans. While RNA-Seq gene expression data provides critical molecular information, it has two major challenges that can affect ML systems and hinder their performance: its large dimensionality and a propensity for class imbalance. This study offers a comprehensive and data-driven comparison of five machine learning classifiers that demonstrate superior performances in multi-class cancer diagnosis using the RNA-Seq-based TCGA dataset with Random Forest, Support Vector Machine (SVM), and Gradient Boosting k-Nearest Neighbors (k-NN) and Multilayer Perceptron (MLP). Mutual information was used to select important features that reduce the dimensions of the data, and then sensitivity analysis showed the successful performance of the method. To overcome these class imbalance problems, the Synthetic Minority Over-sampling Technique (SMOTE) method was used. The model performance was evaluated using a comprehensive testing framework that integrated 5-fold cross-validation with several evaluation metrics such as balanced accuracy, precision, recall, and F1-score, and confusion matrices and ROC curves. The research group performed ablation studies to better understand the individual effects of correct feature selection and SMOTE on their process. These outcomes show that the MLP classifier reached the highest balanced accuracy (0.9917) after resolving methodological avarice. The assessment demonstrates that SMOTE has a significant impact on improving classification results of minority classes due to its positive effects on recall and F1 score metrics. This study provides insights into selected gene features associated with known oncogenic pathways and, therefore, advanced biological interpretations per se of the obtained results. The work develops a framework that is reproducible and allows multi-metric evaluation of ML-based cancer classification, but exposes two key issues leading to data leakage and validation failures. Our research shows how the promise of AI-driven precision medicine can help build approaches to automated and accurate cancer typing.
Peter Makieu, Sahr Foday, Alfred Santigie Turay· Journal of Artificial Intell...· 0 citations