Skip to content

Epigenetic profile drives accurate survival prediction in breast cancer via a multi-omics machine learning model

· 0 citations · 49 references

TL;DR

This study demonstrates that multi-omics integration via machine learning enhances survival prediction and reveals actionable biomarkers in breast cancer and outperformed mutation-based models.

View source

Similar papers

Open access Aug 2026

Identifying multi-omics biomarkers for ovarian cancer survival estimation

Ovarian cancer is among the deadliest gynecologic malignancies, and its molecular heterogeneity limits accurate prognostic stratification. Although multi-omics approaches have improved predictive modeling, many prioritize predictive performance over biological interpretability, limiting their clinical translation. We developed an interpretable three-stage machine learning framework integrating mRNA, microRNA, DNA methylation, copy number variation, and protein expression data from The Cancer Genome Atlas. Hierarchical feature selection was combined with a weighted ensemble of ElasticNet, ridge regression, support vector regression, XGBoost, and random forest models to estimate overall survival time in patients with ovarian cancer. Multi-omics integration outperformed every single-modality model, achieving a Pearson correlation of 0.752, a concordance index of 0.779, and a mean absolute error of 8.57 months between estimated and observed survival time, compared with 0.48 for the best single modality. The framework identified a 20-biomarker signature dominated by tumor-associated macrophage and complement genes. In an independent survival analysis, VSIG4 and CD163 remained significant after false discovery rate correction, and the signature raised the concordance index over clinical covariates alone from 0.615 to 0.686Enrichment analysis implicated PI3K-Akt, MAPK, focal adhesion, hypoxia, apoptosis, and p53 signaling pathways. This framework couples improved prognostic estimation with biological interpretability supporting multi-omics biomarker discovery in ovarian cancer.

Kosar Fateh, S. Sathipati · 0 citations
Open access Jul 2026

The single-cell atlas of programmed cell death signature: A machine learning-based prognostic framework in breast cancer.

Breast cancer remains a leading cause of cancer-related mortality in women, and current prognostic models are suboptimal. The transcriptomic role of programmed cell death (PCD) in breast cancer progression is not fully understood. Here, we integrated single-cell RNA sequencing data from breast tumors with nine bulk transcriptomic cohorts to systematically analyze 19 PCD modalities. Using a machine learning framework incorporating 14 algorithms, we constructed a prognostic signature, with a ridge regression-based PCD riskscore showing optimal performance and being further integrated into a clinical nomogram. Functional roles of key genes were validated through in vitro and in vivo experiments. We identified a prognostic signature comprising 26 core PCD genes, which effectively stratified patients into distinct risk groups and robustly predicted overall survival. Single-cell analyses revealed that a high PCD risk core was associated with an immunosuppressive tumor microenvironment and reduced immune checkpoint expression, whereas low-risk patients showed greater sensitivity to targeted therapies. Among the signature genes, PDIA4 was consistently overexpressed in 50 paired breast cancer tissues, and its knockdown markedly inhibited tumor growth and malignant phenotypes. This study establishes a novel PCD-based prognostic signature for breast cancer and identifies PDIA4 as a functionally important oncogene.

Jixuan Jiang, Yiwen Wang, Yuju Huang et al. · 0 citations
Open access Jul 2026

Three ct-miRNA signature predicts disease outcome in early breast cancer women.

BACKGROUND Breast cancer is still a leading cause of tumor mortality in women. Indeed, despite advancements in early diagnosis, molecular profiling and novel therapeutic approaches, the disease outcome is still not always predictable. This evidence underlines the need for validated biomarkers to predict recurrences, to personalize both disease monitoring and tailored therapies. And to this aim, miRNAs have shown promising applications as circulating biomarkers. Starting from plasma samples collected from women with early-stage breast cancer at the time of diagnosis, we explored the expression of ct-miRNAs to define a molecular signature predictive of recurrence. METHODS Two independent cohorts of plasma samples were retrospectively and prospectively collected at Fondazione IRCCS Istituto Nazionale dei Tumori di Milano (INT) for a total of 203 patients. Ct-miRNAs were previously profiled by using the OpenArray Human microRNA panel (OA) (Thermo Fisher Scientific). Relapse-free survival (RFS) was analyzed using Cox regression models adjusted for cohort to assess associations with clinicopathological variables and circulating miRNA levels. Models performance was evaluated using c-statistics (95% Confidence interval) and an internal validation was performed with bootstrap resamples. Clinicopathological variables were added to the signature in multivariate models to evaluate their independent prognostic value. RESULTS We identified a three ct-miRNA (miR-125b, miR-26b-3p and miR-532-5p) signature associated with disease outcome, with a hazard ratio of 2.803 (95% CI, 1.721-4.565). The c-statistic was 0.72 (95%, CI 0.62-0.83) and was confirmed by the resampling procedure, with a c-median bootstrap statistic of 0.73 (IQR, 0.65-0.81). The 3-circulating miRNA signature retained its significant prognostic performance with respect to RFS even after the inclusion of clinicopathological variables in multivariate models. Considering both ct-miRNA signature and Nottingham Prognostic Index (NPI), the c-statistic of the bivariate model was equal to 0.80 (95% CI, 0.70; 0.90). CONCLUSION We identified a three ct-miRNA prognostic signature in early breast cancer women. Considering the accessibility and stability of ct-miRNAs, this signature might improve the recurrence risk prediction identifying who, despite the early diagnosis, might need a more intensive screening or secondary prevention strategies.

G. Cosentino, Mara Lecchi, M. Giussani et al. · 0 citations
Open access Aug 2026

Predicting Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer Using Multi-Omics and Machine Learning

Pathological complete response (pCR) to neoadjuvant chemotherapy (NAC) in breast cancer remains a clinically important endpoint, but accurate prediction before treatment is challenging. We developed an attention-based multi-omics framework that integrates pretreatment genomics, transcriptomics, proteomics, epigenomics, and clinical variables to predict pCR in early-stage breast cancer. The model was trained on the I-SPY2 neoadjuvant cohort and externally evaluated using The Cancer Genome Atlas Breast Cancer and independent NAC datasets. Performance was assessed using discrimination, calibration, and subtype-specific analyses, while explainability was examined using SHAP-based feature importance and pathway enrichment testing. In the I-SPY2 test set, the multi-omics model achieved an area under the receiver operating characteristic curve of 0.81 and outperformed clinical-only and single-omics baselines across subtypes. Improvements were most apparent in triple-negative and HER2-positive disease. The model showed acceptable calibration and maintained performance in external and transfer analyses, in which higher predicted risk scores were associated with poorer recurrence-related outcomes. Explainability analyses identified proliferation, immune activity, and PI3K/AKT signaling as major contributors to prediction. These findings indicate that integrating pretreatment multi-omics data with clinical variables improves prediction of NAC response while producing interpretable outputs. Further prospective validation is required before clinical application.

J. Fakoya, Catherine Falayi, M. Ajinaja · 0 citations
Conference Jul 2026

Explainable Multi-Omic Machine Learning Framework for Predicting Drug Response in Breast Cancer

Accurate prediction of drug sensitivity in cancer cell lines is vital for precision oncology and patient-specific therapies. However, many computational approaches fail to integrate multi-modal biological and chemical features and often struggle with high-dimensional, imbalanced pharmacogenomic data, limiting predictive accuracy and interpretability. To address these challenges, we developed a machine learning framework that integrates pharmacogenomic profiles-including mutation status, copy number alterations, and microsatellite instabil-ity-with molecular fingerprints and descriptors of 85 anticancer drugs, generated using PaDEL from SMILES strings. Data from 40 breast cancer cell lines in the Genomics of Drug Sensitivity in Cancer (GDSC) dataset were employed. A threestage feature selection strategy combining Boruta, mRMR, and XGBoost was applied to reduce drug feature dimensionality while retaining 130 cell line features. Multiple models were trained, and LightGBM, optimized with grid search, class weighting, and 3-fold cross-validation, demonstrated superior performance in handling severe class imbalance (233 sensitive vs. 3167 resistant samples). LightGBM achieved training AUROC $=0.9455$, AUPRC $\boldsymbol{=} \mathbf{0. 5 1 4 8}$, Accuracy $\boldsymbol{=} \mathbf{0. 8 4 1 5}$, F1-score = 0.4481, Recall = 0.9409, and MCC = 0.4732, underscoring its suitability for sparse biomedical datasets. Model interpretation with SHapley Additive exPlanations (SHAP) highlighted BRCA-related features, identifying cnaBRCA25 (not mutated) as a resistance marker and cnaBRCA47 (mutated) as a context-dependent biomarker, consistent with their roles in DNA repair pathways. Overall, this framework demonstrates the value of multi-modal integration and interpretable machine learning in pharmacogenomics. While results are promising, validation on larger and independent cohorts is essential to establish clinical relevance.

D. Kumari, Aiman, Sakshi Singh et al. · 0 citations
Open access Aug 2026

DSAI-11 AN OPTIMIZED MACHINE LEARNING MODEL FOR OVERALL SURVIVAL PREDICTION IN BRAIN METASTASIS PATIENTS USING GENOMIC MUTATION AND COPY NUMBER FEATURES

Abstract Background Cancer progression and patient survival are influenced by both tumor-intrinsic and microenvironmental factors, including the ability of tumor cells to disseminate and colonize distant organs. Organ-specific metastases, particularly brain metastases (BM), exhibit distinct tumor–microenvironment interactions, therapeutic responses, and clinical outcomes. Integrating metastatic genomic alterations into survival modeling is essential for improving prognostic accuracy. Here, we present an optimized machine learning framework leveraging genomic mutations and copy number variations to predict overall survival (OS) in BM patients. Methods We implemented a rigorous machine learning pipeline for survival prediction. The dataset was randomly divided into training (70%) and independent test (30%) cohorts. Feature selection, model training, and hyperparameter optimization were performed exclusively within the training set. Prognostic features were initially identified using univariable Cox regression (p < 0.05) and refined using machine learning–based selection, retaining features consistently selected across multiple models. Hyperparameters were optimized via 3-fold cross-validation. Model performance was evaluated using the concordance index (C-index) and time-dependent AUC, while Kaplan–Meier analysis assessed risk stratification. Results The cohort comprised 381 BM patients, primarily from lung cancer (51.1%), followed by melanoma (15.2%) and breast cancer (7.9%). Key prognostic features included recurrent single-nucleotide variants in genes such as PTPRT, ARID1A, PREX2, and FAT1. Ridge regression demonstrated the best performance, achieving a C-index of 0.70 in the test cohort. Time-dependent analyses showed AUCs of 0.642, 0.711, and 0.729 at 1, 2, and 3 years, respectively. The model achieved significant risk stratification (HR = 3.45, p < 0.001), with clear separation between predicted risk groups. Conclusions We developed a robust machine learning framework integrating genomic mutations and copy number alterations to predict survival in BM patients. The model demonstrated stable performance and effective risk stratification in an independent cohort, supporting its potential clinical utility for prognostic assessment and precision oncology applications.

M. I. Ali, Z. Majeed, Peng Li et al. · 0 citations