The development and validation of prognostic and predictive biomarkers in breast cancer is limited by the availability of well-annotated datasets linking tumor molecular features to treatment response and survival outcomes. To address this need, we generated an extensive mouse models dataset comprised of 26 immunocompetent mammary tumor models spanning diverse genetic backgrounds, epithelial-mesenchymal states, the basal-luminal axis, and distinct immune microenvironments. For each model, survival was measured under no treatment, immune checkpoint inhibition (ICI), and carboplatin/paclitaxel chemotherapy, and RNA-sequencing was performed on baseline tumors and on 7-day on-treatment samples for both regimens. Baseline murine tumor gene expression features were used to train a machine learning Elastic Net model that predicted survival outcomes on multiple human breast cancer datasets with performance comparable to that of existing prognostic assays. Models trained for ICI benefit, using either the untreated or 7-day ICI treated samples, predicted ICI benefit on human ICI treated datasets, with the 7-day treated tumor model showing better performance. A predictor of carboplatin/paclitaxel response developed from the murine mammary tumor data performed well in mice but did not generalize to human chemotherapy cohorts. Finally, comparison of multiple computational approaches, including XGBoost, random forests, and support vector regression, showed that all methods successfully predicted survival outcomes, with Elastic Net offering the best performance and interpretability. These results indicate conserved cancer biology between mouse and human tumors for prognosis and ICI response and establish a large preclinical dataset with linked phenotypic and genomic data as a resource for biomarker discovery.
Matthew D. Sutcliffe, Kevin R Mott, Tulay Yilmaz-Swenson et al.· Cancer Research· 0 citations
Background: Population-scale molecular profiling integrated into routine healthcare could accelerate biomarker discovery, validation, and implementation, but the feasibility and sustainability of such an approach have rarely been demonstrated prospectively. The Sweden Cancerome Analysis Network - Breast (SCAN-B) Initiative was established to integrate prospective molecular profiling with population-based breast cancer care and create an infrastructure for translating molecular discoveries into clinical practice (ClinicalTrials.gov identifier NCT02306096). Methods: We evaluated the first 10 full calendar years of SCAN-B, encompassing patients with primary invasive breast cancer enrolled between August 30, 2010 and December 31, 2020. Enrollment and biospecimen collection were compared with all eligible breast cancer diagnoses in participating hospitals to assess population coverage and representativeness. Clinicopathological characteristics, treatments, recurrence-free survival, overall survival, RNA-sequencing-based molecular subtypes and risk-of-recurrence, and somatic mutations were evaluated. We additionally report the translation of SCAN-B molecular profiling from the research setting into routine clinical diagnostics. Results: Among 16,381 estimated eligible breast cancer diagnoses, 13,940 patients (85.1%) were prospectively enrolled across participating Swedish hospitals. Baseline blood samples were obtained from 98.4% of enrolled patients and tumor specimens from 71.1%; 9,323 tumors (94.0% of submitted tumor specimens) underwent RNA-sequencing. The enrolled cohort was broadly representative of the underlying breast cancer population across major clinicopathological characteristics. Integration of longitudinal clinical data with molecular profiling enabled characterization of real-world treatment patterns, long-term outcomes, molecular subtypes, risk-of-recurrence, and the somatic mutational landscape in this population-based cohort. Building on prospective real-time RNA-sequencing and subsequent development and validation of single-sample molecular subtype and risk-of-recurrence predictors, the SCAN-B workflow was transferred into routine clinical molecular diagnostics in Sk[a]ne and Blekinge in 2021. Through January 2026, more than 3,000 patients had received clinical RNA-sequencing-based molecular subtype and risk-of-recurrence reports, while prospective SCAN-B enrollment and transfer of samples and molecular data into the research infrastructure continued. Patient enrollment continues prospectively, with over 23,000 patients accrued as of January 2026. Conclusions: A prospective, population-based molecular profiling program can be integrated into routine breast cancer care at scale while maintaining high population coverage and representativeness. Over more than a decade, SCAN-B progressed from prospective biosampling and molecular profiling through biomarker development and validation to implementation of RNA sequencing-based testing in routine healthcare. This model establishes a continuous framework linking population-based molecular research, biomarker discovery and validation, and clinical implementation, and provides a strategy for integrating precision oncology research with routine cancer care.
L. Saal, H. Dalal, P. Meng et al.· medRxiv· 0 citations