Deep learning-based cross-attention fusion of multimodal MRI for survival prediction and risk stratification in IDH-wildtype glioblastoma: a multicenter study
TL;DR
The CAF-derived risk score offers prognostic information complementary to routine clinical variables, representing a promising noninvasive tool for individualized risk stratification when molecular profiling is incomplete or unavailable; these findings warrant prospective external validation before clinical use.
Abstract
Background Glioblastoma (GBM) exhibits profound molecular and spatial heterogeneity, complicating prognostic evaluations. While multiparametric MRI provides crucial multidimensional biological information, conventional end-to-end deep learning integration strategies, such as early or late fusion, often fail to capture complex nonlinear cross-modal interactions. We aimed to systematically evaluate a cross-attention fusion (CAF) architecture for GBM survival prediction and quantify its incremental prognostic value relative to existing clinical tools. Methods In this multicenter retrospective study, 386 adults with IDH-wildtype, WHO grade 4 GBM were assembled from an institutional cohort (n = 226), the Chinese Glioma Genome Atlas (CGGA, n = 62), and The Cancer Genome Atlas (TCGA, n = 98). Using a unified 3D ResNet-18 backbone, we compared single-modality models, early fusion, late fusion, and CAF on preoperative T1-weighted, contrast-enhanced T1-weighted (T1CE), and T2-weighted MRI, and integrated the resulting deep learning risk score with routine clinical variables through multivariable Cox regression. Performance was assessed using Harrell’s C-index, time-dependent AUC, and decision curve analysis. Results CAF showed numerically higher, more consistent C-index trends than early fusion, late fusion, and single-modality models (pooled C-index 0.629, 95% CI 0.594–0.664), although pairwise differences in time-dependent AUC were not statistically significant. Integrating clinical variables raised the pooled C-index to 0.691 (95% CI 0.660–0.721) in the treatment-era model, with comparable performance across the three cohorts (Local 0.688; CGGA 0.716; TCGA 0.689); a pre-treatment configuration excluding adjuvant therapy yielded a pooled C-index of 0.642. Under leave-one-cohort-out external validation, the combined model retained significant risk stratification in all held-out cohorts (C-index 0.63–0.71; all log-rank P < 0.01), albeit with attenuated discrimination. The deep learning risk score remained independent after multivariable adjustment (HR 1.41 per SD, 95% CI 1.26–1.57; P < 0.001). Kaplan–Meier analysis confirmed significant high- versus low-risk separation in all cohorts, and decision curve analysis showed greater net benefit than clinical-only and deep-learning-only models. Conclusion The CAF-derived risk score offers prognostic information complementary to routine clinical variables, representing a promising noninvasive tool for individualized risk stratification when molecular profiling is incomplete or unavailable; these findings warrant prospective external validation before clinical use.