Skip to content
Open access

Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis

Sep 2026 · Journal of Artificial Intelligence in Bioinformatics · 0 citations · 24 references

Abstract

Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-range spatial dependencies that characterize the heart across the cardiac cycle. In this work, we present a fully supervised framework that couples a Vision Transformer (ViT) encoder with a convolutional decoder to segment the left ventricle (LV), right ventricle (RV), and myocardium (Myo) on the Automated Cardiac Diagnosis Challenge (ACDC) dataset. The transformerbackbone models global context across all image patches via multi-head self-attention, while skip connections preserve the fine spatial detail required for accurate boundary recovery. From the predicted masks, we derive ten interpretable morphological indices and train a supervised ensemble classifier to assign each subject to one of five diagnostic phenotypes. The proposed encoder attains a mean Dice similarity coefficient of 0.918 across the three structures, and the downstream classifier reaches a mean cross-validation accuracy of 93.3%. Feature-importance analysis identifies the RV--LV volume ratio, LV volume, and myocardial-thickness variability as the most discriminative descriptors, consistent with established clinical reasoning. The results indicate that transformer-based encoders provide a competitive and interpretable basis for integrated cardiac segmentation and diagnosis.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.