Skip to content
Open access

MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation

Aug 2026 · IEEE journal of biomedical and health informatics · 0 citations · 43 references
Computer Science

TL;DR

To overcome the architectural mismatch inherent in existing hybrid networks, a novel Scaling Feature Pyramid (SFP) is proposed, which effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring that global contextual information is effectively leveraged.

Abstract

Automatic cardiac image segmentation is pivotal for diagnosing and treating cardiac diseases. In this work, we introduce MCSeg, a volumetric transformer-based network tailored for multi-modal cardiac segmentation. To overcome the architectural mismatch inherent in existing hybrid networks, we propose a novel Scaling Feature Pyramid (SFP). Unlike conventional skip connections, the SFP effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring that global contextual information is effectively leveraged. For the training paradigm, the ViT encoder first undergoes self-supervised pre-training via masked image modeling. Subsequently, the network is fine-tuned on downstream tasks, during which a regional mutual information (RMI) loss is integrated to improve boundary segmentation accuracy. In experiments, MCSeg consistently outperforms eleven SOTA methods on CT dataset ImageCHD, multi-modal dataset MM-WHS, MRI dataset HVSMR-2.0 and MSD Heart, highlighting the effectiveness of our MCSeg for multi-modal cardiac segmentation tasks. Furthermore, MCSeg's superior performance in few-shot experiment showcases its significant potential in adapting to limited data scenarios. Codes and pre-trained ViT-B weights are open-sourced at https://openi.pcl.ac.cn/OpenMedIA/MCSeg

Read PDF

Similar papers

Open access Aug 2026

TransCat: a hybrid CNN-transformer network with KAN for medical image segmentation

TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.

Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al. · 0 citations
Open access Sep 2026

Global Self-Attention for Cardiac MRI: Unified Segmentation and Interpretable Pathology Diagnosis

Accurate delineation of cardiac structures from cine magnetic resonance imaging (MRI) is essential for quantitative assessment of ventricular function and for diagnosing cardiomyopathies. Convolutional encoders, although highly effective, capture context within a limited receptive field and may underrepresent the long-...

Faizan Ahmad, Muhammad Arif Anwar · 0 citations
Open access 2026

A Hybrid Vision Transformer and U-Net Framework for Automated Cardiac MRI Segmentation and Abnormality Detection

Cardiovascular diseases remain a leading cause of mortality worldwide, and accurate segmentation of cardiac structures from MRI is critical for clinical diagnosis. We propose a hybrid framework that integrates a Vision Transformer encoder with a U-Net decoder for automated cardiac MRI segmentation and abnormality detec...

Prachi Khune · 0 citations
Open access Aug 2026

Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data

The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.

M. Balakrishnan, K. Ananthi, S. R. et al. · 0 citations
Aug 2026

MSCT-Trans: A Multi-scale Convolutional Neural Network Token Transformer for Interpretable Ultrasound Image Classification.

Multi-Scale CNN Token Transformer (MSCT-Trans), a lightweight and interpretable hybrid architecture for general-purpose ultrasound image classification, is proposed, which consistently outperformed CNN and Transformer baselines across accuracy, macro-F1 and area under the receiver operating characteristic curve, partic...

Mohsin Furkh Dar, Sayima Mukhtar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.