MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation
To overcome the architectural mismatch inherent in existing hybrid networks, a novel Scaling Feature Pyramid (SFP) is proposed, which effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring th...