Skip to content
Conference

Transferring SAM-pretrained 2D ViTs for semi-supervised 3D medical image segmentation

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 1435119 - 1435119-8 · 0 citations · 22 references
Engineering

TL;DR

Experimental results show that the approach outperforms the state-of-the-art on the Pancreas-CT dataset by a large margin and enables rapid transfer learning from 2D-pretrained models to 3D medical tasks with few labeled data, making it especially valuable for rare disease diagnosis.

Abstract

The scarcity of labeled data has made semi-supervised learning essential for medical image segmentation. Recently, Vision Transformers (ViT), pre-trained on large-scale 2D natural images have shown remarkable performance in 2D image segmentation tasks. It is natural to transfer the learned knowledge in ViT to the data-limited semi-supervised medical image segmentation task. However, directly applying ViT faces several challenges, including adapting 2D-pretrained models to 3D medical data and addressing performance degradation in ViT when trained on small datasets. To tackle these challenges, this paper proposes a method for medical image segmentation. The method leverages the strengths of both ViT and Convolutional Neural Networks (CNN) via a hybrid architecture. The CNN-based encoder and decoder play a projector and tokenizer role for the ViT, while the architecture of ViT is fully retained to preserve the knowledge in the pre-trained model derived from SAM (Segment Anything Model) as much as possible. In addition, pseudo-labeling serves as the core guidance for our method to learn from unlabeled data. Experimental results show that our approach outperforms the state-of-the-art on the Pancreas-CT dataset by a large margin and enables rapid transfer learning from 2D-pretrained models to 3D medical tasks with few labeled data, making it especially valuable for rare disease diagnosis.

View source

Similar papers

Conference Aug 2026

A Unified Self-Supervised CNN–Transformer Architecture for Joint Image Enhancement and Segmentation Under Label Scarcity

Visual analysis is essential in practice, and image enhancement and segmentation are fundamental elements of it, but most deep learning models address these problems separately and use large annotated datasets. The paper introduces a single self-supervised hybrid CNN Transformer model that performs joint image enhancem...

E. Punitha, R. Geetha · 0 citations
Open access Aug 2026

MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation

To overcome the architectural mismatch inherent in existing hybrid networks, a novel Scaling Feature Pyramid (SFP) is proposed, which effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring th...

Zhi-Yu Ye, Hai-Rong Zheng, Tong Zhang · 0 citations
Sep 2026

Dynamic adaptation and boundary-aware learning for semi-supervised medical image segmentation.

DABAL is proposed, a semi-supervised framework designed to improve both supervision reliability and contour localization and introduces a Dynamic-static Domain Adaptive Adapter (DDAA) into the Segment Anything Model (SAM) encoder to preserve stable structural priors while providing input-dependent feature compensation...

Wei-Yan Zeng, Zhi-Ming Cheng, Bin Lin et al. · 0 citations
Sep 2026

LMDAU-Net: An Effective Lightweight Multi-scale Deformation Aggregation U-Net for Skin Lesion Segmentation.

Automatic skin lesion segmentation is a pivotal problem in the medical domain and an indispensable component in the computer-aided diagnosis program. Most convolutional neural network-based segmentation algorithms have demonstrated promising performance due to their ability to encode detail and semantic features effici...

Jun-Han Hu, Ming Liu, Jing Yang et al. · 0 citations
Preprint Sep 2026

BruNet: A Cross-Domain Transfer Framework for Bruise Segmentation

Segmenting bruises is a challenging task in medical imaging due to limited data and annotations, diffuse boundaries, and highly variable appearance. In this work, we propose BruNet, a segmentation framework that combines a ViT-based visual encoder (a self-supervised DINOv3 or a pretrained LingBot-Vision backbone) with...

Qi-Ming Wang, Richard J. Motley, E. E. Obi et al. · 0 citations
Sep 2026

EPPNet: Edge Prototype Purification with Auxiliary Supervision for Few-Shot Medical Image Segmentation.

Medical image segmentation plays a pivotal role in computer-aided diagnosis. However, the scarcity of annotated data severely hinders the deployment of deep learning models. Few-shot learning (FSL) is designed to achieve rapid adaptation to unseen classes using limited labeled samples, among which prototype-based metho...

Wen-Jie Meng, Kai Liu, Minghui Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.