Aug 2026· 1 citation· ⚡ 1 influential· 27 references
Computer Science
TL;DR
This work systematically examines how MedSAM generalizes across diverse medical imaging benchmarks, with six adaptation strategies: full-model and encoder-only LoRA, shallow and deep visual prompt tuning (VPT), and decoder-only and full fine-tuning, and concludes that robust MedSAM adaptation requires the combined consideration of prompt noise exposure, domain shift, and representation preservation.
Abstract
Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically examines how MedSAM generalizes across diverse medical imaging benchmarks, with six adaptation strategies: full-model and encoder-only LoRA, shallow and deep visual prompt tuning (VPT), and decoder-only and full fine-tuning. Models are trained on the International Skin Imaging Collaboration Challenge (ISIC 2018) dataset and evaluated under clean and increasingly noisy prompts on IN and Out-of-Distribution (OOD) datasets: close-OOD PH2 (dermoscopy), far-OOD BUSI (Breast Ultrasound Images Dataset) and CBIS-DDSM (Curated Breast Imaging Subset of the Digital Database for Screening Mammography). We show that adaptation improves performance on IN and close-OOD data but often reduces performance on far-OOD data. Full fine-tuning provides the best tradeoff, while encoder-only LoRA is the strongest parameter-efficient alternative, outperforming standard LoRA and VPT under far-OOD shifts. Using Centered Kernel Alignment (CKA), we show that far-OOD degradation is strongly associated with drift in decoder representations, whereas encoder similarity alone does not explain robustness. This suggests encoder-only LoRA provides stronger robustness than standard LoRA by adapting the encoder to distribution shift in visual features, while preserving the decoder pathway. We further show that random 0-100 pixel jitter on prompts produces more robust and better performing models. We thus conclude that robust MedSAM adaptation requires the combined consideration of prompt noise exposure, domain shift, and representation preservation. We release our code: https://github.com/ImSounic/medsam-vpt
Medical image segmenters often get worse when sites, scanner vendors, or protocols change. Continual test-time adaptation (CTTA) addresses this problem without target labels, but it can be impossible to update a model on a non-stationary stream and can lead to a lot of errors. We examine a more reasonable and meaningfu...
Visual prompt tuning (VPT) efficiently adapts foundation models for medical image classification but remains dependent on large labeled datasets. To overcome this, we explore semi-supervised learning (SSL) for label-efficient VPT. In VPT, while the backbone is well-regularized by large-scale pre-training to exhibit sta...
Jie Xu, Qiushi Yang, Wu-Tong Li et al.· Medical Image Analysis· 0 citations
Results show that forgetting depends not only on how much the model changes, but also on which parts of the model are allowed to change, which means that forgetting still increases as more blocks are trained and remains severe when the full backbone is updated.
Amal Saqib, Tausifa Jan Saleem, N. Saeed et al.· 0 citations
Despite the rapid progress of deep neural networks in visual recognition, their adoption in high-risk medical applications remains limited due to reliability and robustness concerns. Models may exploit spurious correlations, particularly in medical imaging, where devices or treatment artifacts often co-occur with patho...
Shenhav Nadir, M. Levi, Eyal Gofer et al.· 0 citations
Multi-dimensional lightweight modules can reconcile segmentation quality with strict computational budgets when adapting video-centric foundation models to 2D clinical data.
Xue-Jia Yuan, Zong-Jian Yang, Yu Guo et al.· IEEE transactions on bio-med...· 0 citations
Text-guided image editors can generate high-fidelity medical deepfakes, challenging the reliability of clinical imagery. Although reasoning-based detectors perform strongly in distribution, they degrade substantially under deployment shift. MedForge-Reasoner, an 8B vision-language model trained with supervised fine-tun...
Zhi-Hui Chen, Meng-Ling Feng· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.