Multi-dimensional lightweight modules can reconcile segmentation quality with strict computational budgets when adapting video-centric foundation models to 2D clinical data.
Abstract
Objective
Foundation models such as MedSAM2 achieve strong zero-shot segmentation, but standard fine-tuning updates 11.73 M parameters, limiting deployment in resource-constrained clinical environments. Existing parameter-efficient fine-tuning (PEFT) methods typically focus on reducing storage cost, with less attention to computational complexity and inference latency. We aim to adapt MedSAM2 to 2D clinical modalities with near-transparent overhead across parameters, FLOPs, and latency.
Methods
We propose PE-MedSAM2, a lightweight adaptation framework organizing four complementary modules along two orthogonal axes (channel vs. spatial; feature enhancement vs. computation reduction): a Low-Rank Adapter (LRA) for domain transfer with 0.002 M parameters; a Parameter-Free Feature Enhancement (PFFE) module that uses fixed multi-scale gradient operators to extract high-frequency spatial priors without learnable parameters; an Ultra-Lightweight Adapter (ULA) that decouples spatial and channel transformations via depthwise separable convolutions; and a Dynamic Sparse Attention (DSA) module that concentrates attention on gradient-guided salient tokens.
Results
Across five RGB- like 2D benchmarks spanning polyps, skin lesions, and cell nuclei, PE-MedSAM2 attains the highest Dice similarity coefficient on four and the best average surface distance on four of five; on a sixth benchmark, chest radiography, it again improves over its MedSAM2 baseline. Relative to MedSAM2, the framework adds only 0.294 M trainable parameters, 0.89 G MACs (+0.7%), and roughly 1.2 ms latency.
Conclusion
Multi-dimensional lightweight modules can reconcile segmentation quality with strict computational budgets when adapting video-centric foundation models to 2D clinical data.
Significance
PE-MedSAM2 enables deployment of foundation-model-based segmentation in resource-constrained clinical settings while maintaining competitive contour fidelity. Code is available at https://github.com/Yexika/PE-MedSAM2.
Medical image segmentation foundation models (MedFMs) perform strongly across diverse imaging modalities, but their large size and computational demands hinder deployment in resource-limited clinical settings. Lightweight fine-tuning is impractical given the high training cost of MedFMs, and the efficiency benefits of...
Peng Huang, Ao-Zhong Zhang, Peng-Hang Yin et al.· Medical Image Analysis· 0 citations
It is suggested that ten annotated cases are sufficient for clinically useful segmentation, effectively reducing bottlenecks for both image annotation and training time.
S. Nagaraju, B. Abrahamsen, Ashkan Moradi et al.· 0 citations
SRWKV is proposed, a shape-guided RWKV (SGR) model for parameter-efficient MedISeg that introduces an SGR block that uses a shape prior predicted from the deepest encoder feature to guide token traversal during decoding, reducing foreground-background interleaving and improving structural coherence during sequence form...
Chun-Li Yu, Yin-Hao Li, Zheng Zhao et al.· IEEE Transactions on Neural...· 0 citations
TvaraNet is pro-posed, an extremely lightweight segmentation network designed to preserve boundary fidelity under strict efficiency constraints and achieves competitive or superior boundary-aware performance compared to heavier architectures.
Sridhatta Jayaram Aithal, Vandana Bharti· Proceedings of the Thirty-Fi...· 0 citations
GeoSFLoRA consistently improves Dice and HD95 on BraTS20, MSD-Prostate, and MSD-Lung, approaching full fine-tuning performance and demonstrating an effective paradigm for 2D-to-3D medical image segmentation.
Qin Hao, Bo-Nian Chen, Shengwei Tian et al.· Proceedings of the Thirty-Fi...· 0 citations
Adapting large pretrained Vision Transformers (ViTs) to medical imaging is challenging because labeled clinical data are limited and full fine-tuning is computationally expensive. This paper presents a multi-seed, multi-modality study of parameter-efficient fine-tuning (PEFT) for low-data binary medical image classific...
Alif Akbar Hafiz, Collyne Collyne, Dzaky Rizha Anargya et al.· International Conferences on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.