Jul 2026· IEEE Transactions on Medical Imaging· Vol PP, pp. 1-1· 0 citations
Medicine
TL;DR
VSS-SAM++ is introduced, a novel dual-branch architecture that combines SAM's foundational visual priors with Vision Mamba's capacity for modeling long-range spatial dependencies and its robustness to domain shifts and scalability across diverse modalities highlights its potential for clinical deployment.
Abstract
The Segment Anything Model (SAM) has demonstrated groundbreaking performance in natural image segmentation, yet its direct application to medical imaging remains suboptimal due to domain shifts in data distributions and the inherent 3D nature of medical data. Although recent SAM-based methods have employed parameter-efficient transfer learning (PETL) to adapt SAM for medical tasks, they often overlook the critical 3D contextual information essential for accurate volumetric segmentation. To address this limitation, we introduce VSS-SAM++, a novel dual-branch architecture that combines SAM's foundational visual priors with Vision Mamba's capacity for modeling long-range spatial dependencies. In this framework, SAM serves as the primary encoder for high-level feature extraction, while a parallel Mamba branch captures cross-slice dependencies in 3D medical volumes. A gated hybrid attention module then dynamically fuses complementary features from both branches, adaptively weighting multi-view representations to minimize feature ambiguity and enhance segmentation precision. Extensive evaluations across nine public CT and MRI datasets demonstrate that VSS-SAM++ outperforms existing methods by 0.2-11.3% in Dice score on multi-organ and lesion segmentation tasks. The framework's robustness to domain shifts and scalability across diverse modalities highlights its potential for clinical deployment.
Three-dimensional medical image segmentation is critical for clinical diagnosis and treatment planning. Medical images typically exhibit heterogeneous characteristics of low inter-slice resolution and high in-plane resolution, leading 3D CNNs to lose fine edge details via cross-slice feature averaging and traditional 2...
Xuan Chen, Shao-Long Chen· Italian National Conference...· 0 citations
This work proposes an enhanced 3D segmentation framework, UAtten-Unetr, designed to improve segmentation accuracy and robustness in complex medical scenarios, and innovatively developed a unified loss function based on bimodal modality-specific Dice constraints and uncertainty regularization, optimized for synchronous...
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anato...
Donghang Lyu, Zi-Chen Zhang, O. Dzyubachyk et al.· 0 citations
This work proposes a conditional Flow Matching (CFM) based framework for efficient 3D medical image segmentation that maintains competitive segmentation accuracy while achieving an inference time of only 1.14 seconds per volume, offering a clear efficiency advantage over existing generative segmentation methods.
Y. Yilihamu, Jian Xue, Chuang Jia et al.· International Conference on...· 0 citations
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
B-MIM is introduced, a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction and suggests that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.
S. González, Karen Sanchez, J. M. Saavedra et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.