F³S: feature fused few-shot segmentation with CLIP guided semantic priors and Sinkhorn attention refinement
Abstract
Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignment. To address these issues, we propose the Feature-Fused Few-shot Segmentation (F³S) framework, which integrates multi-scale feature fusion with CLIP-driven semantic information. By fusing complementary features from PVTv2, RegNetZ, and MobileViTv2 backbones, F³S captures rich spatial and contextual representations. It further employs dual CLIP priors, text-guided Grad-CAM maps, and visual similarity maps refined with attention mechanisms and sinkhorn normalization to enhance semantic alignment and spatial coverage. A lightweight, transformer-based meta-decoder efficiently synthesizes these multi-modal features for precise mask prediction. We have validated the effectiveness of F³S on PASCAL-5i and COCO-20i, where our method achieved superior mIoU over CNN and transformer-based approaches. The source code is available at https://github.com/shahrozajmal95/F3S.git.