Skip to content
Conference Open access

F³S: feature fused few-shot segmentation with CLIP guided semantic priors and Sinkhorn attention refinement

Sep 2026 · International Conference on Image Processing and Pattern Recognition (IC-IPPR 2026) · 0 citations

Abstract

Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignment. To address these issues, we propose the Feature-Fused Few-shot Segmentation (F³S) framework, which integrates multi-scale feature fusion with CLIP-driven semantic information. By fusing complementary features from PVTv2, RegNetZ, and MobileViTv2 backbones, F³S captures rich spatial and contextual representations. It further employs dual CLIP priors, text-guided Grad-CAM maps, and visual similarity maps refined with attention mechanisms and sinkhorn normalization to enhance semantic alignment and spatial coverage. A lightweight, transformer-based meta-decoder efficiently synthesizes these multi-modal features for precise mask prediction. We have validated the effectiveness of F³S on PASCAL-5i and COCO-20i, where our method achieved superior mIoU over CNN and transformer-based approaches. The source code is available at https://github.com/shahrozajmal95/F3S.git.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.