The results show that IPS-Seg achieves a favorable trade-off between segmentation accuracy and computational efficiency while benefiting consistently from the proposed pseudo-label generation strategy.
Abstract
Unmanned aerial vehicle (UAV) target segmentation remains challenging due to the small size of objects, appearance variations, cluttered backgrounds, and the scarcity of densely annotated data. These factors hinder the performance and practical deployment of lightweight segmentation models in real-world UAV applications. To address this problem, this paper investigates the use of SAM3 (Segment Anything Model 3) as a pseudo-label generator for training compact segmentation networks. Specifically, two supervision paradigms are explored: (i) direct pseudo-supervision using unaltered SAM3-generated masks, and (ii) a refinement strategy that re-applies SAM3 to localized image patches for improved mask quality. Based on these paradigms, a two-stage SAM3-guided pseudo-label generation framework is proposed. In the first stage, SAM3 generates coarse masks for initial object localization. The localized regions are subsequently cropped into patches and processed by SAM3 again to generate fine masks with accurate object boundaries and discard false positives. The resulting coarse and fine masks are then used as pseudo-labels to optimize a lightweight network, termed IPS-Seg, which consists of three components: an IdentityFormer backbone for feature extraction, an Atrous Spatial Pyramid Pooling module for multi-scale context aggregation, and a PixelShuffle-based decoder for spatial resolution recovery. Extensive experiments under multiple supervision settings demonstrate the effectiveness of the proposed framework. The results show that IPS-Seg achieves a favorable trade-off between segmentation accuracy and computational efficiency while benefiting consistently from the proposed pseudo-label generation strategy. These findings highlight the potential of large-scale foundation models as annotation sources for training compact task-specific segmentation networks in low-label vision domains.
The first systematic zero-shot evaluation of SAM 2 for aerial building segmentation is presented, establishing SAM 2 as a viable tool for rapid building mapping while highlighting where domain adaptation remains necessary.
Bingning Xiong, Mingyu Ou· Journal of image processing...· 0 citations
Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually similar background structures interfere with component localization, causing missed instances and false positives. We propose a model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content. We integrate this module with Grounded-SAM and SAM 3, yielding SD-Grounded-SAM and SD-SAM 3. SD-Grounded-SAM further applies geometric and depth-aware box refinement before mask generation, while SD-SAM 3 relies on SAM 3's internal setup. On TOW-300, a dataset of 340 communication-tower UAV images, our strategy improves both baselines: SD-SAM 3 achieves the strongest instance-segmentation performance, while SD-Grounded-SAM produces fewer false positives. Ablations confirm complementary gains from saliency, depth, and box refinement, improving robustness in cluttered scenes.
Small object detection in UAV remote sensing imagery plays a crucial role in applications such as infrastructure inspection, disaster assessment, and precision agriculture, where targets of interest frequently occupy fewer than 32×32 pixels under large ground sampling distance variation and complex cluttered backgrounds. Existing methods still face three main challenges in UAV small-object detection: fine-grained detail loss caused by repeated downsampling, feature inconsistency during cross-scale fusion, and unstable boundary regression in densely distributed aerial scenes. To address these issues, this paper proposes HSAR-DETR, a detection framework that jointly improves hierarchical feature representation, cross-scale refinement, and geometry-aware localization. Specifically, a Hierarchical Enhancement Network (HENet) is introduced to preserve shallow spatial details while strengthening deep semantic-context representation. A Dual-Stream Feature Refinement module (DSFR) is designed at the P4-to-P3 fusion stage, combining spatial-domain structural modeling with frequency-domain phase refinement to improve cross-scale feature consistency. A Coordinate-Guided Adaptive Convolution module (CGAC) is further deployed before the detection head, converting coordinate-guided offset magnitudes into modulation weights for adaptive feature recalibration and improved localization stability. In addition, a conventional high-resolution P2 detection branch is incorporated to enhance small-object representation. Experimental results on the VisDrone, RSOD, and TinyPerson datasets demonstrate improved detection performance. On the VisDrone validation set, HSAR-DETR achieves 50.8% mAP50 and 31.4% mAP50:95, outperforming the RT-DETR baseline by 4.2 and 3.0 percentage points, respectively.
With the increasing deployment of autonomous inspection platforms [e.g., unmanned aerial vehicles (UAVs) and robots], the segmentation performance of learning-based road inspection methods is facing severe challenges due to motion blur. In response, an innovative image segmentation framework is proposed in this study that treats motion deblurring as a pretext task for self-supervised learning. The proposed framework uses a shared encoder to process both motion deblurring and segmentation models. Structural features from the low-level visual task are effectively transferred to the high-level semantic task without introducing additional data. Consequently, the efficiency, performance, and robustness of the segmentation model on motion-blurred images are significantly enhanced. Crucially, this single-model approach eliminates the need for a separate, computationally expensive restoration step, making it particularly efficient for resource-constrained edge devices compared to traditional multistage pipelines. This framework proves adaptable to datasets of varying complexities and network architectures. On a simple motion-blurred crack segmentation dataset captured from UAVs, the CNN-based DeblurGAN enhanced by this framework improved the crack Intersection over Union (IoU) from 49.39% to 70.51%, closely matching the performance achieved using sharp training data (71.98%). Furthermore, on a complex dataset captured from terrestrial vehicles, the proposed framework successfully prevented the “model collapse” observed in the Transformer-based Restormer. The augmented model achieved a mean F1 and mean IoU of 88.19% and 80.65%, respectively, surpassing the next best-performing ConvNext (87.34% and 79.53%). Future studies will explore refining the network architecture to further integrate deblurring and segmentation characteristics and extending this self-supervised paradigm to unified perception frameworks robust against a wider spectrum of real-world visual degradations.
Ye Liu, Jinyan Feng, Bowen Du et al.· Journal of computing in civi...· 0 citations
Experimental results demonstrate that the adapted SAM2 model achieves stable segmentation under moderate environmental variability, while degrading under severe visibility loss, consistent across model scales and input resolutions.
Bindusara Nagathihalli Lokesh, Laura Camila Duran Vergara, Hans-Gerd Maas et al.· The International Archives o...· 1 citation
Domain-specific fine-tuning, coupled with the proposed PPA and NPC frameworks, successfully mitigates the limitations of SAM 2 in agricultural remote sensing and provides a robust methodology for automated, high-precision land segmentation.
Yayang Setia Budi, Fardan Al Jihad, Nurjannah Syakrani et al.· Journal of Information Syste...· 0 citations