Skip to content
Review Open access

Comparative Performance Analysis of Mainstream Deep Learning and Vision Foundation Models for Small-Sample Vegetation Segmentation in High-Resolution Remote Sensing Imagery

Jul 2026 · Remote Sensing · 0 citations · 21 references

Abstract

Accurate extraction of vegetation information from high-resolution remote sensing (RS) imagery is crucial for efficient urban ecological environment monitoring and land use management. However, due to the high cost of manual annotation in remote sensing imagery and the complex textural variations and spectral confusion exhibited by vegetation under different terrains and lighting conditions, precise vegetation segmentation under small-sample conditions remains a significant challenge. Using the Nanjing Zijinshan region as a case study, this research conducts a systematic comparison of eight representative models within a unified high-resolution remote sensing small-sample experimental framework to address these complexity challenges. We fine-tuned and systematically compared the recently prominent “Segment Anything Model” (SAM) series (including SAM2-Tiny, SAM2.1-Tiny, MobileSAM, and MobileSAMV2), along with classic fully supervised models (U-Net, DeepLabV3+), open-vocabulary segmentation models (SegEarth-OV), and instance segmentation models (YOLO11s-seg), helping clarify the performance boundaries and applicable conditions of different technical paradigms in vegetation segmentation. Experimental results highlight the distinctive performance characteristics of these models. Notably, fine-tuned vision foundation models (such as SAM2.1-Tiny and SAM2-Tiny) demonstrated superior segmentation performance and cross-dataset generalization capabilities, with SAM2.1-Tiny achieving the highest mean Intersection over Union (mIoU; 0.7821) on the Zijinshan dataset, a 5.5% improvement over the classic U-Net model; SAM2-Tiny also maintained the most stable generalization performance in cross-dataset testing on LoveDA, Potsdam, and Vaihingen. In contrast, zero-shot SegEarth-OV and instance segmentation model YOLO11s-seg showed relatively lower performance in current semantic segmentation tasks, revealing the application boundaries of different paradigms. Beyond these findings, to further leverage unlabeled temporal imagery and break through small-sample constraints, we propose an innovative Cross-Temporal Pseudo-Label Self-Training (CT-PLST) strategy, which successfully improved SAM2-Tiny’s mIoU from 0.7776 to 0.7888 (+1.44%), providing a low-cost efficiency enhancement solution for remote sensing segmentation under scarce annotation conditions. To promote reproducible research in remote sensing and computer vision, we publicly release the fine-tuned models, related comparative experiment code, and a high-resolution remote sensing vegetation dataset covering multi-temporal scenarios; access details are provided in the Data Availability Statement. The findings of this study, combined with the proposed CT-PLST strategy and the high-precision segmentation results achieved by vision foundation models, can strongly support tracking analysis of vegetation cover changes, urban heat island effect assessment, and exploration of ecosystem dynamic evolution. Meanwhile, these achievements also provide valuable theoretical guidance and engineering references for practitioners and researchers in finding lightweight segmentation models suitable for specific image characteristics and computational cost constraints in practical applications such as rapid disaster risk assessment or forestry resource surveys.

Read PDF