Skip to content
Review Open access

WetVeg-2mm: An Ultra-High-Resolution UAV Dataset for Riparian Vegetation Semantic Segmentation

Aug 2026 · Remote Sensing · Vol 18, pp. 2508 · 0 citations · 40 references

Abstract

Fine-grained mapping of riparian vegetation is important for ecological monitoring, invasive species control, and ecosystem restoration. However, riparian plant communities often exhibit fragmented patches, broad transition zones and high visual similarity among classes, making stable species-level segmentation difficult from conventional satellite imagery or lower-resolution UAV imagery. To address this gap, we present WetVeg-2mm, an ultra-high-resolution UAV dataset for fine-grained riparian vegetation semantic segmentation. Built from UAV surveys over a representative riparian section of the Jiuzhou River in Guangxi, China, the dataset provides 2054 image chips (1024 × 1024) with pixel-level annotations at 2 mm ground sampling distance. It contains 17 semantic classes in total, including 14 representative wetland plant classes, such as Colocasia, Eichhornia and Phragmites, together with water, bareland and background. Five baseline models, namely U-Net, Attention U-Net, DeepLabV3+, PSPNet and SegFormer, were evaluated using per-class IoU, mIoU, mDice, PA, Precision and Recall. Across all evaluated baseline settings, SegFormer with ImageNet pretraining achieved the best overall performance, with 76.23% mIoU, 86.01% mDice, 85.17% PA, 87.30% Recall and 85.42% Precision on the test set. Overall, WetVeg-2mm provides a reproducible and challenging benchmark for fine-grained riparian vegetation semantic segmentation.

Read PDF

Similar papers

Review Jul 2026

EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping

The developed system, named EcoVision, establishes a practical foundation for scalable, high-resolution salt marsh monitoring, demonstrating how AI-driven workflows can translate pixel-level predictions into ecologically interpretable metrics.

I. Onyenonachi, Peter J. Lawerance, Nadia Kanwal · 0 citations
Open access Aug 2026

High-Resolution Mapping of Farmland Shelterbelts in an Oasis Agricultural Region Using GF-2 Imagery and Semantic Segmentation

Farmland shelterbelts are important linear vegetation infrastructures in oasis agricultural landscapes. Their accurate extraction is essential for shelterbelt inventory and farmland management, but remains challenging because shelterbelts are narrow, elongated, locally discontinuous, and spectrally similar to croplands, orchards, roadside vegetation, bare soil, and irrigation-related features. This study developed a GF-2-based deep learning workflow for farmland shelterbelt extraction in the 11th Regiment of Alar City, Xinjiang, China. Four representative semantic segmentation models, namely U-Net, U-Net with scSE attention, U-Net++, and DeepLabV3+, were trained and evaluated using four-band GF-2 optical imagery under a unified experimental setting. Model performance was assessed using Precision, Recall, F1-score, Intersection over Union (IoU), overall accuracy, and Kappa coefficient. Patch-level statistical comparison and visual interpretation were further conducted to examine performance differences, shelterbelt continuity, boundary integrity, omission errors, and background confusion. The results showed that U-Net achieved the best overall performance, with a Precision of 94.58%, Recall of 94.77%, F1-score of 94.67%, IoU of 89.88%, overall accuracy of 99.71%, and Kappa coefficient of 0.9452. Compared with U-Net with scSE attention, U-Net++, and DeepLabV3+, U-Net better preserved the continuity and boundary integrity of narrow shelterbelts in regular field-boundary networks. The other models showed varying degrees of omission, boundary fragmentation, or confusion with spectrally similar agricultural objects. The best-performing U-Net model was then applied to the complete study area, and the extracted shelterbelt area was approximately 6.5 km2, accounting for about 4.37% of the cultivated land area. These results indicate that GF-2 optical imagery combined with semantic segmentation can support fine-scale farmland shelterbelt mapping in oasis agricultural landscapes. They also show that model evaluation for narrow linear vegetation features should consider not only pixel-level accuracy but also spatial continuity, boundary integrity, and typical error patterns. The proposed workflow provides a practical reference for GF-2-based farmland shelterbelt inventory, high-resolution linear vegetation mapping, and shelterbelt monitoring in arid oasis agricultural landscapes.

Ying Xu, Ping Lv, Zhuo Zhang et al. · 0 citations
Open access Aug 2026

Disentangling Spectrally Similar Urban Vegetation via Semantic Segmentation-Guided Object Analysis and Multi-Periodic Phenological Features

Fine-grained classification of urban green spaces (UGSs) is important for urban ecological assessment and management but remains challenging because of spectral similarity among vegetation types and inaccurate object delineation in complex urban environments. This study proposes a pixel-to-object framework that combines semantic segmentation-guided object construction with multi-periodic phenological modeling. A semantic green-space mask derived from 0.27 m very-high-resolution imagery constrains superpixel segmentation to generate spatially coherent, boundary-aware green space object-level patches (GSOPs). Pixel-level temporal representations are then derived from Sentinel-2 normalized difference vegetation index (NDVI) time series using TimesNet, aggregated into GSOP-level phenological features, and combined with spatial attributes to classify urban trees, grasslands, and farmlands. Applied to the built-up area of Chengdu, China, the framework achieved an overall accuracy of 91.6%, with F1-scores of 92.5%, 91.9%, and 87.6% for urban trees, grasslands, and farmlands, respectively. Ablation experiments showed that removing phenological features reduced overall accuracy by 13.1 percentage points and decreased the F1-scores of grasslands and farmlands by 16.0 and 23.0 percentage points, respectively. These results demonstrate that semantically constrained object delineation and phenological information jointly reduce boundary fragmentation and improve the discrimination of spectrally similar urban vegetation types.

Chenglong Zhu, Xi Cheng, Tao Liu et al. · 0 citations
Open access Jul 2026

Pixel-based vegetation mapping at class-level from UAV multispectral imagery: application in an alpine lake ecosystem

Abstract. Vegetation mapping in alpine environments is essential for monitoring ecosystem dynamics and climate change impacts, yet remains challenging when using very high-resolution UAV imagery under limited labeled data. This study proposes a data-centric, pixel- based classification framework for class-level vegetation mapping using multispectral UAV data acquired in an alpine study area. The approach prioritizes improving data representation rather than increasing model complexity. To address label scarcity, a feature-rich dataset was constructed by integrating spectral information, vegetation indices, and lightweight spatial descriptors to enhance class separability. Classification was performed using XGBoost, which is well suited for multispectral tabular data and robust under imbalanced conditions. The results show consistent classification performance across vegetation types and demonstrate the effectiveness of dataset enrichment under limited supervision, highlighting the importance of feature representation in data-scarce scenarios.

M. Elahi, Alessandra Spadaro, F. Matrone et al. · 0 citations
Review Open access Jul 2026

Comparative Performance Analysis of Mainstream Deep Learning and Vision Foundation Models for Small-Sample Vegetation Segmentation in High-Resolution Remote Sensing Imagery

Accurate extraction of vegetation information from high-resolution remote sensing (RS) imagery is crucial for efficient urban ecological environment monitoring and land use management. However, due to the high cost of manual annotation in remote sensing imagery and the complex textural variations and spectral confusion exhibited by vegetation under different terrains and lighting conditions, precise vegetation segmentation under small-sample conditions remains a significant challenge. Using the Nanjing Zijinshan region as a case study, this research conducts a systematic comparison of eight representative models within a unified high-resolution remote sensing small-sample experimental framework to address these complexity challenges. We fine-tuned and systematically compared the recently prominent “Segment Anything Model” (SAM) series (including SAM2-Tiny, SAM2.1-Tiny, MobileSAM, and MobileSAMV2), along with classic fully supervised models (U-Net, DeepLabV3+), open-vocabulary segmentation models (SegEarth-OV), and instance segmentation models (YOLO11s-seg), helping clarify the performance boundaries and applicable conditions of different technical paradigms in vegetation segmentation. Experimental results highlight the distinctive performance characteristics of these models. Notably, fine-tuned vision foundation models (such as SAM2.1-Tiny and SAM2-Tiny) demonstrated superior segmentation performance and cross-dataset generalization capabilities, with SAM2.1-Tiny achieving the highest mean Intersection over Union (mIoU; 0.7821) on the Zijinshan dataset, a 5.5% improvement over the classic U-Net model; SAM2-Tiny also maintained the most stable generalization performance in cross-dataset testing on LoveDA, Potsdam, and Vaihingen. In contrast, zero-shot SegEarth-OV and instance segmentation model YOLO11s-seg showed relatively lower performance in current semantic segmentation tasks, revealing the application boundaries of different paradigms. Beyond these findings, to further leverage unlabeled temporal imagery and break through small-sample constraints, we propose an innovative Cross-Temporal Pseudo-Label Self-Training (CT-PLST) strategy, which successfully improved SAM2-Tiny’s mIoU from 0.7776 to 0.7888 (+1.44%), providing a low-cost efficiency enhancement solution for remote sensing segmentation under scarce annotation conditions. To promote reproducible research in remote sensing and computer vision, we publicly release the fine-tuned models, related comparative experiment code, and a high-resolution remote sensing vegetation dataset covering multi-temporal scenarios; access details are provided in the Data Availability Statement. The findings of this study, combined with the proposed CT-PLST strategy and the high-precision segmentation results achieved by vision foundation models, can strongly support tracking analysis of vegetation cover changes, urban heat island effect assessment, and exploration of ecosystem dynamic evolution. Meanwhile, these achievements also provide valuable theoretical guidance and engineering references for practitioners and researchers in finding lightweight segmentation models suitable for specific image characteristics and computational cost constraints in practical applications such as rapid disaster risk assessment or forestry resource surveys.

Le Hu, Fuquan Zhang · 0 citations