The results establish zero-shot NAS as a computationally efficient paradigm for large-scale Earth observation segmentation as a training-free strategy for semantic segmentation in Earth observation.
Abstract
Large-scale forest monitoring from Sentinel-2 imagery is constrained by the high computa- tional cost of deep-learning model selection on multi-terabyte datasets. This work evaluates zero-shot neural architecture search as a training-free strategy for semantic segmentation in Earth observation. An ensemble of proxy metrics (SynFlow, Fisher Information, Gradient Norm) is applied to a search space of 8,640 Attention U-Net variants, enabling efficient architectural pruning without full training. Validation on a 4.5 TB national Sentinel-2 L1C dataset (Romania) demonstrates a strong rank correlation ( $\rho =0.80$ ) between zero-shot scores and trained performance for the candidate configurations evaluated during HPO. The selected architecture achieves a pixel-wise F1 score on reconstructed maps of spatially disjoint holdout tiles of 0.87, outperforming U-Net, DeepLabV3Plus, and SegFormer. In addition, a systematic analysis of reconstruction and thresholding shows that adaptive thresholds (Yen, Entropy) improve segmentation consistency over fixed heuristics. Overall, the results establish zero-shot NAS as a computationally efficient paradigm for large-scale Earth observation segmentation.
The first systematic zero-shot evaluation of SAM 2 for aerial building segmentation is presented, establishing SAM 2 as a viable tool for rapid building mapping while highlighting where domain adaptation remains necessary.
Bingning Xiong, Mingyu Ou· Journal of image processing...· 0 citations
This study proposes a zero-shot burned area mapping approach based on the Segment Anything Model (SAM) using Sentinel-2 data and demonstrates that SAM can serve as a powerful, scalable, and low-cost framework for zero-shot environmental monitoring and automatic burned area detection, particularly in data-scarce or time-critical post-fire assessment scenarios.
Sentinel-2 imagery offers open access, global coverage, and frequent revisit times, making it attractive for practical building mapping at scale; however, its native 10m resolution makes building vs non-building classification challenging, particularly for small or sub-pixel buildings, and performance can vary with both seasonality and the heterogeneity of built-up environments. This paper introduces a Sentinel-2 building-detection framework designed to systematically quantify these effects and to support more formalised, practice-oriented model selection. We construct a dedicated multi-temporal Sentinel-2 dataset over the Warsaw region and derive binary ground-truth masks by rasterising official Polish topographic database (BDOT10k) building footprints onto the Sentinel-2 pixel grid. Using two established convolutional segmentation backbones (U-Net and DeepLabV3+), we first perform scene-specific fine-tuning to select a robust architecture and identify the best monthly models for L1C and L2A products separately. We then conduct cross-temporal inference by applying each best monthly model to all scenes, enabling an assessment of (i) which months provide favourable training and inference conditions, (ii) how performance transfers between seasons, (iii) the impact of processing level, and (iv) how these effects differ across built-up typologies. Based on these results, we provide practical guidance for routine Sentinel-2 building classification under varying acquisition periods and settlement characteristics.
M. Romaszewski, Kamil Drejer, Katarzyna Kołodziej et al.· 0 citations
Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imbalance, feature complexity, and atmospheric interference. In this paper, we build on the foundational U-Net image segmentation model to develop a quantum-hybrid solution in hopes of more effectively modeling the high-dimensional spectral feature space of the Sen2Fire dataset. We inject a variational quantum circuit in the bottleneck portion of U-Net, specifically the QuFeX and QB-Net ansatzes. We test a classical Feature Pyramid Network (FPN) for further comparative analysis of the model, and we also explore classical improvements to the U-Net model and its training process, including a compression of parameters, alternative loss functions, and uniform mixing of input data. Our primary finding is that under matched conditions, both QB-Net (with an $F_1$ score of 31.18) and QuFeX ($F_1 = 30.79$) outperformed the classical U-Net baseline results ($F_1 = 28.71$). Additionally, the classical FPN achieved a comparable score of 31.13. A crucial finding was that data mixing removed a significant domain shift between the geographically-separated train and test sets, which boosted the classical FPN $F_1$ score to 39.76. We validate the architecture's robustness and generalizability to the wildfire detection problem via cross-dataset transfer on the California Burned Areas (CaBuAr) dataset. Overall, we find that quantum machine learning has potential to provide an advantage in the problem of wildfire image segmentation, and further experiments will continue to validate and expand upon this finding.
Jaiman Munshi, Tanvi Tewary, Sawyer Bloom et al.· 0 citations
Active wildfire mapping from satellite imagery is challenging due to the sparse and highly imbalanced nature of fire pixels, especially in early-stage or low-density fire observations. This work investigates the use of multispectral Landsat-8 imagery for active-fire segmentation under multi-scale wildfire size conditions. We propose a data-driven protocol to characterize fire-region size distributions through connected-component analysis and an interquartile range criterion, enabling the evaluation of model robustness across different local fire-region densities. Three segmentation architectures, U-Net, DeepLabV3+, and SegFormer, are evaluated under different SWIR-based spectral configurations. Results show that U-Net achieves the strongest robustness across the evaluated conditions, SegFormer provides competitive performance, and DeepLabV3+ tends to produce conservative predictions with reduced recall. Across architectures, SWIR2 consistently achieves the strongest or near-best results, highlighting its importance for active-fire segmentation in Landsat-8 imagery. These findings suggest that both spectral band selection and architectural design are critical for robust satellite-based active wildfire mapping trained on low active fire-pixel density images.
Matheus F. Kovaleski, C. Premebida, J. Paulo· 0 citations
Street View Images (SVI) are high-resolution, geo-referenced panoramas that capture real-world environments. Integration of Artificial Intelligence (AI) with SVI enables automated analysis for a range of urban applications including object detection, semantic segmentation, text recognition, scene understanding, and socioeconomic prediction. More than 25 recent AI based SVI studies covering applications in crime prediction, building attribute classification, sidewalk inventory, land price estimation, and environmental monitoring were screened for the survey. Across domains, deep learning architectures such as ResNet, ConvNeXt, and Vision Transformers consistently outperformed traditional machine learning models, with reported accuracies up to 94% for classification and R2 values between 0.62–0.83 for prediction tasks. SVI augmented with other data sources like satellite imagery and Global Information System (GIS), enhanced the model performance and contextual understanding. The Findings highlight the dominance of convolutional and transformer-based networks, emerging interest in graph neural networks, and the need for generalized models and diverse datasets to advance SVI research.
Ranjani A, J. C, V. V et al.· 2026 7th International Confe...· 0 citations