Jul 2026· Journal of Information Systems Engineering and Business Intelligence· 0 citations· 27 references
TL;DR
Domain-specific fine-tuning, coupled with the proposed PPA and NPC frameworks, successfully mitigates the limitations of SAM 2 in agricultural remote sensing and provides a robust methodology for automated, high-precision land segmentation.
Abstract
Background: The Segment Anything Model 2 (SAM 2) represents a state-of-the-art foundation model for object segmentation; however, its application to satellite-based agricultural mapping faces significant challenges. Standard SAM 2 architectures often struggle with the spectral ambiguity of fragmented tropical landscapes and the domain gap inherent in remote sensing imagery. Furthermore, the model’s interactive nature requires precise spatial guidance, making it sensitive to both the location and density of input prompts, which limits its scalability for automated large-scale monitoring.
Objective: This study aims to (1) analyze the impact of domain-specific fine-tuning combined with automated Point Prompt Augmentation (PPA) and Negative Prompt Calibration (NPC) on segmentation accuracy; (2) evaluate the performance of four SAM 2 variants (Tiny, Small, Base+, and Large) to identify the optimal backbone for agricultural tasks; and (3) determine the optimal prompt density for both positive and negative points.
Methods: The SAM 2 variants were fine-tuned using the LoveDA satellite dataset. Evaluation was conducted through an automated pipeline comparing two initialization strategies: Largest Agricultural Area (LAA) Centroid and random placement. The study implemented PPA to strategically increase positive prompt density and NPC to suppress "mask leakage" into irrigation infrastructure. Performance was quantified using mean Intersection over Union (mIoU) and Jaccard & F-measure (J&F) metrics.
Results: The Small variant emerged as the superior backbone, achieving a peak mIoU of 0.7255 and J&F of 0.7734, representing a significant improvement over the pretrained baseline. The results indicate that the LAA Centroid strategy provides a more stable spatial anchor, while the integration of three positive and three negative points optimized the boundary alignment. The Small variant maintained a high computational efficiency with an average inference time of 2.62 minutes.
Conclusion: Domain-specific fine-tuning, coupled with the proposed PPA and NPC frameworks, successfully mitigates the limitations of SAM 2 in agricultural remote sensing. This research provides a robust methodology for automated, high-precision land segmentation, bridging the gap between foundation models and specialized geographic information systems.
Keywords: Agriculture Segmentation, Satellite Imagery, Segment Anything Model 2, Fine-tuning, Point Prompt Augmentation, Negative Prompt Calibration
The first systematic zero-shot evaluation of SAM 2 for aerial building segmentation is presented, establishing SAM 2 as a viable tool for rapid building mapping while highlighting where domain adaptation remains necessary.
Bingning Xiong, Mingyu Ou· Journal of image processing...· 0 citations
The results demonstrate that AB-SAM provides a practical parameter-efficient framework for automated, hint-free landslide segmentation, although further evaluation across additional regions, sensors, and landslide-size distributions remains necessary.
Abstract. Accurate segmentation of individual tree crowns (ITCs) from remote-sensing imagery is essential for forest monitoring and ecological analysis, yet remains challenging due to overlapping canopies and structural variability. The Segment Anything Model (SAM) shows strong generalization capabilities but requires effective prompting and domain adaptation for remote sensing applications. In this study, we investigate a lightweight fine-tuning strategy using Low-Rank Adaptation (LoRA) to adapt SAM for ITC segmentation on the BAMFORESTS dataset. The impact of different prompting strategies is evaluated, including manually annotated point and bounding box prompts, as well as automatically generated bounding boxes derived from a pre-trained tree detector. SAM is fine-tuned with instance-level ITC masks, enabling prompt-aware segmentation of multiple tree crowns per image. Performance is assessed before and after fine-tuning using standard instance segmentation metrics, including IoU and F1-score. Results show that LoRA-based adaptation improves mask delineation and robustness to prompt variability, with bounding box prompts consistently outperforming point-based inputs. Automatically generated prompts enable a fully automated workflow, although their effectiveness depends on detection quality. Evaluation on an independent validation site with manually annotated ITC labels shows that the fine-tuned LoRA-SAM model achieves performance comparable to manual annotations, while significantly reducing annotation effort. These findings highlight the importance of prompt design in adapting foundation models for remote sensing tasks and demonstrate that parameter-efficient fine-tuning provides a practical pathway toward scalable ITC segmentation.
Rewanth Ravindran, Janik Steier, S. Karam et al.· The International Archives o...· 0 citations
Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segmentation of remote sensing (RS) imagery offers a promising solution; yet, residential PV systems remain challenging targets because of their small size and sparse distribution, resulting in severe target-background imbalance. Vision-language foundation models (FMs) provide a data-efficient paradigm through prompt-based semantic and spatial guidance, but the relative contribution of different prompt types remains unclear. We systematically evaluate SAM3 for small-scale PV segmentation in RS imagery by comparing textual, geometric, and hybrid prompting, under varying supervision levels, training strategies, spatial resolutions, and imaging conditions. Multi-temporal aerial imagery from a large off-grid rural region serves as a study site, with findings validated across three additional datasets. Prompting strategy emerged as the dominant factor governing model behavior. Textual prompting consistently produced the lowest performance and showed the greatest sensitivity to supervision and imaging conditions. In contrast, spatial guidance substantially improved both segmentation accuracy and robustness. Hybrid prompting achieved the highest accuracy and stability, indicating that semantic and spatial guidance provide complementary information. Most performance gains were achieved with only a few hundred annotated samples, demonstrating strong data efficiency. Transfer learning had limited overall impact, with only modest improvements observed for textual prompting under limited supervision. Overall, our findings establish prompting strategy as a key determinant of SAM3 adaptation, robustness, and generalization, highlighting the potential of promptable FMs for scalable PV mapping in data-constrained off-grid regions.
Roni Blushtein-Livnon, T. Svoray, Osher Rafaeli et al.· 0 citations
Remote sensing image semantic segmentation (RSISS) has attracted significant attention due to the growing demand for fine-grained land cover information. The Segment Anything Model (SAM), proposed as a foundation vision model, offers strong segmentation performance and generalization capabilities for RSISS tasks. However, existing SAM-based approaches face two limitations: (1) Insufficient adaptation of SAM's features to the diverse characteristics of land cover types. (2) Semantic ambiguity at object boundaries, which hinders accurate delineation. To address these limitations, we propose Frequency and Edge-guided SAM (FE-SAM), a scalable and efficient framework for RSISS. Specifically, we introduce a Frequency-Modulated Adapter (FMA) that adaptively decomposes and modulates frequency-domain features based on the input data. It selectively enhances informative high- and low-frequency components corresponding to different land cover types. Furthermore, to improve SAM's ability to capture fine-grained details, we design EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image. Extensive experiments on three benchmark datasets demonstrate that FE-SAM outperforms state-of-the-art methods. The source codes are available at: https://github.com/oucailab/FE-SAM.
Feng Gao, Zizhe Pan, Haoting Wang et al.· IEEE Transactions on Geoscie...· 0 citations