Skip to content

SAGE: A sampling-aware global evaluation benchmark for species distribution modeling

Sep 2026 · 0 citations · 90 references
Computer Science Biology Mathematics

TL;DR

A Sampling-Aware Global Evaluation (SAGE) benchmark is introduced, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species, and an evaluation framework that groups species based on two properties, sampling effort and relative prevalence is proposed.

Abstract

Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale, averaging performance hides substantial species-level variability, particularly for rare species, often of greatest conservation concern. Records are also strongly biased, making occurrence counts misleading. Accounting for these factors is essential for a reliable and informative evaluation of multi-species SDMs. Here, we introduce a Sampling-Aware Global Evaluation (SAGE) benchmark, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species. We propose an evaluation framework that groups species based on two properties, sampling effort and relative prevalence, which describe how densely a species'range is sampled and how frequently the species is recorded. Evaluating single-species SDMs and multi-species DeepSDMs, we find that Random Forests and DeepSDMs perform best overall, but neither dominates: DeepSDMs outperform single-species SDMs for infrequently recorded species while offering no consistent advantage for well-sampled ones. Crucially, this advantage emerges only when established bias-correction practices, such as spatial thinning and reweighting, are carried over to the deep-learning setting. SAGE helps identify the species and data conditions for which a given approach is beneficial, thereby supporting the development of more transparent and ecologically credible SDMs. Data and code: https://earens.github.io/sage/

View source

Similar papers

Open access Sep 2026

Probabilistic species distributions from nine large-scale gridded atlases over five decades

Information on the long-term dynamics of species distributions is essential for assessing global biodiversity change and its causes, and for informed conservation decisions. A valuable source of such historical information are atlases. They provide spatial standardisation, extensive geographic coverage, temporal replic...

F. Grattarola, Gabriel Ortega-Solís, Carmen D. Soria et al. · 0 citations
Open access Sep 2026

Occurrence-based models reveal greater spatial details in tropical species richness

Species richness maps are essential for biodiversity assessment and conservation planning, but their accuracy is limited by uneven sampling effort and large data gaps, especially in tropical regions. Here, we combined the Uniform Sampling from Sampling Effort (USSE) framework with deep neural networks to generate occur...

Ubirajara Oliveira, B. Soares-Filho · 0 citations
#machine learning Review Sep 2026

Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling

This work introduces Active Continuous-Score Occupancy Modeling (ACORN), a method that incorporates ML predictions into occupancy models and strategically selects samples for expert review that are maximally informative for downstream ecological analysis.

T. Haucke, Lauren Harrell, Justin Kay et al. · 0 citations
Review Open access Oct 2026

Species distribution models of ESA-listed corals highlight niche differences among taxa that may guide recovery planning

Understanding the biogeography of living marine resources is essential for informed management, particularly in data-limited environments. Here, we use two corals listed under the U.S. Endangered Species Act (ESA) around Tutuila, American Samoa, to evaluate a transferable workflow that combines presence-only occurr...

Kisei R. Tanaka, Thomas A. Oliver, C. Couch et al. · 0 citations
Review Open access Aug 2026

Closing the biodiversity observation-to-action loop

Citizen science observations are abundant, but conservation requires turning uneven records into reliable predictions and directing new surveys to where information is missing. We developed a biodiversity platform for Japan that is updated monthly and integrates 2.32 million records to predict 8,297 species across seve...

Keiji Yamaguchi, Kei Uchida, Masayoshi K. Hiraiwa et al. · 0 citations
Open access Sep 2026

Modelling spatial deficits in citizen science recording: a case study from UK butterflies

Citizen science data are increasingly used to infer species distributions and monitor biodiversity at large spatial scales. However, recording effort is often uneven across space, potentially biasing inferences about species’ distributions. We analysed 51,045 butterfly records submitted via the iRecord Butterflies app...

Ming-Rui Li, R. ffrench-Constant, Richard Fox et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.