Skip to content
Preprint

Beyond Natural-Image Foundation Models: Benchmarking Satellite Pretraining for Ophthalmic Image Analysis

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

Satellite imagery is proposed as a novel pretraining domain for MedVFM development and benchmarking, motivated by its closer visual alignment with medical data and its freedom from the privacy constraints that limit medical datasets.

Abstract

Vision Foundation Models (VFMs) have emerged as a promising approach in medical imaging, producing broadly applicable systems that can be efficiently adapted across diverse imaging modalities, anatomical regions, and clinical tasks. However, VFMs require extensive training data, and their progress in medical image analysis is constrained by limited data availability, privacy concerns, and high development costs. To alleviate these constraints, medical VFMs (MedVFMs) are often built upon weights from generalist models pretrained on vast amounts of publicly available natural images, introducing a substantial distribution shift for medical task adaptation. To address this, we propose satellite imagery as a novel pretraining domain for MedVFM development and benchmarking, motivated by its closer visual alignment with medical data and its freedom from the privacy constraints that limit medical datasets. Across multiple ophthalmic imaging modalities, we compare DINOv3-SAT493m pretrained on 493 million satellite images against DINOv3-LVD1689m pretrained on 1.7 billion natural images, together with two medical specialist baselines: DINOv3-RETFound and MAE-RETFound. Our experiments show that satellite imagery is a stronger pretraining source than natural images for ophthalmic tasks, particularly on en face vascular-rich modalities. On several tasks, satellite pretraining matches or exceeds the medical specialists on high-resolution en face inputs, despite using no medical data.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is part...

David S. Restrepo, Chen-Wei Wu, L. Nakayama et al. · 0 citations
Open access Aug 2026

Knowledge Distillation from Medical Vision Foundation Models to Lightweight Networks for Edge Deployment: A Comparative Study

Medical vision foundation models pretrained on large-scale domain-specific data achieve strong clinical performance, but their computational cost precludes deployment on resource-constrained edge devices. We investigate whether the domain-specific representations of a medical foundation model transfer more effectively...

Shu-Wei Liu · 0 citations
Sep 2026

MedCure: Medical Data Curation for Efficient Vision–Language Pretraining

Medical foundation models (FMs) typically rely on contrastive vision-language pretraining over large-scale datasets, yet such datasets often exhibit substantial heterogeneity in data quality and demand extensive computational resources. Recent studies suggest that data quality matters more than data quantity, but how t...

Chong Wang, Feng-Bei Liu, Yu-Yuan Liu et al. · 0 citations
Sep 2026

Generalized post-training quantization for medical image segmentation foundation model.

Medical image segmentation foundation models (MedFMs) perform strongly across diverse imaging modalities, but their large size and computational demands hinder deployment in resource-limited clinical settings. Lightweight fine-tuning is impractical given the high training cost of MedFMs, and the efficiency benefits of...

Peng Huang, Ao-Zhong Zhang, Peng-Hang Yin et al. · 0 citations
Preprint Sep 2026

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretraining paradigm uses only RGB images, leaving readily available complementary signals, such as depth m...

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

EyeVQA is introduced, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs, and fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol are benchmarked.

Gu-Jie Shao, Zi-Xun Xie, Xue-Chun Xing et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.