The results show that tailoring augmentation strategies to the characteristics of retinal images plays a critical role in improving performance, and even under constrained settings, lightweight SSL frameworks can learn transferable representations that reduce dependence on large annotated datasets and achieve competitive results.
Abstract
Despite the growing number of public datasets, annotated medical images remain scarce. Supervised learning methods achieve strong performance on many benchmarks, however require large amounts of labeled data, which are costly and time-consuming to obtain in the medical domain. To address this limitation, contrastive self-supervised learning (SSL) has emerged as a promising alternative for learning useful representations from unlabeled data. In this work, we investigate two SSL frameworks, SimSiam and SimCLR, for retinal disease classification from fundus images. We focus on understanding how augmentation strategies and training parameters influence representation learning under resource-constrained settings. Given limited data and computational capacity, we explore the feasibility of training SSL models with small batch sizes incorporated with retinal-specific augmentation techniques. Through a series of experiments, we assess the quality of learned representations via linear evaluation and fine-tuning across downstream tasks, including multi-disease classification and diabetic retinopathy grading. Our results show that tailoring augmentation strategies to the characteristics of retinal images plays a critical role in improving performance. Even under constrained settings, lightweight SSL frameworks can learn transferable representations that reduce dependence on large annotated datasets and achieve competitive results.
Vision Transformers (ViTs) have demonstrated strong performance in medical image analysis due to their ability to model long-range dependencies through self-attention mechanisms. However, training such models typically requires large annotated datasets, which are often difficult to obtain in medical imaging. To address...
Abdulkarem Almshnanah, Rehab M. Duwairi· Journal of Intelligent &...· 0 citations
Purpose To develop and evaluate an unsupervised domain adaptation (UDA) framework for glaucoma classification from fundus images that improves the generalizability of deep learning (DL) models across heterogeneous imaging characteristics and clinical settings. Methods We developed an adversarial UDA framework that adap...
Homa Rashidisabet, R. V. Paul Chan, T. Vajaranant et al.· Translational Vision Science...· 0 citations
FusionEye-Net combines strong internal performance with substantially improved cross-domain generalization through lightweight adaptation, underscoring the critical role of external validation and domain-aware fine-tuning in formulating trustworthy, adaptive AI-assisted decision-support tools that can safely augment au...
Ali M. Duhaim, A. M. Al-Bakry· Discover Artificial Intellig...· 0 citations
This study presents a patch-based contrastive learning framework (Patch-SimCLR) designed to improve the generalizability and calibration of diabetic retinopathy classification across heterogeneous fundus datasets. By extracting overlapping, and non-overlapping patches and leveraging contrastive pretraining, the model l...
Usman Ali, A. A. Imam, R. Apong· International Conference on...· 0 citations
This study introduces a multi-dimensional framework to systematically evaluate the generalization and computational efficiency of deep learning models for retinal vessel segmentation. Using the DRIVE and STARE datasets, this research evaluates the performance of SegNet, U-Net, and DeepLab across in-domain, cross-domain...