Self-supervised vision transformers for intelligent OCT-based retinal disease classification and severity assessment
Abstract
Vision Transformers (ViTs) have demonstrated strong performance in medical image analysis due to their ability to model long-range dependencies through self-attention mechanisms. However, training such models typically requires large annotated datasets, which are often difficult to obtain in medical imaging. To address this challenge, this study investigates the integration of self-supervised learning (SSL) with transformer-based architectures for retinal image analysis. We propose a computational intelligence framework that combines the DINO self-supervised learning method with Vision Transformers to learn robust feature representations from unlabeled optical coherence tomography (OCT) scans. The model is first pretrained on a large collection of unlabeled images and subsequently fine-tuned on a smaller expert-annotated dataset for three diagnostic tasks: binary classification, multiclass retinal disease identification, and disease severity assessment. Experimental results demonstrate that SSL-based models can effectively learn informative features from unlabeled data. The ViT-SSL model achieved accuracies of 95.0% for binary classification, 77.7% for multiclass classification, and 82.5% for severity assessment, while the CNN-SSL counterpart achieved 94.9%, 87.2%, and 88.9%, respectively. Additional evaluation on a public OCT dataset confirms the generalizability of the learned representations. The results highlight the potential of self-supervised transformer-based frameworks for intelligent and data-efficient retinal disease analysis in medical imaging systems.