Retinal image analysis for clinical translation: From deep learning to foundation models, generalization, and trustworthy deployment.
TL;DR
This review examines retinal image analysis from the perspective of clinical translation rather than benchmark-oriented model comparison, arguing that the next stage of retinal AI requires transferable representations, robust external evaluation, trustworthy uncertainty handling, and demonstrated value within real-world care pathways.
Abstract
This review examines retinal image analysis from the perspective of clinical translation rather than benchmark-oriented model comparison. Instead of organizing prior work solely by disease category or model family, we synthesize recent advances through four connected dimensions: imaging modality, task taxonomy, methodological paradigm, and translational bottleneck. We compare major tasks, including classification, detection, segmentation, grading, progression prediction, and treatment-response assessment, across fundus photography, OCT/OCTA, and angiographic imaging. We further review the roles of CNNs, U-Net variants, 3D models, Transformers, graph-based methods, hybrid architectures, foundation models, self-supervised pretraining, and multimodal learning under different data conditions and clinical constraints. Beyond technical progress, we analyze why many high-performing systems still fail to translate reliably into practice, highlighting challenges related to distribution shift, label inconsistency, limited external validation, image-quality control, calibration, fairness, and workflow integration. We argue that the next stage of retinal AI requires transferable representations, robust external evaluation, trustworthy uncertainty handling, and demonstrated value within real-world care pathways, rather than incremental architectural novelty alone.