Knowledge Distillation from Medical Vision Foundation Models to Lightweight Networks for Edge Deployment: A Comparative Study
Abstract
Medical vision foundation models pretrained on large-scale domain-specific data achieve strong clinical performance, but their computational cost precludes deployment on resource-constrained edge devices. We investigate whether the domain-specific representations of a medical foundation model transfer more effectively through knowledge distillation (KD) than generic ImageNet-pretrained features. We compare a domain-specific foundation model (RETFound, ViT-Large, 307M parameters) against a general-purpose pretrained model (ConvNeXt-Base, 89M) as KD teachers for two lightweight students (MobileNetV3-Small, 2.5M; EfficientNet-Lite0, 4.7M) on two clinical benchmarks – HAM10000 (dermoscopy, 7 classes) and APTOS-2019 (diabetic retinopathy, 5 classes) – over a grid of distillation temperature and loss weight, with INT8 quantization and Grad-CAM analysis. Because RETFound is pretrained on retinal images, APTOS- 2019 is in-domain for it whereas HAM10000 (dermoscopy) constitutes an out-of-domain test. On balanced accuracy, the appropriate metric for these class-imbalanced datasets, ConvNeXt-Base yields the stronger student in all four settings (e.g., HAM10000 MobileNetV3: 0.826 vs. 0.763) – including on APTOS-2019, where RETFound’s retinal pretraining is in-domain – despite fewer parameters and no domain-specific pretraining; RETFound distillation can even fall below the no-distillation baseline. RETFound requires higher temperatures for effective transfer, whereas ConvNeXt is robust across a broad range. INT8 quan- tization preserves accuracy within 0.3% while compressing MobileNetV3-Small to 4.2 MB. These results indicate that a teacher’s fine-tuned quality, rather than domain-specific pretraining, predicts distillation success, and offer practical guidance for edge deployment of medical image classifiers.