Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification
Abstract
Deep neural networks (DNNs) have achieved remarkable success in medical image classification, yet their performance remains sensitive to dataset size. Knowledge distillation (KD) alleviates this issue by transferring knowledge from a high-capacity teacher to a lightweight student. However, conventional KD relies on a static teacher, while adaptive teacher updating may improve performance on the student-learning data while reducing retention of knowledge acquired during teacher pretraining. To address these limitations, we propose an Anti-forgetting Adaptive Teacher-driven Knowledge Distillation framework (A2T-KD), which aims to balance teacher adaptation and pretraining-knowledge retention. The proposed framework integrates three modules: MITR for cross-epoch representation consistency, DSDO for prediction-space decoupling and class discriminability, and SGKD for feature- and logit-level knowledge transfer. Across nine medical imaging datasets, A2T-KD achieved higher mean values than the fixed-teacher Vanilla KD baseline in 30 of 36 dataset–metric comparisons. It also exhibited the lowest pretraining-set ACC degradation among the evaluated teacher-update baselines on all nine datasets, supporting the intended balance between teacher adaptation and pretraining-knowledge retention under the evaluated settings.