A continual learning method based on a learnable knowledge transfer network: preliminary exploration
Abstract
Continual learning requires a model to retain knowledge of old tasks while sequentially learning new tasks, but standard neural networks typically suffer from catastrophic forgetting in this setting. To address this challenge, a new method is proposed based on a learnable knowledge transfer network. Specifically, a transfer network is trained to compress the knowledge of a multi-branch model into a single-branch model, effectively extending the application of knowledge distillation in continual learning. Experimental results show that our approach is feasible: in single-task distillation scenarios on Split MNIST, the compressed single-branch network maintains over 90% accuracy. This confirms the existence of a learnable structure in weight space that our transfer network can exploit. Preliminary multi-task experiments are also conducted to analyze the core challenges and explore potential directions for future improvement. Overall, our method provides a fresh perspective by treating the distillation process itself as a meta-learning problem.