Distilling Encoder-Decoder and Decoder-Only Models into BERT for NLP Tasks
Large pretrained language models achieve strong performance on natural language processing tasks but are costly to deploy. Existing distillation methods in this domain almost exclusively assume architectural homogeneity between teacher and student, leaving cross-architecture transfer, well studied in computer vision, l...