Distilling Encoder-Decoder and Decoder-Only Models into BERT for NLP Tasks
Abstract
Large pretrained language models achieve strong performance on natural language processing tasks but are costly to deploy. Existing distillation methods in this domain almost exclusively assume architectural homogeneity between teacher and student, leaving cross-architecture transfer, well studied in computer vision, largely unexplored in natural language processing, where teacher and student differ in fundamental mechanisms such as causal versus bidirectional attention. We present an empirical study of cross-architecture knowledge distillation, fixing a bidirectional encoder student and varying the teacher across encoder-decoder and decoder-only models, with a same-architecture large encoder teacher as baseline. Using vanilla logit-level distillation, we evaluate on three General Language Understanding Evaluation benchmark tasks for sentiment, inference, and paraphrase detection. We find that while distillation efficacy is highly task-dependent, cross-architecture students generally preserve accuracy and F1 within a narrow margin of the same-architecture baseline, while inheriting the lightweight inference profile of the student and yielding substantial reductions in memory, latency, and energy relative to their teachers. These results indicate that output-level knowledge distillation transfers task-discriminative knowledge across fundamental architectural divides.