Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Distilling Encoder-Decoder and Decoder-Only Models into BERT for NLP Tasks

Large pretrained language models achieve strong performance on natural language processing tasks but are costly to deploy. Existing distillation methods in this domain almost exclusively assume architectural homogeneity between teacher and student, leaving cross-architecture transfer, well studied in computer vision, l...

Kristopher Nathanael, Danielson, Jeremy Auriel Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.