Flexible Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on LLMs and Beyond.
Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, as they assume a consistent semantic correspondence across logit dimensions, limiti...