Preprint
Aug 2026
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
This work analyzes what tokens are learned when tokenization is jointly optimized with language modeling, and finds tokenizer-free approaches optimize for contextual and computational efficiency rather than strict morphological structure, resulting in fundamentally different yet effective vocabularies for downstream NLP.
Saketh Reddy Vemula, Parameswari Krishnamurthy
· 0 citations