THE SINDHI TOKEN TAX: ORTHOGRAPHIC PENALTY AND TOKENIZATION INEQUITY BETWEEN ENGLISH AND SINDHI IN LARGE LANGUAGE MODELS
Large language models (LLMs) do not read words; they read tokens, and the tokenizer decides how many tokens a language needs to say the same thing. This paper introduces the notion of a Sindhi token tax and tests it empirically for Pakistan’s English–Sindhi context. Using 1,004 parallel sentences from the FLORES-200 be...