Whose Tokenizer? An Empirical Review of When and How Much Tokenizer Choice Matters for Multilingual Language Models
Clara Meister ⋅ Amit Moryossef ⋅ Tiago Pimentel ⋅ Antoine Bosselut
Abstract
For some tokenizer properties, creating a multilingual vocabulary under fixed size constraints is a zero-sum game. For example, compression in one language typically comes at the cost of compression in another. But do the same tokenizer properties matter equally across languages? We show that for the performance of a multilingual language model, various tokenizer intrinsic properties matter for different languages to a different extent. We train 1.27B parameters multilingual language models that differ only in their tokenizer (41 combinations), with architecture, data, and optimization held fixed, and measure per-language bits-per-byte (BPB) on the 31 natural languages in the training mixture. We report three main findings. First, the across-tokenizer variation in a language's BPB is larger for languages with less training data: the standard deviation of per-language BPB across the tokenizers correlates negatively with training-data share ($\rho = -0.58$). The across-tokenizer BPB range is $0.02$ to $0.03$ for high-resource languages such as Russian and German and $0.18$ for the low-resource language Bengali. Second, which intrinsic metric predicts a language's BPB also depends on the language's resource level, writing system, and morphological type: the within-language coefficient relating Renyi efficiency to BPB is $-0.016$ at low resource and $+0.002$ (not significant) at high resource. Lastly, no single intrinsic metric provides a tokenizer ranking consistent with downstream BPB across all languages, so ranking them requires a small set of information-theoretic metrics rather than any one; an aggregate or English-only selection criterion raises BPB for nearly every language relative to per-language selection. This set orders held-out 1B tokenizers at pairwise accuracy $0.80$ (validation BPB) and $0.78$ (FLORES) with no model trained (leave-one-tokenizer-out. The practical implication is that tokenizer evaluation for multilingual models should be per language, or at least per language type: an aggregate or English-only criterion is least informative for the low-resource languages where the choice of tokenizer matters most.
Successful Page Load