What Makes an Encoding Good for Language Modeling?
Abstract
What are the ideal atomic units of text for language modeling? The de facto approach is to convert all text into UTF-8 bytes prior to tokenization. While UTF-8 encoding expresses all Unicode code points within a limited vocabulary, it also imposes systematic disadvantages on many languages. We propose the Aether framework: a hybrid setting that leverages script-specific encodings for target scripts while maintaining the universal expressivity of UTF-8 bytes. Aether tokenizers allow us to improve encoding representations along axes such as compression and linguistic information, while also enabling comparisons of encodings across tokenizers. We train a comprehensive sweep of models, using language-specific character encodings, Unicode variants, decompositional encodings, and our own custom encodings to represent target scripts, on Korean and Chinese, which have large character sets and multiple pre-existing encoding standards. We find that the compression and vocabulary size of an encoding are strongly correlated with language modeling performance in both languages, and that linguistic information and structural regularities in the encoding are beneficial. Intentional encoding choices can lead to more equitable, efficient, and performant language models for non-Latin scripts.