More Than Words: Compositional Tokenization for Efficient Language Models
Abstract
Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as “On the table.” is usually produced as four separate predictions for the preposition (on), article (the), noun (table), and period (.), affecting the overall context length, and accordingly—inference costs. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a single lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled from-scratch pretraining at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves downstream performance by 1.2 points relative to standard BPE under matched training compute. More broadly, our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.