Objective vs. Search: Decomposing What Makes a Good Tokeniser
Abstract
The two dominant tokenisation algorithms used in modern language models, byte-pair encoding (BPE) and UnigramLM, differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these axes, making it unclear whether observed performance differences arise from what is being optimised or from how it is optimised. We disentangle these factors by introducing two new tokenisers that complete this 2×2 design space: BottomUpLL, a bottom-up likelihood-based tokeniser, and TopDownComp, a top-down compression-based tokeniser. We train language models using all four tokenisers across model scales (100M–1B parameters), vocabulary sizes (8k, 32k, and 128k), and both English-only and multilingual settings. Evaluating these models on bits-per-byte and BLiMP, together with a range of intrinsic tokenisation metrics, we find that the search procedure—not the optimisation objective—is the dominant factor governing downstream performance: bottom-up tokenisers consistently achieve lower bits-per-byte across settings, while the choice of objective mainly matters in the small-vocabulary regime. These findings hold across model sizes and languages, providing new insight into the relative roles of optimisation objectives and search procedures in tokeniser design and offering practical guidance for constructing more effective tokenisers.