Reframing Language Segmentation: Generalizing Past Tokenizers
Abstract
We introduce an interface that enables a more general class of segmentation priors --- of which tokenization is one instance --- to be fed into transformer backbones. At small scale, we show that the interface not only trains better models under GPT-2's BPE tokenizer prior adapted to it, but that more-general priors such as regular expressions used for pretokenization train better still. Each results in lower bits per byte (bpb) at matched non-embedding parameters, under both data-matched (D) and compute-matched (F) accounting. This new class of priors lets each transformer step carry more bytes --- shortening the global sequence for a given text --- even without a trained vocabulary. Our motivation is the long-context regime: because attention is quadratic in sequence length, a shorter sequence yields substantial FLOP savings downstream of pretraining --- long-context RL and, above all, inference.