BBT: BPE-Guided Byte Transformer
Abstract
Byte Pair Encoding (BPE) tokenizers shorten language model input sequences, while standard BPE Transformers typically build token representations through vocabulary lookup. This lookup represents each token with an independent vector, so byte-level structure inside tokens is not reflected in vocabulary-related parameters. In small-scale models, a large vocabulary also makes the token embedding / LM head matrix account for a substantial fraction of parameters. To address this, we propose BPE-Guided Byte Transformer (BBT), a model that uses BPE for segmentation while replacing vocabulary lookup with byte-level encoding and decoding. BBT uses a ByteEncoder to produce byte hidden states and to extract the states at segment boundaries as segment representations. The SegmentTransformer performs causal modeling over the resulting sequence of segment representations. A ByteDecoder then generates bytes in the next segment conditioned on these SegmentTransformer outputs. This preserves BPE's sequence compression effect while shifting the input/output parameterization from a vocabulary-sized token embedding / LM head matrix to byte-level encoder/decoder modules. On a subset of FineWeb with 47.3 GB of raw text bytes, we compare BBT with standard BPE Transformers whose segment-level Transformer stacks are matched at hidden dimensions 1024, 2048, and 3072. Across these paired runs, BBT reduces total parameters by 20.2-40.4% and achieves lower test bits per byte (BPB) by 2.7-6.6%. Beyond BPB, auxiliary evaluations show competitive general task performance, along with improvements in completion behavior and fine-grained string robustness.