Byte-Level Tokenizer Transfer with Auxiliary Objectives
Abstract
Tokenization is a foundational component of large language models (LLMs), but standard subword methods struggle with non-standard orthographies, low-resource languages, and linguistic noise. In this work in progress, we explore tokenizer transfer to address these limitations, replacing a pre-trained model's tokenizer with a more adaptable alternative without full retraining. Building on the Bolmo architecture, we extend its byte-level encoder with auxiliary losses targeting three properties: noise robustness, cross-lingual token alignability, and morphological accuracy. For noise robustness, we either reconstruct the original input from the byte-level encoder via an auxiliary decoder, or train the boundary predictor to produce consistent segmentations across clean and noisy inputs. For token alignability, we apply contrastive learning on parallel data to bring representations of aligned tokens closer across languages, together with a language identification task. For morphological accuracy, we gradually transition the boundary predictor from mimicking subword segmentation to predicting reference morphological boundaries. These objectives aim to produce a tokenizer that is better suited for multilingual and low-resource settings, without the prohibitive cost of full-model pre-training, and more robust to linguistic variations.