Comparative Study of Tokenizer Replacement, Embedding Initialization, and Continued Pretraining for LLM Language Adaptation
Abstract
Adapting pretrained large language models to less represented languages requires modifying their tokenizers to improve segmentation efficiency, which introduces the challenge of aligning the model's embedding space with the newly introduced vocabulary. Using Qwen3.5-0.8B-Base, we investigate strategies for constructing a Polish-oriented tokenizer by either modifying the original vocabulary or training a new one from scratch. We evaluate five embedding initialization methods and architectural optimization techniques during continued pretraining, assessing both intrinsic tokenizer characteristics and downstream cross-lingual benchmark performance. Our results demonstrate that extending the original tokenizer paired with block expansion provides the optimal balance between target-language adaptation and source-language retention. Furthermore, we find that the effectiveness of embedding initialization methods is highly dependent on how the tokenizer is prepared, and that freezing-based methods can improve both training and evaluation results.