Writing-System-Level Tokenizer Adaptation for Byte-Level BPE
Abstract
Inserting new tokens into a byte-level BPE merge table is structurally broken: in our Nemotron-3 setting, heuristic merge splits conflict with BPE's greedy execution and 65% of inserted tokens become unreachable, a failure we formalize as the merge ordering problem. Balde et al. report a similar 64% on LLaMA-2-7B for domain vocabulary adaptation, so this is not language-specific. We resolve it structurally with BPE-guided insertion, which registers each donor merge against the decomposition reachable under the target tokenizer's existing merge ordering, yielding 100% token integrity at construction time without runtime modification to BPE inference. We embed this in a fixed-vocabulary writing-system-level pipeline (script-aware removal, target-script base reconstruction, BPE-guided insertion) that holds vocabulary size constant. As empirical evidence, two Ukrainian adaptations of open-weight models reduce fertility by 25--37% while preserving tokenizer-level behavior on English and the evaluated protected languages. Fertility reduction itself is the expected consequence of any reallocation; the structural contribution is that the reallocation is reachable at all in byte-level BPE. On the target artifacts, BPE-guided insertion leaves 0% unreachable inserted tokens, while archived pre-guided builds leave 65% unreachable on Nemotron-3 and 85% on GPT-OSS. As a separate stress test of the construction, the same procedure leaves 0% unreachable inserted tokens on GPT-2, Qwen3, and Aya byte-level BPE tokenizers, while a naive append baseline leaves 34--99% unreachable.