Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP
Abstract
Subword regularization methods such as BPE dropout are typically applied only during fine-tuning, while pretraining is usually done with deterministic tokenization. This creates a potential word segmentation mismatch between pretraining and fine-tuning. We investigate whether applying BPE dropout during pretraining improves downstream performance in low-resource NLP. We trained monolingual and bilingual BERT models on downsampled subsets of English, German, French, Spanish, Kiswahili, and isiXhosa, and evaluated them on XNLI, PAWS-X, PAN-X, and MasakhaNER 2.0. Across these tasks, the best results are typically obtained when BPE dropout is applied during both pretraining and fine-tuning, whereas applying BPE dropout only during fine-tuning can underperform deterministic tokenization in smaller-data settings. As the amount of fine-tuning data increases, this disadvantage diminishes. We further find that the benefits of pretraining-time BPE dropout are largest when either pretraining or fine-tuning data is scarce, which suggests that its regularization and data augmentation effects might be particularly useful in resource-constrained settings. The benefits of BPE dropout have often been attributed to better compositional representations, especially for rare words. To examine this, we also measured morphological boundary alignment under BPE dropout and found only modest improvements in expected alignment, while better-aligned segmentations remained rare. This suggests that fine-tuning alone may provide limited exposure to such segmentations, whereas applying BPE dropout during pretraining increases cumulative exposure before downstream adaptation. Our morphologically aligned fine-tuning intervention further supports this pattern, as models pretrained without BPE dropout benefit more from targeted aligned segmentations than models already exposed to BPE dropout during pretraining. Overall, these findings suggest that exposure to better-aligned segmentations may contribute to the downstream benefits observed when applying BPE dropout during pretraining.