Fewer Transformer Positions, Same Subword Evidence: Word-Level Fusion for Multilingual Tokenization
Abstract
Subword tokenization is a robust default for multilingual language models, but it fragments the productive word forms that many morphologically rich and underrepresented languages use heavily. This fragmentation is not only a modeling issue: it can create unequal compute and access costs for speakers of Global South languages. We argue that tokenizers should be evaluated as \emph{interfaces}: they trade off preservation of lower-level evidence against the number of transformer positions exposed to the model. We study \emph{compositional tokenization interfaces}, which preserve subword or morphological evidence while presenting a word-level computational unit to later layers. The paper combines two evidence streams. First, controlled language-modeling experiments replace lookup embeddings for productive forms with single-token compositional embeddings derived from morphological parts. These experiments show consistent gains for Turkish and selective transfer to Hindi, German, Finnish, and Telugu. Second, pretrained fine-tuning experiments insert a word-level fusion layer between a frozen multilingual encoder embedding layer and its transformer stack, then adapt the model with LoRA on natural language inference. The strongest pretrained result is additive fusion: Hindi reaches 0.824 test accuracy versus 0.833 for the standard subword baseline while reducing the average effective sequence length from 49.12 subword tokens to 31.69 word-level units, a 1.55x compression; Turkish reaches 0.799 versus 0.809 while compressing from 41.85 to 20.92 positions, a 2.00x compression. Destructive Hindi ablations collapse to chance when word boundaries are removed or only the first subword is kept, while omitting word-level position re-indexing reduces accuracy to 0.733. The resulting claim is narrow but useful: simple additive compositional interfaces can retain nearly all pretrained fine-tuning performance while substantially reducing sequence length, and the gain depends on meaningful word grouping rather than arbitrary compression.