Compression, Lexical Alignment, and Segmentation Similarity of Tokenizers Across Programming Languages
Sanghyeon R Joo
Abstract
Tokenization and its downstream effects have been extensively studied under different natural languages but not so much under different programming languages. This paper studies tokenizer behavior of a panel of widely used pretrained tokenizers and code-specific trained tokenizers across eight languages. We measure compression efficiency, alignment with lexical units, and segmentation similarity across tokenizers. Our results highlight the tradeoff between compression and lexical alignment as well as a surprising amount of similarity across language-specific tokenizers.
Successful Page Load