What Do We Talk About When We Talk About Tokenization?
Abstract
Tokenization appears across a large and heterogeneous body of NLP literature, most often as a design choice rather than the primary focus. We identify more than 800 such papers published over the last decade and use an LLM-based methodology to survey them at a scale that would not be feasible through manual review. We analyze four dimensions: the languages studied, the evaluation tasks, the intrinsic metrics used, and the granularity of the tokenization units considered. We find that machine translation is the most frequently studied when tokenization is discussed. Despite the multilingual focus of much of this work, a strong English and Eurocentric bias persists. Subword tokenization dominates in practice despite sustained research interest in characters, bytes, and morphemes. We release the extracted dataset to support further analysis.