A Call for an Open Tokenizer Benchmark
Marco Cognetta
Abstract
Tokenizer research lags behind other areas of the language modeling stack due in part to the difficulty of: 1) comparing across models with different tokenizers, 2) quickly and accurately measuring the impact of a tokenizer change, 3) a lack of predictive metrics and scaling laws for tokenizers, and 4) a lack of standardized public benchmarks. Here, we hope to start the discussion and formalization of tokenization-targeting language modeling community benchmarks \textit{à la} the NanoGPT Speed/Slowrun.
Successful Page Load