Is Tokenizer Cognitive Plausibility just Fertility?
Abstract
Subword tokenizers are increasingly evaluated using “cognitive plausibility” metrics derived from lexical decision experiments, where reaction times are interpreted as evidence that tokenization approximates human linguistic units. A prominent example is chunkability, defined as 1− tokens-per-character (fertility), which is intended to capture segmentation quality independent of model compression. We re-examine this assumption by evaluating 633 tokenizers (BPE, Unigram, SuperBPE, and Morfessor variants) across eight languages using large-scale lexical decision psycholinguistic studies, including English, Dutch, French, Spanish, Chinese, Korean, Hebrew, and Malay. For each item, we model reaction times using nested regressions controlling for frequency and word length, and compare the incremental explanatory power of fertility, chunkability, and boundary-based segmentation metrics such as MorphScore and a new metric called Boundary–Surprisal Alignment (BSA). We show that “cognitive plausibility” signals used to evaluate subword tokenizers are largely driven by compression (fertility / tokens-per-character) rather than segmentation quality. Across 633 tokenizers and 8 languages, the popular metric chunkability is an exact affine transform of fertility and therefore adds no independent explanatory power for lexical decision times. Fertility (compression) consistently predicts lexical decision times across languages, with strongest effects in Chinese and Korean, while chunkability adds no residual variance once fertility is included. Third, boundary structure is empirically real: tokenizers differ systematically in morphological and surprisal-aligned segmentation (e.g., Morfessor > Unigram > BPE), yet these differences are nearly orthogonal to reaction times after controlling for compression. Across all settings, boundary-based metrics (including BSA) explain negligible additional variance in lexical decision performance once fertility is accounted for. These findings suggest that lexical decision studies primarily measure sensitivity to token count rather than segmentation quality, and that correlations previously attributed to “cognitive plausibility” largely reflect compression effects.