When Tokenization is Secretly Output Supervision
Tanja Baeumel ⋅ Josef v Genabith ⋅ Simon Ostermann
Abstract
Tokenization in language models is universally treated as an input preprocessing decision. In this position paper, we argue that this framing is incomplete: in autoregressive models, tokenizer granularity determines what the model must resolve in a single forward pass, and therefore the supervision signal it receives. This affects both the difficulty of the learning problem and the model internal representations that emerge. We develop the $\textit{minimal computation hypothesis}$: Autoregressive models represent what their output supervision requires, and have no gradient pressure to represent more. We test this in a controlled experiment on numeric reasoning with a novel decoupling of input from output tokenization. We show that differences in performance and model internals are induced by $\textit{output}$ tokenization, not $\textit{input}$ tokenization. We also survey 120 recent *CL papers on numeric reasoning and find that only about 10% report numeric tokenization of the evaluated models, while 69% compare models trained under different numeric tokenization, and thus supervision, regimes without reporting tokenization strategies. Conclusions attributed to numeric reasoning ability may thus partly reflect differences in task definition. While a significant body of prior work documents that tokenization affects model performance across domains, there is no principled account of $\textit{why}$ yet. We argue that framing tokenization as output supervision provides that account and discuss implications for comparability of cross-tokenizer model evaluations. We argue that framing tokenization as output supervision provides that account: performance differences may partly reflect differences in the learning problem induced by tokenization.
Successful Page Load