Reading a Word Two Ways: A Learned Gate over Tokenization Layouts for Frozen Language Models
Abstract
A language model reads text through one tokenizer, and that single choice shapes what it can and cannot do. Byte-pair tokenizers merge common letter groups into single tokens, which hides the individual characters and makes simple tasks like counting letters surprisingly hard. One fix is to show the model both views of a prompt at once, the normal subword view and a character-by-character view, without retraining the model. We study this on two frozen Qwen2.5 models and find that no single way of combining the two views works best everywhere. The best layout changes from one task to another and even from one model size to another. We then train a small router that reads the raw prompt and picks a layout for it. The router beats every fixed layout on both models, recovers about a third of the gap to a per-example oracle, and uses fewer tokens than always showing both views. It also learns a different policy for each model, matching the layout preferences we measured by hand. The useful question is not which tokenization is best, but when to use which, and a model can learn the answer from task outcomes alone.