Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Zhenyu Zhao ⋅ Sander Land ⋅ Daniel M Bikel ⋅ Waseem Alshikh
Abstract
Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that reasoning tokens split into two functional types: low-entropy $\textit{structural}$ tokens (recurring phrases that scaffold the reasoning process) and higher-entropy $\textit{organic}$ tokens (problem-specific content that drives toward a solution). This asymmetry motivates a simple, model-agnostic compression pipeline: apply cross-word BPE merges on a model's own reasoning traces to derive $\textit{supertokens}$ that capture frequent structural patterns, then teach the model to adopt them via supervised fine-tuning. Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by $8.1\%$ on average; under a TOST equivalence analysis at a $\pm 2$~pp margin, $2/15$ model--benchmark cells pass equivalence, $11/15$ are inconclusive (predominantly AIME at $N=30$), and $2/15$ fail equivalence with a real accuracy loss on DeepSeek-R1-Distill-Llama-70B. Beyond compression, learned supertokens often align with interpretable reasoning moves such as backtracking, verification, and strategy shifts. This enables a compact structural analysis of reasoning traces: correct traces show more recovery and verification patterns, while incorrect traces show more repeated hedging and unresolved counterarguments. We release the full pipeline as open-source code.
Successful Page Load