Beyond Metric Stability: Interface-Preserving Token Halting and Fusion
Abstract
Adaptive token-level inference can substantially reduce the compute, latency, and energy cost of language-model deployment. As an emblematic example, QuickSilver reports up to a 39.6% reduction in floating-point operations with at most a 0.2 increase in perplexity by combining token halting, cache skipping, token fusion, and entropy-guided precision allocation. Yet the same framework also provides full-processing overrides, halting blocklists, minimum-depth rules, fusion exemptions, and precision override masks for tokens whose semantic roles are not safely captured by its scalar gates. We argue that these safeguards are not peripheral edge-case patches: they approximate a structural condition missing from scalar stability. Using a category-theoretic formalization of tokens, we distinguish metric stability from interface-preserving stability. Scalar criteria propose identifications among token traces; halting and fusion then induce quotient-like maps. Such a map is admissible relative to a downstream interface only when the interface factors, exactly or approximately, through the quotient. A minimal negation example and a richer engineering example show why a token can become locally stable precisely because its compositional role has been resolved, while remaining indispensable to later computation. We conclude by framing approximate factorization error as a diagnostic and by identifying the open challenges that would have to be solved for interface-sensitive token compression to become practical.