Direct Multi-Token Decoding
Xuan Luo ⋅ Weizhi Wang ⋅ Xifeng Yan
Abstract
Recent studies suggest that pre-trained large language models (LLMs) might develop distinct functional roles across their layers: early layers focus on understanding the input context, middle layers handle task-specific processing, and late layers map abstract representations to output tokens. Inspired by the intuition that humans can produce an extended sequence of speech from a single act of reading and thinking, we hypothesize that a single pass through the early and middle layers could provide sufficient information to support the decoding of multiple future tokens. Accordingly, we propose Direct Multi-Token Decoding (DMTD), which performs only one full forward pass for every $n$ tokens and reuses the late layers to decode multiple subsequent tokens. Unlike speculative decoding, our method introduces no additional parameters, auxiliary routines, or post-generation verification. With minimal training overhead, a fine-tuned Qwen3-4B model utilizing DMTD achieves up to a $2\times$ speedup with marginal performance loss. Moreover, our scaling analysis indicates that the performance of DMTD could be further improved with increased training.
Successful Page Load