SPEED: Specialized Position Experts for Efficient Speculative Decoding
Abstract
Speculative decoding accelerates Large Language Model (LLM) inference by using a small draft model to predict multiple tokens, and a large target model to verify these tokens in parallel. Recent studies leverage features of target and draft models to enhance draft accuracy, but suffer from the degrading accuracy of draft tokens at later positions, due to accumulated deviation in draft model-generated features. Worse still, features of inconsistent levels of deviation provide a noisy training signal at training stage. In this paper, we propose SPEED, applying specialized position experts to draft tokens at preassigned positions. Position experts substantially improve acceptance rate at later draft positions, as each expert only needs to focus on handling a certain level of feature deviation. Experiment results on six datasets demonstrate that SPEED effectively improves over baselines on average acceptance length and speed-up ratio. Our codebase is available at https://github.com/speed-speculative-decoding/speed.