Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Yuqian Fu ⋅ Haohuan Huang ⋅ Kaiwen Jiang ⋅ Jiacai Liu ⋅ Zhuo Jiang ⋅ Yuanheng Zhu ⋅ Dongbin Zhao
Abstract
On-policy distillation (OPD) has emerged as a promising paradigm for LLM post-training, offering an efficient alternative to reinforcement learning by leveraging a teacher model during student rollouts. However, the dominant sampled-token formulation suffers from fundamental limitations: it reduces distribution matching to a single-token signal and degrades in quality as rollout prefixes drift outside the teacher's training support. We revisit OPD from both estimator-theoretic and practical perspectives. Theoretically, we show that token-level OPD is biased with respect to sequence-level reverse-KL minimization, while admitting a substantially tighter worst-case variance bound; a controlled synthetic study further confirms that stronger future-reward coupling increases gradient variance and destabilizes training. Empirically, we identify three concrete failure modes of sampled-token OPD: imbalanced token-level supervision, unreliable teacher guidance on out-of-distribution prefixes, and tokenizer/special-token mismatch; and address each through teacher top-$K$ local support matching, implemented via truncated reverse-KL with top-$p$ rollout sampling and special-token masking. On both single-task mathematical reasoning and multi-task benchmarks spanning agentic and mathematical settings, our objective achieves more stable optimization and a **+14\%** performance gain over standard sampled-token OPD baselines, establishing a principled and practical recipe for on-policy distillation at scale.
Successful Page Load