Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
Abstract
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, especially in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed replay within larger training pipelines---exploration, experience restructuring, or mixed-policy optimization---making the contribution of replay itself difficult to isolate. We ask a more focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two explicit decisions: (1) Headroom ranks stored groups by remaining learning value, (2) Drift gates them by compatibility with the current policy. The fresh on-policy stream is left untouched, and no auxiliary generation or training machinery is added. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, this single intervention outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32. In Agentic Search, where environment interaction dominates cost, it delivers comparable quality at materially lower wall-clock time.