Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
Abstract
Reinforcement learning for large language model reasoning often suffers from signal collapse: under uniform small-group sampling, difficult prompts frequently produce uninformative all-fail or all-success groups, causing useful gradients to disappear. We show that this failure is largely a statistical artifact of undersampling rather than an inherent limitation of the model. This suggests that sampling budget should be allocated adaptively across prompts according to their pass rates. We formalize this idea through a family of non-linear RL objectives, which induce difficulty-aware weighting and recover prior log-objective analyses as a special case. Based on this view, we propose Reinforce-Ada, with an estimation-based variant and a model-free sequential variant that adaptively allocate rollouts to recover informative signals. Across multiple models and math reasoning benchmarks, Reinforce-Ada consistently improves over standard GRPO baselines, recovering much of the benefit of large-group training at substantially lower cost.