SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
Abstract
Group-based reinforcement learning objectives such as GRPO allocate learning signal poorly across prompt difficulty: group normalization induces a divergent weighting on easy prompts, wasting gradient budget on well-solved examples. We introduce Softmax Advantage Group Estimation (SAGE), a drop-in alternative that replaces z-score-normalized group advantages with temperature-scaled softmax advantages, keeping weights bounded regardless of prompt difficulty. We prove that SAGE exactly optimizes a population-level objective under binary rewards; for general scalar rewards, the same softmax weighting recovers a centered reward-augmented maximum likelihood (RAML) target, which we implement in practice via PPO-style trust-region optimization. Empirically, SAGE achieves strong results on verifiable tasks, improving from 30.0\% to 51.8\% on DeepMath, and is especially effective when only weak rewards are available: it improves a 1.5B instruction-tuned model from 35.0\% to 68.0\% on Poetry using only lightweight text-similarity rewards, outperforming supervised fine-tuning and several preference and distillation baselines across both settings.