A Comedy of Estimators: On KL Regularization in RL Training of LLMs
Abstract
The reverse Kullback-Leibler (KL) divergence between the trained policy and a reference policy is commonly used as a regularizer when training large language models (LLMs) with reinforcement learning (RL). Since computing the exact sequence-level KL divergence is intractable, practical algorithms use sample-based estimators computed from on-policy rollouts, and incorporate them in the reward or the loss. Despite its ubiquity, the interaction between estimator choice, position of the regularization and downstream performance is poorly understood. Recent work also identifies that certain implementation choices can result in biased gradients. We further analyze these practices and examine the gradients of several estimator configurations, showing how the choices shape gradient bias. We substantiate these observations with empirical evidence by RL fine-tuning LLMs with different KL configurations, and evaluating their performance on both in- and out-of-distribution tasks. In on-policy settings, we find that biased-gradient estimator configurations can cause training instabilities, while unbiased-gradient configurations lead to better performance on in-domain as well as out-of-domain tasks. We also observe that KL regularization can help stabilize asynchronous training. Overall, our findings provide useful takeaways for using KL-regularized objectives during RL post-training of LLMs.