In-Context Learning as Implicit Policy Gradient
Abstract
Recent work has shown that Large Language Models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. However, the theoretical underpinnings of this phenomenon remain unclear. In this paper, we establish a formal connection between score-conditioned In-Context Learning (ICL) and policy gradient methods. We first provide a constructive proof showing that self-attention mechanisms can implement reward-weighted aggregation, which is structurally equivalent to the REINFORCE algorithm. Furthermore, we show that the parameter-fixed nature of ICL implicitly enforces a bounded distribution shift from the reference model, analogous to KL-regularized policy optimization. We validate our theory through extensive experiments across multiple LLMs, demonstrating that LLMs indeed use score information to shift output distributions toward high-scoring examples, and attention weights correlate strongly with example scores.