ExpRL: Exploratory RL for LLM Mid-Training
Abstract
Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success depends critically on the coverage present in the base model. In practice, models are often primed for RL through mid-training on curated reasoning traces that teach useful primitive skills such as decomposition, verification, or self-correction. Although effective, this strategy requires manually specifying what the model should learn, and it remains unclear whether such primitive coverage is enough for much harder problems, which require combining these skills into broader solution techniques. We study a more automated approach to mid-training for diverse problem-solving techniques using large corpora of human-written question-answer data. Rather than treating reference solutions as targets to imitate, we use them to define dense rewards for on-policy RL. Our method, ExpRL, uses an LLM judge to compare sampled reasoning traces against reference solutions and assign process-level and outcome-level rewards. This enables RL-based mid-training to reinforce partial progress and useful intermediate behaviors across a broad range of problems, building coverage over techniques that sparse rewards alone often fail to discover. On math reasoning tasks, ExpRL yields better RL priming and a stronger initialization for subsequent sparse reward RL. The resulting policies also transfer better, improving downstream performance on both the original task distribution and a proof domain.