Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond
Abstract
The reasoning paradigm, where language models reason before answering, has enabled breakthroughs on tasks such as mathematical problem-solving. While current tooling for reasoning is built around next-token prediction trained models, recent works introduce an alternative choice: masked diffusion language models (MDLMs). MDLMs are trained to in-fill positions in randomly masked sequences. We find that the in-filling capacity of MDLMs provides new ways to prompt, sample and post-train for reasoning. First, we propose reasoning-as-infilling, a prompting technique where tokens are pre-filled to explicitly delimit reasoning and answer regions. Reasoning-as-infilling with MDLMs enables new possibilities for post-training. Given question-answer pairs, we can generate high-quality reasoning traces conditioned on answers. We show that on GSM8k, fine-tuning LLaDA- 8B-Base on posterior reasoning traces provides a similar improvement as human-written traces. Additionally, given an answer, we find that MDLMs can score their reasoning process at intermediate steps, providing intermediate rewards more correlated with correctness than a specialized model. At inference-time, reasoning-as-infilling enables measuring answer uncertainty at intermediate steps and exiting when the model is certain, providing a 4× acceleration when combined with parallel decoding. Our results show that the MDLM training objective provides promising benefits for reasoning tasks.