Reasoning with Image Generation
Abstract
Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to generate intermediate reasoning steps. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content. We propose ReImaGin, which leverages image generation models as a flexible visual reasoning mechanism in multimodal LLMs: unlike fixed-function tools, they accept natural language commands and can perform open-ended visual transformations, like removing an occlusion or annotating a trajectory. Across six visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25.0 points, demonstrating the advantage of flexible, generative visual reasoning.