Ill-Defined Math: Benchmarking LLM Reasoning Beyond Well-Defined Problems
Abstract
Ill-defined problems are ubiquitous in science and engineering, where solutions may fail to exist, be non-unique, or lack meaningful interpretation. Despite recent progress of large language models (LLMs) on math reasoning, their ability to reason about ill-defined problems remains poorly understood. A key obstacle is the lack of appropriate benchmarks and evaluation protocols. Existing benchmarks for unsolvable problems focus on shallow commonsense issues and rely on inflexible and simplistic evaluation schemes, primarily measuring refusal behavior rather than reasoning capability. To address this gap, we introduce Ill-Defined Math (IDM), a benchmark of 1,300 ill-defined mathematical problems whose ill-definedness arises from deep, fundamental flaws in their mathematical capabilities. IDM includes a test set of 300 manually constructed problems that are rigorously verified by PhD-level experts, and a larger training set constructed via an automatic generation pipeline, guided by few-shot examples from the test set. To address the lack of versatile evaluation protocols, we further propose a fine-grained, three-stage LLM judge framework that evaluates responses from three complementary perspectives: final answer statements, issue recognition, and issue fixing. Evaluations on recent LLMs reveal substantial performance degradation compared to well-posed counterparts, and a clear gap between recognizing issues and handling them correctly: even when models recognize ill-definedness, they often continue to produce definitive answers without fixing the underlying issues, suggesting a tendency to satisfy user expectations rather than respond accurately. These findings highlight the roles of both reasoning limitations and sycophancy behavior for LLM reasoning under ill-definedness.