CreativityBench: Evaluating Creative Reasoning via Affordance-Based Tool Repurposing
Abstract
Recent advances in large language models (LLMs) have enabled strong performance in reasoning and environment-interaction, yet their ability for creative problem solving remains underexplored. One core form of such creativity is tool use: repurposing objects by reasoning about their affordances and attributes rather than relying on canonical usage. To study this capability, we introduce CreativityBench, a benchmark for affordance-based creative tool use. It is built on a large-scale affordance knowledge base containing 4K entities and 150K+ annotations linking objects, parts, attributes, and uses, from which we generate 14K grounded tasks requiring non-obvious but physically plausible solutions under constraints. Evaluation across 10 closed and open-source LLMs reveal a consistent gap: while models often select plausible objects, they fail to identify the correct parts, affordances, and underlying physical mechanism. Moreover, general reasoning ability does not reliably transfer to creative affordance discovery, gains from scaling quickly saturate, and Chain-of-Thought yields limited benefit. These findings highlight creative tool use as a missing dimension in current LLMs capabilities, and position CreativityBench as a systematic testbed for advancing creative intelligence in future LLMs and embodied agents.