BigCodeArena: A Platform for Executable Code Generation Evaluation
Abstract
Crowdsourced evaluation platforms such as Chatbot Arena enable real-time human assessment of model responses, but evaluating LLM-generated code remains challenging due to long outputs and the need to reason about execution behavior. We introduce BigCodeArena, an open human evaluation platform for code generation with an on-the-fly execution backend. BigCodeArena executes LLM-generated code and allows humans to interact directly with execution processes and outcomes. We collect over 14K code-centric conversations across 10 widely used LLMs, spanning 10 programming languages and 8 execution environments, including more than 4.7K multi-turn samples with pairwise human preferences. Our analysis reveals fine-grained preferences across tasks, languages, and frameworks. With the collected data, we construct two benchmarks: BigCodeReward, which measures alignment between reward models and human preferences, and AutoCodeArena, an automated Elo-style benchmark for code generation. Results show that execution feedback improves preference modeling, and that proprietary models such as GPT-5 and Claude-4 variants remain strong performers.