Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
Abstract
We introduce the Visual Aesthetic Benchmark (VAB), a new benchmark designed to evaluate whether multimodal large language models (MLLMs) can make expert-aligned aesthetic judgments. Existing image aesthetics datasets reduce this problem to predicting a scalar score for individual images. Through a controlled human study with expert annotators, we show that score-derived rankings do not faithfully capture experts' preference orders, yielding 42 percentage points lower inter-annotator agreement than direct comparative ranking. Instead of scoring, VAB formulates aesthetic evaluation as comparative selection over candidate sets with matched subject matter, isolating differences in execution quality rather than content. VAB comprises 400 tasks and 1,195 images spanning three visual domains and 24 topics, with ground-truth labels derived from the consensus of 10 independent expert judges. Each task is evaluated under three random permutations of candidate order. We evaluate 20 frontier MLLMs and 6 visual quality reward models. The best model achieves only 26.5\% on the strictest metric, far below the 68.9\% of human experts, with severe sensitivity to candidate ordering and sharp degradation as set size grows. We further show that fine-tuning a 35B-parameter model on 2,000 expert-annotated examples enables it to approach the performance of a model over 10× larger, highlighting the transfer value of expert comparative supervision. VAB establishes a reliable benchmark for aesthetic evaluation and provides a concrete foundation on expert-aligned multimodal aesthetic judgment.