ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Abstract
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a button set) that are short enough to diagnose and rich enough to capture the burdens of modern interfaces. We present ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs. ComponentBench is organized around a library-agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified tasks across widely used component libraries, paired with cleaned human reference trajectories that enable evaluation of both task success and interaction efficiency. Evaluated across seven models, including GPT-5.4, Gemini 3 Flash, GPT-5.4 mini, GPT-5 mini, Gemini 3.1 Flash-Lite, Qwen3-VL-235B, and UI-TARS-1.5-7B, and four observation and action spaces (e.g., GUI-based and coordinate-based representations), ComponentBench shows that these design choices critically impact performance. Varying the observation or action space can shift task success rates by over 30% within a single model, causing GPT-5 mini's performance to degrade from 87.0% with Browser-Use to 49.0% with coordinate-only pixel. Moreover, even the fastest agents remain 3.7× slower than human references, and components involving spatial manipulation that is trivial for humans continue to challenge current agents.