Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
Abstract
GUI grounding, the task of mapping natural-language instructions to pixel coordinates, is crucial for autonomous computer-use agents (CUA) yet remains difficult for current VLMs. The core bottleneck is reliable patch-to-pixel mapping: current approaches generate coordinates as text tokens directly from visual features, forcing the model to infer complex position-to-pixel mappings implicitly. This implicit regression is unstable and breaks when extrapolating to high-resolution displays unseen during training. We address this with RULER tokens, explicit coordinate markers that share positional embeddings with image patches, letting the model reference known positions and adjust within a bounded range rather than regress coordinates from scratch. We also propose Interleaved MRoPE (I-MRoPE), which corrects a frequency imbalance in standard positional encodings so that width and height dimensions receive equal representational capacity. Experiments on ScreenSpot, ScreenSpot-V2, and ScreenSpot-Pro show consistent grounding improvements, with the largest gains on high-resolution displays unseen during training (+2.6% on ScreenSpot-Pro), achieved with less than 1% additional tokens.