iOSWorld: A Realistic iOS Environment for Benchmarking Phone Agents with Personalized User Identity and Memory
Abstract
Existing benchmarks for phone agents evaluate them in impersonal environments that ignore a basic fact about real devices: each one reflects its owner's identity, history, and preferences. We introduce iOSWorld, the first dynamic iPhone simulator benchmark built around a persistent user identity that spans 26 purpose-built iOS apps. These apps contain interconnected data including transaction histories, messaging threads, travel records, social relationships, and financial activity. iOSWorld includes 128 tasks across three categories of increasing difficulty: single-app tasks (26), multi-app tasks (58), and memory and personalization tasks (44), which require agents to infer implicit user preferences, routines, and constraints from distributed evidence across apps. We evaluate five frontier computer-use models under both vision-only and structured XML input settings. The best model achieves 39% overall accuracy, with sharp drops on multi-app and personalization tasks: current agents struggle to maintain context and reason across applications over a user's digital footprint. We release iOSWorld as an open-source benchmark, including all apps, seed data, tasks, rubrics, and evaluation code, to support reproducible research on personalized phone agents.