BrowseSafe: Understanding and Detecting Prompt Injection Within AI Browser Agents
Abstract
The integration of AI agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified prompt injection as a critical attack vector for web agents, yet why detection fails in real-world environments, which exhibit substantially greater noise, distractor elements, and semantic complexity than the simple, single-line injections on which existing defenses are trained and evaluated, remains poorly understood. We conduct a large-scale empirical study of prompt injection detectability in browser agents, evaluating over 20 open- and closed-weight models on BrowseSafe-Bench, a dataset of 14,719 annotated samples grounded in real-world browser agent usage data. Our study reveals that detection accuracy degrades significantly in the presence of real-world complexity, with systematic gaps across attack visibility, linguistic sophistication, and language diversity. Motivated by these findings, we present BrowseSafe, a multi-layered defense comprising architectural and model-based components that enforces explicit trust boundaries on tool outputs and runs a fine-tuned detection classifier in parallel with agent inference, achieving an F1 score of up to 0.904 with low latency of less than 1 second, demonstrating that the identified gaps are actionable targets for practical defense. BrowseSafe-Bench and the BrowseSafe model will be publicly released to support future research.