Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
Abstract
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. On ToM-SB scenarios with a shared attacker-defender universe, we find that a strong frontier model like Gemini3-Pro struggles with ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge even when prompted to reason about the attacker’s beliefs. To close this gap, we train models on ToM-SB using reinforcement learning, testing both fooling and ToM rewards. Notably, we first find a bidirectionally emergent relationship between ToM and attacker-fooling: rewarding fooling success alone improves ToM, and rewarding ToM alone improves fooling. Across four attackers with different strengths, six defender methods, and both in-distribution and OOD evaluation, we find that gains in ToM and attacker-fooling are well-correlated, highlighting belief modeling as a key driver of success on ToM-SB. Combining both ToM and fooling rewards into unified AI Double Agents yields the strongest fooling and ToM performance, outperforming Gemini3-Pro with ToM prompting on hard scenarios. We also show that ToM-SB and AI Double Agents can be extended to stronger attackers, demonstrating generalization to OOD settings and the upgradability of our task.