Attractor States Emerge in Multi-Turn LLM Conversations
Abstract
Large language models (LLMs) are increasingly deployed in multi-agent pipelines, yet the long-term dynamics of autonomous model-to-model interactions remain poorly understood. We investigate whether open-ended LLM discussions exhibit attractor-like behavior: a tendency to converge toward stable regions of behavior space regardless of initial conditions. We analyze dyadic debates between agents initialized with opposing stances on controversial topics, comparing self-play and mixed-play across different model families. Through per-turn stance tracking and PCA trajectory analysis of embedding spaces, we find that interaction dynamics stabilize rapidly and consistently. Stances typically neutralize within five turns, while token entropy decreases monotonically as conversations settle into a progressively narrower vocabulary. Critically, we identify asymmetric influence in mixed-play scenarios, where non-Claude agents gravitate toward the stylistic and semantic markers of Claude-based models. Our results demonstrate that multi-agent discussions do not wander freely; instead, they settle into model-conditioned local attractors shaped more by model identity than by topical content. These findings suggest that deployed multi-agent systems may converge toward predictable but biased conversational regimes, posing new challenges for safety and behavioral diversity.