Misalignment Contagion: Can a Misaligned Minority Shift Aligned Agents in Multi-Agent LLM Deliberation?
Abstract
Multi-agent deliberation is often proposed as a way to improve the safety and reliability of AI systems. We show that it can also create a channel through which misalignment spreads. In particular, a single emergently misaligned agent, obtained through narrow fine-tuning without adversarial intent, can shift the privately elicited post-deliberation stances of an otherwise aligned majority through ordinary structured debate. We refer to this phenomenon as misalignment contagion. To study it, we use a three-stage protocol designed to distinguish public conformity during discussion from persistent post-discussion change under private re-elicitation. Across 27.9k trials spanning five normative and safety-relevant datasets and four communication topologies, we find that the effect is not well described as surface compliance alone. In many settings, shifted stances persist when agents are queried again in a fresh non-social context, with the strongest persistence arising in chain topologies. We also find that, within our tested settings, increasing the size of the misaligned minority does not increase contagion, while placing a misaligned agent in a structurally central position increases the depth of persistence. More generally, topology shapes not only how much agents move during deliberation, but whether that movement survives private re-elicitation. Finally, prompting aligned agents to be more rigid does not reduce susceptibility and can instead increase it. These results show that component-level alignment does not by itself guarantee aligned behavior at the level of the interacting system.