Understanding Primacy Effects in Large Language Models with Sparse Autoencoders
Abstract
The primacy effect is a well-established cognitive phenomenon in which information presented early in a sequence is often weighted more strongly than information presented later. Recent work suggests that large language models (LLMs) exhibit analogous behavior across a range of settings, yet the representation-level mechanisms underlying this bias remain underexplored. In this work, we provide a representation-level account of primacy effects in LLMs using sparse autoencoders (SAEs). We propose the Dedicated Primacy Subspace (DPS) hypothesis: LLMs appear to preferentially recruit a particular representational subspace to encode the semantics of the first demonstration, which subsequently modulates the processing of later context. Using few-shot classification as our testbed, we show that this mechanism has two complementary forms: DPS-Affirmative features, which encode support for the first-demonstration label, and DPS-Contrastive features, which activate when subsequent inputs contradict that initial label, thereby emphasizing competing alternatives. We further show that a subset of these features exhibits primacy-sensitive causal effects under intervention, providing direct support for the proposed mechanism. Taken together, these results provide a concrete representation-level account of how primacy effects manifest in LLMs.