How Conversational Pressure Alters Belief in Large Language Models: A Hidden State Boundary Perspective
Abstract
Large language models (LLMs) often give the correct answer at the start of a factual dialogue, but may give up that answer under continued conversational pressure. We do not view this simply as a prediction error. Instead, we study it as a process in which the dialogue context gradually weakens the model's internal belief. By tracking the full hidden state trajectory of LLMs across multi-turn dialogues under pressure, we find that this change is not just a brief fluctuation at the output layer. Even before the answer changes in later turns, the hidden states already show a gradual movement toward a learnable hidden state boundary. Multi-view CNN analysis further shows that the signal related to this boundary is concentrated in a sparse internal subspace within a small number of key dimensions and in the middle to late layers. Boundary-aligned intervention in hidden state space then shows that perturbations along the boundary normal are most effective at strengthening or weakening the model’s internal belief in its answer, while tangent perturbations are much less effective. This pattern is consistent across different model families. Overall, we show that continued conversational pressure gradually weakens an LLM's internal belief, and that this process appears in hidden state space as a gradual shift toward a boundary, eventually resulting in a change in the model’s answer.