Identifying Introspection From the Inside
Abstract
Language models increasingly make claims about themselves that are both consequential and difficult to verify from behavior alone. How, then, can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to adopt fictitious personas that make decisions according to latent linear preference functions. Sustained fine-tuning on the decision task alone can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First, we ask: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? To study this, we conduct weight ablations and frozen-layer experiments, which together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Our second research question asks: can the structural differences we have observed be used to distinguish faithful models from unfaithful ones? Using attribution patching scores on weights, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks. Cross-task causal patching confirms this: in faithful models, patching a small fraction of adapter weights into the base model recaptures substantially more of the original model's behavior---a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.