How social context shapes value-related content in Large Language Model outputs
Abstract
We study whether Large Language Models (LLMs) can reliably express and recognize value-conditioned behavior. To this end, we introduce the Generator-Inquisitor framework, in which one model generates text conditioned on a target value and another infers the underlying value from that text. Using Schwartz's four value dimensions—Openness to Change, Conservation, Self-Enhancement, and Self-Transcendence—instantiated across Fiske's four relational domains—Communal Sharing, Equality Matching, Authority Ranking, and Market Pricing—we evaluate six state-of-the-art LLMs across five experiments. Value recognition accuracy is generally high but varies systematically across value dimensions and relational domains, with Self-Enhancement showing the hardest misalignment, especially in the Communal Sharing context. Sentiment analyses further revealed higher positivity in generated texts for Self-Enhancement-vs-other values. Cross-model evaluation showed negligible decrease in accuracy when Generator and Inquisitor are from two different model families, and a perturbation experiment confirmed that recognition relies on value-relevant content rather than superficial lexical cues. Together, these results show that value representations in LLMs are not abstract model attributes but emerge through relational contexts. Like for humans, this observation calls for contextualized approach to the study of value alignment in machines.