Uncovering the computational ingredients that support human-like conceptual representations in large language models
Abstract
The ability to translate diverse input patterns into structured behavior has been thought to rest on learning robust representations of concepts. The rapid advancement of transformer-based large language models (LLMs) has produced a diversity of computational ingredients — architectures, fine-tuning methods, and training datasets among others — yet it remains unclear which are most crucial for developing human-like representations. Further, most current benchmarks are ill-suited to measuring \textit{representational alignment}, making their scores unreliable for assessing whether LLMs are progressing as cognitive models. We address these limitations by evaluating over 77 models on a triplet similarity task — a method well established in cognitive science for measuring conceptual representations — using concepts from the THINGS database. We find that instruction fine-tuning and larger attention head dimensionality are among the strongest predictors of human alignment, while activation function choice, multimodal pretraining, and parameter size have limited bearing. Correlations between alignment scores and existing benchmark scores reveal that while some benchmarks (e.g., BigBenchHard) better capture representational alignment than others (e.g., MUSR), none fully accounts for the variance in human-model alignment, demonstrating their insufficiency for measuring human-AI alignment. Taken together, our findings highlight the essential computational ingredients for advancing LLMs as models of human conceptual representation and address a key gap in LLM evaluation.