Bottom-Up Interpretability of Pretraining Dynamics via Loss Curve Decomposition
Abstract
Average loss curves during language model pretraining obscure diverse learning dynamics across individual training data instances. Gaining a more granular understanding of these dynamics currently requires tracking intermediate checkpoint performance on pre-defined tasks, or computing higher-order gradients with respect to each data instance. To instead enable a scalable bottom-up decomposition of language model training dynamics, we propose identifying data-driven learning domains using non-negative matrix factorization (NMF) to group training data based solely on their instance-wise losses, without any gradient computation. Applied to 3.2k checkpoints across 21 configurations of the Pythia model suite, our method automatically recovers distinctive training dynamics for different training data clusters. The resulting learning domains are consistent across model sizes and random seeds, while being interpretable as, e.g., code, mathematics, repetitive patterns, and non-English texts. The corresponding domain-wise loss curves further reveal potential causes for low-performing training runs that do not surface in the aggregate loss alone.