A Spectral Transport Mechanism Underlying the Training Dynamics of Large Language Models
Yizhou Zhang ⋅ Weichen Wu ⋅ Lun Du ⋅ Zhengjie Miao
Abstract
Large language models exhibit remarkably regular loss trajectories during training, often following approximate power-law decay over long optimization intervals. The dynamical origin of this phenomenon remains poorly understood. We derive an exact transport--dissipation equation governing the spectral distribution of prediction error under gradient flow, with the irreducible entropy cleanly separated via a KL-divergence-based error observable. The equation is exact; what requires empirical input is the behavior of the spectral flux $\mathcal{F}$ during training. Analyzing public LLM checkpoints via stochastic Lanczos probes, we find that $\mathcal{F}$ is weak but nonzero and directional throughout training. Because $\mathcal{F}$ remains weak, perturbation theory yields power-law loss decay as the leading-order solution over any finite window. The first-order correction predicts a systematic, monotonically growing undershoot relative to any early power-law extrapolation---a deviation small in absolute terms, but precisely what the framework directs attention to. We confirm this prediction directly on LLMs with various model size: at every held-out checkpoint set, the observed loss falls below the extrapolated power law, with deviation onset governed by the data budget rather than model capacity.
Successful Page Load