LF²AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models
Abstract
Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging. Perceptual modalities are finer-grained and less semantically dense than text, making it unclear how functions learned during text pretraining can be reused. We study this problem through the lens of a layerwise abstraction-refinement dynamic observed in transformer LMs: representations first become more abstract and compositional, then are refined into representations predictive of fine-grained structure. This perspective suggests that, to adapt an LM to finer-grained modalities requires: (i) allocating additional fine-to-coarse processing at the input and coarse-to-fine processing at the output, consistent with late modality fusion and an output-side analogue we term late fission; and (ii) allowing the output predictor to preserve input-dependent selective access to both high-level semantic structure and low-level perceptual detail, motivating our use of attention residuals in fission. Across models ranging from 135M to 2B parameters and adapted to text-as-images and speech, these components increase feature abstraction, strengthen cross-modal alignment, enable such alignment to emerge at smaller compute budgets, improve preservation of text-like predictive structure, and yield better performance on image and speech versions of language understanding and reasoning benchmarks. Attention residuals also induce sparse, interpretable use of deep backbone layers, enabling early-exit decoding with a 1.9× generation speedup. Together, our results suggest that multimodal adaptation improves when it respects the layerwise organization learned during text pretraining.