Reinforcement Routing for Mixtures of LoRAs in Parameter-Efficient LLM Finetuning
Ruizhong Qiu ⋅ Hanqing Zeng ⋅ Yinglong Xia ⋅ Yiwen Meng ⋅ Ren Chen ⋅ Jiarui Feng ⋅ Dongqi Fu ⋅ Qifan Wang ⋅ Jiayi Liu ⋅ Jun Xiao ⋅ Xiangjun Fan ⋅ Benyu Zhang ⋅ Hong Li ⋅ Zhining Liu ⋅ Hyunsik Yoo ⋅ Zhichen Zeng ⋅ Tianxin Wei ⋅ Hanghang Tong
Abstract
Low-rank adapters (LoRAs) are a parameter-efficient finetuning technique that injects trainable low-rank matrices into pretrained models to adapt them to new tasks. Mixture-of-LoRAs models expand neural networks efficiently by routing each layer input to a small subset of specialized LoRAs of the layer. Existing Mixture-of-LoRAs routers assign a learned routing weight to each LoRA to enable end-to-end training of the router. Despite their empirical promise, we discover, both theoretically and empirically, that the routing weights often collapse to only one LoRA even when we activate $k>1$ LoRAs during finetuning. When one LoRA has a dominantly large weight, then the computation of the other $k-1$ LoRAs are essentially **wasted** because using $k>1$ would have similar accuracy to $k=1$. This essentially limits the number of effective LoRAs and thus severely hinders the realized expressive power of existing Mixture-of-LoRAs models. How can we address this critical weakness? In this work, we attribute this weakness to the nature of learnable routing weights and rethink the fundamental design of the router. To address this critical issue, we propose a simple yet effective router design that we call *Reinforcement Routing for Mixtures of LoRAs* (ReMix). Our key idea is using **non-learnable** routing weights to ensure all active LoRAs to be equally effective, with no single LoRA dominating the routing weights. However, such non-learnable routing weights make it infeasible to directly train routers via gradient descent. In response, we further propose an unbiased gradient estimator for the router and employ the reinforce leave-one-out (RLOO) technique to reduce the variance of the estimator. Our gradient estimator also enables to scale up training compute to boost the predictive performance of our ReMix. Extensive experiments demonstrate that our proposed ReMix significantly outperforms state-of-the-art parameter-efficient finetuning methods under a small number of activated parameters.
Successful Page Load