Accelerating Federated LoRA Fine-Tuning for LLMs with Adaptive Ranks at Network Edge
Abstract
The rapid development of large language models (LLMs) has played a crucial role in advancing artificial intelligence. Pretrained LLMs can be adapted to various downstream tasks through fine-tuning. To alleviate the intensive resource requirements of fine-tuning and safeguard data privacy, researchers have integrated Low-Rank Adaptation (LoRA) with Federated Learning (FL), giving rise to a new framework known as Federated LoRA fine-tuning. However, due to heterogeneous computation and communication capacities across devices, federated LoRA fine-tuning still suffers from the straggler problem when using fixed LoRA ranks. In this work, we propose a method named FLFAR, which accelerates federated LoRA fine-tuning with adaptive ranks. Specifically, at the beginning of each global round, the LoRA rank of each device is adapted according to its computational and communication capacity. As a result, devices can complete fine-tuning nearly simultaneously, effectively reducing synchronization delays. Extensive experiments show that, compared with baseline methods, FLFAR reduces federated LoRA fine-tuning time by up to 40% while maintaining comparable accuracy.