FedSubMuon is proposed, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces that keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices.
Abstract
Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces. This design keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices. We further introduce FedSubMuon-GT, an accuracy-oriented extension that uses projected gradients to adapt tracked subspace bases toward task-relevant gradient directions. Experiments on instruction tuning and mathematical reasoning show that FedSubMuon-GT achieves the best overall accuracy on four of five dataset-model pairs, while FedSubMuon performs best under all matched communication budgets. On Dolly-15K, the closest communication baseline requires 5.5 times and 1.4 times more total communication on Llama-1B and Qwen-4B, respectively.
FraQ, an efficient coordinate-space recompression method for federated LoRA, is proposed, an efficient coordinate-space recompression method for federated LoRA that achieves accuracy close to uncompressed baselines while substantially reducing downlink communication with low server-side recompression overhead.
L-shaped SFT is presented, a split fine-tuning framework that removes the need for continuous client participation and introduces one-shot SFT, in which clients upload activations once and then go offline while the server continues optimization over cached representations.
The rapid development of large language models (LLMs) has played a crucial role in advancing artificial intelligence. Pretrained LLMs can be adapted to various downstream tasks through fine-tuning. To alleviate the intensive resource requirements of fine-tuning and safeguard data privacy, researchers have integrated Lo...
Federated fine-tuning of large language models with low-rank adaptation reduces per-client trainable parameters, but client-to-server communication remains the dominant cost. Existing accounting for federated LoRA protocols omits the asymmetric transition round when a protocol changes aggregation mode, and reports savi...
FedGSA, a geometry-consistent aggregation framework for differentially private federated LoRA, is proposed and it is proved that FedGSA incurs no additional privacy loss beyond client-side DP training and establishes its convergence under standard assumptions.
Le-Le Zheng, Rui Hu, Tao Zhang et al.· 0 citations
Large Language Models (LLMs) have achieved remarkable success in NLP tasks, but fine-tuning them on resource-constrained mobile devices remains challenging due to prohibitive memory and computation requirements. Federated Learning (FL) enables privacy-preserving distributed fine-tuning, yet conventional approaches, inc...
He Sun, Jinrui Zhou, Li Li et al.· Proceedings of the 32nd ACM...· 0 citations