This paper shows that the proposed method ensures a bounded Byzantine influence on both distillation gradients and individual client private gradients after cross-modality fusion, thereby enabling stable local optimization for honest clients under Byzantine distillation.
Abstract
This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. In the proposed method, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. It then filters unreliable clients, assigns reliability-based weights to the retained clients, and constructs a teacher for each type of knowledge. Finally, the corresponding distillation gradients are validated using a supervised gradient computed from private data. Conflicting prediction and boundary gradients are removed, and conflicting relation gradients are suppressed before the final model update. We prove the convergence of the proposed method by showing stable local optimization for honest clients under Byzantine distillation. Particularly, we show that our method ensures a bounded Byzantine influence on both distillation gradients and individual client private gradients after cross-modality fusion, thereby enabling stable local optimization for honest clienunder Byzantine distillation. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that the proposed method improves the prediction accuracy of heterogeneous models of clients under non-IID data and Byzantine attacks. As the booming demands of federated learning in decentralized environments such as edge computing and mission-oriented UAV collaborations, our method has a great potential for adoption of DFL in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions.
This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation and history-based update trend prediction, rather than purely aggregating client models as in the existing work. In R-DPFL...
Federated Continual Learning (FCL) enables distributed clients to collaboratively learn a sequence of tasks while preserving data privacy and mitigating catastrophic forgetting. However, most existing FCL methods rely on the assumption that all clients share an identical model architecture, which is impractical in real...
Pei-Yi Zeng, Shu-Ming Yang, Jia-Wei Liao et al.· 2026 12th International Conf...· 0 citations
This work proposes Class-wise Reliability-Aware Distillation (CRAD), which, per class, first discards teachers that disagree with the peer consensus and then takes a weighted average of the rest, weighting each teacher by its per-class reliability (precision, or inverse variance).
Baraa Bilbeisi, Meng-Chen Fan, Bao-Cheng Geng et al.· 1 citation
Due to distributed data and privacy concerns, federated learning(FL) is a promising approach to learn global models from distributed data, with personalized federated learning (PFL) being a key enabler for customized services in future 6G networks. Federated Distillation (FD) is a classic communication-efficient PFL pa...
This study delves into large PFMs adaptation in the resource-constrained federated learning environment, and proposes an innovative framework, namely ADRAP, which alternates between large model distillation and resource-adaptive pruning, with guaranteed convergence.
Xiao Zhang, Yang-Yang Wang, Xing-Yu Sun et al.· IEEE Transactions on Pattern...· 0 citations
Federated Reinforcement Learning (FedRL) improves sample efficiency while preserving privacy; however, most existing studies assume homogeneous agents or utilize public datasets for knowledge distillation to address agent heterogeneity, which limits its applicability in real-world heterogeneous scenarios. Knowledge dis...
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
Microsoft Research Blog· microsoft.comAug 20, 2026
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.