Skip to content
Open access

Adaptive Dual-Mode Distillation for Robust and Communication-Efficient Federated Learning Under Statistical and Model Heterogeneity

Sep 2025 · IEEE Access · Vol 14, pp. 145907-145925 · 0 citations · 84 references
Computer Science

Abstract

The growing volume of data from smart devices offers significant potential for machine learning, yet privacy concerns hinder centralized use. Federated Learning (FL) has emerged as a promising decentralized learning (DL) approach enabling the use of distributed data without compromising privacy. However, practical deployments face three major challenges: 1) strong assumptions of homogeneous model architectures across clients despite diverse computational capabilities and application requirements, 2) severe performance degradation under statistically heterogeneous (non-IID) data distributions, and 3) lack of effective, communication-efficient incentive mechanisms for client participation. We propose a unified confidence-aware distillation framework that leverages unlabeled public data to support robust knowledge aggregation under statistical and model heterogeneity. The framework operates in three complementary modes: DL-SH, which mitigates statistical heterogeneity through confidence-weighted distillation; DL-MH, which extends the same framework to fully heterogeneous client models and label spaces via schema-based mapping and masking; and I-DL-MH, an incentive mechanism that enables clients to distill updated global knowledge with negligible communication overhead. Across all modes, client contributions are adaptively weighted using confidence estimates, enabling aggregation without requiring model homogeneity in a single distillation round. We evaluate the proposed methods across multiple model architectures, benchmark datasets, and both IID and non-IID data distributions. Results demonstrate that our methods consistently outperform representative baselines while substantially reducing communication costs and improving resource efficiency. In particular, DL-SH improves global accuracy by up to 153% over the standard FL baseline under extreme non-IID settings, while I-DL-MH yields up to a 225% client-level improvement after a single round of post-distillation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.