Scalable Data Engineering Architectures for Federated Learning in Decentralized Cloud Environments
Abstract
The growing adoption of Federated Learning (FL) is reshaping the way machine learning models are trained across distributed, privacy-sensitive datasets. However, the scalable and efficient orchestration of data engineering pipelines in decentralized cloud environments remains a significant challenge. This paper presents a comprehensive architectural framework for scalable data engineering tailored for FL in heterogeneous and resource-constrained environments. By integrating modern distributed computing paradigms, such as Kubernetes-based orchestration, edge-aware data preprocessing, and secure federated communication, we propose a modular architecture that addresses data heterogeneity, scalability, and compliance. A case study in a healthcare IoT scenario validates the performance and flexibility of the proposed system. Our work serves as a blueprint for deploying robust FL systems in real-world decentralized cloud ecosystems.