Scalable Data Engineering Architectures for Federated Learning in Decentralized Cloud Environments
The growing adoption of Federated Learning (FL) is reshaping the way machine learning models are trained across distributed, privacy-sensitive datasets. However, the scalable and efficient orchestration of data engineering pipelines in decentralized cloud environments remains a significant challenge. This paper presents a comprehensive architectural framework for scalable data engineering tailored for FL in heterogeneous and resource-constrained environments. By integrating modern distributed computing paradigms, such as Kubernetes-based orchestration, edge-aware data preprocessing, and secure federated communication, we propose a modular architecture that addresses data heterogeneity, scalability, and compliance. A case study in a healthcare IoT scenario validates the performance and flexibility of the proposed system. Our work serves as a blueprint for deploying robust FL systems in real-world decentralized cloud ecosystems.