Privacy-Preserving Feature Engineering for Federated Learning Analytics
Federated learning (FL) offers a decentralized approach to machine learning that preserves data privacy by training models locally across distributed devices. However, the feature engineering process—an essential step in improving model performance—often remains centralized or privacy-invasive, risking sensitive data exposure. This paper proposes a comprehensive framework for privacy-preserving feature engineering (PPFE) within federated learning analytics. We explore techniques such as homomorphic encryption, differential privacy, and secure multi-party computation to enable robust, privacy-safe feature selection, transformation, and extraction across clients. Our framework includes both vertical and horizontal FL settings and evaluates the trade-offs between privacy, utility, and communication overhead. Experimental results on real-world datasets demonstrate that our PPFE methods can significantly improve model performance without compromising data privacy. This work contributes towards building a more secure and efficient FL pipeline that ensures end-to-end data confidentiality.