Income Prediction Using Dimensionality Reduction Analysis and Machine Learning Models
Abstract
Predictive analytics has become an essential component of modern data-driven decision-making across industries. One key application of predictive analytics is income prediction. Machine learning models are developed to classify individuals based on their income level using demographic, educational, and employment-related features. Income prediction helps monitor inequality trends, evaluate labor market dynamics, and understand demographic gaps in earning power. This paper illustrates no feature selection, Uniform Manifold Approximation and Projection (UMAP), Neighbourhood Component Analysis (NCA), and Partial Least Squares Discriminant Analysis (PLSDA) for dimensionality reduction and using the results to train ten models including support vector machine (SVM), logistic regression (LR), K-nearest neighbour (KNN), gaussian naïve bayes, Decision tree (DT), random forest (RF), adaptive boosting (ADA), bagging, stacking, and voting to predict individuals’ income. Additionally, SMOTE and one-hot encoding are used in data preprocessing. Results show that no feature selection has the best performance across all dimensionality reduction techniques, with the Random Forest classifier achieving the best, with an accuracy rate of 89.91% and an ROC-AUC of 96.89%. SHAP analysis confirms that “age” is the dominant predictive feature in all the models. The findings highlight that retaining original features maximizes performance for the dataset, providing a reliable framework for income prediction and guiding dimensionality reduction choices in similar tasks.