Skip to content
Open access

Machine Learning-Based Predictive Analytics: Investigating the Influence of Feature Selection and Model Optimization on Prediction Accuracy

Sep 2026 · Journal of Global Social Transformation · Vol 2, pp. 313-326 · 0 citations · 31 references

TL;DR

Overall, the findings demonstrate that integrating feature selection with model optimization substantially enhances classification performance, with the optimized Random Forest emerging as the most accurate and balanced model for the evaluated dataset.

Abstract

This study comparatively evaluated the predictive performance of machine learning classification models and examined whether feature selection and model optimization improved classification accuracy and overall predictive effectiveness. A quantitative computational research design was employed using a structured dataset containing the predictor variables required for classification, which was preprocessed and divided into training and testing subsets before model development. Four machine learning algorithms, namely Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine (SVM), were trained under baseline, feature-selection, model-optimization, and combined feature-selection-plus-optimization configurations. Feature selection was performed using SelectKBest and Recursive Feature Elimination (RFE), while model optimization was applied to improve the predictive configurations of the algorithms. Model performance was evaluated using accuracy, precision, recall, and F1-score, allowing comparative assessment of classification effectiveness. The results showed that feature selection consistently improved predictive accuracy, with RFE producing stronger improvements than SelectKBest. Model optimization also enhanced performance, with accuracy improvements of 5.50% for Logistic Regression, 6.20% for Decision Tree, 6.10% for Random Forest, and 5.40% for SVM compared with their respective baseline models. When feature selection and optimization were combined, mean accuracy increased from 83.25% in the baseline configuration to 93.00%, demonstrating a substantial improvement in overall predictive performance. Under the combined approach, Logistic Regression achieved 92.30% accuracy, Decision Tree achieved 90.40%, SVM achieved 94.10%, and Random Forest achieved the highest accuracy of 95.20%, with precision, recall, and F1-score all reaching 0.95. Overall, the findings demonstrate that integrating feature selection with model optimization substantially enhances classification performance, with the optimized Random Forest emerging as the most accurate and balanced model for the evaluated dataset.

Read PDF

Similar papers

Open access Sep 2026

Intelligent Predictive Analytics using Machine Learning: A Comparative Evaluation of Classification Algorithms for High-Dimensional Data

This study comparatively evaluated the predictive performance of selected machine-learning classification algorithms for high-dimensional data using a quantitative computational design-and-evaluation methodology. The analysis involved data preprocessing, feature processing, model development, hyperparameter optimizatio...

Muhammad Awais, Muhammad Haad, Hameed Hussain et al. · 0 citations
Open access Sep 2026

Advanced Machine Learning Approaches for Predictive Analytics: Comparative Evaluation of Classification Performance in High-Dimensional Data

This study examined advanced machine learning approaches for predictive analytics by comparatively evaluating the classification performance of Support Vector Machine (SVM), Random Forest, and Gradient Boosting in high-dimensional data. A high-dimensional dataset containing multiple observations and a large number of f...

Imtiaz Ali, M. Ali, Shahrukh Nawaz · 0 citations
Review Open access Aug 2026

Hybrid Machine Learning Models for Predictive Analytics and Intelligent Decision-Making in Complex Data Environments

Logistic Regression, Random Forest, soft voting, and weighted voting using a publicly available educational dataset comprising 4,424 student records, 36 predictors, and three outcome classes provided accurate, interpretable, and decision-oriented predictions, although external validation and prospective evaluation are...

Haleeful Jud · 0 citations
2026

A Study Report on Comparative Performance Evaluation of Machine Learning Algorithms for Structured Data Classification

This study presents a comprehensive comparative evaluation of three widely used machine learning algorithms: Logistic Regression, Decision Tree, and Random Forest, for classification tasks, and demonstrates that Random Forest achieves superior generalization performance compared to the other models.

Badhur Ammulya, G. Nagalakshmi · 0 citations
Open access Sep 2026

A Comparative Study of Feature Selection Techniques for Customer Churn Prediction

This paper investigates the effectiveness of feature selection techniques in optimizing supervised machine learning pipelines for customer churn prediction using the publicly available Customer Churn Dataset from Kaggle. Feature selection plays a crucial role in enhancing model interpretability and generalization by...

M. C. Opara · 0 citations
Open access Sep 2026

Student Performance Prediction Using Machine Learning: A Comparative Analysis of Decision Tree, Random Forest, and Optimized XGBoost

Student performance prediction has become an important application of Machine Learning in educational data mining, enabling institutions to identify academically weak students at an early stage and provide timely academic support. Accurate prediction of student performance helps educators implement personalized learnin...

K. S. Sangeetha, Tulasi Miryala · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.