Intelligent Data Management in IoT: A Machine Learning-Based Approach
Abstract
The rapid proliferation of Internet of Things (IoT) devices has resulted in massive, heterogeneous, and high-velocity data streams, making accurate data management and device classification increasingly challenging—especially under class imbalance. This study proposes a lightweight and scalable classification pipeline that integrates the Random Forest (RF) classifier with the Synthetic Minority Over-sampling Technique (SMOTE) to improve minority-class recognition in IoT device identification tasks. Experiments were conducted using the IoT Device Identification dataset (UCI repository) with a reproducible train–test protocol and standard evaluation metrics (accuracy, precision, recall, and F1-score). The proposed SMOTE+RF approach achieved an overall accuracy of 96% and macro-level precision/recall/F1-score of 96%/96%/96%, demonstrating consistent performance across most device classes while improving the classification quality of underrepresented classes compared to the non-oversampled baseline. These results highlight the practical value of balancing strategies for reliable IoT data management and support the deployment of interpretable ML-based classification models in resource-constrained IoT ecosystems.