ML-Based SMS Messaging Spam Detection: Impacts of Text Feature Extraction Techniques
Abstract
Spam detection on SMS messaging has not received as much attention from researchers recently as the spam detection studies on emails or social media platforms. However, spam SMS messaging can be more intrusive, annoying, and harmful. Thus, detecting and filtering spam SMS messages is becoming a priority that saves human productivity. This work aims to introduce a dynamic, accurate, and efficient machine learning-based spam detection technique for SMS messaging. It aims at protecting users and businesses from spam SMS attacks. It aims to detect and identify suspicious messages that contain promotional, misleading, irrelevant, or harmful content. It primarily aims to test and evaluate the impact of feature extraction methods on the performance of machine-learning-based spam detection. Several text feature extraction techniques have been used and tested, including classical, statistical, contextual, and advanced embedding techniques. An extensive set of experiments has been presented on benchmark datasets in this field. From the comparative study, we can infer that all investigated feature extraction techniques have achieved high accuracy (90%+) on the in-domain dataset. However, their performance decreased when they were tested on the out-of-domain dataset (70%+). The advanced embedding techniques achieved the best performance across both datasets compared to the other tested feature extraction models.