Performance of Classical Machine Learning Algorithms in Three-Class Sentiment Classification of Indonesian DANA E-Wallet Reviews
Abstract
Automated sentiment classification of e-wallet reviews can support service monitoring, yet reported performance may be inflated by duplicate texts, conflicting labels, and test-set-driven model selection. This study compares Multinomial Naive Bayes (MNB), linear Support Vector Machine (SVM), and Logistic Regression (LR) for three-class sentiment classification of Indonesian DANA e-wallet reviews under a leakage-controlled design. From 50,000 public Kaggle reviews, empty normalized texts, conflicting-label groups, and duplicates were excluded after text normalization, leaving 29,658 unique reviews. The data were divided into an 80% development set and a 20% held-out test set. TF-IDF features were fitted within five-fold stratified cross-validation, and hyperparameters were selected by macro F1. SVM achieved the highest cross-validation macro F1 (0.767±0.008), held-out accuracy (0.801), and held-out macro F1 (0.770), followed by LR (0.790; 0.767) and MNB (0.766; 0.738). SVM outperformed LR in paired accuracy (McNemar p=0.00022), although the 95% bootstrap interval for their macro-F1 difference included zero. NEUTRAL remained the most difficult class (F1=0.617), and balanced class weighting raised its recall from 0.453 to 0.596 with a slight accuracy decrease. These findings show that normalization-aware auditing, leakage-controlled model selection, and class-wise evaluation materially affect conclusions drawn from imbalanced review data.