Improving Sentiment Classification Performance Using Pseudo-Labeling with Naive Bayes and Random Forest
Overall, TF-IDF outperformed Count Vectorizer, and larger threshold values yielded more consistent performance improvements across datasets, though lower values offered greater potential for gains on large, diverse datasets, which suggest pseudo-labeling is a viable method for incorporating unlabeled data.