Skip to content
Review Open access

Ensemble Learning for Multi-Source Multi-Domain Sentiment Analysis

2026 · International journal of research and scientific innovation · 0 citations

TL;DR

Findings show that ensembling does not guarantee improvement over an already-strong individual classifier, and that domain-aware evaluation is essential in multi-source sentiment analysis, and that domain-aware evaluation is essential in multi-source sentiment analysis.

Abstract

Sentiment analysis systems are typically trained and evaluated on a single data source and domain, limiting their reliability when applied across sources that differ in vocabulary, review length, and writing style. This paper presents an ensemble learning framework for multi-source, multi-domain sentiment classification evaluated across three structurally distinct sources: Twitter (short, informal), TripAdvisor (long, first-person), and Rotten Tomatoes (short, editorial). A unified pipeline combining text normalization with a TF-IDF-weighted Word2Vec (Skip-gram, 200-dim) representation was used to train six base classifiers (Naive Bayes, KNN, Logistic Regression, SVM, Decision Tree, MLP) and four ensembles (Bagging, Boosting, Stacking, Voting) under an identical 5-fold cross-validation protocol. On a held-out, domain-stratified test set of 2,360 samples, SVM achieved the best overall performance (83.05% accuracy, 82.28% F1), narrowly outperforming Stacking (82.58%) and Voting (81.78%). A domain-wise breakdown revealed that TripAdvisor (88.13% F1) was classified far more reliably than Rotten Tomatoes (78.41%) or Twitter (76.27%), a gap associated with review length rather than dataset size. These findings show that ensembling does not guarantee improvement over an already-strong individual classifier, and that domain-aware evaluation is essential in multi-source sentiment analysis.

Read PDF

Similar papers

Sep 2026

Deep Learning-based Multi-Class Sentiment Classification from Social Media Comments using LSTM Architecture

This paper presents a lightweight sentiment classification model based on Long Short-Term Memory networks, developed as a foundational text-analysis component for future multimodal emotion recognition systems, and provides a reproducible and computationally efficient baseline suitable for integration into broader multi...

Munmun Kakkar, Hemant Patidar · 0 citations
Review Open access Sep 2026

Automated Sentiment Analysis of Hindi Text using Machine Learning Techniques: A Lightweight and Scalable Framework for Regional Language NLP

This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages by proposing a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual dat...

Satyapal Singh, Jarnail Singh, D. S · 0 citations
Open access 2026

Comparative Analysis of Language Models for Sentiment Classification

Comparing and analysing the performance of several machine learning algorithms on fine-grained sentiment classification problems to examine their suitability and shortcomings for use as models in sentiment analysis suggests large language models perform significantly worse on the 28-class classification task in zero-sh...

Shangjiafeng Guo · 0 citations
Open access Aug 2026

Linguistically Informed Machine Learning for Gujarati–English Code-Mixed Sentiment Classification: A Comparative Study of Feature Fusion Strategies

Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.

Chirag D. Shah, Shailesh A. Chaudhari · 0 citations
#small language model Open access Sep 2026

An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data

Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performan...

Bhumit Peshavariya, S. Nahar · 0 citations
Open access Sep 2026

Sentiment classification of stock forum text based on a parameter-decoupled ERNIE-Transformer architecture

To address the challenges of diverse domain-specific terminology, highly colloquial expressions, and limited annotated samples in sentiment analysis of stock forum texts, this study proposes an ERNIE-Transformer sentiment classification model that integrates ERNIE and Transformer architectures. First, a systematic data...

Xiu-Mei Li, Fei Chen, Wen-Chao Ling et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.