Skip to content
Review Open access

Automated Sentiment Analysis of Hindi Text using Machine Learning Techniques: A Lightweight and Scalable Framework for Regional Language NLP

Sep 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages by proposing a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual data.

Abstract

The paper presents a lightweight yet effective machine learning-based framework for automated sentiment classification of Hindi textual data. This proposed work addresses the persistent challenges of data sparsity, linguistic diversity, and limited annotated resources that hinder sentiment analysis in regional Indian languages. Here methodology encompasses a systematic pipeline comprising data preprocessing, Term Frequency–Inverse Document Frequency (TF-IDF) feature extraction with unigram and bigram representations, and supervised classification using Multinomial Naive Bayes (MNB) and Logistic Regression (LR) algorithms. Experiments on the IIT Patna Movie Reviews Hindi Sentiment Analysis dataset (2,480 training, 310 validation, and 310 test samples spanning the negative, neutral, and positive classes) demonstrate that Logistic Regression substantially outperforms the Naive Bayes baseline, achieving 88.19% training accuracy and 54.84% test accuracy against 67.82% and 43.23% for MNB, together with a higher micro-averaged Receiver Operating Characteristic – Area Under the Curve (ROC–AUC) (0.742 vs. 0.653). Class-wise analysis shows that positive sentiment is the easiest to detect (LR F1 = 0.63), while the neutral class remains the most challenging (LR F1 = 0.39). Proposed framework offers a computationally efficient, interpretable, and scalable solution for Hindi sentiment analysis, with direct applicability to social media monitoring, customer feedback analysis, and opinion mining in regional language ecosystems. Experimental codes is made available as a fully executable Google Colab notebook to ensure reproducibility and facilitate future research extensions.

Read PDF

Similar papers

Review Open access Sep 2026

Improving Sentiment Classification Performance Using Pseudo-Labeling with Naive Bayes and Random Forest

Overall, TF-IDF outperformed Count Vectorizer, and larger threshold values yielded more consistent performance improvements across datasets, though lower values offered greater potential for gains on large, diverse datasets, which suggest pseudo-labeling is a viable method for incorporating unlabeled data.

Arvidion Havas Oktavian, A. Aribowo · 0 citations
Review Open access Sep 2026

Machine Learning-Based Sentiment Classification of Reviews from Indonesian Mobile Applications Using TF-IDF

Comparisons of classical machine learning algorithms and the effectiveness of class-weighted learning in improving minority-class recognition in Indonesian mobile application reviews demonstrate that the highest overall accuracy does not necessarily indicate the most balanced classifier under class imbalance.

Tuti Handayani, Sri Mardiyati · 0 citations
Open access Sep 2026

Comparison of Naive Bayes, SVM, and Logistic Regression for Sentiment Analysis of the Makan Bergizi Gratis Program

The Makan Bergizi Gratis Program (MBG) became one of the widely discussed public issues on platform X and generated diverse responses from users. These responses included supportive, critical, and neutral opinions, making sentiment analysis relevant for understanding public opinion toward the program. This study compar...

Andika, Julio Cesar Alessandro, Adjie Perkasa Tarigan et al. · 0 citations
Open access 2026

Comparative Analysis of Language Models for Sentiment Classification

Comparing and analysing the performance of several machine learning algorithms on fine-grained sentiment classification problems to examine their suitability and shortcomings for use as models in sentiment analysis suggests large language models perform significantly worse on the 28-class classification task in zero-sh...

Shangjiafeng Guo · 0 citations
Sep 2026

Deep Learning-based Multi-Class Sentiment Classification from Social Media Comments using LSTM Architecture

This paper presents a lightweight sentiment classification model based on Long Short-Term Memory networks, developed as a foundational text-analysis component for future multimodal emotion recognition systems, and provides a reproducible and computationally efficient baseline suitable for integration into broader multi...

Munmun Kakkar, Hemant Patidar · 0 citations
#small language model Open access Sep 2026

An Empirical Benchmarking of Traditional Machine Learning and DistilBERT-Based Zero-Shot Hierarchical Sentiment Analysis on Large-Scale Twitter Data

Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performan...

Bhumit Peshavariya, S. Nahar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.