This work shows that specialization can outweigh scale in the case of small, low-resourced dialects and highlights the significance of investments into gathering resources for them.
Students' opinions play a pivotal role in the formulation of successful educational policies. However, the computational analysis of the Moroccan Arabic Dialect remains widely viewed as problematic due to a focus on complex morphology and the absence of specific datasets. To close this gap, we propose an end-to-end system for educational sentiment mining, which utilizes EduDarijaBERT—a transformer model explicitly fine-tuned on this academic domain. In the development of the system, we initially constructed and systematically assessed a specifically curated corpus of 12,223 social media comments. To assess the model's reliability, we evaluated the framework on an independent test set. The proposed EduDarijaBERT system achieved an accuracy of 83.15% on this dataset, demonstrating strong generalization capabilities for educational sentiment mining. In addition to these quantitative metrics, the system facilitates a qualitative interpretation of sentiment categories, providing practical insights revealing that administrative challenges are among the major contributors to negative feedback, while active academic support is the key to positive interaction. This study provides universities with a stable, easily implementable system to track student feedback and fuel data-driven innovations.
Abderrahim Ait Ichou, Ali Ouacha· International Journal of Adv...· 0 citations
Arabic sentiment analysis (ASA) has received increasing attention due to the rich availability of Arabic content across online platforms. However, the Arabic language presents several issues including rich morphology, dialectal diversity, and complex grammatical structures. This survey introduced a comprehensive systematic review of ASA research from 2018 to 2025, analyzing over 70 peer-reviewed studies. We examine three main methodological approaches:(1) traditional machine learning approaches, (2) deep learning techniques, and (3) transformer-based architectures including AraBERT, MARBERT, and CAMeLBERT and emerging aspect-based sentiment analysis (ABSA) research. Additionally, we provide a novel comparative analysis of a set of publicly available datasets and NLP tools. We identify four critical challenges: data scarcity, code-switching, dialectal generalization and evaluation standardization. Finally, we propose future directions including cross-lingual transfer learning, multimodal sentiment analysis, domain-specific ASA applications. The survey provides researchers with a structured roadmap for improving Arabic sentiment analysis systems.
Ola Adnan Altiti· International journal of com...· 0 citations
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al.· Buana Information Technology...· 0 citations
A bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text is introduced, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days.
Ahmed Amine Aliane, H. Aliane, N. Semmar· 0 citations
This study introduces a novel, publicly available balanced Algerian Arabic Aspect Based Sentiment Analysis (ABSA) dataset consisting of 11,338 comments written in the Algerian dialect. The dataset follows the semantic evaluation 2016 annotation guidelines and includes both explicit and implicit aspects within the telecommunications domain. Furthermore, the study proposes a unified End-to-End (E2E) framework for Arabic ABSA based on the newly introduced dataset and the Arabic ABSA hotels dataset, where aspect term extraction and sentiment classification are integrated into a single sequence labeling task. Transfer learning was leveraged by fine-tuning the Arabic Bidirectional Encoder Representations from Transformers (AraBERT) model, which was further enhanced with Bidirectional Gated Recurrent Units (BiGRU) and a Softmax output layer, forming the fine-tuned AraBERT-BiGRU-Softmax model. Using the unified E2E ABSA approach, the proposed model achieved an overall accuracy of 88.07% on our dataset, along with a macro precision of 72.96%, a macro recall of 60.79%, and a macro F1-score of 65.48% across all labels. When excluding the O label, the model obtained a micro-precision of 54.94%, a micro-recall of 53.85%, and a micro-F1 score of 54.39%. Evaluated on the Arabic ABSA hotel reviews dataset, the model obtained an overall accuracy of 92.91%, with a macro precision of 57.79%, a macro recall of 47.04%, and a macro F1-score of 50.69% across all labels. In addition, when excluding the O label, it reached a micro-precision of 64.74%, a micro-recall of 55.74%, and a micro-F1 score of 59.90%.
Fatiha Tebbani, C. Kara-Mohamed, A. Hamdi-Cherif· Journal of King Saud Univers...· 0 citations