Jul 2026· Buana Information Technology and Computer Sciences (BIT and CS)· 0 citations· 14 references
TL;DR
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.
Abstract
This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language. Unlike prior works that often focus on sentiment analysis or use unbalanced datasets, this research utilizes a balanced dataset consisting of 2,500 lyric segments from five regional languages: Javanese, Sundanese, Batak, Minangkabau, and Banjarese. A comprehensive preprocessing pipeline is applied, including case folding, text cleaning, tokenization, stopword removal, stemming, sequence padding, and label encoding to transform textual data into numerical representations. The model is evaluated using 5-fold cross-validation to ensure robustness and generalization across different data partitions. Experimental results show that the proposed model achieves an accuracy of 95.24%, precision of 95.36%, recall of 95.24%, and F1-score of 95.26%, indicating strong and consistent performance. These findings demonstrate that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages, enabling accurate classification despite similarities in vocabulary and structure. Furthermore, this study contributes to the advancement of natural language processing for low-resource languages and highlights the potential of deep learning approaches in supporting the digital preservation and automatic organization of Indonesian regional cultural content.
It can be concluded that the Bi-LSTM model is effective for automatic text classification of educational questions and has strong potential for further development in technology-based question grouping systems.
The integration of the IndoBERT-BiLSTM architecture with SHAP is demonstrated to deliver accurate and explainable Indonesian sentiment analysis, which effectively bridges the gap between deep learning performance and decision transparency without compromising classification accuracy.
A. Widiyatmoko, A. Nugroho, Muhammad Nurul Firdaus· Journal of Electrical Engine...· 0 citations
The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.
Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al.· Informatica· 0 citations
A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.
Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
This study integrates parallel multi-kernel word-level convolutional features into conventional and hybrid deep learning models for Arabic text analysis tasks, providing a systematic within-study assessment of model sensitivity to architecture, preprocessing, and learning-rate selection.
Ahmed I.Taloba, George Samy Rady, Khaled F. Hussain· International Journal of Adv...· 0 citations
This study analyzes classification performance through four critical dimensions: model architecture, temporal data shift, source-origin bias, and training data scale, and reveals distinct learning behaviors across model families: transformer models benefit increasingly from larger training sets, whereas the strongest traditional and deep learning baselines remain competitive throughout the evaluated range.
B. Chau, D. Duong, Phuoc Tran· IEEE Access· 0 citations