Skip to content
Open access

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

Jul 2026 · Buana Information Technology and Computer Sciences (BIT and CS) · 0 citations · 14 references

TL;DR

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Abstract

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language. Unlike prior works that often focus on sentiment analysis or use unbalanced datasets, this research utilizes a balanced dataset consisting of 2,500 lyric segments from five regional languages: Javanese, Sundanese, Batak, Minangkabau, and Banjarese. A comprehensive preprocessing pipeline is applied, including case folding, text cleaning, tokenization, stopword removal, stemming, sequence padding, and label encoding to transform textual data into numerical representations. The model is evaluated using 5-fold cross-validation to ensure robustness and generalization across different data partitions. Experimental results show that the proposed model achieves an accuracy of 95.24%, precision of 95.36%, recall of 95.24%, and F1-score of 95.26%, indicating strong and consistent performance. These findings demonstrate that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages, enabling accurate classification despite similarities in vocabulary and structure. Furthermore, this study contributes to the advancement of natural language processing for low-resource languages and highlights the potential of deep learning approaches in supporting the digital preservation and automatic organization of Indonesian regional cultural content.

Read PDF

Similar papers

Open access Jul 2026

Implementation of a Bi-LSTM Model for Automatic Text Classification of Mathematics, Science, and Indonesian Language Questions

It can be concluded that the Bi-LSTM model is effective for automatic text classification of educational questions and has strong potential for further development in technology-based question grouping systems.

Mochamad soffan Muslim, Aviv Yuniar Rahman, Rangga Pahlevi · 0 citations
Open access Aug 2026

Implementation of the IndoBERT-LSTM Model for Indonesian Sentiment Analysis withan Explainable AI Approach Using SHAP

The integration of the IndoBERT-BiLSTM architecture with SHAP is demonstrated to deliver accurate and explainable Indonesian sentiment analysis, which effectively bridges the gap between deep learning performance and decision transparency without compromising classification accuracy.

A. Widiyatmoko, A. Nugroho, Muhammad Nurul Firdaus · 0 citations
Open access Aug 2026

Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Hanan Fawzy, Ahmad Salah, Heba El-Fiqi et al. · 0 citations
Review Open access Jul 2026

A Hybrid VADER–IndoBERT Framework for Robust Sentiment Analysis of Long and Ambiguous Indonesian Texts

A Hybrid VADER–IndoBERT framework designed to improve sentiment classification robustness on complex Indonesian texts is introduced, demonstrating the superiority of Transformer-based architectures in capturing long-range dependencies and handling ambiguous sentiment cues.

Margareta Valencia Suci Handayani, R. S. Basuki, Muljono et al. · 0 citations
Open access 2026

Arabic News Text Classification Using Deep Learning Models with Dynamic N-grams

This study integrates parallel multi-kernel word-level convolutional features into conventional and hybrid deep learning models for Arabic text analysis tasks, providing a systematic within-study assessment of model sensitivity to architecture, preprocessing, and learning-rate selection.

Ahmed I.Taloba, George Samy Rady, Khaled F. Hussain · 0 citations
Open access 2026

A Large-Scale Vietnamese News Dataset for Text Classification: Construction and Evaluation

This study analyzes classification performance through four critical dimensions: model architecture, temporal data shift, source-origin bias, and training data scale, and reveals distinct learning behaviors across model families: transformer models benefit increasingly from larger training sets, whereas the strongest traditional and deep learning baselines remain competitive throughout the evaluated range.

B. Chau, D. Duong, Phuoc Tran · 0 citations