Aug 2026· JOURNAL OF APPLIED INFORMATICS AND COMPUTING· 0 citations· 34 references
TL;DR
It is demonstrated that integrating contextual language representations with sequential modeling, supported by an efficient LLM-assisted labeling strategy and class imbalance handling, provides an effective approach for Indonesian cyberbullying detection and offers a practical solution for large-scale social media content moderation.
Abstract
Cyberbullying on social media platforms, particularly X (formerly Twitter), has become a serious issue that negatively affects users' mental health and well-being. Automatic cyberbullying detection in Indonesian remains challenging due to the widespread use of informal language, slang, abbreviations, and highly imbalanced class distributions. This study proposes a hybrid deep learning model that integrates IndoBERTweet with a Bidirectional Gated Recurrent Unit (BiGRU) to improve cyberbullying detection performance on Indonesian tweets. A dataset of Indonesian tweets was collected from X and annotated using a multi-stage dual large language model (LLM) labeling strategy to reduce the time and effort required for manual annotation while maintaining label consistency. To address class imbalance, this study investigates the effectiveness of Focal Loss and label distribution modification through multiple experimental scenarios. The proposed approach was evaluated using accuracy, precision, recall, and F1-score. The best performance was achieved by combining Focal Loss with a modified four-class label configuration consisting of Rude and Vulgar Words, Sexual Harassment, Body Shaming and Hate Speech, and Non-Cyberbullying. This configuration obtained an accuracy of 0.93, precision of 0.90, recall of 0.90, and F1-score of 0.90. These findings demonstrate that integrating contextual language representations with sequential modeling, supported by an efficient LLM-assisted labeling strategy and class imbalance handling, provides an effective approach for Indonesian cyberbullying detection and offers a practical solution for large-scale social media content moderation.
The effectiveness of suicide prevention in Indonesia is severely hindered by significant underreporting and social stigma, leading at-risk individuals to express their distress on social media platforms such as Twitter. However, detecting these signals is computationally challenging due to the informal nature of Indone...
Wiyan Herra Herviana, R. Kusumaningrum, B. Surarso· Jurnal Teknik Informatika (J...· 0 citations
Cyberbullying detection in social media remains a challenging task due to noisy textual content, contextual ambiguity, and the rapidly evolving nature of online language. This paper proposes a domain-specific transformer ensemble framework for cyberbullying detection that leverages fine-tuned Twitter-RoBERTa models. Th...
The rise of hate comments on social media, especially during politically sensitive periods such as Indonesia’s 2024 election has increased the urgency of automated cyberbullying detection. This study aims to evaluate and compare the performance of two Indonesian-language NLP models IndoBERT and Cendol in classifying ha...
Nancy Olivia Syahanifa, Kartika Dwi Maharani, Anggara Budiyanto et al.· Journal of Measurements Elec...· 0 citations
The experimental results demonstrate that the BiLSTM model with the RMSProp optimizer is effective for detecting cyberbullying in bilingual Indonesian and English texts.
A sophisticated framework that combines Long Short-Term Memory (LSTM) networks with Natural Language Processing (NLP) techniques is suggested to enhance cyberbullying detection in online communication.
Verganti Sreelatha, M. U. Farooq· International Journal of AI...· 0 citations