Skip to content
Open access

Hybrid BERT-XGBoost Framework for Early Detection and Classification of Online Cyberbullying across Social Media

2026 · Journal of Cyber Security and Risk Auditing · Vol 2026, pp. 281-305 · 0 citations

TL;DR

A Hybrid BERT–XGBoost model to detect cyberbullying in short social media texts, which combines the strengths of both models, provides an interpretable, practical and category-aware solution for early detection of cyberbullying.

Abstract

Cyberbullying via social media is a constant digital safety issue because the content can be widely shared and openly visible and can have a negative impact on users before it is removed by manual moderation. Current detection models are mostly based on shallow lexical features or transformer-only classifiers, resulting in low-level accuracy and explainability. This study introduces a Hybrid BERT–XGBoost model to detect cyberbullying in short social media texts, which combines the strengths of both models. The contextual sentence embeddings are extracted using BERT and the auxiliary linguistic and behavioral features are extracted in parallel, such as sentiment polarity, profanity score, punctuation intensity, capitalization ratio, hashtag usage, mention count, emoji frequency, and post length. XGBoost is used for the classification of the fused representation. The model is tested on stratified training, validation, and testing splits, compared to a baseline model, ablated, tested with macro-F1, weighted-F1, ROC-AUC, early detection recall, and grouped explainability. The proposed framework achieved 96.18% accuracy, 96.05% macro-F1, 96.16% weighted-F1, 95.88% early detection recall, and 98.42% macro-AUC. It performs better than the BERT + Dense baseline, which obtained 94.31% accuracy and 94.08% macro-F1 score, demonstrating the advantage of fusion of contextual and auxiliary features. The framework provides an interpretable, practical and category-aware solution for early detection of cyberbullying, but further research is needed to validate the framework in multiple languages, modalities and in conversations.

Read PDF

Similar papers

Open access Sep 2026

Robust Binary Cyberbullying Detection on Social Media Using a Multi-Class Guided Transformer Ensemble

Cyberbullying detection in social media remains a challenging task due to noisy textual content, contextual ambiguity, and the rapidly evolving nature of online language. This paper proposes a domain-specific transformer ensemble framework for cyberbullying detection that leverages fine-tuned Twitter-RoBERTa models. Th...

Shahad Ghazi Abd, Bashar Taleb Hameed, Ali Hussein Fadhel · 0 citations
Open access Aug 2026

Indonesian Cyberbullying Detection Using IndoBERTweet-BiGRU Model on Class-Imbalanced X (Twitter) Data

It is demonstrated that integrating contextual language representations with sequential modeling, supported by an efficient LLM-assisted labeling strategy and class imbalance handling, provides an effective approach for Indonesian cyberbullying detection and offers a practical solution for large-scale social media cont...

F. Nafiah, Aviolla Terza Damaliana, K. M. Hindrayani · 0 citations
Review Open access Sep 2026

Advances in AI-Based Methods for Cyberbullying Detection on the X Platform: A Systematic Review

Cyberbullying on social media platforms presents a serious and growing concern due to its psychological and social consequences, particularly among young users. The X platform, known for its high user activity and informal language, has become a prominent site for the spread of online abuse. In response, AI technologie...

K. H. Al-Omari, Raghad Saleh Al-Ghamdi, Sumaya Hassan Al-Suhaimi Al-Suhaimi et al. · 0 citations
Open access Aug 2026

Performance Evaluation of Quantum Machine Learning Models for Cyberbullying Detection

Cyberbullying on social media has become a serious issue that affects users' mental well-being, safety, and online engagement. Detecting cyberbullying automatically is challenging due to informal writing styles, short messages, and the presence of subtle abusive patterns that often overlap with normal communication. Th...

Vandana P. Pawar, P. M. Yawalkar · 0 citations
Sep 2026

Chinese cyberbullying detection via multi-feature prompt learning

Most existing research on cyberbullying detection reduces the task to comment-level hate speech classification, overlooking its collective, event-driven, and dynamic nature. To address this limitation, we propose CDMPL, a multi-feature prompt learning framework for Chinese cyberbullying incident detection. CDMPL integr...

Xin Zou, Yi Zhu, Ye Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.