Robust Binary Cyberbullying Detection on Social Media Using a Multi-Class Guided Transformer Ensemble
Abstract
Cyberbullying detection in social media remains a challenging task due to noisy textual content, contextual ambiguity, and the rapidly evolving nature of online language. This paper proposes a domain-specific transformer ensemble framework for cyberbullying detection that leverages fine-tuned Twitter-RoBERTa models. The proposed framework introduces a multi-class guided binary classification strategy in which the transformer model is initially trained on fine-grained cyberbullying categories and subsequently aggregated into binary predictions during inference. To improve robustness and generalization, we employ a weighted ensemble combining K-Fold fine-tuning, RoBERTa-large optimization, and word-dropout augmentation. We conducted experiments on the English-language Cyberbullying Classification Dataset, a publicly available Twitter dataset containing 47,692 annotated tweets across six cyberbullying categories. The proposed weighted ensemble achieved 91.06% accuracy and 90.64% F1-score, outperforming recent streaming machine learning approaches. Comprehensive evaluation through ablation analysis, cross-validation, statistical significance testing, and error analysis further demonstrates the framework's effectiveness and stability. The results confirm that domain-specific transformer fine-tuning combined with heterogeneous ensemble learning provides a robust and scalable solution for automated cyberbullying detection in social media environments