Skip to content

The Enforcement and Feasibility of Hate Speech Moderation on Twitter

Apr 2026 · arXiv.org · Vol abs/2604.12289 · 2 citations · 57 references
Computer Science

Abstract

Online hate speech is associated with harms ranging from deteriorating mental health to violence, yet how consistently platforms moderate hate, and whether enforcement is feasible at scale, remain poorly understood. We audit hate speech moderation on Twitter (now X) using 540,000 tweets annotated by trained native speakers, representative of a full day on the platform. Five months after posting, 80% of hateful tweets, including violent ones, remained online. Removal was only marginally more likely than for non-hateful tweets, far below scams or adult content, and insensitive to severity and reach. Automated detection could not reliably classify hate but ranked it highly, enabling human triage. Simulating this workflow, current staffing curbed little exposure, yet substantial reductions proved financially feasible, far below applicable regulatory fines. Persistent hate reflects resource allocation, not technical limits.

View source

Similar papers

Open access Aug 2026

Detecting and Categorizing Islamophobic Hate Speech on Spanish Twitter/X

Islamophobic hate speech on social media includes overt toxicity and contextual frames that associate Muslims with invasion, criminality, cultural threat, or racialized exclusion. This study examines 21,326 Spanish-language tweets about Islam and Muslims posted in Spain between 2010 and 2022. We combine human annotatio...

William González-Baquero, Carlos Arcila Calderón, Javier J. Amores et al. · 0 citations
Open access Sep 2026

From Words to Harm: Hate Speech and Societal Violence in U.S. Presidential Campaigns

Political speeches are typically viewed as mobilizing democratic participation, yet research shows elite political rhetoric—particularly when it contains hate speech—may contribute to societal harm. Our study extends this line of research by investigating the association between hate speech and societal violence from a...

Zheng-Yi Liang, Joel Sandoval-Valdez, Jaeho Cho et al. · 0 citations
#human-computer interacti... Preprint Sep 2026

Reclaiming the social in social media

This work argues that the harms commonly attributed to social media arise not from technologies supporting social connection but from their implementation within the attention economy, and proposes an architecture that does, making the attention economy not just discouraged but structurally impossible.

Luca Benn, M. Merz, D. Grüning · 0 citations
Open access Aug 2026

The Performance Analysis of Hate Comments in Cyberbullying Cases Based on IndoBERT and Cendol

The rise of hate comments on social media, especially during politically sensitive periods such as Indonesia’s 2024 election has increased the urgency of automated cyberbullying detection. This study aims to evaluate and compare the performance of two Indonesian-language NLP models IndoBERT and Cendol in classifying ha...

Nancy Olivia Syahanifa, Kartika Dwi Maharani, Anggara Budiyanto et al. · 0 citations
Open access Sep 2026

What drives hateful tweets to go viral? Threat perception and Twitter engagement: The case of the COVID ‐19 pandemic

Drawing on Intergroup Threat Theory as an analytical framework, this study analyses pandemic‐driven Sinophobic tweets to examine how threat‐laden language influences user engagement through computational text analysis. Findings showed that those conveying ideological threats generated higher engagement, whereas those...

Hyejin Kim, Thyago Mota, Sanga Song · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.