Online hate speech is associated with harms ranging from deteriorating mental health to violence, yet how consistently platforms moderate hate, and whether enforcement is feasible at scale, remain poorly understood. We audit hate speech moderation on Twitter (now X) using 540,000 tweets annotated by trained native speakers, representative of a full day on the platform. Five months after posting, 80% of hateful tweets, including violent ones, remained online. Removal was only marginally more likely than for non-hateful tweets, far below scams or adult content, and insensitive to severity and reach. Automated detection could not reliably classify hate but ranked it highly, enabling human triage. Simulating this workflow, current staffing curbed little exposure, yet substantial reductions proved financially feasible, far below applicable regulatory fines. Persistent hate reflects resource allocation, not technical limits.
Islamophobic hate speech on social media includes overt toxicity and contextual frames that associate Muslims with invasion, criminality, cultural threat, or racialized exclusion. This study examines 21,326 Spanish-language tweets about Islam and Muslims posted in Spain between 2010 and 2022. We combine human annotatio...
William González-Baquero, Carlos Arcila Calderón, Javier J. Amores et al.· Central European Journal of...· 0 citations
Political speeches are typically viewed as mobilizing democratic participation, yet research shows elite political rhetoric—particularly when it contains hate speech—may contribute to societal harm. Our study extends this line of research by investigating the association between hate speech and societal violence from a...
Zheng-Yi Liang, Joel Sandoval-Valdez, Jaeho Cho et al.· International Journal of Com...· 0 citations
This work argues that the harms commonly attributed to social media arise not from technologies supporting social connection but from their implementation within the attention economy, and proposes an architecture that does, making the attention economy not just discouraged but structurally impossible.
The rise of hate comments on social media, especially during politically sensitive periods such as Indonesia’s 2024 election has increased the urgency of automated cyberbullying detection. This study aims to evaluate and compare the performance of two Indonesian-language NLP models IndoBERT and Cendol in classifying ha...
Nancy Olivia Syahanifa, Kartika Dwi Maharani, Anggara Budiyanto et al.· Journal of Measurements Elec...· 0 citations
Drawing on Intergroup Threat Theory as an analytical framework, this study analyses pandemic‐driven Sinophobic tweets to examine how threat‐laden language influences user engagement through computational text analysis. Findings showed that those conveying ideological threats generated higher engagement, whereas those...
Hyejin Kim, Thyago Mota, Sanga Song· Asian Journal of Social Psyc...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.