IRAIVI: Agentic Framework for Tamil Hate Speech Detection and Counter-Speech Generation Toward Responsible AI Intervention
Abstract
Online gender-based hate speech affects women, men, and LGBTQIA+ communities and is difficult to moderate in low-resource Tamil because social media content often contains transliteration, code-mixing, dialectal variation, sarcasm, and culturally specific abusive expressions. This paper presents IRAIVI (Intelligent Responsible Agentic AI for Identifying Hate Speech, Verifying Legal References, and Intervention), a CrewAI-based prompt-driven agentic framework that orchestrates multiple agents, tasks, and a sequential workflow for hate speech detection, target identification, hate category classification, legal knowledge retrieval, and Tamil counter-speech generation. The framework integrates Tree-of-Thought (ToT) reasoning, Retrieval-Augmented Generation (RAG), Genetic Algorithm (GA)-based prompt optimization, and human-in-the-loop (HITL) escalation for uncertain cases. A Qwen-based self-instruction pipeline expanded 200 seed examples into 5,747 Alpaca-style instruction-input-output triplets. The dataset was collected from publicly available Tamil YouTube comments using the YouTube API and contained no personally identifiable information therefore, institutional ethics approval was not required. The proposed three-level annotation scheme includes hate or non-hate classification, target identification, and hate category (implicit or explicit), achieving strong annotation agreement with Cohen’s kappa (k = 0.87). Experiments were conducted using GPT-3.5, Gemini, Mistral-7B, Qwen-4B, Phi-2, and TinyLLaMA. GPT-3.5 achieved the highest classification F1-score of 0.91, while Qwen-4B achieved 0.85 among the evaluated open-source models. The ToT + GA configuration achieved a counter-speech quality score of 0.86 and a BERTScore of 0.92, demonstrating the effectiveness of the proposed framework for responsible Tamil hate speech detection and intervention.