தமிழ் வட்டார மொழிகளை AI மொழி மாதிரிகள் கையாளும் திறன்: கொங்கு வட்டாரத் தமிழை மையமாகக் கொண்ட ஆய்வு
Abstract
While Large Language Models (LLMs) demonstrate significant capabilities in standard Tamil, the extent to which they accurately understand and translate Kongu regional Tamil—spoken in the Kongu region including Coimbatore, Tiruppur, Erode, Salem, and Karur—remains an insufficiently explored area. This article examines the dialect-handling capabilities of AI language models, focusing on Kongu Tamil. For this purpose, a small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary. Using this dataset, zero-shot and few-shot experiments were conducted with a publicly accessible Large Language Model (Claude), and the responses were evaluated based on the author's linguistic knowledge, with errors categorized accordingly. Additionally, citing results from previously published studies comparing various LLMs (Gemini, ChatGPT, Claude), this paper proposes a comprehensive methodological plan for a complete multi-model comparative study for Kongu Tamil, along with a human expert evaluation protocol. Preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions. This study underscores the need to build large-scale, human-expert-verified benchmark datasets for Tamil dialects.