ArabCulture-Dialogue is introduced, a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both MSA and each country’s respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics, to address the performance gap between MSA and Arabic dialects.
This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.
Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al.· 0 citations
ALMIEYAR is introduced, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models, including the first published benchmark for Ahwazi Arabic.
Omid Ghahroodi, Anas Madkoor, Dima Faris Alsaudi et al.· 0 citations
Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective translation requires modeling not only semantic content but also dialectal variation, conversational context, speaker and addressee characteristics, and sociolinguistic a...
Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy et al.· 1 citation
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level...
Priyanka Dasari, Yuvrajsinh Bodana, V. Mujadia et al.· 0 citations
This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoRA adapter on NileChat-3B using structured system/user prompts that condition gen...
N. Esmaeil, Fathima Rena, S. Subhash et al.· 1 citation
These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation, and the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks is presented.
Minh Tran, C. Trinh, T. Lê et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.