Skip to content
Review

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

Aug 2026 · 0 citations · 34 references
Computer Science

Abstract

Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers'ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.

View source

Similar papers

Open access Jul 2026

When Morphology Indexes Prestige: Arabic-English Code-mixing as Linguistic and Social Practice in Saudi Arabia

This study explores the intersection of linguistic form and social meaning within Arabic-English code-mixing in relation to the morphological patterns that arise in mixed speech and their association with prestige and modern identity. Based on naturally occurring spoken and digital data and questionnaires from educated Arabic-English speakers in Saudi Arabia, the analysis reveals that morphological adaptation strategies include the introduction of English lexical items into Arabic morphological patterns, affixal incorporation, and the creation of hybrid lexical forms, whose recurrent patterns provide evidence of systematic linguistic innovation, rather than mere borrowing. These morphological adaptations are analyzed through the lens of indexicality and social meaning (Silverstein, 2003; Eckert, 2008) that correlate language choice with social aspiration, education, and symbolic capital. The occurrence of English insertions in the data is often related to prestige, global orientation and identification with a modern lifestyle, although these meanings are contextually inferred, and Arabic morphology is used as a sign of authenticity and local identity. The study argues that morphological choices in code-mixing are socially motivated, and they serve as a site for negotiating status, identity, and belonging. By drawing on morphological and sociolinguistic approaches, this study illustrates how form and meaning converge to shape prestige meanings in contemporary Arabic-English speech. 

W. Alshammari · 0 citations
Open access Aug 2026

Syllable Structure Variation between Two Iraqi Arabic Dialects: A Comparative Optimality-Theoretic Investigation

This study is an Optimality-Theoretic comparison of two major dialects of Iraqi Arabic: Baghdadi Gilit Arabic (BGD) and Basri Arabic (BAS), focusing on syllable structure variation. The data were collected from 20 native speakers (10 for each dialect), representing both rural and urban backgrounds, as participants in recorded naturalistic conversations and in targeted elicitation tasks. The analysis examined a large set of more than 200 phonological tokens in the framework of Optimality Theory (OT), in particular, the ranking of the constraints for dealing with complex onsets, the resolution of triconsonantal clusters, and the Gahawa Syndrome. Key findings reveal that complex onsets are allowed in both dialects, but that they have different strategies for resolving the triconsonantal clusters and different characteristics in the occurrence and phonological conditioning of the Gahawa Syndrome.

Israa Ali Kareem Al-Aaydi · 0 citations
Preprint Aug 2026

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.

Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas et al. · 0 citations
Open access Jul 2026

Exploration of Arabic Collocation Patterns in the Indonesia Al-Youm News Corpus

This study investigates the structural and semantic characteristics of Arabic collocations in the Indonesia Al-Youm digital news corpus, a corpus compiled from Arabic-language online news texts reporting Indonesian social, political, cultural, and economic issues. Using Sketch Engine’s Multiword Term extraction feature, the study identifies and classifies the top 100 recurrent multiword units according to their morphosyntactic structures, grammatical functions, and semantic domains. The findings show that Arabic journalistic discourse in this corpus is dominated by nominal constructions, particularly idafah, noun–adjective patterns, and Masdar-based formations. Semantically, the collocations cluster around six domains: media identity and digital technology, politics and diplomacy, social protection and administration, economy and development, tourism, culture, and geography, and proper names. The novelty of this study lies in its corpus-based mapping of Arabic collocation patterns within an Indonesian digital news context, an area that remains underexplored in Arabic corpus linguistics. The results contribute to the description of Arabic journalistic phraseology and provide empirical input for Arabic for Specific Purposes, especially journalistic Arabic and media translation instruction.

I. Romadhon, Muassomah Muassomah, Zakiyah Arifa · 0 citations
Open access Aug 2026

DIALECTS IN TURKEY AND NEIGHBOURING COUNTRIES BETWEEN ROOTS AND BRANCHES: A COMPARATIVE STUDY IN THE LIGHT OF THE ARABIC LINGUISTIC HERITAGE AND ISLAMIC STUDIES

This study compares the Arabic dialects of Southeastern Turkey (Anatolian Arabic) with the neighbouring Arabic dialects of Syria and Iraq in terms of their linguistic relationship. These varieties, which have been lying on the periphery of Arabic and have been cut off from the Arabic heartlands since 1923, are identified with the sedentary Mesopotamian (qəltu) branch and their contact-induced changes under intensive contact with Turkish and Kurdish are investigated. The data collected in the field are compared to the well-known corpora of Syria and Iraq in Mardin, Hatay, Urfa and Siirt. The results reveal that the Anatolian Arabic still has some archaic features (like the voiceless uvular stop [q] and the perfective suffix –tu) which are not present in neighbouring dialects, and some new features (like the SOV word order and the gender neutralisation). The sociolinguistic analysis shows that remote dialects like Tillo are facing a serious language shift by the younger generation. The results of this study, when considered in the light of the Arabic linguistic heritage (al-turāth al-lughawī), reveal that the preserved features come under the rubric of the articulatory categories central to the Qurʾānic recitation (tajwīd), while the religious field has been one of the means of their preservation. It concludes that Anatolian Arabic is a unique peripheral Arabic and that the documentation of the same preserves a common Islamic linguistic heritage.

Mahmud Şuş · 0 citations