Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas et al.· 0 citations
The findings suggest that the impact of GenAI use is present in various contexts, highlighting the need for instructional guidance on how students should use GenAI as a learning aid, and insights for other instructors that wish to integrate GenAI tools into computing curricula.
Valeria Ramirez Osorio, Ido Ben Haim, Ahmed Ashraf et al.· Annual Conference on Innovat...· 1 citation