These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation, and the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks is presented.
Abstract
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.
This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.
Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al.· 0 citations
Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly targe...
Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed et al.· 0 citations
Vietnamese dialect normalization transforms regional linguistic variants into standard Vietnamese, thereby improving the robustness of downstream natural language processing systems. However, models trained only on dialect-to-standard pairs often suffer from over-normalization, in which already-standard inputs are unne...
Y. Duong, Hao Nguyen, T. Tran et al.· International Conference on...· 0 citations
Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this"dialect tax"across the natural language processing pipeline. Using parallel English dialect corpora that hold meanin...
Elle Michelle Yang, Mark Chen, Jerry Tworek et al.· 0 citations
Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice quest...
Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi et al.· 0 citations