Skip to content
Preprint

How Robust Are LLMs to Vietnamese Dialects?

Aug 2026 · 0 citations · 11 references
Computer Science

TL;DR

These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation, and the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks is presented.

Abstract

Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.

Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al. · 0 citations
#natural language process... Preprint Sep 2026

5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly targe...

Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed et al. · 0 citations
Conference Aug 2026

VietNorm: Vietnamese Dialect Normalization

Vietnamese dialect normalization transforms regional linguistic variants into standard Vietnamese, thereby improving the robustness of downstream natural language processing systems. However, models trained only on dialect-to-standard pairs often suffer from over-normalization, in which already-standard inputs are unne...

Y. Duong, Hao Nguyen, T. Tran et al. · 0 citations
Preprint Aug 2026

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this"dialect tax"across the natural language processing pipeline. Using parallel English dialect corpora that hold meanin...

Elle Michelle Yang, Mark Chen, Jerry Tworek et al. · 0 citations
#natural language process... Preprint Sep 2026

To What Extent Do Large Language Models Understand Bangla Idioms?

Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice quest...

Mousumi Akter, Md. Faiyaz Abdullah Sayeedi, Nurul Labib Sayeedi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.