VietNorm: Vietnamese Dialect Normalization
Abstract
Vietnamese dialect normalization transforms regional linguistic variants into standard Vietnamese, thereby improving the robustness of downstream natural language processing systems. However, models trained only on dialect-to-standard pairs often suffer from over-normalization, in which already-standard inputs are unnecessarily rewritten and useful stylistic cues may be lost. To address this issue, we introduce VietNorm, a simple yet effective framework based on Identity Data Augmentation (IDA). VietNorm augments model training for BARTpho, ViT5, and Vietnamese-correction-v2 with standard-to-standard identity pairs. Experiments on the ViDia2Std benchmark show that VietNorm consistently improves the original baselines; in particular, BARTpho-syllable gains up to 5.93 BLEU points. These results indicate that identity mapping helps mitigate over-normalization while preserving dialect normalization quality.