Skip to content
Conference

Data Saturation in Low-Resource TTS Fine-Tuning

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 24 references

Abstract

When fine-tuning large language models, the assumption is that more data means better results. This principle is often extended to text-to-speech (TTS) fine-tuning, yet remains underexplored, particularly for low-resource languages where high-quality data is difficult to obtain. In this work, the "more data is better" phenomenon is investigated for Turkish TTS using XTTS v2 by incrementally increasing training data. Speech quality is evaluated using multiple metrics, including UTMOS, NISQA, and an LLM-based TTS evaluation framework. For the LLM-based evaluation, a multimodal language model (Gemini) was prompted to assess Turkish-specific pronunciation, naturalness, and synthesis artifacts on a calibrated 1-10 scale. Results suggest that for TTS fine-tuning on non-mainstream languages, modest data investments may achieve near-optimal quality.

View source

Similar papers

Open access Aug 2026

Speech-to-text model comparison using XLS-R, XLSR-53, and Wav2Vec 2.0

Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.

Juan Hebert, Amalia Zahra · 0 citations
Conference Jul 2026

QMOS: Qwen-Based MOS Prediction for TTS

Text-to-speech systems are improving fast, but measuring how natural they sound still requires expensive human listening tests. Existing automatic methods struggle to generalise well across different datasets. We present QMOS, a MOS prediction framework that extracts hierarchical speech quality features from a frozen Qwen2-Audio large audio language model and combines them with layer-weighted WavLM SSL representations through a learned cross-attention fusion. This design captures both high-level semantic naturalness and lowlevel acoustic distortions in a unified model. On two standard benchmarks, SOMOS and BVCC, QMOS achieves competitive system-level SRCC of 0.923 and 0.921 on SOMOS and BVCC, respectively, while using no system-ID conditioning or listener embeddings. Cross-domain evaluation yields a competitive SRCC of 0.702, suggesting the learned representations generalise across acoustic domains.

Sri Ravi Sastry Kolluru, Charan Devarakonda, S. Radhe et al. · 0 citations
Conference Jul 2026

Fine-Tuning Text-to-Speech Models with Turkish Speech Data

Developing text-to-speech (TTS) systems for a language with limited accessible speech data such as Turkish remains a challenge. This study describes a process for creating a Turkish text-to-speech system using web-scraping data to train deep learning models. The data collection approach is based on transcribing Turkish audiobook content from YouTube and converting it into a usable dataset using normalization, piece segmentation, and human annotation methods. In this study, the performances of fine-tuning KaniTTS and Dia voice models are compared with the performance of Elevenlabs voice clone. It has been observed that fine-tuned voice models with limited resources gained the ability to synthesize at the level of commercial based API voice model.

Hüseyin Çakmak, Kuzey Arar, F. B. Tek · 0 citations
Preprint Jul 2026

Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

It is suggested that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling, in self-supervised fine-tuning.

Wangjin Zhou, Yizhou Zhang, Yichi Wang et al. · 0 citations
Review Open access 2026

A data-centric approach to performance improvement in under-resourced ASR: The case of Dënë Sųłıné

This paper presents a study focused on advancing Automatic Speech Recognition (ASR) for the under-resourced language Dënë Su ˛ łıné through data-centric approaches. We explore multiple strategies to enhance the quality of training data—both audio recordings and tran-scriptions—to address the challenges posed by mixed-quality datasets. Our experiments investigate which data preparation techniques most effectively improve ASR performance in this context. Our findings show that reducing spelling variants of the same lexeme in the corpus significantly improves model generalization, resulting in a substantial increase in recognition accuracy. Additionally, we demonstrate that increasing manually reviewed transcriptions consistently improves word and character error rates, while audio enhancement slightly reduces performance, highlighting the complex trade-offs in low-resource ASR development.

Olga Kriukova, O. Lovick, Antti Arppe · 0 citations