Aug 2026· 5 citations· ⚡ 1 influential· 44 references
Computer Science
TL;DR
A frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space and improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple.
Abstract
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlled voice design, and speech editing, but remains susceptible to error accumulation during autoregressive generation. Existing solutions often require additional semantic modules, multi-stage tokenizer training pipelines, or complex autoregressive architectures. In this work, we propose FireRedTTS3, a simple yet effective speech generation and editing framework that mitigates error accumulation at the representation level. Specifically, we leverage a frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space. This improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple. FireRedTTS3 provides two variants: FireRedTTS3-Base for multilingual and multi-dialect zero-shot voice cloning, and FireRedTTS3-Instruct for unified voice cloning, instruction-controlled voice design, and speech editing. Experiments show that FireRedTTS3-Base achieves the best average speech intelligibility and speaker similarity among compared systems on Seed-TTS-Eval and MiniMax-MLS-Test, while FireRedTTS3-Instruct outperforms competing systems on InstructTTSEval and Ming-Freeform-Audio-Edit. These results demonstrate that semantically enriched continuous speech representations, combined with a simple architecture, enable stable, controllable, and high-fidelity speech generation and editing. Code and models are available at https://github.com/FireRedTeam/FireRedTTS3.
FireRedAudio is introduced, a general-purpose audio language model with a shared 9B-parameter LLM that achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantia...
Fei-Yu Shen, Fenglong Xie, Junjie Li et al.· 3 citations· ⚡1
This work proposes Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling and is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss.
Pei-Jie Chen, Zhuanling Zha, Zhipeng Nie et al.· 0 citations
AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.
Zi-Yang Ma, Zhi-Kang Niu, Wen-Ming Tu et al.· 2 citations· ⚡1
Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned featur...
Xiao-Yu Yang, Arthur Hinsvark, Antonios Alexos et al.· 0 citations
This work presents TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU, and is designed primarily for English and German, with additional multilingual support.
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning,...
Jian Chen, Zhang You, M. Vinton· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.