Skip to content

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Sep 2026 · 2 citations · ⚡ 1 influential · 89 references
Computer Science

TL;DR

AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context, is introduced and leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing is demonstrated.

Abstract

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.

View source

Similar papers

#natural language process... Preprint Sep 2026

AURA: Uncertainty-Routed Activation Editing for Acoustic Grounding in Speech Foundation Models

Attention encoder-decoder (AED) Speech Foundation Models achieve strong ASR performance but can generate acoustically unsupported text when inputs contain no speech, weak acoustic evidence, or unreliable transcription. We propose AURA: Activation-editing with Uncertainty-Routed Adaptation, an ultra-efficient representa...

Natarajan Balaji Shankar, Zilai Wang, Zi-Han Wang et al. · 0 citations
Preprint Sep 2026

Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

A unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration is proposed, highlighting post-training as a practical approach to extending existing speech synthesis models.

Lian-Ru Gao, Yu-Jie Guo, Yong Qin · 0 citations
#natural language process... Preprint Sep 2026

DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning...

Ya-Yue Deng, Ding-Dong Wang, Yuxuan Hu et al. · 0 citations
Preprint Aug 2026

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

FireRedAudio is introduced, a general-purpose audio language model with a shared 9B-parameter LLM that achieves competitive or leading performance in audio understanding and multilingual ASR, strong content accuracy and speaker preservation in zero-shot TTS, leading instruction following in Instruct TTS, and substantia...

Fei-Yu Shen, Fenglong Xie, Junjie Li et al. · 3 citations · ⚡1
#artificial intelligence Preprint Sep 2026

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-alig...

Yi-Zhao Li, Pu-Sen Gao, Ming Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.