FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
A frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space and improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple.