Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can faithfully realize. While recent reasoning m...