Voice Cloning: A Survey of Zero-Shot and Controllable Speech Synthesis
Abstract
This survey examines zero-shot voice cloning through the linked views of representation, generation, control, and deployment. Rather than treating recent systems as isolated milestones, we organize the literature around the technical decisions that shape modern speaker-conditioned synthesis: neural codec design, acoustic-token modeling, autoregressive and non-autoregressive decoding, diffusion and flow-matching objectives, multilingual conditioning, streaming constraints, preference alignment, controllability, and safety. We review representative systems including YourTTS, VALL-E, Voicebox, F5-TTS, MaskGCT, Seed-TTS, CosyVoice 2, MiniMax-Speech, GLM-TTS, Spark-TTS, and Qwen3-TTS, and we compare how these systems trade off naturalness, speaker similarity, latency, controllability, and deployment risk. The survey is unified by a single frame, which is a deployment-oriented reading of the codec language model era. That frame is delivered through four contributions. We provide a four-axis taxonomy across representation, generation, control, and deployment. We provide a qualitative architectural comparison of representative systems along axes that headline metrics obscure. We provide a structured treatment of evaluation reliability that maps automatic and human metrics to the quality dimensions they actually capture. We provide a consolidated safety and governance perspective that couples technical mechanisms with dataset licensing and a pre-deployment checklist. Three field-level shifts recur across these views. The first is the move from waveform or spectrogram prediction toward discrete acoustic-token generation. The second is the transition from offline quality-first models toward streaming and conversational architectures. The third is the growing need for trustworthy evaluation, watermarking, consent-aware deployment, and demographic robustness. We conclude by identifying open problems in benchmark standardization, speaker-similarity assessment, multilingual low-resource performance, preference-aligned synthesis, and secure real-world use.