Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget and apply latent adversarial post-training inspired by Flow2GAN for few-step generation.
Abstract
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.
This work proposes FlowSep2, a text-conditioned flow-matching generative model for LASS, which learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query.
Yiitan Yuan, Xu-Bo Liu, Haohe Liu et al.· 0 citations
Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. Howev...
Run-Wu Shi, Kai Li, Yu-Jin Wang et al.· 0 citations
CuteTTS is presented, a compact continuous-autoregressive TTS system that reconciles high-fidelity generation with the latency demands of real-time interaction and introduces guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student.
Yu-Qian Zhang, Yao Shi, Kexin Huang et al.· 0 citations
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and appl...
Tiernon Riesenmy, Zhang You, Gautam Bhattacharya et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.