Skip to content
Preprint

Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget and apply latent adversarial post-training inspired by Flow2GAN for few-step generation.

Abstract

Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.

View source

Similar papers

Preprint Aug 2026

FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation

This work proposes FlowSep2, a text-conditioned flow-matching generative model for LASS, which learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query.

Yiitan Yuan, Xu-Bo Liu, Haohe Liu et al. · 0 citations
Preprint Sep 2026

UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation

Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. Howev...

Run-Wu Shi, Kai Li, Yu-Jin Wang et al. · 0 citations
Preprint Aug 2026

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

CuteTTS is presented, a compact continuous-autoregressive TTS system that reconciles high-fidelity generation with the latency demands of real-time interaction and introduces guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student.

Yu-Qian Zhang, Yao Shi, Kexin Huang et al. · 0 citations
Preprint Sep 2026

PLACE: Positional Latent Adaptation via Conditioned Embeddings for Binaural Audio Generation

We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and appl...

Tiernon Riesenmy, Zhang You, Gautam Bhattacharya et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.