This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregate...
Fei Liu, Yang Ai, Xiao-Hang Jiang et al.· 0 citations
Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time...
An EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels, which achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.
Chen-Yuan Ning, Yang Ai, Hui-Peng Du et al.· 0 citations
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an en...
Ye-Xin Lu, Xin Wang, Yang Ai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.