Recent improvements in voice-cloning speech generation systems raise concerns about misuse by malicious actors to impersonate others and spread misinformation. Detecting such tampering is difficult, since deepfakes in the wild may be created by different generative models in a wide range of languages. Furthermore, the...
Yuan Tseng, Aishwarya Fursule, Andrew Zijun Ma et al.· 0 citations
Multi-user voice agents must track who said what across dialogue sessions. Text LLMs are attractive backbones for such agents, but transcripts alone do not expose acoustic speaker identity, leaving the model without a persistent reference for linking information to speakers across sessions. We address this gap by intro...
Run-Qiu Xu, Zhi-Sheng Zheng, David Harwath· 0 citations
This paper presents a scalable method that measures consonant contribution using acoustic masking, and relates MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load.
Eunjung Yeo, Kwanghee Choi, K. Kothadia et al.· 0 citations
Multiparty Bench is introduced, the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts and assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness.
Yi-Jen Shih, S. Kuan, Guan-Ting Lin et al.· 2 citations· ⚡1
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR ar...
Zhi-Sheng Zheng, Xiaohang Sun, Zhu Liu et al.· 0 citations
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imp...
Yi-Jen Shih, Pu-Yuan Peng, Abdel-rahman Mohamed et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.