Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. This paper introduces"cached LLM probability retrieval,"which involves querying a local teacher LLM offline to obtain...
This work proposes CLAP stacking, a lightweight method that combines multiple CLAP embedding spaces that improves SRCC from 0.5521 for the best single MS-CLAP pipeline to 0.5934 and reduces MSE from 3.2500 to 2.9418.
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, makin...
Pengcheng Wang, Sheng Li, Ji-Yi Li et al.· 0 citations
D DuplexGen is presented, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics and produces conversational dynamics closer to the real-dialogue reference distribution than the stitching baselines evaluated in this study.
Pengcheng Wang, Sheng Li, JiyiLi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.