Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis...
Lu-Jia Bao, Qian Chen, Luyao Cheng et al.· 1 citation
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...
Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al.· 2 citations
This work proposes LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration, and introduces a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints.
Tianrui Pan, Qinglin Zhang, Chong Deng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.