Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a datas...
Nishit Anand, Jia-Qi Su, Ke Chen et al.· 1 citation
D DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding, and shows that even the best voice agents leave substantial room for improvement on all 3 axes.
It is shown that steering vectors can be sparsified by up to 85-96% while retaining most performance, and that different steering methodologies agree on a subset of important dimensions, and that different steering methodologies agree on a subset of important dimensions.
Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha· arXiv.org· 6 citations
VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjec...
This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the eval...
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.