Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whisper-Tiny and Whisper-Small with EnCodec...
Song-ha Jo, Sehyun Lee, Soyoon Kim et al.· 0 citations
OmniEvaluator connects existing inference engines and curated evaluation libraries at a higher level, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface.
Ho-Dong Lee, Sanghee Park, Dohoon Ryu et al.· 0 citations
OmniACBench, a benchmark for evaluating context-grounded acoustic control in omni-modal models, is introduced and three common failure modes are identified-weak direct control, failed implicit inference, and failed multimodal grounding-providing insights for developing models that can verbalize responses effectively.
Seunghee Kim, B. Park, Kyudan Jung et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.