This work introduces Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning and reveals that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
Abstract
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.
The objective of this paper is to improve and analyze auditory reasoning in large audio–language models for audio question answering (AQA), where a model must infer the correct answer from acoustic evidence and textual answer options. Although reinforcement learning (RL) with verifiable rewards has recently improved re...
In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...
Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al.· 0 citations
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-im...
Yu-Xiang Wang, Sheng-Bo Cai, Ying Shen et al.· 0 citations
Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The chall...
Xing-Ming Shui, Da-Peng Chen, Bo-Wei Liu et al.· 0 citations
SCoPE is introduced, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other, and derives an exact condition for when this competition removes an FCA in a two-label fit.
J. Jeong, Junho Yoon, Hyunju Kim et al.· 0 citations
STAG is introduced, to the authors' knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs, and provides behavioral support for the faithfulness and selectivity of the explanations.
Lucia Cascone, V. Fraenza, Michele Nappi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.