Skip to content

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.20166 · 1 citation · 37 references
Computer Science

TL;DR

This work introduces Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning and reveals that increasingly fine-grained auditory descriptions emerge naturally from game pressure.

Abstract

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.

View source

Similar papers

Open access 2026

Enhancing Auditory Reasoning in Large Audio–Language Models Via Supervised Fine-Tuning and Reinforcement Learning With Verifiable Rewards

The objective of this paper is to improve and analyze auditory reasoning in large audio–language models for audio question answering (AQA), where a model must infer the correct answer from acoustic evidence and textual answer options. Although reinforcement learning (RL) with verifiable rewards has recently improved re...

Jian Wang, Sheng-Yan Hao · 0 citations
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations
#machine learning Preprint Sep 2026

EvoAudio: Recursive Self-Improvement for Audio Understanding

Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-im...

Yu-Xiang Wang, Sheng-Bo Cai, Ying Shen et al. · 0 citations
Preprint Sep 2026

OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The chall...

Xing-Ming Shui, Da-Peng Chen, Bo-Wei Liu et al. · 0 citations
Preprint Aug 2026

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

SCoPE is introduced, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other, and derives an exact condition for when this competition removes an FCA in a two-label fit.

J. Jeong, Junho Yoon, Hyunju Kim et al. · 0 citations
#machine learning Preprint Sep 2026

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

STAG is introduced, to the authors' knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs, and provides behavioral support for the faithfulness and selectivity of the explanations.

Lucia Cascone, V. Fraenza, Michele Nappi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.