Skip to content

Audio Active Learning With Noisy Labels

2026 · IEEE Transactions on Audio, Speech, and Language Processing · Vol 34, pp. 4001-4014 · 0 citations · 69 references

Abstract

Audio annotation is particularly costly and prone to errors due to the temporal nature and semantic ambiguity in audio perception. Active learning (AL) addresses this by iteratively selecting the most informative samples from an unlabeled pool for expert labeling, thereby maximizing model performance with minimal annotation effort. In the standard continual fine-tuning framework widely adopted in audio AL, the model trained in each cycle serves as the initialization for the next, preserving accumulated knowledge to enhance both sample selection and model accuracy. However, we discover a critical limitation when annotation noise is inevitably introduced: this default continual fine-tuning approach becomes susceptible to Primacy Bias — a phenomenon where early-learned patterns persistently influence subsequent learning. Our experiments show that this bias causes the model to overfit noise more rapidly in later cycles. To our knowledge, this represents the first comprehensive study specifically addressing Active Learning with Noisy Labels (ALNL) in the audio domain. To address this issue, we introduce a re-initialization strategy for ALNL scenarios. Our experiments demonstrate that periodically resetting model parameters preserves the model’s ability to learn from clean samples. Furthermore, we propose the Self-Purify Active Learning (SPAL) method, which dynamically identifies potential label noise via training loss modeling and supports either human-in-the-loop correction or automated label refurbishment. Extensive experiments on underwater acoustics, general audio, and speech datasets demonstrate the effectiveness of our framework against label noise in AL scenarios.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Sample-Conditioned Representation Selection for Audio Few-Shot Learning

Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propos...

Feng-Rui Liu, Ning-Xin Shen, Yi Li et al. · 0 citations
Preprint Aug 2026

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

SCoPE is introduced, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other, and derives an exact condition for when this competition removes an FCA in a two-label fit.

J. Jeong, Junho Yoon, Hyunju Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Tracing Audio Grounding and Answer Selection in Audio LLMs

Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve...

Hyebin Cho, Suho Yoo, Jihoo Jung et al. · 0 citations
Preprint Sep 2026

Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...

Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al. · 0 citations
Preprint Aug 2026

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.

A. Shukla, R. Thakur, Aryan Das et al. · 0 citations
Aug 2026

Certainty-Aware Partition and Sufficient Utilization of Noisy Samples

CAPSUN is proposed, a robust framework that improves the precision of clean sample selection and mitigates distribution bias through alignment among subsets, and designs a distribution alignment module to adjust the class distribution contrast of labeled and unlabeled subsets to mitigate class distribution discrepancie...

Jia Zhang, Gao-Xia Jiang, Sen-Yu Hou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.