This work proposes CLAP stacking, a lightweight method that combines multiple CLAP embedding spaces that improves SRCC from 0.5521 for the best single MS-CLAP pipeline to 0.5934 and reduces MSE from 3.2500 to 2.9418.
FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction, is introduced, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction that substantially outperforms all baselines and remains robust under real ASR transcript...
Mithilesh Vaidya, Stephen W. Bailey, Sumukh Badam et al.· 0 citations
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...
Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al.· 0 citations
Mizar, a 159.3M-parameter ALM, is introduced, a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM that surpasses the previous best-performing ALM below 200M parameters on all three benchmarks.
Kai-Yang Li, Shaobo Han, Yue Tian et al.· 0 citations
Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, unde...
Y. El Kheir, Xin Wang, Wan-Ying Ge et al.· 1 citation
MuLA-Bench exposes conditional failure patterns that a single long-context score does not capture, and evaluates ten audio-language models and conducts pooled diagnostics on a fixed eight-model cohort.
Ze-Yu Yang, Xin-Yu Zhang, Zi-Bo Bi et al.· 1 citation
This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.
Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.