Skip to content

Stacking Complementary CLAP Embeddings for Improving Text-Audio Alignment Correspondence Scoring

· 0 citations · 19 references

TL;DR

This work proposes CLAP stacking, a lightweight method that combines multiple CLAP embedding spaces that improves SRCC from 0.5521 for the best single MS-CLAP pipeline to 0.5934 and reduces MSE from 3.2500 to 2.9418.

View source

Similar papers

#machine learning Preprint Sep 2026

FuseAlign: Forced Alignment in the Wild

FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction, is introduced, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction that substantially outperforms all baselines and remains robust under real ASR transcript...

Mithilesh Vaidya, Stephen W. Bailey, Sumukh Badam et al. · 0 citations
Preprint Sep 2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of...

Le-Jun Min, Jun-Yu Dai, Rui-Chen Zheng et al. · 0 citations
#natural language process... Preprint Sep 2026

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Mizar, a 159.3M-parameter ALM, is introduced, a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM that surpasses the previous best-performing ALM below 200M parameters on all three benchmarks.

Kai-Yang Li, Shaobo Han, Yue Tian et al. · 0 citations
Preprint Aug 2026

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, unde...

Y. El Kheir, Xin Wang, Wan-Ying Ge et al. · 1 citation
#small language model Preprint Sep 2026

Towards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval

This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.

Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.