Skip to content

Towards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval

Sep 2026 · 0 citations · 25 references
Engineering

TL;DR

This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.

Abstract

Audio deepfake forensics is moving beyond a simple real-or-fake verdict toward attribution: which system generated the audio clip? Most source-attribution methods cast this as closed-set classification, so they cannot name a generator that was absent from training, a gap that widens with every newly released text-to-speech (TTS) system. We instead frame attribution as cross-modal retrieval: each generator is described in natural language, and a clip is attributed by retrieving the description closest to it in a shared audio-text embedding space. Adding a new system then takes nothing more than writing its description, with no retraining and no new classifier head. Our model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms. We evaluate on MLAAD v9 (140 TTS models, 51 languages) under 10-fold leave-models-out cross-validation. For generators it has never encountered before, the model reaches a model-level mean reciprocal rank (MRR) of 58.4%. Even when the correct model is not identified, the audio clip is often matched to systems that share the true generator's vocoder, acoustic model, or architecture. Because the same embedding space also answers natural-language attribute queries, one set of descriptions covers both open-set attribution and attribute-level forensic profiling.

View source

Similar papers

Preprint Aug 2026

Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, unde...

Y. El Kheir, Xin Wang, Wan-Ying Ge et al. · 1 citation
Preprint Sep 2026

Tracing and Relearning Detection Evidence in Text-to-Speech Systems

Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel ca...

Eunji Shin, Kyudan Jung, Jihwan Kim et al. · 0 citations
#machine learning Preprint Sep 2026

AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models

Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and q...

Pooneh Mousavi, Amir Ivry, M. Ravanelli et al. · 0 citations
#machine learning Preprint Sep 2026

Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

Voice-cloning fraud increasingly relies on surgical injection: a genuine conversation in which only one or two sentences are replaced by synthetic speech. Utterance-level deepfake detectors emit a single real/fake label per clip and cannot report where the synthetic speech lies. We formalise this as Temporal Deepfake L...

Soumyadeep Roy · 0 citations
#natural language process... Preprint Aug 2026

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

This work presents the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels, and assesses fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder mod...

Meruyert Aristombayeva, Jason Samuel Lucas, Chaewan Chun et al. · 1 citation

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.