Towards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval
This model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms.
Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu et al.
· 0 citations