It is shown that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks and learn representations which encode words with reasonable fidelity independently of local phonetic content.
Abstract
Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness have largely focused on probing their representations'ability to discriminate phonemes and words. However, discriminative ability for words need not imply specialized representation of words per se. Good discrimination of words may be explained by good encoding of word form (phonemes) rather than form-independent word representations encoding identity or syntactic/semantic properties. By partialling out phoneme information using residualization, we show that, in later layers, HuBERT and wav2vec 2.0 do in general learn representations which encode words with reasonable fidelity independently of local phonetic content. We show that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks.
This work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.
Gabriel Pirlogeanu, Dan Oneata, H. Cucu et al.· 0 citations
This work evaluates two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER).
It is observed that the correlation between speakers'L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models.
Tingyu Cheng, L. Clemmensen, Sneha Das· 0 citations
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on ea...
Diego Torres, Axel Roebel, Nicolas Obin· 0 citations
This work proposes Agentic-GER, an LLM-based agent for terminology correction in long-form speech, which uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses.
Yan-Qiao Zhu, Wu-Peng Wang, Zhifu Gao et al.· 0 citations
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, clos...
Maureen de Seyssel, Jie Chi, Zakaria Aldeneh· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.