Skip to content

Do speech foundation models really learn words?

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

It is shown that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks and learn representations which encode words with reasonable fidelity independently of local phonetic content.

Abstract

Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness have largely focused on probing their representations'ability to discriminate phonemes and words. However, discriminative ability for words need not imply specialized representation of words per se. Good discrimination of words may be explained by good encoding of word form (phonemes) rather than form-independent word representations encoding identity or syntactic/semantic properties. By partialling out phoneme information using residualization, we show that, in later layers, HuBERT and wav2vec 2.0 do in general learn representations which encode words with reasonable fidelity independently of local phonetic content. We show that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks.

View source

Similar papers

#machine learning Preprint Sep 2026

A Comprehensive Study of Content Representations for Speech Synthesis

Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on ea...

Diego Torres, Axel Roebel, Nicolas Obin · 0 citations
#natural language process... Preprint Sep 2026

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, clos...

Maureen de Seyssel, Jie Chi, Zakaria Aldeneh · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.