Skip to content
Preprint

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Aug 2026 · 0 citations · 33 references
Computer Science Engineering

TL;DR

Preliminary analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation, indicating that phonological information from synchronized audio can be partially transferred to articulatory models.

Abstract

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.

View source

Similar papers

Open access Sep 2026

Visual speech enhances phoneme separability in human superior temporal gyrus

Using intracranial recordings from human superior temporal gyrus, it is shown that congruent visual speech does not simply amplify all speech-related information, instead, visual input selectively enhances the separability of confusable phoneme-level representations and improves word identity decoding, while leaving th...

Yi-Ke Li, Iain DeWitt, Jonathan R. Brennan et al. · 0 citations
#natural language process... Preprint Oct 2026

Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis

Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) acros...

Abner Hernandez, T. A. Vergara, Andreas K. Maier et al. · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Preprint Sep 2026

ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...

B. Halpern, Thomas B. Tienkamp, D. Abur et al. · 0 citations
Preprint Aug 2026

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

A linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing is introduced and suggests that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition i...

Sophia Riaz, Haoze Zheng, Amos Roche et al. · 0 citations
Preprint Sep 2026

Articulatory Source-Filter TTS: Physically Grounded Control through Vocal Tract Kinematics

Modern neural text-to-speech (TTS) systems achieve remarkable acoustic fidelity but act as black boxes, offering little interpretable control over the vocal tract filter. We propose a controllable source-filter TTS architecture grounded in articulatory kinematics. An Acoustic-to-Articulatory Inversion (AAI) model, enha...

Jesuraja Bandekar, Shinji Watanabe, P. Ghosh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.