This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.
Abstract
A major limitation in the speech sciences is access to naturalistic data in experimental settings and the difficulty of translating laboratory designs to real-world contexts. Researchers studying speech production or perception often rely on in-lab recordings, which constrain the sociolinguistic contexts examined and limit ecological validity. The Toolkit for Acoustic–Phonetic Analysis (TAPA) is an open-source pipeline that automates the acquisition, transcription, speaker diarization, forced alignment, and per-segment acoustic analysis of naturalistic, single/multi-speaker audio. The current release supports vowel formant extraction, stop voice onset time, and fricative spectral moments. We demonstrate TAPA on the 2016 U.S. presidential debate, extracting nearly 33,000 segments from a 90-min recording, and validate each measurement type against hand-coded annotation. Vowel formant agreement with expert measurements was high (F1 r = 0.89, F2 r = 0.87). Stop VOT showed reliable aggregate means but poor per-token agreement (r = − 0.04) because of a training–deployment mismatch in the neural VOT classifier. Fricative spectral standard deviation agreed strongly with hand-coded values overall (r = 0.80), and center of gravity agreed strongly for sibilants (/s/ r = 0.87, /ʃ/ r = 0.94), while non-sibilant moments diverged systematically. These findings suggest that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts’ supervision.
Japanese vowel duration is a key phonetic feature for distinguishing lexical meaning and stylistic register, but its realization patterns across different speech contexts remain insufficiently quantified. This study constructs a bimodal corpus involving 20 native Japanese speakers, including 50 minutes of spontaneous c...
Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic"diluted"accent: a phonetic profile attempting to represent all regional distrib...
Pedro H. L. Leite, Pedro Benevenuto Valadares, L. Biscainho· 0 citations
Automated Speaking Assessment of non-native speech must effectively evaluate prosody, including speech rhythm, to align with human perception. However, commonly employed rhythm metrics rely on segmental duration, requiring an additional alignment step, which is error-prone in non-native speech containing disfluencies a...
Advancement of zero-shot text-to-speech synthesis is currently hindered for low-resource languages by the scarcity of large-scale, high-fidelity speech datasets. Traditional alignment-based methods require rare verbatim transcripts, while standard in-the-wild pipelines often rely on single-model automatic speech recogn...
Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al.· 0 citations
This study investigates the role of prosodic-acoustic features in distinguishing two inland varieties of Brazilian Portuguese, Paraíba (PB) and São Paulo (SP), and examines how these features influence automatic speech recognition (ASR) performance. Grounded in sociophonetic theory and dialect-aware speech modeling, th...
Leônidas Silva, L. Tenani, João Marcelo Monte· Cadernos de Linguística· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.