Skip to content

Forced alignment of Istro-Venetian speech via neural versus non-neural models

Aug 2026 · Journal of the Acoustical Society of America · 0 citations

Abstract

Forced aligners are an essential tool to facilitate access to the word and phone level of speech data, but they rely on acoustic models, pronunciation dictionaries, and complete word-level transcriptions that may not exist for low-resource languages. Alternatively, word-level alignments can be extracted from neural network-based tools for automatic speech recognition like Wav2Vec-BERT 2.0, avoiding the need to create bespoke acoustic models and dictionaries. This paper compares the ability of the Montreal Forced Aligner (MFA) and Wav2Vec-BERT 2.0 to produce word-level alignments for the low-resource variety Istro-Venetian, spoken in Croatia. Results indicate that MFA, which uses Gaussian mixture models with hidden Markov models, significantly outperforms Wav2Vec-BERT 2.0 on forced alignment. However, we show that Wav2Vec-BERT 2.0 can produce transcriptions with a word error rate of only 14.4%. Neural models can thus aid research on low-resource languages by creating transcriptions that can be used as input to forced aligners, and with sufficient data they may also perform acceptably on forced alignment.

View source