Skip to content

A Dataset for Evaluating ASR on Specialized Vocabulary

2026 · International Conference on Language Resources and Evaluation · pp. 470-480 · 0 citations · 38 references
Computer Science

TL;DR

A linguistically curated bilingual dataset comprising 13,846 utterances distributed across synthetic and literature-derived subsets, with OOV rates reaching up to 100%, and a diagnostic evaluation framework that partitions recognition performance into Biased Word Error Rate (B-WER), which targets domain-specific jargon, and Unbiased Word Error Rate (U-WER), which focuses on general vocabulary is introduced.

View source

Similar papers

#natural language process... Preprint Sep 2026

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an...

Hritika Sharma, Thibault Bañeras-Roux, Alessandra Pinto et al. · 0 citations
Preprint Aug 2026

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

The results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil et al. · 0 citations
#natural language process... Preprint Sep 2026

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often...

Shivam Singh, Aditya Yadavalli, Catherine Arnett et al. · 0 citations
Preprint Sep 2026

Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval

Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issu...

Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi et al. · 0 citations
#natural language process... Preprint Sep 2026

mu-bench: A Multilingual Utterance Transcription Benchmark

Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an...

Andrea Li, Soham Ray · 0 citations
#natural language process... Preprint Sep 2026

Benchmarking Automatic Speech Recognition Tools for Iberian Languages

Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Por...

Fernando López, Pablo Gómez, David Solans et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.