Skip to content

AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth

Sep 2026 · 0 citations · 25 references
Computer Science Engineering

Abstract

Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets'rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.

View source

Similar papers

#machine learning Preprint Sep 2026

Do Audio Language Models Hear and Read Distinctive Features Alike?

Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members'mean representations. Av...

Yuanhao Chen, Peter Chin · 0 citations
#natural language process... Preprint Sep 2026

mu-bench: A Multilingual Utterance Transcription Benchmark

Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an...

Andrea Li, Soham Ray · 0 citations
Preprint Aug 2026

Towards Quantifying Benchmark Optimization in ASR Models

This work presents a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript, and indicates that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transc...

Theo Lebryk, David Ayllón, Alice Baird et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

SpeechCritic is introduced, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons, and shows that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish.

Ming-Yue Huo, Shivam Mehta, Bhavin Jawade et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.