Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets'rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members'mean representations. Av...
Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an...
MuLA-Bench exposes conditional failure patterns that a single long-context score does not capture, and evaluates ten audio-language models and conducts pooled diagnostics on a fixed eight-model cohort.
Ze-Yu Yang, Xin-Yu Zhang, Zi-Bo Bi et al.· 1 citation
This work presents a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript, and indicates that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transc...
Theo Lebryk, David Ayllón, Alice Baird et al.· 0 citations
Modeling language use as a joint distribution over meanings, contexts, and utterances, upper bounds are derived on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance.
SpeechCritic is introduced, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons, and shows that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish.
Ming-Yue Huo, Shivam Mehta, Bhavin Jawade et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.