CoLMbo-SV, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports, substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an e...
Massa Baali, Sarthak Bisht, Zi-Yue Qiu et al.· 0 citations
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference...
Sheng-Tse Lin, Si-Yuan Zhai, Chien-Liang Kuo et al.· 0 citations
Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse, is introduced.
A mechanistic analysis of paralinguistic information in four open source models using the Expresso dataset with controlled speaking styles identifies a gap between what models encode and what they use, highlighting a key limitation in current audio language models.