CoLMbo-SV, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports, substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an e...
Massa Baali, Sarthak Bisht, Zi-Yue Qiu et al.· 0 citations
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the...
Wei Chu, Yuan-Zhe Dong, Ke Tan et al.· 0 citations
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference...
Sheng-Tse Lin, Si-Yuan Zhai, Chien-Liang Kuo et al.· 0 citations
Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse, is introduced.
Massa Baali, Bhiksha Raj· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.