Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syllabic rhythm, and turn interaction through eight interpretab...
S. Satish, Erica Cooper, Patrícia Schmidtová et al.· 0 citations
This work evaluates state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias.
Maike Zufle, Patrícia Schmidtová, Vilém Zouhar et al.· 0 citations
This position paper draws on AAC as a setting where speech AI failures are most visible and their stakes highest, alongside other underserved speakers - people who stutter, multilingual speakers, and non-binary and transgender users.
Maria Teleki, Kimi Wenzel, Anna Seo Gyeong Choi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.