Exposing Weaknesses in Emotion Recognition in Conversations
An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.
Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al.
· 0 citations