Skip to content
Open access

Fine-Grained Pose-Aware Visual Fusion for Emotion Recognition in Conversational Video Streams

Jul 2026 · Mathematics · 0 citations · 36 references

Abstract

Despite the rapid advancement of Emotion Recognition in Conversation (ERC), prevailing systems that primarily integrate language and speech exhibit substantial performance disparities on underrepresented emotion classes (e.g., Fear, Disgust). This study investigates whether fine-grained non-verbal visual modalities (facial action units, hand gestures, and body pose) can effectively mitigate these biases. We propose a multi-stream fusion architecture combining language, speech, and engineered pose-aware visual features, trained with class-imbalance-aware objectives. Experiments on MELD demonstrate that hybrid pose augmentation improves F1 on the least frequent classes: Fear +9.19%, Disgust +6.07%, Sadness +6.46%. We achieve an overall weighted F1 of 68.79%, competitive with recent state-of-the-art systems while uniquely targeting minority-class debiasing. These results establish fine-grained body language as a critical debiasing signal, recovering accuracy on the subtle expressions that text and speech alone fail to capture.

Read PDF