Skip to content
Open access

Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition

Aug 2026 · Algorithms · Vol 19, pp. 669 · 0 citations · 23 references

TL;DR

A large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability and adaptively fuses predictions from the instruction-tuned VLM and the distillation module for final emotion recognition.

Abstract

Visual emotion recognition plays a critical role in human–computer interaction and mental health applications. Although existing Vision–Language Models (VLMs) alleviate the limitations of conventional vision models in high-level semantic understanding, they still face three main challenges: limited emotional semantic understanding, insufficient visual emotional perception capability, and high computational costs when deploying both models simultaneously. To address these issues, a large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability. Furthermore, we transfer the visual emotion discrimination knowledge of a conventional vision model into the VLM using a distillation module while keeping the VLM frozen during training, which reduces the computational costs. Following that, we design a fusion and prediction module that adaptively fuses predictions from the instruction-tuned VLM and the distillation module for final emotion recognition. The experimental results on the Abstract, ArtPhoto, Emotion6, and FI datasets demonstrate that VERLADF achieves recognition accuracies of 36.71%, 52.38%, 74.73%, and 79.69%, respectively, significantly outperforming many methods in the literature and demonstrating the effectiveness of the proposed framework.

Read PDF

Similar papers

Preprint Aug 2026

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained wit...

Ahmed-Shehab Khan, Zhiyuan Li, Yan Tong · 0 citations
#human-computer interacti... Preprint Aug 2026

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce One...

Jiahao Huang, Zheng Lian, Jingyi Zhang et al. · 0 citations
Open access Aug 2026

HYBRID DEEP LEARNING MODEL FOR EMOTION RECOGNITION SYSTEM WITH PROPOSED WDCRF TECHNIQUE

Facial Emotion Recognition (FER) is widely used for many applications and Deep Learning (DL) has significantly advanced performance of FER, yet many models fail to capture temporal coherence and realistic emotional transitions in visual sequences. This study introduces a hybrid of Residual Network 18-Layer (ResNet-18),...

Durgesh Kumar Kotangle, H. S. Hota · 0 citations
Sep 2026

CLMER: A Framework for Contrastive Learning-Based Multimodal Emotion Recognition.

Emotion recognition plays a crucial role in human-computer interaction and affective computing, yet its effectiveness is limited by the difficulty of integrating heterogeneous modalities with fundamentally different structures, such as physiological signals and visual data. In this article, we propose CLMER, a contrast...

Shuang Niu, Jian He, Yu Liang et al. · 0 citations
Open access Aug 2026

An Attention-Guided Framework for Feature-Level and Decision-Level Fusion in Multimodal Emotion Recognition

A comparative analysis of four multimodal configurations of early Fusion without attention, early Fusion with attention, late Fusion without attention, and late Fusion complemented by attention provides a methodological framework that can be used to develop more effective and understandable MER systems using a systemat...

Chintan Chatterjee, Brijesh Bhatt · 0 citations
Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.