This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment that incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources.
Abstract
Audio-visual emotion recognition (AVER) is central to affective computing systems that require reliable, real-time interpretation of human emotions. However, many existing multimodal models treat feature learning and deployment efficiency separately, limiting their ability to preserve hierarchical facial relationships, capture long-range speech dynamics, and operate with low latency in distributed settings. This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment. The visual stream employs a CNN–Capsule branch to retain spatial hierarchies in facial expressions, while the audio stream uses a CNN–Transformer branch to learn local spectral patterns and long-range temporal dependencies from speech. A cross-modal Transformer fusion module integrates complementary emotional cues, and a latency-aware task-allocation mechanism allocates preprocessing, inference, and training-related operations across edge, fog, and cloud layers according to workload, node capacity, and communication delay. Unlike approaches that optimize multimodal representation learning and distributed deployment as separate problems, the proposed framework adopts a deployment-aware co-design in which spatial visual representation, temporal acoustic modeling, multimodal interaction, and deterministic latency-aware task allocation are coordinated within a unified processing pipeline. The framework is evaluated on RAVDESS, CREMA-D, and SAVEE using a subject-independent protocol. Experimental results show an average accuracy of 91.5%, an F1-score of 90.7%, an MCC of 0.894, and an AUC of 0.950. The framework further incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources. Physical-device deployment and comprehensive resource profiling remain subjects for future validation.
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.· Signal, Image and Video Proc...· 0 citations
Prolonged tracking of affective states empowers caregivers, educators, and medical staff to identify initial signs of distress. This is particularly vital for vulnerable individuals, such as the elderly and children, who may struggle to verbally express their emotional needs. This paper proposes a compact audio--visual...
Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal...
Ali Amini, Sarmad Maqsood, Irfan Abbas et al.· Proceedings of the 28th Inte...· 0 citations
Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency....
Wamika Jha, Mea Wang, U. Alim et al.· Proceedings of the 28th Inte...· 0 citations
Music-assisted therapy relies on the accurate perception of a user's affective state to deliver contextually appropriate acoustic stimuli. However, existing emotion-driven retrieval methods frequently encounter representational mismatches between raw multimodal behavioral signals and therapeutic music indexing, which s...
Chuanfan Guo, Xiang Gao, Qianqian Wang et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.