Skip to content
Open access

Latency-Aware Hybrid Transformer–Capsule Network for Audio-Visual Emotion Recognition in Edge–Fog–Cloud Environments

Jul 2026 · Algorithms · Vol 19, pp. 626 · 0 citations · 20 references

TL;DR

This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment that incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources.

Abstract

Audio-visual emotion recognition (AVER) is central to affective computing systems that require reliable, real-time interpretation of human emotions. However, many existing multimodal models treat feature learning and deployment efficiency separately, limiting their ability to preserve hierarchical facial relationships, capture long-range speech dynamics, and operate with low latency in distributed settings. This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment. The visual stream employs a CNN–Capsule branch to retain spatial hierarchies in facial expressions, while the audio stream uses a CNN–Transformer branch to learn local spectral patterns and long-range temporal dependencies from speech. A cross-modal Transformer fusion module integrates complementary emotional cues, and a latency-aware task-allocation mechanism allocates preprocessing, inference, and training-related operations across edge, fog, and cloud layers according to workload, node capacity, and communication delay. Unlike approaches that optimize multimodal representation learning and distributed deployment as separate problems, the proposed framework adopts a deployment-aware co-design in which spatial visual representation, temporal acoustic modeling, multimodal interaction, and deterministic latency-aware task allocation are coordinated within a unified processing pipeline. The framework is evaluated on RAVDESS, CREMA-D, and SAVEE using a subject-independent protocol. Experimental results show an average accuracy of 91.5%, an F1-score of 90.7%, an MCC of 0.894, and an AUC of 0.950. The framework further incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources. Physical-device deployment and comprehensive resource profiling remain subjects for future validation.

Read PDF

Similar papers

Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations
Open access Oct 2026

A Context-Aware Multimodal Transformer for the Affective Monitoring of Human Cognitive Behavior Using Facial and Speech Emotion Recognition

Prolonged tracking of affective states empowers caregivers, educators, and medical staff to identify initial signs of distress. This is particularly vital for vulnerable individuals, such as the elderly and children, who may struggle to verbally express their emotional needs. This paper proposes a compact audio--visual...

Dhanraj, Arun Biradar · 0 citations
Open access Aug 2026

Lightweight and robust audio-visual emotion recognition via multi-scale mamba temporal modeling and quality-aware expert fusion

Experiments demonstrate that the proposed lightweight audio-visual emotion recognition framework achieves competitive recognition accuracy with significantly fewer parameters and lower computational cost.

Tianxing Zhang, Hadi Affendy Bin Dahlan, Fadhilah Rosdi et al. · 0 citations
Book Open access Oct 2026

Multi-Scale Spatiotemporal EEG and Self-Supervised Audio Fusion: A Mixture-of-Experts Approach to Continuous Affect

Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal...

Ali Amini, Sarmad Maqsood, Irfan Abbas et al. · 0 citations
Book Open access Oct 2026

Light-ED: Lightweight Multimodal Emotion Detection using Enhanced EfficientNet

Emotion recognition plays a key role in affective computing and human–computer interaction, where understanding emotions from multimodal signals such as facial expressions and speech remains challenging. Most existing methods treat data fusion and classification as separate stages, limiting performance and efficiency....

Wamika Jha, Mea Wang, U. Alim et al. · 0 citations
Conference Aug 2026

Emotion-driven music retrieval via cross-modal audio-visual representation learning

Music-assisted therapy relies on the accurate perception of a user's affective state to deliver contextually appropriate acoustic stimuli. However, existing emotion-driven retrieval methods frequently encounter representational mismatches between raw multimodal behavioral signals and therapeutic music indexing, which s...

Chuanfan Guo, Xiang Gao, Qianqian Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.