Aug 2026· American Journal of Artificial Intelligence· Vol 10, pp. 198-208· 0 citations· 19 references
TL;DR
These findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets.
Abstract
Emotion recognition is a core component of affective computing, enabling intelligent systems to interpret human emotional states across critical applications such as healthcare, online education, and human–computer interaction. Early unimodal approaches relying solely on facial expressions, speech, or text have proven insufficient due to noise, cultural variability, and signal ambiguity, prompting a decisive shift toward multimodal integration. This systematic review, conducted following the PRISMA framework, examines this transition by analyzing 89 peer-reviewed studies selected from an initial pool of 160. The objective is to synthesize current methodologies, compare performance across modalities, and identify persistent technical and ethical barriers. Our findings reveal that multimodal systems, which fuse visual, acoustic, linguistic, and physiological signals, consistently outperform unimodal counterparts, achieving accuracy levels above 85% on benchmark datasets. Deep learning architectures particularly convolutional networks for spatial features, recurrent networks for temporal dependencies, and transformer-based models enhanced with attention mechanisms dominate the field, enabling effective dynamic weighting and fusion of heterogeneous data streams. Despite these advances, several challenges impede real-world deployment. Cross-subject and cross-session variability degrades generalizability, while data scarcity and the lack of large-scale, annotated multimodal corpora constrain model training. Computational complexity, especially in transformer-based fusion, limits edge-device feasibility, and ethical concerns surrounding privacy, demographic bias, and model interpretability remain unresolved. Future research must prioritize scalable and lightweight architectures, inclusive and culturally diverse dataset curation, and explainable AI frameworks that build user trust. Ultimately, transitioning these systems from laboratory prototypes to ethically sound, practical applications will require close interdisciplinary collaboration among computer scientists, psychologists, and ethicists, ensuring that emotion recognition technologies are not only accurate but also fair, transparent, and accessible across diverse real-world settings.
Emotion recognition plays a critical role in affective computing systems that aim to understand human behavior through observable signals. While uni-modal approaches based on audio, facial expressions, or body language provide useful cues, their performance often degrades under real-world conditions due to noise, occlusions, and modality-specific limitations. This paper presents a multi-modal emotion detection framework that integrates audio, facial, and body-language information using a learnable weights strategy. Each modality is processed by an independently trained deep learning model tailored to its feature characteristics, and predictions are fused at the score level with robustness to missing or unreliable inputs. Emotions are inferred over successive non-overlapping 3-second temporal windows, enabling time-resolved analysis of emotional dynamics. Experiments conducted on a custom multimodal dataset demonstrate that the proposed fusion model achieves an overall accuracy of 77% and provides stable per-class performance across diverse emotions. In addition to static classification, the system produces continuous emotion probability spectra and predicted emotion timelines, offering interpretable insights into emotional transitions over time. The proposed approach is particularly suited for therapeutic and behavioral assessment scenarios where robustness and temporal interpretability are essential.
K. Joshitha, Gaja Dhanush, Gudipati Thrishal et al.· International Conference on...· 0 citations
Emotion recognition has emerged as a critical component in the development of intelligent systems capable of understanding and responding to human affective states across applications such as healthcare, human-computer interaction and smart environments. Despite substantial progress, achieving accurate and generalizable emotion recognition remains challenging due to the complex, subjective and multimodal nature of human emotions. This review presents a comparative analysis of recent studies employing machine learning, deep learning, hybrid and ensemble approaches across diverse data modalities, including speech, facial expressions and physiological signals. The analysis highlights that deep learning and hybrid models significantly enhance feature representation by capturing spatial and temporal dependencies, while ensemble techniques improve classification robustness and stability. Furthermore, multimodal approaches consistently outperform unimodal systems by integrating complementary emotional cues from heterogeneous data sources. However, several limitations persist, including small and imbalanced datasets, limited generalization across diverse populations, noise in physiological signals and high computational complexity that restricts realtime deployment. Based on these observations, this study identifies key research gaps and emphasizes the need for scalable, lightweight and generalized multimodal emotion recognition frameworks. Future research should incorporate multimodal fusion framework that effectively integrate heterogeneous data modalities, robust feature learning strategies and computationally efficient architectures with improved accuracy, adaptability and practical applicability.
Jasmine George, M. D.· International Conference on...· 0 citations
Emotion recognition has emerged as a pivotal research frontier at the intersection of affective computing, signal processing, and human–computer interaction (HCI), driven by the growing demand for systems capable of perceiving, interpreting, and responding to human affective states in real time. This review synthesises the state of the art in deep learning-based emotion recognition, examining the multidisciplinary convergence of computer vision, speech processing, physiological signal analysis, and natural language understanding that underlies contemporary affect-sensing pipelines. The paper traces the evolution from handcrafted feature engineering to end-to-end representation learning, critically analysing convolutional neural networks, recurrent architectures, attention mechanisms, transformer-based models, and multimodal fusion strategies that have redefined benchmark performance on datasets such as IEMOCAP, CREMA-D, AffectNet, DEAP, and MELD. A structured taxonomy is proposed to classify existing approaches along the axes of modality (unimodal versus multimodal), learning paradigm (supervised, self-supervised, and few-shot), and application context (embodied agents, adaptive tutoring, mental health monitoring, and automotive safety). Architectural blueprints, mathematical formulations of loss functions and fusion operators, and comparative performance tables are presented to ground the discussion in empirical evidence. The review further examines real-world deployment through a case study of affect-aware conversational agents, surveys the software and hardware ecosystem supporting reproducible research, and evaluates recognition accuracy, latency, and robustness across cross-corpus and cross-cultural conditions. Persistent challenges—including data scarcity, annotation subjectivity, domain shift, and the fundamental ambiguity of emotional expression—are discussed alongside ethical considerations surrounding consent, bias, and affective privacy. The review concludes by outlining promising directions, including foundation-model-driven affective reasoning, neuro-symbolic integration, and personalised continual learning, offering a roadmap for researchers seeking to advance emotionally intelligent human–computer interaction systems.
Dhrub Kumar, Dr. Mohini Mittal· International Journal of Res...· 0 citations
Affective computing has emerged as a cornerstone of human-computer interaction (HCI), healthcare analytics, and digital mental health monitoring. Traditional emotion recognition frameworks heavily rely on unimodal architectures—analyzing either text logs, facial expressions, or acoustic patterns in isolation. However, unimodal systems are inherently prone to environmental noise, semantic ambiguities, and cross-channel context blindness, which restrict their real-world reliability. This paper presents a systematic review of contemporary advancements in automated mood swing analysis, focusing on the evolution from handcrafted unimodal classifiers to deep-learning-driven multimodal architectures. We dissect the structural components of feature extraction across linguistic, visual, and acoustic domains, evaluate Early, Late, and Hybrid fusion mechanics, and analyze the deployment bottlenecks in transitioning from complex, resource-intensive models to lightweight web-based frameworks. Finally, we highlight critical gaps in current literature, particularly regarding the handling of cross-modal emotional inconsistencies and real-world framework deployments.
Mayuri More, Shridevi Amol Nandi· International journal of res...· 0 citations