Skip to content
Review

A Comprehensive Review of Multimodal Facial State Analysis: Tasks, Methods, and Resources

Sep 2026 · 0 citations · 55 references
Computer Science

TL;DR

This survey reviews core tasks, representative methods, and datasets in multimodal facial state analysis, focusing on facial expression recognition, AU detection, and face-based soft biometric estimation, and emphasizing the unique value of language in providing contextual semantics, enhancing reasoning, and generating explanations.

Abstract

Facial state analysis plays a crucial role in understanding human expressions, psychological modeling, and human computer interaction. Traditional unimodal vision-based methods are often limited by environmental sensitivity and weak interpretability. Multimodal facial state analysis addresses these issues by integrating complementary cues from visual, audio, textual, physiological, and other related modalities. This survey emphasizes two key aspects: on one hand, multimodal learning enables contextual semantic understanding for improved facial state reasoning and leverages interpretable language generation to enhance model explainability; on the other hand, multi-task learning allows simultaneous analysis of expressions, action units (AUs), and face-based soft biometrics (e.g., age, gender), effectively capturing fine-grained expressions and improving cross-scene generalization. This survey reviews core tasks, representative methods, and datasets in multimodal facial state analysis, focusing on facial expression recognition, AU detection, and face-based soft biometric estimation, and emphasizing the unique value of language in providing contextual semantics, enhancing reasoning, and generating explanations. The survey aims to provide an up-to-date overview of the literature and to highlight future research directions for multimodal, interpretable, and multi-task adaptive facial state analysis.

View source

Similar papers

Open access Sep 2026

Multimodal Emotion Recognition System Using Facial Expressions and Speech Analysis

Human communication involves emotion as an important issue. Because of emotion, people express and understand feelings, attitudes, and actions. However, effective automatic emotion recognition is still difficult to achieve when based on one single cue, as for instance facial expressions or voice may both present weak o...

S. G., N. S · 0 citations
Open access Aug 2026

An Attention-Enhanced ConvNeXtTiny Model for Robust Facial Expression Recognition

Facial Expression Recognition (FER) plays an important role in affective computing and human–computer interaction by enabling automated interpretation of human emotional states from facial images. Despite recent advances in deep learning, reliable FER remains challenging because of variations in facial appearance, illu...

Manisha B. Thombare, S. Gumaste · 0 citations
#computer vision Preprint Sep 2026

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face vi...

Birgit Nierula, Karam Tomotaki-Dawoud, M. Akguel et al. · 1 citation
#graph neural networks Open access Aug 2026

Facial expression recognition using a non-exclusive learning search-Fossa optimization algorithm with a convolutional neural network

A non-exclusive learning search-Fossa optimization algorithm integrated with a convolutional neural network (NELS-FOA-CNN) is proposed to select the most relevant features for accurate FER, demonstrating improved performance compared with the baseline GCN.

Aswini Vadladi, Kavitha Valasa, Sruthi Kandukuri et al. · 0 citations
Open access Sep 2026

Model Capacity for Small-Sample Facial Expression Recognition

A stability-oriented convolutional study that rethinks model capacity for seven-class small-sample FER and combines stratified partitioning, grayscale normalization, compact VGG-style representation learning, global average pooling, label smoothing, dropout, L2 regularization, and momentum-based optimization to control...

Yang-Xuan Xie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.