Skip to content

Toward Trustworthy Dynamic Facial Expression Recognition via Information Bottleneck Modeling

2026 · IEEE Transactions on Information Forensics and Security · Vol 21, pp. 6985-6999 · 1 citation · 59 references
Computer Science

Abstract

Due to the presence of semantic ambiguity among similar expression categories and the inherent imbalance in spatio-temporal feature intensities, dynamic facial expression recognition (DFER) in the wild poses significant challenges for building trustworthy and robust systems. These factors often lead to inconsistent feature representations and unreliable decision boundaries, which hinder the model’s ability to perform stable and accurate recognition under uncertainty, and further pose a serious safety hazard, e.g., misdiagnosis of depression. To tackle these challenges, we propose a novel adaptive framework, Semantic-Aware Facial Expression Recognition framework (SAFE), which is developed from an Information Bottleneck (IB)-inspired perspective to improve the robustness and prediction reliability of DFER in complex, unconstrained scenarios. Specifically, we first design a Temporal-aware Augmentation Module (TAM) to introduce structurally perturbed yet temporally coherent training samples, effectively mitigating spatio-temporal feature imbalance. Then, to ensure stable long-range modeling under temporal variation, we introduce the Spatio-temporal Modeling Module (STM) with a sparsity-aware state-space fusion gate. Furthermore, an Ambiguity-aware Calibration Loss (ACL) is formulated to dynamically refine decision boundaries by focusing on confusing and underrepresented categories, improving the model’s resilience to distributional skew and semantic uncertainty. Extensive experiments on two large-scale in-the-wild DFER benchmarks, DFEW and FERV39k, demonstrate that SAFE consistently outperforms state-of-the-art methods across multiple metrics, particularly under ambiguous and imbalanced conditions. These results validate the effectiveness of our approach in promoting more robust and stable expression recognition, which is important for trustworthy DFER in real-world environments. Codes are released at https://github.com/QIcita/SAFE_DFER

View source

Similar papers

Open access Sep 2026

PCD-Net: prior-context collaborative learning for robust facial expression recognition under noisy labels

Noisy annotations and class imbalance commonly coexist in real-world facial expression recognition (FER), making hard but correctly labeled minority-class samples difficult to distinguish from unreliable training samples. Existing noisy FER approaches mainly improve robustness through sample selection or re-weighting...

Chang-Shuang Wang, Wen-Zhong Yang, Ya-Bo Yin et al. · 0 citations
Open access Aug 2026

Transformer With Decoupled Self-Attention Regularization for Age-Unbiased Facial Expression Recognition

A novel bias-mitigation method that decouples and regularizes age- and emotion-related components within the self-attention mechanism of Transformer to reduce age-related bias and enhance age-invariant emotion separation is proposed.

Jaeil Park, Sung-Bae Cho · 0 citations
Open access

Integrating Domain Knowledge for Robust and Interpretable Deep Neural Networks in Facial Affect Recognition

Facial affect recognition is a key component of human-centered AI, enabling systems to respond appropriately to human nonverbal signals. The deployment of deep learning for facial affect recognition is critically hindered by vulnerability to data-driven biases and a lack of transparency. This thesis addresses these cha...

Ines Rieger · 0 citations
Sep 2026

GCA-ODN: A global context-aware dropout network for joint facial landmark detection and emotion recognition under occlusion.

We propose the Global Context-Aware Dropout Network (GCA-ODN), a CNN-based, computationally practical neural architecture for joint facial landmark detection (FLD) and facial expression recognition (FER) under partial facial occlusion. GCA-ODN learns a shared embedding that encodes facial geometry and affective cues, i...

Muhammad Sadiq · 0 citations
Open access Sep 2026

Model Capacity for Small-Sample Facial Expression Recognition

A stability-oriented convolutional study that rethinks model capacity for seven-class small-sample FER and combines stratified partitioning, grayscale normalization, compact VGG-style representation learning, global average pooling, label smoothing, dropout, L2 regularization, and momentum-based optimization to control...

Yang-Xuan Xie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.