The Prior-aware Sparse Transformer (PAST) is proposed, a novel Transformer-based architecture integrating spatial and semantic prior knowledge synergistically into a sparse attention mechanism, enabling linear-complexity processing of large amounts of fine-grained features.
Abstract
Micro-expression recognition (MER) has a lot of applications in lie detection, education, healthcare, etc., as involuntary micro-expressions (MEs) may provide subtle facial cues associated with affective responses. With the development of deep learning, many studies have recently employed Vision Transformers (ViTs) to investigate MER, since ViTs show promising performance in various visual domains due to their excellent local–global modeling ability. However, such methods confront two fundamental challenges: First, fine-grained visual features are needed to capture the subtle facial movements of MEs, which ViTs relatively fall short on due to coarse patch resolution constrained by their quadratic complexity. Second, the data-intensive nature of ViTs impedes effective learning given the limited scale of ME data. To overcome the aforementioned limitations of using ViTs for MER, we propose the Prior-aware Sparse Transformer (PAST), a novel Transformer-based architecture integrating spatial and semantic prior knowledge synergistically into a sparse attention mechanism, enabling linear-complexity processing of large amounts of fine-grained features. Specifically, we first designed an extraction algorithm to generate a representative set of motion-intensive Principal Anchors, which are used to guide the model’s focus on biologically critical regions during sampling. Second, we introduced the Semantic Dictionary, which was trained with a carefully designed self-contrastive loss to embed task-invariant discriminative semantics of the anchors. Such global semantics further modulate patch sampling and attention weighting in the sparse attention procedure, achieving better training performance with limited ME data. Extensive evaluations on MEGC and CD6ME protocols demonstrate state-of-the-art performance, validating PAST’s efficacy for MER.
A lightweight FER model that combines a truncated MobileNetV2 backbone with a patch-based local feature extraction module and a channel-attention refinement module and a channel-attention refinement module, followed by a compact classifier is proposed, outperforming many existing lightweight models.
Anh Viet Vu, Khanh Gia Pham, Nam Quy Tran et al.· Journal of Ambient Intellige...· 0 citations
Micro-expression recognition (MER) is a challenging step in various multimedia applications, such as media understanding and human–computer interaction, as it can reveal genuine human emotions. However, traditional MER often overlooks person-specific facial nuances, limiting generalization and personalized adaptability...
Rui-Qi Wang, Ke-Rong Li, Jiateng Liu et al.· Italian National Conference...· 0 citations
: Micro-expression recognition (MER) is a challenging task because micro-expressions are extremely short in duration, weak in intensity, and often distributed over subtle local facial regions. Existing methods either rely on handcrafted descriptors with limited representation capacity or focus on single-stream deep mod...
Zishi Li, Xiao-Dong Huang· Computers, Materials & C...· 0 citations
It is hypothesized that the FER task does not necessarily require all facial information to correctly interpret emotional states, as specific regions such as the eyes, the mouth, and parts of the cheeks carry discriminative information that can be sufficient to recognize emotions.
Aya Manel Zitouni, Aicha Zenakhri, Karim Haroun et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.