Geometry and mask aware vision transformer for masked face recognition in unconstrained scenarios
Abstract
The concealed facial features make it difficult to identify masked faces. The features are further distorted and deteriorated when masked faces are combined with low resolution and pose variation in unrestricted contexts. Pose variation and low-resolution circumstances, along with face masks, are not assessed for current masked face recognition systems. We developed a transformer-based model a mask-aware geometry-guided vision transformer (MGViT), to address these issues. First, a learnable geometry-guided patch weighting (LGW) is used in the proposed model to suppress occluded regions and concentrate on the key face regions. Second, a mask-aware feature adaptor is created to improve the domain embeddings between masked and unmasked faces. Following that, by combining Identity and Consistency Loss functions to align identity, a strong consistency learning is integrated. Experiments with various challenges are carried out on the various masked face datasets. The proposed model, MGViT, performs well in identifying and verifying low-quality and cross-pose masked faces. Additionally, the model achieves 94.70%, 95.25%, 87.67%, and 78.56% accuracy on RMFRD, Masked LFW, Masked Multi-PIE, and Masked LR, respectively, outperforming the various state-of-the-art methods and previously proposed methods. Facial identity identification in unrestricted real-world environments may benefit from this model.