Through comprehensive evaluations and analysis, it is shown that LFD effectively captures shared features and artifacts across different lensless imaging devices, making it a valuable dataset for advancing lensless face recognition.
Abstract
Face recognition is a ubiquitously used computer vision task that has a wide range of applications ranging from everyday smartphone biometrics to high-stakes security systems. Most face recognition systems rely on traditional cameras, which often suffer from limitations such as bulky form factors, high costs, and limited privacy protection. To address these limitations, lensless cameras have emerged as an alternative. Lensless cameras use thin optical encoders, enabling smaller size, lower cost, and greater design flexibility. These cameras are typically paired with reconstruction algorithms that convert raw captures into recognizable images. However, reconstructed images often contain artifacts, and the reconstruction methods struggle to generalize well to real-world conditions. Furthermore, existing face datasets do not account for the artifacts present in lensless images. To address this issue, we introduce the Lensless Face Dataset (LFD). LFD comprises 21,080 lensless raw measurements, reconstructions, and standard images of faces captured under diverse lighting, angle, and distance. Our key contributions are: (1) Real-world lensless face data: LFD focuses on capturing a diverse face dataset with varying levels of artifacts introduced under different environments; (2) In-the-wild captures: 4,976 images are captured in outdoor settings with varying intensities of natural light and different background patterns; (3) Multiple lensless devices: LFD includes face images collected from three different types of lensless cameras, each with a unique optical encoder. We use this hardware diversity to demonstrate generalization across different lensless cameras. Through comprehensive evaluations and analysis, we show that LFD effectively captures shared features and artifacts across different lensless imaging devices, making it a valuable dataset for advancing lensless face recognition.
Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks. With the rising emphasis on data privacy and responsible AI, face De-ID has emerged as an active research area spanning computer vision and privacy-preserving communities. Early approaches, and many contemporary ones, operate in the digital domain by modifying pixel-level or appearance-level features through post-capture processing. Recent advances extend face De-ID beyond post-processing by integrating privacy mechanisms directly into sensors during image acquisition, bridging sensing systems and downstream vision algorithms. In parallel, physical-domain methods explore wearable accessories and materials that conceal identity information in real-world environments prior to capture. In this survey, we present the first unified overview that spans the full data acquisition pipeline, encompassing the physical, sensor, and digital domains. Through this domain-centric lens, we systematically analyze current methodologies, technical progress, and the distinct challenges inherent to each stage. We then review and organize existing evaluation protocols, examining current practices and highlighting the critical need for standardized, comprehensive benchmarks. Finally, we identify key open problems and outline emerging research directions to guide future work in this rapidly evolving field. To support ongoing research, we maintain a project page that organizes relevant literature with collected datasets and open source code: https://github.com/CV-AC/Awesome-FaceDe-ID.
In forensic environments, automated identification of perpetrators is difficult due to pose changes, changes in light, occlusion, and lack of labeled data. This paper presents ForensicNet, a lightweight deep learning framework for forensic face recognition that enhances attention. The suggested model combines the MobileNetV2 backbone with Convolutional Block Attention Modules (CBAM) to improve the learning of discriminative features while maintaining computational speed. A two-phase transfer learning strategy with adaptive layer unfreezing is used to improve domain adaptation and reduce overfitting. This study used publicly available datasets such as LFW and SCFace, with 15,000 facial images spanning 68 identity classes. The proposed model outperforms baseline architectures such as AlexNet, ResNet-50, and MobileNetV2, with an accuracy of 92.4%, a precision of 90.8%, and a recall of 89.5%. Additionally, the framework requires only 2.1 GFLOPs per inference, and hence can be used in real-time forensic surveillance applications.
S. N. J., L. B. T.· Engineering, Technology &...· 0 citations
Missing person identification using surveillance imagery remains a challenging problem due to adverse visual conditions, limited availability of reference images, and strict privacy constraints surrounding real-world data. This paper presents a hybrid deep-learning framework for missing person identification that combines YOLO-based person detection with deep face embedding models, specifically FaceNet and ArcFace. To address ethical and privacy limitations, a manually curated composite synthetic dataset is constructed by combining publicly available crowd and in-the-wild face datasets with additional curated images to realistically emulate CCTV conditions, including low illumination, occlusions, visually similar individuals, accessories, and clothing-matched decoys, while restricting each identity to only three to four reference images. The proposed system is evaluated across five YOLO variants (v8–v12) and a wide range of cosine similarity thresholds to analyze detection sensitivity, false positive behavior, and overall identification robustness. Extensive experiments conducted on 780 group images demonstrate that the YOLO–ArcFace pipeline achieves superior performance, reaching a peak identification accuracy of 97.50% with near-zero false positives, while the YOLO–FaceNet pipeline attains a maximum accuracy of 95.42% at an optimized operating threshold but exhibits higher susceptibility to false matches. Threshold–accuracy analysis and operating point comparison reveal that ArcFace embeddings provide stronger inter-class separation and greater stability under surveillance-specific distortions. The results highlight the importance of discriminative embedding models in safety-critical identification tasks and demonstrate that reliable missing person detection is achievable even under constrained and privacy-preserving settings, with future extensions proposed toward continuous video-based tracking using multi-object tracking algorithms.
G. S. Rakshika, Babyrani Waikhom, U. Muthaiah· International Conference Com...· 0 citations
The increasing use of vlog videos on social media creates privacy risks because third-party faces are often unintentionally recorded and distributed without consent. Existing face blurring approaches generally apply uniform anonymization to all detected faces and do not provide an identity-selective mechanism that keeps the content creator visible while blurring other individuals. This study develops PRIVA, a desktop-based selective face blurring application that runs locally without an external AI server. The proposed pipeline integrates YOLOv8n-Face-960 for face detection, MobileFaceNet for face recognition using 512-dimensional embeddings, and Deep SORT for maintaining identity consistency across video frames. Face enrollment is performed through guided multi-pose webcam capture, while video evaluation is conducted on extracted YOLO analysis frames from five real vlog-like test videos. YOLOv8n-Face-960 achieved an overall detection precision of 95.02%, recall of 89.32%, and F1-score of 92.09%. The baseline comparison showed that YOLOv8n-Face-960 achieved a higher mean detection F1-score than MTCNN, while MobileFaceNet provided a smaller and faster recognition model than FaceNet for CPU-based local inference. For correctly detected face instances, PRIVA achieved a system precision of 99.45%, recall of 98.70%, F1-score of 99.08%, and accuracy of 98.50% in determining whether faces should be blurred or kept visible. Processing performance testing showed an average analysis speed of 4.83 FPS, average export speed of 70.05 FPS, and average processing ratio of approximately 2.40 times the original video duration. These results indicate that PRIVA can support practical local identity-selective face blurring for video privacy protection, although detection robustness remains important under low-light, crowded, distant, or partially occluded face conditions.
Muhammad Satrio, Mohammad Nasucha· SinkrOn· 0 citations
A survey of demographic attribute estimation from periocular images, covering publicly available datasets, methodological trends from handcrafted descriptors to deep learning architectures, and the state of the art in gender, age, and ethnicity prediction is provided.
F. Alonso-Fernandez, Kevin Hernandez-Diaz, J. Bigun· 0 citations
Face recognition in real-world and uncontrolled environments is greatly affected when the face is partially covered by masks, glasses, scarves, hair, or due to pose-related self-occlusion. Although three-dimensional (3D) face recognition is generally more robust to lighting changes and moderate pose variations, its performance still reduces when important facial regions are heavily covered. To overcome this problem, this study presents an occlusion-aware hybrid biometric framework for reliable 3D face recognition. The proposed method combines reconstructed 3D shape features, deep texture features, and additional biometric cues using an adaptive weighted fusion approach. An occlusion detection and generative reconstruction module is used to recover missing facial regions before feature extraction, and an attention mechanism reduces the impact of unreliable areas while focusing on important facial features. Extensive experiments on benchmark datasets, such as Bosphorus, BU-3DFE, and FRGC v2.0, show that the proposed framework performs better than strong single-modality and traditional multimodal methods. The proposed system reaches an accuracy of up to 98.7%, even in partial occlusions, and significantly reduces the Equal Error Rate (EER), demonstrating its effectiveness and suitability for real-world biometric authentication applications.
M. L. Gangadhar, A. S. Raju, C. R. Roopashree· Engineering, Technology &...· 0 citations