Face recognition systems are widely used in surveillance, biometric authentication, access control, and digital identity verification; however, supervision sensitivity, evaluation stability, and performance consistency across datasets remain insufficiently understood. This study investigates the behavior of convolutional, transformer-based, and hybrid face recognition architectures under both Softmax and ArcFace supervision using five-fold subject-disjoint cross-validation on the Labeled Faces in the Wild (LFW) and FAGEv2 datasets. ResNet50, MobileNetV3, DeiT-Small, and a Hybrid multi-branch architecture integrating complementary convolutional and transformer feature representations were evaluated using Top-1 identification accuracy, Area Under the ROC Curve (AUC), Equal Error Rate (EER), computational complexity, and fold-level statistical analysis. Experimental results revealed substantial supervision sensitivity across architectures and datasets. On the LFW dataset, Hybrid-Softmax achieved the highest Top-1 identification accuracy (62.4%), while DeiT-Small-Softmax achieved the strongest verification performance with an AUC of 0.905 and EER of 0.159. On the FAGEv2 dataset, Hybrid-Softmax and DeiT-Small-Softmax achieved the highest identification accuracy (38.0%), while Hybrid-Softmax achieved the strongest verification performance with an AUC of 0.825 and EER of 0.251. Fold-level analyses demonstrated that the effect of ArcFace supervision varied across architectures and datasets, with consistent improvements observed for some convolutional architectures but not for transformer-based or hybrid models. Cross-dataset evaluation further revealed changes in model ranking and supervision behavior, indicating that comparative performance is strongly influenced by dataset characteristics and evaluation conditions. The findings demonstrate that additive angular margin supervision does not universally outperform conventional Softmax optimization and highlight the importance of multi-dataset benchmarking, fold-level evaluation, and supervision sensitivity analysis for robust and reproducible face recognition benchmarking.
Andisani Nemavhola, C. Chibaya, Serestina Viriri· Frontiers in Artificial Inte...· 0 citations
Face analysis systems are widely used in security, authentication, and public-sector applications; however, demographic bias and the statistical reliability of reported performance remain key concerns. Many studies rely on aggregate accuracy without quantifying subgroup disparities or uncertainty, potentially overstating model fairness. This study presents a statistically grounded evaluation of demographic bias in face attribute classification across three representative architectures, ResNet50, MobileNetV3, and a vision transformer (DeiT), using the FairFace and UTKFace datasets. Subgroup analysis is conducted across race and gender, incorporating disparity indices, bootstrap confidence intervals, and inferential statistical testing with effect size analysis. The evaluation uses an embedding-based nearest-neighbor approach to examine representation-level behavior consistently across models. Results show that race-based disparities are substantially larger than gender-based disparities across both datasets. On FairFace, race disparity gaps range from 0.1124 to 0.1266, while on UTKFace they increase significantly to 0.4726–0.4944, with large effect sizes (Cohen's d>1). In contrast, gender disparities remain smaller, with gaps between 0.0280 and 0.0582 on FairFace and 0.0194–0.0326 on UTKFace, and correspondingly small effect sizes (d < 0.13). Despite modest differences in overall accuracy across models, subgroup disparities remain statistically significant across all architectures. These findings emphasize the importance of subgroup-level evaluation, uncertainty quantification, and statistical validation for reliable fairness assessment in face analysis systems.
Andisani Nemavhola, Serestina Viriri, C. Chibaya· Frontiers in Artificial Inte...· 0 citations