Skip to content
Conference

RCF-Net: Degradation-Aware Hybrid CNN–Transformer for Child Face Identification in Surveillance

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 1209-1218 · 0 citations · 21 references

Abstract

Child face identification from surveillance video remains difficult because facial crops are frequently low-resolution, blurred, partially occluded, and captured under unstable illumination. Age-related facial variation further increases the difficulty of maintaining discriminative identity embeddings for children. This paper presents RCF-Net, a degradation-aware hybrid CNN–Transformer architecture that combines surveillance-oriented image degradation, dual-branch local/global feature extraction, and learnable cross-attention fusion. MTCNN is used for face detection and alignment, ArcFace supervision is used for discriminative embedding learning, and DeepSORT can optionally be integrated to improve temporal identity consistency in video streams. To address deployment concerns raised by surveillance use, the revised framework also specifies age-progression handling, latency-aware scheduling for live video, multi-camera scaling, and adversarial/spoof-risk safeguards. Experiments are conducted using public face datasets, namely VGGFace2, CASIA-WebFace, CelebA, and IMDB-WIKI, with child-oriented filtering and synthetic surveillance degradations. Compared with representative CNN, transformer, and hybrid baselines, RCF-Net achieves the best overall accuracy of 91.4% and yields the strongest robustness under low-resolution, blur, and occlusion stress tests. The results indicate that explicit degradation modeling and local-global feature fusion are complementary for surveillance-oriented child face identification.

View source