Facial aging simulation has become an essential instrument in various fields, and some of them are forensic medicine, cosmetic surgery, entertainment, and mental health. Within the scope of criminal investigation, age progression approaches help in forecasting how missing persons might look like in the future. This paper presents the proposed Hybrid Age GAN framework, termed MP-CGAN (Multi-Phase Conditional Generative Adversarial Network), developed for generating realistic, progressive facial aging. The proposed model is built upon a conditional GAN backbone, incorporating identity-preserving representation, age- and gender-conditioned feature modulation via Adaptive Instance Normalization, residual aging blocks for gradual structural transformation, and self-attention mechanisms targeting aging-sensitive facial regions. A multi-scale discriminator operating at three resolutions (128×128, 64×64, 32×32) enforces realism at both global and fine-grained levels. In contrast to approaches that focus solely on architectural design, this research primarily addresses training stability as a fundamental factor in achieving realistic aging. A multi-phase progressive training strategy, supported by N_CRITIC update scheduling and label smoothing, was adopted to systematically resolve the discriminator collapse problem D_DEAD. The model was trained on the UTK Face dataset and used 5,500 balanced face images covering 11 age groups, including both genders, over 100 training epochs. Experimental results demonstrate strong performance with a Frechet Inception Distance (FID) of 27.78, Structural Similarity Index (SSIM) of 0.6580, Peak Signal-to-Noise Ratio (PSNR) of 19.72 dB, Cosine Identity Similarity (CSIM) of 0.7818, and Mean Absolute Error (MAE) of 7.09 years at the age bucket classification level.
Amina Taha Alazwe, Y. Mohammad· IEEE Jordan Conference on Ap...· 0 citations
Scene text detection and recognition in multilingual environments remains challenging, particularly for morphologically complex scripts such as Arabic. This paper presents an end-to-end deep learning system for detecting and recognizing Arabic and English text in natural scene images. The detection pipeline combines a Swin Transformer Tiny backbone pretrained on ImageNet-22K, a Feature Pyramid Network for multi-scale fusion, and a DBNet++ head, achieving an F1-score of 94.52% on a strictly held-out 2,000-image ICDAR 2019 MLT test set. For recognition, PARSeq with permutation language modeling is trained on 656,868 cropped word samples (85/15 split), reaching 89.50% word accuracy and 94.13% character accuracy on an 865-character bilingual charset. The system includes a vertical-projection word-segmentation fallback and a smart RTL/LTR ordering algorithm. Ablations show FPN contributes +4.6 pp and DBNet++ +3.22 pp F1. Comparisons with prior baselines are reported as non-comparable references. The contribution is a fully reproducible bilingual Arabic-English pipeline.
Taher Ali Mahmood, Y. Mohammad· IEEE Jordan Conference on Ap...· 0 citations