Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.
Driver distraction is a major road-safety concern that requires reliable and efficient in-vehicle monitoring systems. The main contribution of this work is a reproducible driver-disjoint and deployment-oriented evaluation framework that jointly examines unseen-driver generalization, lightweight model benchmarking, explainability, calibration, and embedded inference. Experiments on the State Farm Distracted Driver Detection dataset show that MobileNetV3-Large provides the best trade-off among the evaluated lightweight models, achieving 88.92% test accuracy, 89.02% balanced accuracy, 88.07% macro-F1, and 97.88% Top-3 accuracy on unseen drivers. Explainable AI analysis indicates that the model mainly focuses on behavior-relevant regions, including the hands, face, phone area, steering wheel, and upper-body posture. For embedded deployment, TensorRT optimization on the Jetson Orin Nano Super achieved 212.94 FPS with 4.67 ms end-to-end latency in FP16 mode. These results demonstrate a practical balance between unseen-driver generalization, interpretability, and real-time embedded inference.
Siham Essahraui, Chaymae Rami, Khalid El Makkaoui et al.· Scientific Reports· 0 citations