Skip to content
Review Open access

A Comprehensive Survey of Generative AI: Applications and Future Directions Across Domains

2026 · IEEE Access · Vol 14, pp. 129914-129941 · 0 citations · 208 references

TL;DR

The review shows that diffusion and autoregressive foundation models increasingly dominate high-fidelity image, language, and multimodal generation, while GANs, VAEs, and flow-based models remain important in data-limited, structured, scientific, and privacy-aware settings.

Abstract

Generative Artificial Intelligence (GenAI) has become a foundational paradigm for learning complex data distributions and synthesizing realistic content across text, image, audio, tabular, scientific, and multimodal data. Despite rapid progress, the literature remains fragmented across model families and application domains, making it difficult to compare methodological choices, evaluation practices, and domain-specific deployment constraints. This survey addresses this gap by reviewing GenAI from both an algorithmic and cross-domain perspective. Following a PRISMA-inspired search and screening procedure, we organize representative and technically relevant studies across major GenAI families, including variational autoencoders, generative adversarial networks, diffusion models, autoregressive models, flow-based models, and recent foundation and multimodal models. We synthesize applications across healthcare, agriculture, manufacturing, transportation, earth and environmental systems, computer science and networks, materials science, finance and economics, education, business and services, and creative industries. The review shows that diffusion and autoregressive foundation models increasingly dominate high-fidelity image, language, and multimodal generation, while GANs, VAEs, and flow-based models remain important in data-limited, structured, scientific, and privacy-aware settings. Beyond summarizing applications, the paper compares domain-specific data structures, evaluation practices, robustness concerns, failure modes, and ethical risks. The survey contributes a unified reviewed abstraction of GenAI workflows, a comparison with prior surveys, a domain-aware synthesis of open gaps, and a future roadmap for trustworthy, validated, and responsible GenAI deployment.

Read PDF

Similar papers

Review Open access 2026

Subject Review: Generative Adversarial Networks from Architectural Foundations to Future Trajectories

An in-depth and up- to-date overview of the GANs environment, principally highlighting the progress made over 2020 and beyond and proposing the idea of hybrid generative systems in the future while emphasizing the oppositional approach's extraordinary and enduring features.

Zahraa Salah Dhaif, Hind Jumaa Serteep · 0 citations
Review Open access Aug 2026

Challenges and opportunities of generative artificial intelligence models in audio/acoustic domain: a comprehensive survey

Generative artificial intelligence (AI) has transformed image and text processing, but its adoption in the audio/acoustic domain remains underexplored due to inherent challenges in modeling long-term temporal dependencies in one-dimensional signals and achieving human-perceptible coherence. This comprehensive survey addresses these gaps by systematically reviewing state-of-the-art generative AI models, including generative adversarial networks (GANs), diffusion/flow-matching models, variational autoencoders (VAEs), recurrent neural networks (RNNs), transformers, and Neural Codec Language Models (Codec LMs), organized around three primary application domains: (1) speech synthesis , encompassing text-to-speech conversion, neural vocoding, voice conversion, and zero-shot voice cloning; (2) music generation , covering both symbolic and acoustic composition, multi-track generation, and style transfer; and (3) general audio synthesis, sound effects, and source separation , including text-to-audio generation, audio restoration and enhancement, data augmentation, and conditional source separation. We provide a detailed taxonomy of architectures, functionalities, comparative strengths, and limitations, supported by common evaluation metrics. Structured comparisons with existing surveys demonstrate that this is the first work to jointly cover all three audio domains and all generative model families, while also providing dedicated evaluation-metric analysis and cross-domain comparative assessments. We further highlight emerging opportunities in AI applications such as healthcare monitoring (e.g., symptom analysis and mental health assessment) and biometric authentication, demonstrating the potential of synthetic audio to address real-life challenges. Through research gap identification, model efficacy comparison, and future direction outlining, this survey serves as a foundational reference for advancing generative AI techniques across the audio domain.

S. Dibbo, Sudip Vhaduri, Chia-Hua Lin · 0 citations
Review Jul 2026

History, Development, and Principles of Representation Learning—An Introductory Survey

This survey deeply explains the basic principles of representation learning, and introduces its practical application cases in various fields, and points out the main limitations of current models and prospects the future research directions.

Zhiyong Wang, Qiang He, Jun Mou et al. · 0 citations
Review Open access 2026

Synthetic Data Quality Evaluation in Generative AI: Current Trends, Challenges, and Future Directions for Social Science Research

The findings indicate that while modern generative models can produce highly realistic and analytically useful datasets, persistent challenges remain, including the lack of standardized benchmarking protocols, utility–privacy trade-offs, privacy leakage risks, bias amplification, limited explainability, and governance concerns.

N. Emran, Ruhaila Maskat, Abdulrazzak Ali · 0 citations
Review Open access Jul 2026

Text-to-Image Generation via Deep Learning: A Comprehensive Review of Models, Architectures, and Future Directions

Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.

Abdussalam Elhanashi, Siham Essahraui, Qinghe Zheng et al. · 0 citations