From Transformers to World Simulators: A Survey on Generative AI Video Production Technology
Abstract
The rise of generative artificial intelligence has triggered a paradigm shift in the field of video production. This survey systematically examines the evolutionary progress of generative AI video production technologies, with a focus on analyzing the transition from CNN architectures to Transformer-dominated frameworks. Adopting a systematic literature review methodology, this study analyzes relevant academic papers and technical reports. Centered on three core research questions, this survey explores: (1) What architectural innovations have enabled the shift from CNN-based to Transformer-based video generation? (2) What are the current capabilities and limitations of cutting-edge models such as Sora? (3) What challenges and future directions lie ahead for this rapidly evolving field? Key findings demonstrate that the self-attention mechanism of Transformers fundamentally addresses the inherent long-range temporal dependency problem of CNNs, reducing temporal consistency error from approximately 42% for CNNs to roughly 15% for Transformers. Nevertheless, substantial challenges persist in multi-modal consistency, computational resource requirements (approximately 10¹⁸ FLOPs for generating a single 10-second video), and ethical concerns. This survey identifies world models and interactive generation as promising research frontiers for the domain.