Skip to content

Author

Yutian Zhang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

From Transformers to World Simulators: A Survey on Generative AI Video Production Technology

The rise of generative artificial intelligence has triggered a paradigm shift in the field of video production. This survey systematically examines the evolutionary progress of generative AI video production technologies, with a focus on analyzing the transition from CNN architectures to Transformer-dominated frameworks. Adopting a systematic literature review methodology, this study analyzes relevant academic papers and technical reports. Centered on three core research questions, this survey explores: (1) What architectural innovations have enabled the shift from CNN-based to Transformer-based video generation? (2) What are the current capabilities and limitations of cutting-edge models such as Sora? (3) What challenges and future directions lie ahead for this rapidly evolving field? Key findings demonstrate that the self-attention mechanism of Transformers fundamentally addresses the inherent long-range temporal dependency problem of CNNs, reducing temporal consistency error from approximately 42% for CNNs to roughly 15% for Transformers. Nevertheless, substantial challenges persist in multi-modal consistency, computational resource requirements (approximately 10¹⁸ FLOPs for generating a single 10-second video), and ethical concerns. This survey identifies world models and interactive generation as promising research frontiers for the domain.

Yutian Zhang · 0 citations
Preprint Aug 2026

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

DECOWAM is introduced, a whole-body world-action model that separates camera ego-motion from base and arm actions through dedicated conditional interfaces and shows that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.

Siyuan Ma, Boshi Zhang, Yutian Zhang et al. · 1 citation