Skip to content
Conference

Scene-Level Planning for Temporally Coherent AI Video Generation Using Multimodal Representations

Aug 2026 · 2026 International Conference on Intelligent Multimedia, Networking, and Security (IMNS) · pp. 1-6 · 0 citations · 12 references

Abstract

Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. This frequently leads to temporally inconsistent or structurally ambiguous outputs. In this work, we investigate whether introducing an explicit scene-planning layer can improve multi-stage video generation. We compare three generation paradigms: direct single-prompt generation, naive prompt decomposition, and a structured scene-planning pipeline. The proposed approach first converts a narrative prompt into a lightweight structured scene representation containing global subject information, visual style constraints, and scene-level descriptions, which is then compiled into scene-conditioned prompts for sequential video generation. A reference-guided continuation mechanism conditions the second scene on the final frame of the first scene to improve cross-scene identity and visual continuity. The framework is generator-agnostic and can operate on modern text-to-video backends without modifying the underlying models. To evaluate generation quality, we adopt a vision-language model (VLM) as an automatic judge that assesses prompt relevance, temporal continuity, aesthetic quality, and narrative clarity across candidate videos. This study provides an empirical investigation of how structured intermediate planning influences narrative video generation and offers a lightweight framework for improving temporal coherence in AI-generated videos.

View source